Skip to content
v2.15.0MIT

An execution posture for substantial agent work: parallel agents in phases and waves, adversarial red-teaming, looping until confident, ruthless simplicity. Fourteen skills: a shared doctrine, wrappers for coding, web design (the gauntlet loop), debugging, auditing, documentation sweeps, deep research, document writing, and project and epic tracking (doctrine-project), plus doctrine-backup, doctrine-handoff, doctrine-resume and doctrine-primer for session continuity across cleared sessions, and doctrine-pane for an interactive terminal session in a pane the user watches and can take over. Composes with (and degrades gracefully without) Matt Pocock's engineering skills, superpowers, codex, ponytail, and Claude Design. Two prerequisites do not degrade, and both refuse rather than falling back to a weaker version of themselves: doctrine-gauntlet needs Playwright, because its critics judge rendered output — without a browser its harness refuses to run rather than report a pass — and doctrine-pane needs herdr, because an interactive session nobody can watch has no purpose. The other twelve skills need nothing installed.

Doctrine Skills

Fourteen skills for Claude Code that put an agent's work through a real gate: every check your project documents, an adversary that did not write the code, and a loop that does not exit until one full pass comes back clean.

The problem

You already get good work out of Claude Code by pushing on it: check that again, did the linter run, are you sure. That works, and you do it differently every time, so what you get back tracks how much attention you had that day. The failures that survive it do not look like failures.

  • The tests pass, the linter is clean, the build is green, and the code is wrong.
  • The context that wrote the code reviews it and finds it reasonable.
  • One defect gets fixed in the one place you were looking at.
  • The fix creates the next bug, and the round that would have found it is over.
  • A long run gets compacted, and the list of what still fails goes with it.

What the doctrine does

Seven rules, applied to every job the hub and its eight wrappers run. Rules 3 to 6 are a loop rather than a sequence, which is the point of rule 6.

  1. Ask first. Questions reach you before work is aimed, and your answers go verbatim to every agent that later judges the result.
  2. Split the work into phases with checkable exit gates, and run independent work in parallel.
  3. Run every check your project documents, not the ones an agent thought of. A check that could not run is recorded as not run, never as clean.
  4. Send an adversary: a reviewer that did not write the code, a different model if you have one installed. Its findings are verified against source first, because adversaries invent things too. Without a second model the run says so, rather than reporting a stronger gate than it ran.
  5. Ask where else. Every blocking finding that survives verification gets one question: what else the same mistake would have touched. The fix goes to the cause, not just the place it showed up.
  6. Loop until one full pass comes back clean with no blocking finding left over from an earlier pass. Two alarms stop the loop and put the decision to you rather than spending your budget without telling you.
  7. Cut what nobody asked for before the pass that certifies the work, so the cut is inside what that pass checked, then deliver by whatever route your project normally ships work, and say whether what shipped is what the clean pass certified.

What it touches. It edits your working tree, the way Claude Code already does. It commits the way your repo does, and where your repo documents no norm the coding workflow commits locally and asks. It pushes or opens a PR only if that is your documented norm or you asked for one. Where it cannot write, it hands you the diff instead. A phase writes its state to a file as it goes, so a run that is interrupted or compacted resumes from disk rather than from whatever the conversation still remembers.

What that catches

Three defects from real work. In each pair the first block is the defect and the second is the fix. The code is real and the identifiers are renamed: no client, product, person or repository is named, so these are mnemonics for the rule beside them, not citations you can follow.

Every check was green

const pairs = (o: Record<string, unknown>) =>
  Object.entries(o).map(([key, value]) => `${words(key)}: ${value === null ? '' : String(value)}`)
const show = (value: unknown) =>
  value === null ? '' : typeof value === 'object' ? JSON.stringify(value) : String(value)

String(x) throws when x is an object whose own toString is not callable, and custom field names are user-supplied. Name a field toString, open that record's history, and the page dies. It dies again on every reopen, because the event is stored.

At that moment: 837 backend tests green, 492 frontend tests green, a clean build, linters green apart from one pre-existing error the change did not cause, and a reviewer that read the whole 2,283-line diff end to end and returned nothing blocking. A different-model adversary read it and also returned nothing. The same adversary, run again on that unchanged revision with no memory of the first pass, found it.

One finding was a class

const globalItem = poolItems.find(g => g.id === item.itemId);
const defaultValue = globalItem?.value ?? 0;
onOverride(item.itemId, defaultValue, item.calcType);
const globalItem = poolItems.find((g) => g.id === itemId);
if (!globalItem) {
  setError('Re-include is unavailable: the pool could not be read. Reload and try again.');
  return;
}
const value = globalItem.value;

?? 0 is the defensive default a linter and a reviewer both read as correct null handling. It is wrong here because this zero is written to the database as a price: when a background reload failed, clicking "Re-include" on a $5,000 benefit saved it as $0, silently.

The different-model adversary found one button. Rule 5 asked what else could reach the same code, and the answer was five other handlers that already did, so the fix went into the shared function they all call rather than into the button the round happened to be looking at.

The fix was where the next bug lived

labels: list[str] = []
seen = {normalize(name) for name in header}
for label in custom_labels:
    if normalize(label) not in seen:
        seen.add(normalize(label))
        labels.append(label)
levels = {normalize(name) for name in LEVEL_COLUMNS}
labels = [label for label in custom_labels if normalize(label) not in levels]

De-duplicating a header against every name already in it is textbook correct, and it fixed the bug that round had reported. But the importer resolves custom labels above ordinary field names, so dropping the duplicate leaves the wrong column standing: export the catalog, import the same file back unedited, and the import silently overwrites a field on every product with an unrelated catalog value.

The first block is not the original code. It is the previous round's repair. Before it, the importer rejected that file outright and named the duplicate header, so the repair replaced a loud refusal with a silent write across the catalog. The next round's different-model adversary caught it by running the real exporter and importer rather than reading them.

How this differs from looping

Running an agent in a loop is not new and this does not claim it. Geoffrey Huntley's Ralph is an unconditional bash loop that pipes a prompt file into the agent over and over. Each pass gets a fresh context and progress lives on disk, a design this borrows rather than improves on.

What differs is what ends the loop.

Ralph's published loop has no exit condition, and the post documents no stopping rule. Clayton Farr's Ralph playbook adds one, an optional iteration cap. Anthropic's own Ralph plugin adds two, an iteration counter or a phrase the working agent emits about itself, and its README says to rely on the counter rather than the phrase.

Here the loop cannot reach a certified exit until a reviewer that did not write the work passes it, on a revision that has actually run, with no blocking findings left over from any earlier pass. It can still end other ways, and the report names which: stopped by you, shipped at an alarm with findings open, or closed with a punch list of what it did not fix. What it never does is present an uncertified ending as a clean one.

Independent review and gating are both documented practice rather than inventions here. Anthropic's own docs describe an adversarial reviewer in a fresh context, a stop hook that blocks a turn until a check passes, and a separate evaluator that keeps working until a goal resolves. What this adds is the combination: the reviewer is the gate, and the gate's conditions include the real run and the findings carried over from before.

Those conditions are also what lets a long run leave you alone. It stops at its own thresholds and comes back to you there, rather than needing you to watch for the end.

Scope. Your words are written down before any work is aimed and handed to every agent that later judges the result, and delivery walks your original request item by item.

Where Ralph is better. Greenfield. Huntley: "There's no way in heck would I use Ralph in an existing code base though, if you try, I'd be interested in hearing what your outcomes are. This works best as a technique for bootstrapping Greenfield, with the expectation you'll get 90% done with it." That is the opposite end of the problem from this one.

One number worth carrying, from Addy Osmani's case for adversarial review: across four review tools on 146 real pull requests, 93.4% of flagged locations were caught by exactly one tool, and none by all four. He draws the conclusion this page draws, that heterogeneity is the point, and attaches a limit this page keeps too: measure it on your own code, because each of those results was specific to a codebase.

Install

In Claude Code:

/plugin marketplace add scottcrosby-securebine/doctrine-skills
/plugin install doctrine@doctrine-skills

If the install message says to, run /reload-plugins or restart. A freshly installed plugin does not always load into the session that installed it.

In Codex CLI:

codex plugin marketplace add scottcrosby-securebine/doctrine-skills
codex plugin add doctrine@doctrine-skills

On Codex the hub reads a Codex reference file for its tools, and runs its red team through Claude. Codex loads no plugin hooks, so the seat panes, the restore after /clear, the context gauge and auto-cycle need one install command, run again after each plugin update:

node ~/.codex/plugins/cache/doctrine-skills/doctrine/<version>/hooks/dctr-codex.mjs install

It writes the hook entries into ~/.codex/hooks.json (or $CODEX_HOME's) and trusts them in config.toml. Where sandbox_workspace_write.network_access is unset it sets it to true, so commands in the workspace-write sandbox can reach herdr, and where agents.max_concurrent_threads_per_session is unset it sets it to 8, so Codex does not refuse a seat past its default limit. A value you set for either is left as it is, comment and all, and the install prints a line saying so. Where either setting's table is written inline (agents = { ... }) or with dotted keys and the setting is unset, it leaves that line as it is and prints how to add the setting by hand. doctrine-pane needs the same install on Codex.

Codex has no setting that turns automatic compaction off. model_auto_compact_token_limit in config.toml can only make it come earlier, and by default Codex compacts at 90% of the model's context window. So for auto-cycle on Codex, set the tier on the record's auto-cycle: on cap <n> tier <tokens or percent> line low enough that the gauge warns, and the handoff that follows the warning finishes, before that point.

Codex fires no hook for a session left idle, a turn ended by an API error, or the exit of a background terminal, so while auto-cycle is on each prompt starts a detached watcher, hooks/dctr-watch.mjs, that follows the turn and hands the auto-cycle hook the event Claude Code would have fired. Idle is a turn whose end is in the session's rollout with no error, no background work and a last message that is not the ready line, and, where the session runs in a herdr pane, herdr reporting that pane idle or done, for 60 seconds. An API error is the error the turn's end carries in the rollout, and a turn whose rollout has not grown for 15 minutes while it is unfinished is handed on as one. A running background terminal is a live process under the session's Codex process whose environment carries CODEX_SESSION_ID, read from /proc. The Stop reads it too and waits while one runs, and the watcher hands the Stop back once the work holding it has finished. A process whose environment cannot be read counts as running, and when that holds a Stop back for 60 seconds the record gets a paused line saying the doctrine could not tell.

After a Codex, Claude Code or herdr update, run node <plugin-root>/hooks/dctr-doctor.mjs, with --host codex or --host claude to check one. It drives each installed host in scratch homes under its own herdr server, never writes ~/.codex or ~/.claude, and prints a row for each host signal the hooks depend on, the areas its drive does not exercise, and a verdict. Exit 0 means every signal it reads holds, 1 means drift, with the row naming each drifted signal, and 2 means it could not run, with the message saying why. A run takes a minute or two and two cheap model turns per host.

Try it on one file

Do not start with a big audit. Point it at one file you suspect: use doctrine-audit on src/whatever.ts (scope is that file only). It asks what counts as a bug before changing anything, runs whatever your repo has written down as a gate, and puts the verdict at the top of its report, ahead of the work:

Gate: clean pass on 4a3e4ca, real run on record. No blockers open.
Gate: shipped at an escalation, round 6, zero clean passes. Open: 2 blocking findings.

Those are two of the forms. A run can also stop for your decision, end blocked on a question you have not answered, or close a prose job with a list of what it did not fix. The wording varies. What the line always does is distinguish a clean pass from an ending without one, and name what is still open.

The fourteen skills

SkillUse it for
doctrineThe shared posture. The eight wrappers invoke it.
doctrine-codeFeatures, specs and tickets.
doctrine-debugAnything broken, throwing, failing or slow.
doctrine-auditBug hunts and deep code audits.
doctrine-docsDocumentation sweeps.
doctrine-writeProposals, briefs, PRDs, reports.
doctrine-researchMulti-source questions needing a fact-checked answer.
doctrine-gauntletWeb design, judged on the rendered page.
doctrine-projectStarting or adopting a repo's project and epic tracking, with a checkable meaning of done.
doctrine-paneAn interactive terminal session in a pane you can take over. Not a wrapper: it does not load the posture.
doctrine-backupSaving the session's memory file for the next session, red-teamed before it is done.
doctrine-handoffA detailed handoff under docs/handoffs/, tied to the plan, so nothing is lost across a cleared session.
doctrine-resumePicking up where the last session left off: state, handoff, drift.
doctrine-primerA cold start: everything doctrine-resume restores, plus a tour of the project.

What it costs

More time and more tokens than a single pass, so point it at work where being wrong is expensive. No mode skips the gate, though doctrine-gauntlet asks which of its two modes you want and the hub proposes a gate, with the reach evidence behind it, for every phase whose exit it sets, and you choose that gate or the full one. No token figures appear here: the runs behind this page did not measure them, and an unsourced number is what this page argues against.

Built on other people's work

These are the disciplines the loop runs on, each published by someone who does this work. The doctrine invokes them by name at runtime rather than copying them, so their updates flow through. All optional, each with a fallback: Matt Pocock's engineering skills for the build and review disciplines; superpowers, by Jesse Vincent and Prime Radiant, for parallel dispatch and worktree isolation; ponytail, by Dietrich Gebert, for the pass that deletes what nobody asked for; writing-clearly-and-concisely, Strunk's rules as a skill, by Josh Thomas via softaworks; OpenAI's codex plugin for the different-model red team; and the gauntlet loop by Matt Shumer, whose method doctrine-gauntlet builds on. One exception to invoking by name: you install Matt Pocock's code-review yourself as a renamed matts-code-review copy, because the original name collides with Claude Code's own /code-review. This plugin does not bundle it.

Requirements

Twelve of the fourteen skills need nothing installed. doctrine-gauntlet judges rendered pages, so it needs a browser, and without one its harness refuses to run rather than report a pass. doctrine-pane needs herdr 0.8.2 or later, and outside herdr it refuses. Neither degrades into a weaker version of itself.

Optional, each with a fallback: Matt Pocock's engineering skills, superpowers, the OpenAI codex plugin on Claude Code, the Claude Code CLI (claude, signed in) on Codex, ponytail, writing-clearly-and-concisely, Claude Code's deep-research workflow, Claude Code's Workflow tool, Claude Design, and herdr for everything other than doctrine-pane and auto-cycle, below. docs/requirements.md has the install commands and says what happens when each is missing.

Auto-cycle is off by default and switched on per phase in its record. Using it needs these set up once: herdr, with the session running in a herdr pane and herdr's Claude integration installed (herdr integration install claude); auto-compaction off, with claude config set -g autoCompactEnabled false (a global config key in ~/.claude.json, not settings.json), since the phase resets by /clear and a compaction mid-phase discards the session's context and the gauge's reading; and the statusline bridge, installed with node <plugin-root>/hooks/dctr-bridge.mjs install and re-run after every plugin update. docs/requirements.md says what each one does, and docs/watching-a-run.md gives the $autocycle sidebar row that shows whether it is cycling or paused.

Watching it work

A doctrine run dispatches a lot of subagents and Claude Code shows you none of them: the terminal goes quiet and an answer appears some minutes later. Run it inside herdr and every subagent gets a live pane you can read while it works. Optional, and nothing here needs it: docs/watching-a-run.md.

Limits

These rules came out of real work, and the failures behind them are ones this author hit. There is no third-party benchmark, and the three examples are anonymized, so you cannot check them yourself. Defects found and not yet fixed are in docs/known-issues.md.

The specification is skills/doctrine/SKILL.md, the file the agent actually loads, so trust it over this page wherever the two disagree. How the work is checked, and what those checks cannot reach, are in docs/how-this-is-tested.md.

License

MIT