Skip to content

jwilger/development-system

v6.12.0AGPL-3.0-or-later

John's consolidated development workflow, engineering standards, worktree, task, review, and agentic-system tooling.

Development System

The single personal development plugin for Codex, also packaged for pi.

pi package (experimental)

This directory is a pi package as well as a Codex plugin (ADR-0018). Install it in place:

pi install ./plugins/development-system

The package loads the skills in skills/ and the adapter in pi/extension.ts. The adapter:

  • registers both MCP servers from mcp.json using the same plugin-root launchers, so binaries are repaired the same way as under Codex;
  • runs bin/development-system session-start --harness pi in the background at session start and shows its warnings as notifications;
  • adds /development-system doctor to rerun that check on demand.

The check warns when a project or user mcp.json entry shadows development-discipline or tiber, and when pi settings disable built-in MCP support. Codex subagents are not yet available in pi, and some skills still describe Codex-specific setup.

Host-local binaries

After installing or upgrading this plugin, download and verify its Rust executables once from the marketplace checkout:

just install-development-system-binaries

On Linux x86_64, the command downloads the exact plugin-version bundle from the repository's GitHub Release, verifies its SHA-256 sidecar and fixed archive layout, and atomically installs tiber and development-discipline-mcp in $PLUGIN_DATA/ai-plugins/development-system/<plugin-version>/<host>/ under Codex. Direct CLI use falls back to $XDG_DATA_HOME, then ~/.local/share. Re-running the command is safe; a newer plugin version is installed alongside previous versions. The plugin root mcp.json declares both servers with plugin-relative launchers. Each launcher repairs a missing matching binary before starting, so consuming projects do not configure MCP servers. Project setup removes only a previously managed Development System MCP block from .codex/config.toml. Signed Tiber operations use the plugin-owned $PLUGIN_DATA/signing-agent-socket file. Outside a portable harness, the launcher uses $XDG_CONFIG_HOME/ai-plugins/development-system/ (or ~/.config). The root installer captures a valid SSH_AUTH_SOCK; update the plugin-owned file with an absolute socket path if your agent moves. The release bundle contains statically linked Linux x86_64 executables, the applicable license, and the exact source tag and commit. To force a locked source build for diagnosis:

just install-development-system-binaries --from-source

The repository's Nix devshell is optional. If just is unavailable, run scripts/install-development-system-binaries.sh from the marketplace checkout.

The installed SessionStart hook and both MCP launchers verify the binaries and atomic installation marker against the plugin manifest version. It repairs missing or stale installations automatically: Linux x86_64 uses the verified release bundle, while hosts without a prebuilt bundle use the locked Cargo build. The setup skill performs the same automatic check before configuring a repository. Manual use is only needed for diagnosis or an intentional source build.

It is inert outside a Git repository or without a valid schema-3 .development-system.toml. Read-only repository inspection remains available in every state. Start configuration through the structured setup.preview tool. It detects the project's manifests, lockfiles, conventional source/test/docs and build paths, Nix devshell, and stack-native test commands, then produces a schema-validated project-specific configuration instead of a generic template. Review its exact scopes and named command catalog, then review the generated project lefthook.yml and its separate pre-commit and pre-push command selections, then explicitly confirm setup.apply. Setup installs both Git hooks with Lefthook but never stages or commits configuration.

Normal Development System work relies on those Git operations for repository verification: git commit triggers the fast pre-commit checks and git push triggers the pre-push checks. Agents must not run the same gate commands separately unless the user explicitly requests a diagnostic run.

Recover a failed checkpoint gate

Version 6.9.5 supports recording a failed pre-commit attempt or a lightweight review that requires remediation without pretending a commit or passing gate exists. Both become canonical failing checkpoints; only the recorded causal repair and immediate fresh testing may follow. Fresh lightweight review and a new normal hook-backed commit are then required before exact verification and authorized delivery. Full terminal review and exact-SHA CI requirements remain unchanged.

For a session stranded by an older plugin, update and reinstall:

codex plugin marketplace upgrade ai-plugins
codex plugin add development-system@ai-plugins
codex plugin list --json

Verify the installed version is 6.9.6 or newer. Restart the harness and resume the work in a fresh thread with its original request, immutable baseline, and checkpoint ID. The startup hook repairs version-matched host binaries; verify both MCP servers use the new plugin root and call workspace-reader.status through the connected Development Discipline MCP. Its final-review attestation must report contract_version >= 2, minimum_clean_iterations >= 3, and durable_pending_assignment_recovery: true before terminal review. If the release is not yet available, build from the updated marketplace checkout into the installed MCP's exact data directory:

plugin_data=$(codex -C /tmp mcp list --json | jq -er '.[] | select(.name == "development-discipline") | .transport.env.PLUGIN_DATA | select(type == "string" and startswith("/"))')
PLUGIN_DATA="$plugin_data" nix develop -c just install-development-system-binaries --from-source

Verify the checkout version matches the installed plugin version, and verify the runtime after restart. A stale explicitly configured project MCP binding requires the separately previewed and approved setup migration.

For new checkpoint transitions, use the bundled transition-local-checkpoint.sh operation API. It derives successors under the authoritative writer lock and returns durable receipts keyed by stable operation IDs. Initialization, focused test outcomes, review and gate outcomes, actual commits, exact verification and retry, local or pushed delivery, CI observations, and terminal outcomes have explicit operations. The recovery command below remains supported for older sessions.

Before changing source, reconcile the retained failed command/review evidence with Git. HEAD must still equal the checkpoint's HEAD: if a commit exists, recover its normal committed transition instead. Keep the real failure evidence in a readable, nonempty file outside the worktree and retain that file for handoffs. Include the failed command or review identity, actual failed checks/findings, exit status when available, and causal diagnosis; avoid secrets. The helper records a bounded path and SHA-256 reference, not raw log content. If evidence is missing, recover it from the original session, or retry the still-pending normal commit with captured output and record the actual result.

Run from the target repository; resolve the installed root rather than using the old session's cached path:

plugin_root=$(codex -C /tmp mcp list --json | jq -er '.[] | select(.name == "development-discipline") | .transport.env.PLUGIN_ROOT')
checkpoint_id=YOUR_CHECKPOINT_ID
checkpoint="$(git rev-parse --path-format=absolute --git-common-dir)/development-system/checkpoints/$checkpoint_id.latest"
generation=$(sed -n 's/^checkpoint-v1 //p' "$checkpoint" | jq -er '.generation + 1')
predecessor=$(sha256sum "$checkpoint" | cut -d ' ' -f 1)
"$plugin_root/scripts/record-checkpoint-failure.sh" \
  "$checkpoint_id" "$generation" "$predecessor" pre-commit-hook \
  'THE ACTUAL FAILED GIT COMMAND' /absolute/path/to/retained-failure-evidence \
  'THE SPECIFIC CAUSAL REPAIR'
cat "$checkpoint"

For a failed review, substitute lightweight-review and the actual review identity/evidence. The helper uses the validated writer's lock and generation/predecessor compare-and-swap. It recomputes the current snapshot, retains the immutable baseline and CI history, clears gate receipts and delivery, and records causal-edit: <repair>. It rejects stale callers, incompatible pending actions, changed HEAD, and missing/empty/in-worktree evidence. It never edits source, stages, unstages, commits, pushes, or replaces the checkpoint directly.

Staged work can remain staged, including newly staged previously untracked files. Staging changes snapshot partitioning; reconcile source identity rather than discarding work to reproduce old hashes. Include identified hook-made changes in the actual failing snapshot. Unexplained source changes or a malformed predecessor remain a recovery hold. Read back and reconcile the published failing record before performing its causal repair. Never manually replace .latest, fabricate evidence, or use the helper to skip testing, review, commit hooks, or delivery gates.

The plugin-wide Development Discipline MCP surface provides bounded repository inspection, deterministic setup, and multi-agent final review. Those services and the Codex hooks are advisory: they do not establish agent identity, isolate project tools, execute project mutations, or deny ordinary host capabilities. setup.apply writes only the previewed repository-local configuration and installs its Git hook launchers. It does not generate privileged agents or profiles and never changes global Codex, marketplace, MCP, shell, or SSH settings.

Within that advisory boundary, final-review state fails closed: risk planning must select a lens, every selected lens and assigned verifier reruns in each iteration, and completion requires at least three consecutive complete finding-free iterations. Findings, malformed results, and material deltas reset the streak; the review-budget ship choice cannot bypass it. The reader status attests the connected final-review protocol, and current workflow guidance refuses to create review state when that attestation is missing or too old, so a newly installed skill cannot silently coordinate through a stale one-pass MCP runtime. The planned Tiber can additionally enforce those receipts at the task boundary when a project opts into [final_review].minimum_clean_reviews: its Git-backed task history then blocks both task completion and trailer-driven delivery until the required clean sequence is current. This gate is available through the same CLI and MCP surface in Codex.

The editor, runner, repository, and diagnostic services are retained as unexposed reusable components for the standalone Tiber harness. Ordinary Codex remains able to inspect, edit, verify, commit, and push while Tiber is being bootstrapped. Tiber, rather than this plugin, will own authoritative identity, isolation, workflow, memory, verification, and delivery.

Codex loads the plugin-root mcp.json and plugin-relative launchers. Project MCP entries and global compatibility overrides are unnecessary. If a signed append cannot reach the agent, update the plugin-owned signing socket path.

The core workflow requires only this plugin. The root installer offers GitHub and CodeRabbit as optional companions. GitHub provides a connector mapping and CodeRabbit provides skills backed by its CLI; neither requires a Development System MCP dependency. Superpowers is omitted because its workflow overlaps. The plugin owns its bundled MCP integrations; user-added MCPs are warned about for compatibility review, not automatically rejected.

The plugin root is the active public surface: its manifests, hooks, launchers, and skills/ directory define the installed development-system plugin and the development-system:<skill> routing names. Directories under components/ retain independently authored source, tests, manifests, and runtime assets that the public plugin owns or wraps. Those component directory and manifest names remain valid internal identities; they are not separate marketplace install targets or public routing labels.

The top-level plugin manifest starts the advisory Development Discipline inspection, setup, and multi-agent final-review surface plus the independent Tiber MCP server. The advisory surface stores final-review transitions on a separate local-only Git-backed EventCore authority and never publishes them to development-workflow; the standalone workflow service retains that remote authority. Development Discipline reads Tiber's unresolved CI hold when evaluating delivery readiness. Tiber is the sole CI-incident and receipt authority on tiber, in addition to publishing task events there. Promptfoo remains an optional MCP owned by the agentic-systems-engineering component and is intentionally excluded from the top-level manifest because projects may disable that capability and must supply the pinned Promptfoo runtime separately. Configure it explicitly when needed; do not install the retained component as a separate marketplace plugin.

Skill descriptions are narrow routing indexes. Detailed workflow context is loaded only after a matching skill routes.

Model routing

The coordinator assigns bounded, standard, or strong capability according to the task's eligibility and risk, then selects a concrete model and supported reasoning setting from current runtime evidence. The plugin does not prescribe model generations or provider-specific effort labels. Availability alone does not establish fitness, and mappings cannot grant tool authority.

Optional string-valued [model_routing] entries in .development-system.toml supply operator preferences. DEVELOPMENT_SYSTEM_MODEL_ROUTING_FILE can select an absolute path to a machine-specific TOML mapping instead. These are coordinator-read preferences, not native spawn configuration; explicitly selected invalid mappings and unconfirmed routes are reported visibly. See the routing skill and mapping contract for keys, precedence, examples, and fallback rules.

Continue required review after a timed checkpoint

A medium-risk review's 75-minute boundary is a progress assessment. When required repairs or fresh independent review remain and no human decision is needed, call final_review.continue_review with the returned state_ref, a stable operation_id, and a concrete rationale. The coordinator records that assessment, retains the original start time and every blocker, and opens another 75-minute window. Three consecutive complete finding-free review rounds remain required. Retry the exact request after a lost response; conflicting reuse of an operation ID fails. Recovery of a legacy escalation hold additionally needs recovery_reference explaining why its recorded dependency is resolved.

After source changes to a completed session, use final_review.reopen with its authoritative reference, stable operation ID, reason, current diff identity, complete changed-file inventory, and fresh shared test evidence. Reopening preserves the pinned baseline and history, clears stale review credit, and requires an independent delta assessment followed by all selected review lenses. A replayed historical reopen receipt explicitly identifies stale assignments; resume current state before launching more reviewers.

For native persistence failures, use the version-matched one-operation replay helper through the host's supported approval mechanism. Its fixed writable paths cover the Git objects, advisory refs and review replica, the exact per-binding SQLite report directory, snapshot objects and private scratch. Source, unrelated refs and configuration stay read-only; IP networking is denied. Signing, publication locks and concurrency checks still apply. The historical narrow reproduction shows that these writes were necessary, without proving every host's original permission failure had the same cause.

Independent rejection evidence now travels with subsequent reviewer packets. The coordinator uses exact stable identity and actual checked source dependencies, so an unrelated documentation change does not discard a valid source-bound resolution. A reviewer may challenge it with explicit new evidence and an explanation of why the prior resolution fails; reopening requires independent adjudication. Unsupported repeats do not create renewed escalation work and still prevent a finding-free round.

Read final_review.yield_report for actual round counts and separate raw versus adjudicated outcomes; use final_review.evidence to inspect the associated records. Source changes compare immutable captured Git trees, with snapshot and tree identities available in the round evidence. Scope-hash or staging changes alone do not prove source changed. Historical hash-derived source claims and missing snapshot evidence report source_changed: null. Missing legacy evidence is reported as unavailable. Reports never infer that a clean round was useless or reduce review requirements. The existing serialized-state size limit remains a limitation for evidence-heavy sessions.