Skip to content

joymin5655/agent-harness

v0.5.13MIT

Portable AI agent harness — curated review/build/test agents, secret-hardening + worktree + plan-gate hooks, and spec/supervise/wrap/verify-completion skills. Install once, use in every project. AI-agnostic core (Claude Code / Codex / Gemini).

Changelog

All notable changes to this project will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

Unreleased

[0.5.13] - 2026-10-02

Added

  • Antigravity native hook adapter (W5-3): adapters/antigravity/adapter.sh and adapter.py translate agy 1.2.12's PreToolUse/PostToolUse/Stop hook JSON (event name from argv, since agy's stdin carries none) into canonical events and run the same core-hook chains as the Codex template (run_command as Bash, write_to_file as Write, replace_file_content and each multi_replace_file_content chunk as Edit, capped at 100). agy treats a PreToolUse {} as a deny and a hook allow was not observed to grant more than ask (unmeasured, so never emitted); a pass-through is {"decision":"ask"}, a core-hook ask becomes force_ask (plain ask defers to the user's allow rules), a core deny is a deny, and any hook failure, timeout, bad JSON, exhausted 25s budget or unverified argument shape is a fail-closed deny. PostToolUse always answers {}. Stop turns a core block into one {"decision":"continue"} guarded by a marker under ${AGENT_STATE_DIR:-$HOME/.agent/state}/antigravity-stop/<conversationId> (6h TTL). With AGENT_ANTIGRAVITY_WORKER=1 every matched tool call is denied without running a hook. GEMINI_API_KEY/GOOGLE_API_KEY are scrubbed from the hooks' environment. Review fixes: send_command_input (text typed into a live shell) is now matched and guarded as Bash; the Stop chain keeps a reserved time slice for brain-capture.py and session-close.sh and reports a gate that overran; a malformed AGENT_ANTIGRAVITY_BUDGET_S falls back to 25 instead of crashing the hook; session-quality-gate.py reads agy's tool_calls transcript shape. core/tests/antigravity-adapter-test.sh (113 checks, includes a drift check against the Codex template) and a new adapter-parity.sh antigravity section cover it.
  • Antigravity plugin install (W5-4): adapters/antigravity/install-plugin.py, plugin.json and hooks.json.template. setup.sh --antigravity installs a plugin folder (${AGENT_ANTIGRAVITY_PLUGIN_DIR:-~/.gemini/config/plugins/agent-harness}) whose hooks.json points at the absolute adapter path. It writes atomically and idempotently, refuses a foreign plugin folder or a framework root containing shell metacharacters, offers --uninstall, --check and --dry-run, and never reads or writes ~/.gemini/config/hooks.json or agy's settings.json. Setup also prints (never applies) API-key and permissions.deny guidance. setup.sh --doctor gained an "antigravity native hooks" check (core/tests/antigravity-native-hooks-test.sh, 86 checks). A marker file keeps a folder ours after an uninstall that left a user file behind, so reinstall and setup.sh --antigravity no longer refuse it, and the setup step warns instead of aborting. The permissions.deny guidance uses agy's documented prefix form, not * globs.
  • Antigravity worker API-key opt-in (W5-2): ANTIGRAVITY_AUTH=apikey makes antigravity-worker.sh read a Gemini API key from the Keychain (service gemini-api-key) and export GEMINI_API_KEY into agy's environment only, never argv or logs. A missing item or security binary exits 2; a missing "modelProvider": "gemini" in agy's settings is a stderr warning, and the worker never edits that file. The keyring stays the default. The worker also exports AGENT_ANTIGRAVITY_WORKER=1 for every dispatch. Review fixes: the worker writes a static workspace deny plugin (.agents/plugins/agent-worker-deny/) before agy starts, exports the API key only after that succeeds and exits 2 otherwise, and its sandbox profile denies writes to ~/.gemini/config/hooks.json, config/plugins/ and settings.json (core/tests/antigravity-worker-test.sh, 95 checks, includes a real sandbox-exec enforcement check).

Changed

  • Antigravity docs and registry brought current with agy 1.2.12 (W5-5): adapters/antigravity/README.md now records the 2026-09-29 probe (soft-deny in denied_actions plus a jetski ... auto-denied stderr notice, exit 3 on status ERROR), the Gemini CLI individual end date 2026-06-18, the API-key opt-in, the native-hook plugin, and the worker threat-model drift (headless 1.2.12 ran echo with no allow rule). docs/runtime-registry.json antigravity: cli_version_measured 1.2.12, measured_on 2026-09-29, hook_events_wired PreToolUse/PostToolUse/Stop. docs/hook-protocol.md gained section 13 and the antigravity ai value; docs/cross-runtime-harness-design.md states what the plugin enforces and what it does not.

Fixed

  • Antigravity lane counted a soft-denied run as success (W5-1). Headless agy soft-denies a tool call it cannot get approval for: the run continues and exits 0 with a stderr notice. antigravity-worker.sh now exits 9 on that notice and 10 when the --output-format json envelope is unparseable or its status is not SUCCESS. It captures agy's output outside the sandbox's writable dir, so a prompt-driven write cannot forge the envelope. antigravity-preflight.sh reports a soft-deny as exit 8 (lane absent), reads the probe token from .response only, and treats authentication required as an auth failure. New core/tests/antigravity-preflight-test.sh.
  • Antigravity soft-deny detection missed agy 1.2.12 (W5-1b). 1.2.12 reports a soft-deny as a non-empty denied_actions array in the json envelope plus a stderr notice with new wording, which the 1.1.14 pattern did not match, so antigravity-worker.sh returned exit 0 for a run that was denied. It now exits 9 on a non-empty denied_actions (single-value stdout only) or a stderr match on permission check failed|denied permission to|auto-denied|cannot prompt for; an empty denied_actions is not a soft-deny, and agy's own nonzero exit codes (including the undocumented 3) pass through unchanged.

[0.5.12] - 2026-10-01

Added

  • Supply-chain scan classes 5–7 (adapted from ECC v2.2 pi/core): core/tests/supply-chain-scan.sh delegates to core/tests/supply-chain-remote.py. Class 5 flags fetch-and-execute (curl … | sh, bash <(curl …), eval "$(curl …)"); class 6 flags unpinned remote runners (npx/npm exec with --yes/--package, bunx, pnpm/yarn dlx, uvx, pipx run); class 7 flags URL hosts in auto-fired hooks and hooks/*.json / .mcp.json that are not in core/tests/supply-chain-allowlist.txt. Threat model documented in rules/policy/security-guards.md (#138).
  • Impact context (idea from Graft's blast radius): core/infra/impact-context.py lists dependents outside the diff and the test files the change reaches, from the existing CodeGraph index. Fail-open (10 s budget, 60-line cap, exit 0). /council-review adds it to the shared review core; /wrap shows it as an advisory pre-flight step (#138).

[0.5.11] - 2026-10-01

Added

  • Codex native hook path (W4-1): adapters/codex/adapter.py's native mode (run_native) translates Codex's native PreToolUse stdin (Claude-shaped: hook_event_name, tool_name, tool_input, tool_use_id, session_id, cwd, transcript_path) into canonical events, splitting apply_patch's patch text (carried in tool_input.command, same field as Bash) into one Write/Edit event per file with absolute paths, and aggregating deny > ask > advisory > allow across files. A canonical ask and any hook failure (non-zero exit other than 2, a timeout, invalid JSON, or a zero-op patch) become a fail-closed deny JSON; non-PreToolUse events stay fail-open by design. adapters/codex/hooks.json.template (merged into ~/.codex/hooks.json by the new adapters/codex/merge-hooks.py, which replaces only Agent-owned entries and leaves other tools' hooks alone) wires SessionStart/UserPromptSubmit/PreToolUse/PostToolUse/Stop/SessionEnd. Hardened after a security review: one 25s check budget per hook across all files of a patch, a 100-file cap, indented patch markers denied as ambiguous, normpath on patch paths, Move to destinations checked with the source file's content (regular files only), a deny for a missing guard, missing python3, or an unknown decision verb, and every per-file event run on PostToolUse (council review). core/tests/codex-native-hooks-test.sh (41 checks) and a new adapter-parity.sh native section cover it; one live codex exec session (codex-cli 0.157.0, 2026-09-27) confirmed the Bash and apply_patch denials end-to-end.
  • Portable plugin manifest + marketplace (W4-2): root plugin.json (agent-plugins.org 1.0.0, extensions.com.openai.hooks → hooks/codex-hooks.json) and .agents/plugins/marketplace.json, so codex plugin marketplace add joymin5655/Agent + codex plugin add agent-harness@agent installs the same hooks and skills — install verified with the codex CLI (skills appear in codex debug prompt-input) on 2026-09-27. core/tests/version-parity.sh now covers the root manifest.

Changed

  • council-escalation-gate.py escape 2 now requires a stated reason. A retry on an already-denied council-scale diff passes only when the code-reviewer dispatch prompt carries council-unavailable: <reason> (≥10 chars after whitespace collapse, not the pasted <placeholder>); the reason is written to security-violations.jsonl. If the diff hash is unavailable, a stated reason alone opens the escape. The deny text used to advertise "re-issue this exact dispatch", and a model was observed retrying reflexively without trying /council-review — a bare identical retry is now denied again, and repeat denials no longer refresh the ledger entry. Tests: 9 new cases in council-escalation-gate-test.sh (32 pass).
  • model-routing-observer records both session ids. session_id now prefers the hook event's runtime session UUID (falls back to AGENT_SESSION_ID), and the env id is kept as agent_session_id, so concurrent sessions in one cwd stay distinguishable. manager-audit.sh --session <id> matches either field, keeping existing filters working.
  • setup.sh --codex now merges hooks.json.template into ~/.codex/hooks.json beside config.toml via merge-hooks.py, symlinks each skills/<name> into ~/.agents/skills, repoints an old ~/bin/codex-bash symlink to legacy/codex-shell-wrap/, and prints a reminder that Codex only enforces a hook after you trust it with /hooks.
  • setup.sh --doctor gained a "codex native hooks" check: WARN when [features] hooks = false or nothing is installed, FAIL on invalid JSON or a moved adapter path, WARN when entries are installed but nothing is trusted yet, PASS otherwise (core/tests/setup-doctor-test.sh).
  • docs/runtime-registry.json, docs/ai-adapters.md, docs/cross-runtime-harness-design.md, and docs/hook-protocol.md brought current with the native path: Codex hook_events_wired, Tier A coverage for Bash + apply_patch + MCP tools, the ask→deny translation now living in the adapter itself, and the fail-open/fail-closed split by adapter in the exit-code table.
  • Claude Code hook manifests brought current with 2.1.282 (W3-1/W3-2/W3-3/W3-4). hooks/hooks.json and adapters/claude-code/settings.json.template: matchers Write|Edit|MultiEdit → Write|Edit, Task|Agent → Agent, ExitPlanMode|Task|Agent → ExitPlanMode|Agent (MultiEdit is no longer a documented tool; Task is a legacy alias of Agent). Wired 6 Claude-only extended events non-canonical to the cross-AI protocol: PostToolUseFailure → circuit-breaker.py, SessionEnd → session-close.sh (timeout: 2), PreModelSwitch/PostModelSwitch → session-tier-observer.py, SubagentStart/SubagentStop → model-routing-observer.py. PermissionRequest deliberately left unwired (different decision schema; exit 2 not honored), and so are WorktreeCreate/WorktreeRemove (a WorktreeCreate hook replaces git's worktree creation and must print the new path — the observer r4-mutex-check.sh would have broken every Claude worktree; wired during W3, unwired before release) — see docs/hook-protocol.md §12. secret-content-scan.py's MCP matcher collapsed from an explicit per-tool pipe-list to per-vendor mcp__<server>__.* wildcards; rubric-commit-judge.sh gained a narrowing "if": "Bash(git commit*)". adapter.sh header now documents the extended events and the no-op-vs-fail-open distinction. docs/hook-protocol.md gained §12; docs/runtime-registry.json claude-code hook_events_wired reflects the full wired list.
  • CI: added plugin-validate job running claude plugin validate --strict . (best-effort CLI install; explicit ::notice:: skip if the CLI never lands — never a silent pass, and not a required check).

Deprecated

  • codex-shell-wrap.sh moved to legacy/codex-shell-wrap/ (W4-3). It remains a fallback only for [features] hooks = false builds, or an admin requirements.toml that allows managed hooks only; adapters/codex/tests/run.sh (T5/T6) still exercises it so the fallback path doesn't rot.

Fixed

  • session-close.sh missed SessionEnd in json.dumps-style input. The event-name sed required :" with no space, so "hook_event_name": "SessionEnd" fell through to the full Stop path (TODO scan, notification, broadcast) inside the 1.5s SessionEnd budget. The extraction now allows whitespace after the colon.
  • circuit-breaker.py answered PostToolUseFailure with hookEventName: "PostToolUse". The advisory now echoes the event it was invoked for.
  • runtime-currency.sh read a future measured_on/checked_on as fresh. A date after today is now a FAIL (typo), not a negative age.
  • codex-template-currency-test.sh denylist lagged the registry (gpt-5\.[234] vs docs/runtime-registry.json's gpt-5\.[2-6]); synced, registry named as SSOT.
  • Deferred council findings (runtime-currency-2026-09).
    • The rubric-commit-judge manifest entry no longer carries "if": "Bash(git commit*)", which missed git -C <dir> commit and rtk git commit. The hook's internal regex now does all of the filtering.
    • circuit-breaker.py serializes its shared state file with flock and writes it atomically (tmp + rename). Before this, concurrent sessions lost failure records.
    • setup.sh --bootstrap with no usable OS package manager (or no Homebrew) now still reaches the PyYAML pip step instead of returning early. It still exits 1.
    • session-tier-observer.py stamps origin on its session-start record, which the W1 log-origin tag had missed.

[0.5.10] - 2026-09-02

Fixed

  • session-quality-gate.py layer 3 fired without end and blamed the wrong session. The unverified-session advisory read the whole dirty work tree while claiming the files "changed this session", and recomputed and reprinted on every Stop: stop_hook_active gates blocking only, and the note is re-injected as model-visible context, so it fed the turn that produced the next Stop. A file left uncommitted by a concurrent session was reported to five unrelated sessions as their own change — 80 firings in 7m22s in one session (median gap 2.4s) and 79 in another, ending only when a verification command happened to land in the sink. The documented exit ("or state explicitly why none applies") was inert: nothing consumed a stated reason, so a session that correctly declined to test someone else's work had none. The diff is now intersected with the files the session actually edited (read off transcript_path; sidechain kept, since a subagent's edit is still the session's work) and the note is suppressed once recorded for the same (session, changed-file set), with blocking keyed separately so enabling AGENT_VERIFY_OBSERVER_BLOCK=1 midway is not swallowed by an earlier advisory. Without a transcript it falls back to the work-tree list and says "attribution unavailable" rather than claiming one. Two ways the intersection could go silently empty are fixed with it: git diff and git ls-files disagree about their base outside the work-tree root (git now runs at the toplevel, diff.relative pinned off), and core.quotePath C-quotes non-ASCII paths (-z throughout). The FIFO/S_ISREG/tail-window hardening moved into one shared traversal instead of being copied per predicate. (#122)

Added

  • top-edit-advisor.py hook (PostToolUse Write/Edit/MultiEdit). Accumulation-time counterpart to model-routing-advisor.py/ model-routing-observer.py, which only see Task/Agent dispatches and so cannot see the leak where the TOP model implements directly and never dispatches at all — a 2026-09-02 audit measured 511 such direct edits across 26 sessions, with a single warning proven insufficient. Repeats a systemMessage advisory every +15 measured main-loop edits (env AGENT_TOP_EDIT_THRESHOLD) instead of warning once and going silent; never caches a not-TOP verdict since the session model can switch mid-session. Advisory only — never blocks. Tests: core/tests/top-edit-advisor-test.sh.
  • Codex quick/deep tier profile installation in setup.sh --codex. install_codex() now installs the existing adapters/codex/{quick,deep}.config.toml.template beside ~/.codex/config.toml, so the tier ladder documented in docs/model-routing.md (codex --profile quick|deep) is wired up by setup instead of requiring a manual copy; doctor check 13's remediation now also points at setup.sh --codex.
  • Review-tier ladder (core/infra/review-tier.sh) + /wrap step 1d integration. Review dispatches are the second-largest routing cost after implementation, so cadence — not just model tier — is now a lever: every diff gets tier 0 (docs-only or ≤AGENT_REVIEW_SKIP_LINES non-risk code lines — skip, self-check only), tier 1 (the common case — one code-reviewer pass at wrap/commit time), or tier 2 (council-scale — /council-review), delegating the tier-2 judgment to the existing council-threshold.sh SSOT instead of re-mirroring its risk-area patterns. agents/code-reviewer.md and docs/model-routing.md updated to the hybrid timing this implies (wrap-time default, immediate only on risk-area paths). Tests: core/tests/review-tier-test.sh.
  • council-escalation-gate.py same-diff-hash escape visibility upgrade. The loop-safety escape (a council-scale diff already denied once is let through on retry) used to be stderr-only; it now also emits a PreToolUse additionalContext advisory (model-routing-advisor.py's emission pattern) so the model/user see that the dispatch skipped review, not just a log line nobody reads. Allow behavior, TTL, and hash binding unchanged.

[0.5.9] - 2026-08-25

Added

  • OpenRouter free advisory lane + purpose launcher set + codex global rules (spec: .agent/plans/free-lanes-and-launchers/). New adapters/openrouter/ worker-lane bridge to OpenRouter :free routes (non-voting advisor-free role, sensitive-cwd guard + per-dispatch retention warning, fail-open on 429, free exact-token preflight; pin nvidia/nemotron-3-super-120b-a12b:free in the adapter-owned tiers file, chosen by live probe over the congested z-ai/glm-5.2:free). /council-review gains --with-free (advisory lane, mirrors --with-grok). New adapters/claude-code/launchers/ purpose launcher set (claude-build/claude-quick/claude-research tier launchers + claude-ox.template OpenRouter gateway launcher absorbed as repo SSOT, personal blocklist externalized to the shared ~/.config/agent-harness/sensitive-paths file) — session-start human allocation, the allowed side of the no-runtime-switching policy. New adapters/codex/AGENTS.global.md.template deployed to ~/.codex/AGENTS.md so manual codex sessions carry the portable harness rules (evidence contract, tier discipline, review-before-done). docs/model-routing.md graduates the "tier/cost-aware task allocation" follow-up to a designed free-lane allocation section (NIM evaluated and deferred on ToS; Groq documented as strongest future candidate); new docs/launchers.md. Tests: core/tests/openrouter-worker-test.sh, adapters/claude-code/tests/launcher-test.sh (stubbed, zero paid calls).

Changed

  • kiro gateway roster drift absorbed (2.19.1). The kiro-cli 2.19.1 roster carries no OpenAI models and no claude-opus-5 (live-probed): kiro-openai backend disabled with a dated disabled_reason (templates kept for revival); second-opinion-review/second-opinion-verify/advisor fallback kiro-openai → null with rationale (a fallback must not silently change the vote's vendor — same rule as third-opinion-review); kiro-anthropic-top re-pinned to claude-sonnet-4.5 (1.3x). adapters/kiro/README.md tier table updated with the roster-recheck rule.

Security

  • Free-lane egress hardening (council + security-reviewer, 2026-08-25). openrouter worker/preflight + ox launcher: sensitive-cwd guard now strips trailing slashes, canonicalizes blocklist entries through pwd -P (symlink parity with the cwd side), fails CLOSED when no guard file resolves (an explicitly set OPENROUTER_SENSITIVE_PATHS_FILE is honored strictly — no silent template fallback), and announces its guard source / FORCE override on stderr. Prompt egress floor: hard byte cap (OPENROUTER_PROMPT_MAX_BYTES, default 256KiB) + credential-shape refusal (private key / AKIA / sk- / ghp_ / xox-) with loud AGENT_OPENROUTER_UNSAFE_PROMPT=1 override. Key-residue windows closed: TERM-first watchdog + TERM/INT/HUP→EXIT trap chain so the key-bearing curl config is scrubbed on timeout, interrupt, and hangup. ox launcher header now documents the isolation tradeoff (deny rules/hooks absent in gateway sessions) and env-token readability. Battery grew to 27 cases incl. a guard-bypass matrix (trailing slash, symlinked entry, fail-closed, /dev/null, egress floor, byte cap).

  • Conditional council auto-escalation — a council-scale diff can no longer be signed off by a solo Claude reviewer. New PreToolUse Task|Agent gate core/hooks/council-escalation-gate.py denies a plain code-reviewer dispatch when core/infra/council-threshold.sh judges the staged diff council-scale (line/file threshold, or a path in a declared risk area) and points the caller at /council-review --staged instead. Every other case is silent: non-dispatch tools, other subagents, small diffs, and any internal error all fail open — a broken gate must not block review, but it says so on stderr and in the audit log rather than failing open silently. Escape hatches, both deliberately narrow:

    • /council-review marks itself active through the gate's own --council-flag set|clear CLI (steps 0.5 and 6) so the council's own internal code-reviewer dispatch isn't denied by the gate that routed the caller there. The flag lives OUTSIDE the reviewed workspace (keyed by project root under ~/.agent/state/council, override AGENT_COUNCIL_STATE_DIR), is TTL-bound (AGENT_COUNCIL_ACTIVE_TTL_S, default 300s) and content-bound to the current diff's hash — so a stale flag, a flag for a different diff, and an in-repo file a reviewed diff could plant all fail to open the gate. Grant-side state outside the workspace follows the trust_tier.py precedent.
    • An identical dispatch re-issued after a denial is let through once, with a warning, so an agent that genuinely cannot run the council isn't trapped in a loop. The ledger entries expire (AGENT_COUNCIL_DENY_TTL_S). /wrap gains an advisory step 1d covering the path the hook can't see — edits made without any Task/Agent dispatch — recommending /council-review before a solo commit rather than aborting. Tests: core/tests/council-escalation-gate-test.sh (24 checks), core/tests/council-threshold-test.sh (19 checks).
  • /worker-setup skill — guided onboarding for the cross-vendor worker lanes (codex, antigravity, grok, kiro): a read-only status sweep of core/infra/backends.json (live jq query, never a hardcoded lane list), a cost-model/tier-allocation briefing before anything installs, then per-lane install → auth → verify with every real round-trip probe announced first (kiro's is billable). Never a dispatcher itself (core/infra/call-worker.sh is) and never a paid probe without explicit user approval. Test: core/tests/worker-setup-skill-test.sh (14 checks: frontmatter, plugin-cache-safe dispatch, vendor coverage, no model IDs, no credential collection).

  • council-review dispatch fixed for plugin-cache installs — skills/council-review/SKILL.md step 3 resolved core/infra/call-worker.sh via ${CLAUDE_PLUGIN_ROOT:-$PWD} instead of a bare cwd-relative path, which silently no-op'd every external lane on a plugin install (docs/claude-plugin-install-lifecycle.md §7 item 1 — now resolved for this one consumer; other consumers named in that item may still carry the gap). A missing dispatcher now names itself and reports every external lane absent rather than guessing a path. Test: core/tests/council-dispatch-path-test.sh (3 checks).

  • setup.sh — ~/bin created unconditionally + doctor PATH row — every install_codex/install_gemini/install_grok/install_antigravity/ install_kiro now calls a shared ensure_home_bin() instead of silently skipping the symlink when ~/bin didn't already exist; a NOTE prints once if ~/bin is not on PATH. New doctor row "worker symlink dir" (PASS when ~/bin exists and is on PATH; WARN naming exactly which half is missing, with the export one-liner) runs just before the existing generic worker-lane PATH sweep. Vendor-CLI-not-initialized NOTEs added to the gemini/grok/antigravity tier-seeding blocks (config dir missing -> told to run the CLI once, then re-run the install flag).

  • install_kiro() + --kiro flag (opt-in, not part of --all/default — kiro is metered/paid, same stance as --grok/--antigravity): seeds every shipped adapters/kiro/*.json.template into ~/.kiro/agents/ (existing user-owned profiles never overwritten), symlinks kiro-preflight, and notes the kiro-cli install command + KIRO_API_KEY auth when the CLI isn't on PATH. Always exits 0. Test: setup-doctor-test.sh (r)/(r2) sections — throwaway-HOME smoke (profiles seeded, pre-existing profile preserved, preflight symlink resolves) + the new doctor row's PASS/WARN branches (11 checks).

  • antigravity-preflight.sh hint bug fixed — the missing-on-PATH message wrongly told users to symlink into ~/bin/grok-worker (copy-paste from the grok adapter); now names ~/bin/antigravity-worker and mentions antigravity-preflight. Regression check added to core/tests/antigravity-worker-test.sh (17th check: no stray grok-worker string anywhere under adapters/antigravity/).

  • codex preflight is now auth-aware — core/infra/backends.json's codex backend switched preflight from ["codex", "--version"] (proves the binary exists, nothing about auth) to ["codex", "login", "status"] + preflight_timeout_s: 15. Measured 2026-08-20 on codex 0.147.0: logged out exits 1 and prints "Not logged in"; logged in exits 0 — a local check, no network call, not billable.

  • Lane cost-model docs — docs/model-routing.md § Cross-vendor lanes gained a "Lane cost models (2026-08-20)" entry (grok/antigravity/codex quota-based, kiro-* metered/paid including its own preflight) as the SSOT /worker-setup cites, plus a note that tier/cost-aware cross-vendor task allocation is a candidate /spec, not designed yet.

  • Council lens split — codex and gemini voting lanes get distinct review perspectives (decided 2026-08-20: perspective diversity over duplicate generalists). Each voting external lane's prompt is now a lens preamble + the shared core: codex (second-opinion-review) reviews implementation correctness (logic, edge cases, error handling, concurrency); gemini/agy (third-opinion-review) reviews architecture & consistency (design/simplification, contract coherence, doc–code drift). A lens states emphasis, not permission — either lane still reports any defect it sees, and cross-lens agreement keeps (strengthens) the ≥2-vendor high-signal rule. Grok's advisory seat stays deliberately unscoped. Framing lives in skills/council-review/SKILL.md step 1; role comments in core/infra/backends.json record the split.

  • Antigravity (agy) review lane restores the seated google reviewer — new adapters/antigravity/ (worker + preflight + tiers template + measured posture README), following the grok adapter pattern. The council's third-opinion-review lane, dead since the gemini CLI's individual OAuth was retired (2026-07), is enabled: true again via agy — Google's official successor — authenticating from the OS keyring (no API key; the GEMINI_API_KEY path is upstream-contradicted). Measured 2026-08-19 (agy 1.1.14): default headless mode fails closed on shell exec and created no file in any probe; worker forbids --dangerously-skip-permissions and runs under a sandbox-exec deny-write/deny-cred-read profile; argv is flags-before--p (flags after are misparsed); tiers differ by model (effort is baked into the ID). Live-verified: preflight exit 0 + one E2E council dispatch returned a real finding with status: complete. Tests: core/tests/antigravity-worker-test.sh (16), setup.sh --antigravity install path.

  • Rate-limit fail-open contract — a vendor quota/rate limit is now a distinct lane condition, not a generic failure: grok-worker.sh classifies the xAI free-tier "usage limit" ending (measured 2026-08-18: message + generic exit 1) as EX_TEMPFAIL (75), call-worker.sh maps 75 to status: rate-limited in the capture (escaping exit code stays 1), and /council-review reports the lane "rate-limited, retry later" and proceeds without it — decided 2026-08-19: the grok lane stays on the free tier and is skipped when exhausted; no upgrade prompt mid-review. Watchdog kills (124/137/143) are excluded from reclassification. Tests: grok battery case (g) — limit→75, plain failure→1, output still streams.

  • Grok (xAI) advisor lane — adapters/grok/ worker-lane bridge (grok-worker.sh stdin→prompt-file + sandbox-exec deny-write + neutral cwd; grok-preflight.sh exact-token probe; tiers template owns the model pin) and registry entry carrying the advisor-third role only — never a gate vote. Measured (0.2.118, 2026-08-19): the CLI's own flags (--tools '', --permission-mode plan, --deny, --disallowed-tools) do NOT block writes; only the OS sandbox does. The sandbox denies reads of credential stores (~/.ssh/~/.aws/…) and narrows writes to the run's WORK_DIR + the CLI session store minus its own tiers file (argv-persistence hole), forwards signals instead of exec so the prompt file never outlives the run, and validates $HOME before building the scheme profile (adapters/grok/README.md).

  • Gemini worker-lane contract restored — gemini-worker.sh tier bridge + gemini-preflight.sh probe + third-opinion-review role (fallback: null for vendor independence). Backend stays enabled: false: re-verified 2026-08-19, a cached oauth credential still gets 401 UNAUTHENTICATED; the registry's re-enable condition is a passing preflight, not a credential file.

  • /council-review skill — multi-vendor review council: Claude code-reviewer agent + codex + gemini in parallel, grok opt-in (--with-grok, advisory only). Synthesis order: mechanical capture status → dedup → citation verification against the actual files (hallucinated findings dropped and counted per lane) → severity re-rating → source tagging (≥2-vendor findings marked high-signal) → disagreement adjudication. False-council guard: an all-external-lanes-absent run must say so.

  • setup.sh --grok + worker-lane symlink installs for gemini/grok, and a generic doctor check: every enabled backend's cmd[0]/preflight[0] must resolve on PATH.

  • Tests: core/tests/gemini-preflight-test.sh (10 checks), core/tests/grok-worker-test.sh (14 checks: argv/prompt bridge + OS-level sandbox write-block, credential-read-deny, tiers-write-deny, unsafe-HOME regressions) — PATH stubs, zero paid calls.

  • Autonomous-loop runner core/infra/loop-run.sh + skills/loop + skills/harness-loop (P2-1 + P2-4 + O-2). The §5 loop's mechanical enforcement in one script — init creates the SQLite goal row (attempts modeled as supervisor-goal.sh waves) plus a JSON state file and the active-loop flag; attempt runs the grader (default grade.sh --base/--target, stubbable via LOOP_RUN_GRADE_CMD) under a background watchdog that TERMs then KILLs at a timeout (600s default), classifies the outcome (keep / discard / crash / timeout), appends exactly one row to .agent/loop/results.tsv via loop-ledger.sh, and prints a LOOP: continue / LOOP: stop(cap|circuit-breaker) verdict — two consecutive harness_score: 0 attempts trip the circuit-breaker, reaching the cap stops the loop, both remove the active flag. Git reset/keep of the candidate commit stays a skill-level step, not this script's job. skills/loop is the generic fresh-context/one-task/resumable protocol (O-2); skills/harness-loop specializes it to this harness's own reviewer prompts with §5.2's 9-step procedure verbatim, numbered, human-only-edited. core/tests/loop-run-test.sh (37 checks, hermetic temp-git-repo fixture) and core/tests/loop-skill-test.sh (13 checks, binds the SKILL prose to the mechanics with RED reorder/removal/token-drop fixtures) drive both. loop-run-test.sh needs sqlite3 + jq (supervisor-goal.sh's own hard dependency) and exits 2 when they are absent, so verify-all's discovered-check SKIP lane tallies it as skipped with its reason rather than as a pass.

  • Autonomous-loop grader core/tests/grade.sh (P2-2 + L-1 impl). The loop (§5) now has its grader: it runs the GATE floor (sanitize-audit / adapter-parity / hook-config-test / post-commit-autosync / gitleaks) and then evaluates each named failure mode in evals/failure-modes.yaml by RE-RUNNING the battery that encodes it — a mode PASSes iff its guard is still green (the hole stays closed), FAILs iff the candidate re-opened it. Output is the L-1 checklist (mode:<id> PASS|FAIL — reason) plus a rollup harness_score: X.Y = (#PASS) − 0.5×(#untrustworthy), the mandatory last line on every path so an empty grep is always a real crash. It fails closed to harness_score: 0 (= discard) on any untrustworthy path — GATE failure, a TARGET-boundary violation (off-target diff), or an unparseable rubric — the operator-chosen baseline verdict. On a clean tree it emits harness_score: 10.0 (10 code-guarded modes PASS; vacuous-parity is GATE-covered — its only guard, adapter-parity, is a GATE battery, and a second checklist path would be unreachable behind the GATE short-circuit, so it is reported N/A on the single GATE path — and review-false-clean is N/A), superseding the retired single-scalar 8.0. It is a loop-time tool (re-runs batteries; ~1 min) and is excluded from verify-all to avoid recursion; grade-test.sh (36 checks) drives it hermetically and gates mode↔guard-map drift, including a ban on GATE batteries in the guard map. The TARGET clean-tree check uses git status --porcelain, so an UNTRACKED tamper file (which the batteries would execute but a committed-range diff cannot see) also refuses the grade (fail-closed).

  • Append-only loop ledger core/infra/loop-ledger.sh (P2-3). The sanctioned writer for .agent/loop/results.tsv (untracked run state): a 5-column schema (commit / harness_score / duration_s / status / description≤80), a status enum (keep|discard|crash|timeout), numeric validation that rejects rather than coerces a malformed score/duration, and free-text sanitization. It only ever appends — the header is written once. Every successful append notarizes the ledger in a sidecar witness (<ledger>.witness = sha256 + line count) that notarizes a PREFIX: the next append REFUSES a ledger that is missing (delete-recreate), shorter than the witness (truncation), or whose first witnessed lines no longer hash to it (rewritten history), while rows beyond the prefix are append-only extension and are accepted — which self-heals the crash window where an append landed but the process died before the witness update (P2-4 kills runs on timeout, so that window is real). Tamper-evident, honestly bounded (a shell can still delete both files; that two-target act is what loop-write-guard escalates during a loop). loop-ledger-test.sh (30 checks).

  • Loop write-ban core/hooks/loop-write-guard.py (L-2). While an improvement loop is active (env AGENT_LOOP_ACTIVE=1 or a flag file), a Write/Edit to the grader/verifier surface (core/tests/, evals/) or a non-append rewrite of the results ledger escalates to ask (not deny — the calibration policy reserves deny for secrets) so a human stays on the loop and a run cannot silently rewrite the code that scores it. Outside a loop the guard is fully inert (zero added friction). Containment uses realpath, so a symlink into the guarded dir cannot dodge it. Wired into the Write|Edit|MultiEdit PreToolUse chain; loop-write-guard-test.sh (31 checks) covers the ask/allow matrix, the symlink-escape case, Bash write-path detection, the delete-recreate/witness escalations, and WHY/FIX tags.

  • B4 follow-ups: gitleaks allowlist regression + 5-surface drift tripwire + a real SKIP lane for discovered checks. core/infra/gitleaks-fire-test.sh proved the DETECTION side of the secret gate; nothing proved the ALLOW side. core/tests/gitleaks-allowlist-test.sh plants each gitleaks.toml placeholder shape (your_*_key, dummy_*, example_*, sk-proj-placeholder…, {{VAR}}, <your_x>, $USER_*_JWT) in realistic config/doc contexts and asserts a clean scan. Adversarial review reshaped it three ways: the control is now per shape rather than one aggregate scan (an aggregate passes on a single finding anywhere, which proved only 1 arm of 8); findings are read from gitleaks' JSON report by rule id, never from its exit status (non-zero means "leaks found" OR "gitleaks errored", and conflating them lets a broken control report ok); and the placeholder value is assembled at runtime so no secret-shaped literal is committed — a committed one would make the repo's own secret-scan depend on the very allowlist arm under test, blocking future tightening. It reports an honest ceiling: measured 2026-08-10, only 2 of 10 shapes are load-bearing (sk-proj-… and Bearer USER_B_TOKEN); the other 8 are caught by no rule with or without the allowlist, so their clean pass is stated as proving nothing rather than counted as coverage. A floor assertion keeps the load-bearing set non-empty. Separately, the enforcement surface (core/tests/, evals/, loop-write-guard.py, pre-tool-guard.sh, loop-ledger.sh, hooks.json, adapter.sh, gitleaks.toml) is declared five times — grade.sh's surface_list pathspec, its guarded_surface_re, loop-write-guard.py's _guarded_dirs/_guarded_files, and that module's GUARDED_TOKENS (hoisted to module level here), which is the only copy gating Bash writes: a file present in the other four but missing there escalates on Write/Edit yet stays freely rewritable via sed -i/redirect. grade-test.sh's new (g4) SURFACE DRIFT extracts all five LIVE (never hand-mirrored) and asserts they cover one surface, (g4-floor) adds a content floor so a synchronized shrink of every copy cannot pass, and (g4-mutation) drives each comparison through its own shared function with a dropped entry to prove the checks are sensitive rather than vacuously equal. verify-all.sh gained a SKIP lane for auto-discovered checks (exit 2 = inapplicable → tallied skipped, reason printed). Without it a battery gated on an optional binary could only signal "cannot run" by exiting 0, which the runner printed as PASS while discarding the battery's own SKIP text — and the CI verify job has no gitleaks installed, so the new allowlist battery would have reported green in CI having asserted nothing. verify-all-test.sh case (9) pins the lane, including that exit 1 is still a hard FAIL.

Fixed

  • Council escape hatch never opened for a session running in a subdirectory — council-escalation-gate.py keyed its state directory on the raw root string, which the --council-flag set CLI (git top-level) and the hook (the PreToolUse event's cwd) spell differently whenever the session's cwd is a subdirectory of the repo, or the checkout is reached through a symlink (macOS /tmp and /var are symlinks — three distinct keys for one repo were measured). The flag was written where the hook never read it, so /council-review step 0.5 was a no-op and the council's own reviewer dispatch was denied. Both paths now canonicalize through canonical_root() (git top-level, then realpath). Regression: gate test section F13 — verified to fail on the pre-fix code, not just pass on the fixed code.
  • /wrap and /verify-completion shelled into core/ by bare relative path — same plugin-cache defect class already fixed for /council-review: bash core/infra/council-threshold.sh and bash core/infra/call-worker.sh resolve against $PWD, which on a plugin install is the user's project, not the harness. Both now resolve via ${CLAUDE_PLUGIN_ROOT:-$PWD}, and /wrap reports a missing script as SKIPPED rather than reading its nonzero exit as "not council-scale". core/tests/council-dispatch-path-test.sh grew a 4th check sweeping every shipped skill for bare invocations of the three council dispatch scripts. The wider "skills shell into bare core/ paths" gap (7 SKILL.md files as of 2026-08-21, mostly core/tests/ helper calls) is NOT fixed here — it stays tracked in docs/claude-plugin-install-lifecycle.md §7 item 1, with a grep to reproduce the current list.

Known limitations (not fixed, tracked here)

  • Linux sandboxing for grok/antigravity worker lanes — both workers use sandbox-exec (macOS-only) for their write-deny/credential-read-deny profile; on Linux they refuse unless the explicit unsandboxed opt-out (GROK_WORKER_ALLOW_UNSANDBOXED=1 / ANTIGRAVITY_WORKER_ALLOW_UNSANDBOXED=1) is set. /worker-setup warns about this rather than solving it — a Linux sandbox equivalent is backlog.
  • Cross-vendor task-allocation routing — /worker-setup step 2 briefs on cost models and a general allocation shape (prefer free/subscription lanes for advisory volume, reserve metered kiro lanes for gate votes) but does not implement any routing logic. A candidate /spec item when this is taken up.
  • Preflight stderr is not redacted at the adapter layer — the skill now reports only a probe's exit code and first stderr line (the auth-failure branch is where a vendor CLI is likeliest to echo a key prefix or account identity), but antigravity-preflight.sh's emit_capture still strips control characters only. call-worker.sh itself is clean: preflight stdout/stderr go to /dev/null and never reach .agent/workers/*.md (only the argv does, via unavailable_reason). Adapter-level redaction of key-shaped tokens is backlog.
  • No CI gate on remote installers in shipped instruction files — the /worker-setup install step is fenced as user-run with an inspect-before-run note, but nothing mechanically stops a future skill from telling a session to pipe a remote script into a shell. A fifth supply-chain-scan.sh class over skills/, agents/, commands/, rules/ would make this a CI failure instead of a review catch.
  • Consumer repos do not inherit a .gitignore for .agent/workers/ — this repo ignores worker captures, but setup.sh --project ships no gitignore, so a consumer running worker lanes can git add -A full reply bodies of a reviewed diff. Surfaced by the 2026-08-20 security review; pre-existing, not introduced here.

[0.5.8] - 2026-08-20

Added

  • Grounded completion gate (P1) — a completion CLAIM ("done, tests pass") is now checkable against evidence, not just self-report. False-success claims are 45-76% of agent failures in the literature (arXiv 2606.09863); independent verification cuts that to ~3%, but LLM-judge-only blocking overtrusts itself (arXiv 2406.07791, 2410.21819) — so this lands as a deterministic hard gate with the judge kept strictly advisory on top. core/infra/completion-verify.py gains two opt-in flags, additive to its existing behavior (a regression case pins the no-flags output unchanged): --require-evidence refutes a claim that changed code but cites zero tests/assertions (new evidence dimension); --diff-base <ref> audits every ADDED line of git diff <ref> for a skip/xfail/it.only(/test.skip/ describe.only(/broad-mock marker, and flags a changed code file outside the claim's declared files/scope (new trajectory dimension; lockfiles/build artifacts are exempt via anchored patterns). New consumer core/infra/completion-gate.sh (docs/gate-registry.md completion-gate row — the first BLOCKING consumer of a completion-verify.py verdict, docs/scoring-convention.md): AGENT_VERIFY_BLOCKING=off|dryrun(default)|block, a liveness canary that fails the gate closed (exit 2) if the verifier cannot REFUTE an obviously-false claim, a JSONL sink written in every reachable mode, and a needs_semantic exit-0 advisory (never a hard block) when a claim carries no deterministically-checkable evidence at all. Wired into /supervise's wave-audit step via --verify-blocking; /verify-completion cross-references it as the blocking counterpart to its own advisory verdict. New failure mode false-done-claim in evals/failure-modes.yaml. Tests: core/tests/completion-verify-test.sh (+22 cases, r1-r9), core/tests/completion-gate-test.sh (new, 41 cases).

Fixed (code-review round, same day)

  • needs_semantic no longer downgrades a real REFUTED — the gate's "claim cites nothing checkable, dispatch a semantic pass instead of blocking" advisory used to trigger off core_total==0 alone, which also fires when --require-evidence legitimately refutes an empty claim against a real code diff; that refutation was silently getting waved through as advisory even in block mode. Now gated on whether the evidence/trajectory dimensions themselves failed, not just on citation count.
  • Claim path now falls back to the repo convention — .agent/claims/<slug>-w<wave-i>.yml first, then .agent/claims/<slug>.yml (skills/verify-completion/SKILL.md, docs/scoring-convention.md); neither existing is reported as a distinct "claim file missing" refutation, not completion-verify.py's generic parse-error text. skills/supervise/SKILL.md and templates/delegation-contract.md now instruct the wave's worker to write the claim file on completion.
  • dist/build/.agent/node_modules scope exemptions are repo-root anchored (^dist/, not (^|/)dist/) — src/dist/real.py no longer rides the build-artifact exemption.
  • Liveness canary checks both sub-checks independently (file-existence AND test-failure refutations both present), not just the top-level verdict — catches a partial death (e.g. a regressed _run() that always returns success) that a file-check alone would still report as REFUTED.
  • AGENT_COMPLETION_VERIFY_BIN is honored only under AGENT_REPRODUCE_TEST=1 — removes the arbitrary-binary-injection surface the seam had on a production block-mode run (security review finding).
  • Empty/unparseable verifier stdout now logs an explicit named refutation instead of a bare, unexplained REFUTED; the record also carries verifier_stderr (truncated) and diff_base_used (ref or null) for post-hoc debugging.
  • _TRAJECTORY_SKIP_PATTERNS gained @unittest.skip/skipIf(/skipUnless(.

Known limitations (not fixed, tracked here per code-review request)

  • TOCTOU on the sink path: resolve_sink's confinement check and the later >> append are two separate operations — a symlink swapped in between them could redirect the write. Not exploitable by the gate's own callers (no untrusted input reaches the sink path outside the env override already confinement-checked), but not hardened against a concurrent attacker with write access to the sink's parent directory. Backlogged, not fixed in this round.

[0.5.7] - 2026-08-02

Release-only version bump: ships the already-merged #101 (decision-time model-routing advisory + telemetry-digest.sh --model), #102 (cross-runtime harness docs), and #103 (circuit-breaker exit-code fix + verify-observer) to plugin consumers. The installed-plugin cache is keyed by manifest version, so without this bump claude plugin update was a no-op and the PreToolUse model-routing advisor never reached live sessions (measured: dispatches still inheriting the session top model four days after #101 merged).

Added

  • Decision-time model-routing advisory hook (core/hooks/model-routing-advisor.py). The 2026-07-11 audit measured 7/7 dispatches silently inheriting the session top model with the post-hoc observer already live — observation alone does not change dispatch behavior. This PreToolUse (Task|Agent) hook adds the decision-time counterpart: when a dispatch is about to inherit top (no registry pin, no call-time model override, and not the intentionally inheriting Plan agent) it injects a one-line additionalContext reminder pointing at the tier table. Advisory only — never blocks, never rewrites the model, always exits 0, swallows all errors (same fail-safe contract as the observer), writes no logs. The "no runtime model-switching" rejection in docs/model-routing.md § What-this-policy-deliberately-does-not-do is refined, not reversed: blocking/auto-switching/classifiers stay rejected; a deterministic decision-time reminder sits inside the line, and the dispatch decision stays with the dispatcher. Registered in hooks/hooks.json, the Claude settings template, docs/gate-registry.md (decision=advise), and core/hooks/README.md; battery core/tests/model-routing-advisor-test.sh (18 checks: warn on unpinned, silent on override/pin/Plan, fail-safe on broken stdin, zero observer-log pollution). Doc counts (hooks 25→26, tests 56→57) updated in §7 live-count declarations and READMEs.
  • Model axis in telemetry (telemetry-digest.sh --model). New mode (same architecture as --gates) summarizing .agent/logs/model-routing.jsonl: verdict distribution (override / pinned_specialist / inherit_top) and tier-multiplier-weighted relative spend (reuses manager-audit's LOW 0.15 / MID 1 / TOP 3.5 and its env seams). Loud SKIP when the log is absent; always exit 0. Adversarial security-lane hardening (1 MED + 2 LOW, same wave): per-record type validation routes malformed log lines (string total_tokens, non-string model/verdict) into the existing skipped_malformed tally instead of a traceback-then-exit-0 false-green, an outer guard turns any analysis failure into the same loud SKIP as the absent-log path, log-derived strings are whitelist-sanitized ([A-Za-z0-9:._-], 80-char cap) in the text report (JSON keeps raw values for machine consumers), and supervise Step 5 writes routing: skip (no routing log) — never clean — when the routing log is absent or empty (manager-audit short-circuits an empty log to lane-clean PASS, so absence must not be transcribed as cleanliness). Battery +16 checks total in telemetry-digest-test.sh (8 model-axis, 3 type-confusion, rendering-fuzz forged-line/raw-ESC probes).
  • M-9 backlog row — per-model harness re-tuning. Cursor's swarm experiment had to drop a frontier model mid-run (emphasis markers taken literally → runaway loop): external evidence that harness prompts need re-validation per model. M-9 adds a prompt-compat smoke row to the tier-promotion procedure (backlog only, no mechanism).
  • M-6 — cost-effective-harness concept doc. docs/concepts/cost-effective-harnesses.md distills the 2026-07 happytlog article + ClaudeDevs thread (intelligence-placement patterns orchestrator/advisor/verifier, the advisor checkpoint finding — mid-run re-ranking beats upfront frontier advice, measured anti-correlation — coordination-cost economics with the small-task inversion, prompt-cache worker-reuse), plus a dated 4-guidance audit table mapping each point to harness state. Backlog rows M-6/M-7 (✅, this PR) and M-8 (open — delegation economics telemetry: spawn-vs-reuse ratio, per-wave delegated volume) added to docs/harness-improvement-plan.md §4.10 with the M-series count declaration updated 5→9 (M-8 from the recovered commit, M-9 added in this PR below); article added to §8 references. (Recovered from stranded branch claude/cost-effective-harness-article, commit 7aaa587, 2026-07-16 — authored pre-0.5.x but never merged; landed here unchanged.)

Changed

  • M-7 — delegation-economics policy wiring. docs/model-routing.md gains an Intelligence placement — the advisor pattern section (three TOP placements by task shape; advisor rule: checkpoints stay TOP and recur mid-run — a single upfront TOP ranking is measured anti-correlated; no new mechanism, stays a documented convention) and two new Floors: a coordination-cost floor (boundary tokens billed ≥2×; below a threshold task size solo TOP is cheaper than orchestration — delegate only when delegated volume dwarfs the handoff) and prompt-cache preservation (route repeat calls to the same worker; fresh spawn re-pays the context write; verifiers always fresh — isolation beats cache). Supervise SKILL.md Model policy gains the placement corollaries paragraph and Step 2b's orchestration rules go four→five (Worker reuse (cache)); delegation-contract.md Wave shaping gains Handoff must pay for itself and Worker reuse over fresh spawns bullets. Enforcement map labels the new rules honestly as conventions (call-time choices are not statically verifiable). docs/README.md index gains model-routing.md (pre-existing omission) and the new concept doc. Docs only, no behavior. (Same recovered commit as M-6 above.)
  • Cursor swarm evidence folded into the delegation-economics canon. docs/concepts/cost-effective-harnesses.md gains a "Planner-output quality is a cost control" section (tier placement dominates the bill — hybrid at ~1/8 of frontier-everywhere at equal final score; ambiguous briefs convert planner savings into worker spend; low-correlation cheap review stacks; per-model re-tuning) with the Cursor source added; docs/model-routing.md Intelligence-placement gains the orchestrator corollary tying the restatement-quality lane and LE-9 into the cost path (relative numbers only, no price constants). LE-9's evidence column cites the same measurement.
  • Supervise Step 5 routing visibility. The manager-audit offer is promoted to an automatic post-run summary of the routing-waste + token-spend lanes, written to RECORD.md's new routing: stub field (full 4-lane audit stays an offer; still non-blocking). supervisor-goal.sh stub gains the field; battery +1 check.

[0.5.6] - 2026-07-28

Added

  • Machine-identity PII guard. After a full-history PII audit (350 commits, gitleaks + targeted scans; zero secrets, two low-sensitivity history-only artifacts accepted as-is by owner decision), the sanitize gate now forward-blocks personal machine identifiers: core/tests/sanitize-audit.sh gains a token group for real macOS home paths (placeholder spellings stay legal) and Apple mDNS device hostnames (shaped to never false-positive on .env.local / settings.local.json); the 3 pre-existing placeholder paths in the template/fixtures were normalized in the same commit. Paired RED-mutation cases extend core/tests/sanitize-audit-test.sh (catch real path, catch hostname, no-false-positive). Rule + git-author noreply-email guidance: rules/public-repo.md § Machine-identity PII, cross-referenced from AGENTS.md §2. Adversarial security-lane hardening (4 MED, same wave): hostname pattern covers macOS's real hyphenated ComputerName forms (Mac-mini/Pro/Studio); the PII tokens moved to a case-sensitive TOKENS_CS group (kills latent false positives on lowercase REST-route /users/ text and imacros-style .local filenames); --range now also scans each commit's metadata (author/committer name+email, subject, body) — the historical leak class was a hostname-bearing author field that diff lines never show; battery extended to 10 checks (hyphenated RED case, FP guards, metadata RED case).
  • Community files. Root CONTRIBUTING.md (verification battery + ground rules), CODE_OF_CONDUCT.md (Contributor Covenant 2.1, contact via issues — no personal email), .github/ISSUE_TEMPLATE/ (bug / feature); the PR template moved to the standard .github/ path.
  • docs/demo.md — three reproducible gate-catch scenarios (denied secret read, REFUTED false-"done" claim, caught PII plant) with real captured outputs and GIF-recording instructions.
  • docs/launch-checklist.md — maintainer-run distribution guide (directory listings, launch copy drafts); all submissions stay manual.

Changed

  • README star-readiness rework (en + ko). Pain-first opening with the install one-liner in the first screen, live CI badge, evidence-cited category comparison table, promoted blind-benchmark result (honest caveat intact), and stale ungated prose counts corrected to live values (agents 2→3, skills 8→9, tests 54→56, hooks measured as 21 wired / 25 scripts). Hero copy and launch-checklist wording keep the enforcement claim honest: secret access and destructive commands hard-deny; the spec gate is disclosed as observation-mode by default with block opt-in.
  • docs/benchmark/landscape.md self-row refresh. The document violated its own refresh cadence and still described this repo pre-0.3 ("no eval suite", 4 CI jobs, 2 agents/4 skills); the self-row, strengths and gap→backlog map now match the live repo (E-1/O-1/T-1..3/L-1/M-4/M-5 shipped; O-2/L-2 open), with a dated self-row verification line.
  • Guard trim (evidence-based, GT series). Executed the docs/freedom-enforcement-calibration-2026-07.md §3c "instrument, then recalibrate" follow-up against 30 days of gate telemetry (new §3e records the verdicts):
    • check-hardcoding.py: hard deny → default dryrun with AGENT_HARDCODING_MODE=off|dryrun|block (block restores the old deny); gains a firing sink (.agent/logs/hardcoding.jsonl, schema 2.0.0, reproduce_test exclusion) closing its UNINSTRUMENTED status; the docstring/error claim that hook-config.yml: hardcoding_patterns[] customizes it — never wired — is now honestly labeled as planned (T-4). Battery rewritten to 21 checks (3 modes, unset-env-defaults-to-dryrun contract, sink schema).
    • session-quality-gate.py: the subjective style scan (inline types, hex colors, console.log) is advisory by default — reported and logged, never blocking; the block decision is owned solely by the objective session.completion_tests layer (P3-1, unweakened). Opt back in with AGENT_QUALITY_STYLE_BLOCK=1. Battery +6 checks (style-only no-block + sink write, opt-in restore, completion-still-blocks).
    • High-firing safety gates (destructive / secrets / verify-bypass) are explicitly unchanged — firing volume is defensive track record, not creativity suppression.

Fixed

  • Adversarial-review hardening (2 lanes, 2026-07-27). code-reviewer (2 MINOR + 1 NIT) and security-reviewer (2 LOW, verdict APPROVE — both demoted gates confirmed to carry no security duty; pre-tool-guard.sh / secret-content-scan.py verified untouched):
    • check-hardcoding: unknown AGENT_HARDCODING_MODE values now warn loudly on stderr and degrade to dryrun (a typo like Block no longer silently weakens an intended hard gate); sink writes are confined to the repo root or the system temp dir (escaping/empty overrides fall back to the default in-repo sink — mirrors the digest's read-side confinement), and the same writer-side confinement was applied to spec-gate.py.
    • session-quality-gate: the block reason's FIX remedies now name only the layer that actually blocked (completion vs opt-in style), instead of always listing style fixes that no longer block.
    • telemetry-digest: reproduce_test-excluded records are now tallied and reported as suppressed instead of silently dropped, so a lingering AGENT_REPRODUCE_TEST=1 in a real session surfaces as an anomaly rather than invisibly blinding the DEAD/FATIGUE audit.
    • gate-registry: hardcoding row's decision column reads deny(opt-in; default=dryrun) so the scannable machine block can't be misread as a live deny.
  • telemetry-digest --gates double-counting. Two gates sharing one sink AND one guard value (secrets-bash vs secrets-content, both guard=secrets in security-violations.jsonl) each reported the union (both fired=2179; actual in-window 1353/826). count_sink now also matches the record's hook field against the registry hook when present (legacy hook-less records keep guard-only matching). Regression battery section (l), 3 checks.

Added

  • supervisor-ask gate registered (registry blind-spot closure). supervisor.py had logged ~440 records / 11 ask-intents in 30 days while absent from docs/gate-registry.md (firing but unreviewable). Ask-family records now carry guard/hook/reproduce_test stamps and the registry gains a supervisor-ask row — mode unchanged (measure first; the digest counts from registration forward since older records lack the stamp).
  • Roadmap v2 (§4.13, growth axes + GT/AP/MC series). docs/harness-improvement-plan.md gains a 5-axis growth plan (guard calibration / low-cost-high-power / beginner onboarding + proposal-based adaptation / plugin interop / model currency) that re-sequences the open backlog (P2, LE, T-4, H-4, W-10, O-2) into waves and adds only three new series: GT (this release's executed trim, 5/5 ✅), AP (proposal-based personalization: local usage aggregation → /tune proposal skill → append-only adopt/rollback ledger → onboarding presets; auto-change stays forbidden), MC (model-currency watch + new-model bench-replay procedure).
  • docs/model-routing.md Currency log. Dated re-verification section: 2026-07-27 entry confirms the generation-neutral reviewer alias pins self-update (no edit needed), the codex adapter profiles match the current lineup, and files the above-top client variant as a watch item gated on bench evidence (MC series) per effort-before-tier-up.

[0.5.5] — 2026-07-23

Added

  • Persona-review skill + orchestrator agent (citizen/user review lane). A new /persona-review <target> skill and persona-review-orchestrator agent seat a panel of grounded Korean citizen personas in front of UX / copy / content and report how ordinary users would react — a user-perspective lens beside code-reviewer (correctness) and security-reviewer (vulnerabilities), replacing neither. Ships a stratified persona catalog (skills/persona-review/personas/catalog.json, N=119) subsampled from the public nvidia/Nemotron-Personas-Korea dataset (CC BY 4.0) via DuckDB httpfs: age_group × province hard-balanced across a 7×17 grid (one persona per cell — every province gets an even seat, not a population-proportional one) plus occupation_group × education_level soft-balancing, with a 1-2-tag review_lens (UX/카피/접근성/신뢰/가격민감) per persona. A reproducible builder (skills/persona-review/scripts/build_catalog.py) and a determinism battery (core/tests/persona-catalog-test.sh: schema, CC-BY attribution, stratification sanity, skill/agent wiring) guard it. Personas are synthetic (no real individuals, no name/contact fields). (Converged from two parallel builders: an initial population-proportional HTTP-paging sampler was superseded by the DuckDB full-dataset hard-balanced one for genuinely even province coverage — the earlier approach left small provinces as few as 2 of 120 seats.)

[0.5.4] — 2026-07-21

Truthfulness repair wave. An external cross-AI audit (2026-07-21) found the docs claiming more than the code delivered and the doctor verifying installability but never installed state. This release makes the claims match the code, the two install paths match each other, and the doctor check the machine it runs on — and adds gates so each repaired drift class cannot silently return (version parity, hook-manifest parity, memory-dump pollution).

Added

  • Version-parity gate (core/tests/version-parity.sh + battery). README (en/ko) badge + status lines, plugin.json, marketplace.json, and the CHANGELOG's latest release heading must all agree on one version; any lag fails the suite (this release repaired a three-file drift: README and marketplace at 0.5.1 vs plugin manifest at 0.5.3).
  • backends.json v2 lane registry (PR #89). Cross-vendor worker lanes gain tiers, enabled/preflight flags, and status frontmatter, replacing the flat single-backend registry.
  • Doctor real-wiring checks (15–18). setup.sh --doctor now verifies the install is actually wired, not just installable: codex wiring (brain MCP [mcp_servers.brain] + codex-shell-wrap.sh in the live config; WARN when absent, FAIL when a wired path is missing on disk), gemini wiring (same policy for ~/.gemini/settings.json — doctor previously had zero Gemini checks; new GEMINI_SETTINGS seam), claude install path (reports whether the plugin or the shell install is live; WARN when both — hooks can run twice — or neither), and brain strict lint over the live store (WARN when the /brain-ingest promotion gate would be blocked). 12 new fixture cases in setup-doctor-test.sh.

Fixed

  • Brain W1 lint measures isolation, not just out-degree. lint.py W1 now fires only for notes with no typed edges in either direction and skips status: seed imports (seeds are declared-unconnected; distillations are status: growing, where the strict gate keeps its teeth). A freshly seeded store no longer fails --strict wholesale — previously 143 warnings made the /brain-ingest precondition (--strict at 0) permanently unsatisfiable.
  • doc-reality prunes .claude/. A live multi-session worktree under .claude/worktrees/ carries another branch's doc snapshots; the gate was scanning them and failing this branch for that branch's paths. .claude/ is untracked runtime state, pruned like .agent/.
  • Hook-manifest parity gate (core/tests/hook-template-parity.sh + battery). hooks/hooks.json (plugin path) and adapters/claude-code/settings.json.template (shell-install path) had silently diverged by six hooks; the template is now a faithful mirror of the manifest (SSOT: hooks/hooks.json) and the gate fails the suite on any future drift (path prefixes normalized; chain order enforced). Shell installs now also wire spec-gate.py, plan-scope-allow.py, model-routing-observer.py, rubric-commit-judge.sh, the WebFetch/MCP secret-content-scan.py matcher, and supervisor.py on UserPromptSubmit (replacing a stale plan-gate.py) — behavior change is observation-only since spec-gate/tdd-guard default to dryrun.
  • Plugin session capture. brain-capture.py added to the plugin Stop chain (hooks/hooks.json) — the 0.5.3 brain feature now actually captures sessions on the plugin install path, not just shell installs.
  • Codex profile templates refreshed to the GPT-5.6 family (PR #88). quick → light lane, deep → top lane at xhigh effort, ahead of the 2026-07-23 legacy-model sunset; model ids verified against the Codex CLI's own models cache, with a currency test (codex-template-currency-test.sh) pinning them.
  • Memory-pollution guard (core/tests/memory-pollution-guard.sh + battery). Fails the suite when an AI-memory plugin's session-context dump (observed injected into AGENTS.md) is present in any committable file — tracked or untracked-unignored — so a personal session log can never reach a public commit. Wired into the /wrap pre-flight gate list.

Docs

  • Docs truthfulness. README (en/ko) no longer claims spec-gate.py / tdd-guard.py "physically block" — both ship in observation mode (off | dryrun | block, default dryrun; only pre-tool-guard.sh always blocks) and the modes are now documented. The "same guardrails no matter which AI" claim is split into decision parity (machine-tested) vs per-runtime event coverage, with an honest Runtime coverage table. docs/architecture.md no longer claims the plan-approval flag has no consumer — both wired consumers (spec-gate.py, plan-scope-allow.py) are now named. Stale README counts corrected (39→48 test scripts, 7→8 skills — brain-ingest added to the catalog).

[0.5.3] — 2026-07-19

Consolidation note. The changelog was not rolled from [0.2.5] (2026-07-08) through the 0.5.3 cut, so entries for the 0.3.x–0.5.3 releases accumulated under Unreleased. They are gathered here at 0.5.3 rather than retro-split into per-version sections, because no version tags or per-release release: commits exist for 0.3.x/0.4.x/0.5.0 to place undated entries reliably (a guess would be fabrication). The inline (vX.Y.Z) / (PR #NN) labels below preserve the finer per-release provenance. Future releases roll Unreleased into a dated section.

Added

  • Candidate scoring + project rubric (PR #79). /spec --score-candidates scores a field of candidate problems/approaches numerically (candidates × dimensions → top-K) instead of a prose pick, mirroring the --interview submode. A domain-neutral project rubric (templates/rubric.yml.template → .agent/rubric.yml) is scored deterministically by core/infra/rubric-score.py (shared verdict schema, refute-by-default) and run per commit by the advisory core/hooks/rubric-commit-judge.sh hook — trust-gated to personal-tier repos only (grader_checks are shell commands and .agent/rubric.yml ships with the repo tree, so a foreign clone's rubric is never auto-executed) — and folded into verify-completion's semantic judge on-demand (the two-layer split that keeps the fresh-spawn judge off the commit path). Follow-up PR #80 hardened the hook's test battery: a .agent/rubric.json fallback (no PyYAML dep) makes the trust-gate tests self-standing, plus a foreign-origin collab case proving the owner-decides-alone branch also blocks auto-execution.
  • sanitize-audit.sh --range mode (PR #78). The domain-neutrality gate can now scan a commit range (PR/push span), catching taint that a single commit adds and a later commit removes — an add-then-remove sequence the per-commit scan misses. CI wires an explicit --range verification step.

Fixed

  • doc-reality.sh no longer scans gitignored .agent/ (PR #79). The gate's find walked .agent/plans/, so a /spec plan naming to-be-created files false-positived as phantom-path drift. .agent is now pruned (matching the gate's own "tracked *.md" intent), with a regression case in doc-reality-test.sh.
  • .gitignore — .agent/plans/ re-entry gap closed (v0.5.2). The runtime-state block enumerated .agent/locks|logs|state|workers/ but not .agent/plans/, so a git add -A could commit per-run records (RECORD/RESTATEMENT/PROPOSALS) — artifacts that demonstrably carry machine-local absolute paths. Audit note for the record: the tracked tree itself was verified clean (sanitize-audit tokens + manual grep — no personal paths shipped); this closes the door, it does not clean up a leak. If committed specs are ever wanted, un-ignoring .agent/plans/**/spec.md is the narrower future carve-out.

Docs

  • sqlite3 + jq declared as goal-mode prerequisites (v0.5.2). core/infra/supervisor-goal.sh hard-requires both (exit 127) but README Prerequisites, setup.sh --doctor, and the supervise skill never said so — the same undeclared-dependency class the jq/telemetry-digest fix already established as a bug (docs/harness-improvement-plan.md P1-5 — jq removed there precisely because a hard dep contradicts doctor's WARN tier). Now: README/README.ko Optional entries, doctor check 14 ("goal-mode deps", WARN-tier — goal-mode is optional), and a prerequisite note on the --goal-mode row in skills/supervise/SKILL.md.

Added

  • docs/concepts/fable-5-prompting.md — frontier-model dispatch guidance (v0.5.2). Distills Anthropic's Fable-5-class prompting guide into 8 harness-mapped rules (effort-before-tier-up backing, anti-wrap-up, evidence-grounded progress claims, boundaries+why, delegation, memory surface, two registers, no reasoning replay). Advisory: cited from the supervise Model policy, the delegation-contract template (evidence-citation
    • anti-wrap-up lines added), verify-completion (artifacts-not-replayed- reasoning hard rule), and docs/model-routing.md. Enforcement lane is backlog (LE-9), not shipped.
  • Agent SDK loop cross-check — docs/loop-engineering-audit-2026-07.md §4 (v0.5.2). Maps the SDK's loop controls (max_turns, budget, effort, permission modes, compaction, subagent isolation, result subtypes, hooks) to harness equivalents. One real gap found: no per-run turn/dispatch cap (token budget only) → LE-8; compaction/scheduling stay intentionally runtime-native. New backlog: LE-8 (dispatch cap), LE-9 (fable-5 prompt audit lane).

Fixed

  • v0.5.1 version bump — re-cut of stale 0.5.0 caches. PR #73 changed shipped content (manager-audit --since, Explore-MID fan-out exception, Step 0 run-start timestamp) without bumping the version, so installed 0.5.0 caches were cut at the earlier commit and silently missed those fixes. The plugin distribution is the git tree keyed by version — a content change without a bump leaves every existing install stale.

Added

  • manager-audit --global — slug-less routing sweep. The audit's routing-waste + token-spend lanes previously required a <plan-slug>, so ad-hoc Explore / execution dispatches outside any /supervise run — the common case for TOP-inherit leaks — were never swept. --global drops the slug and run scope, skips the slug-scoped lanes (restatement-quality, role-compliance), and reports inherit-top / floor / fan-out leaks across the whole model-routing.jsonl. Still after-the-fact, still WARN-only, still exit 0 (no runtime switching). Covered by 8 new cases in manager-audit-test.sh.
  • /manager-audit — meta-audit of the supervisor (v0.5.0). Four lanes answering what the supervisor cannot be trusted to answer about itself: restatement-quality (intake restatement exists, six sections filled, measurable criteria, scope-drift candidates), routing-waste (TOP-inherit leaks, verify/judge MID-floor violations, fan-out not at LOW), token-spend (relative dispatch cost = tokens × tier multiplier — LOW 0.15 / MID 1 / TOP 3.5, midpoints of the docs/model-routing.md relative ranges; still no price constants in the repo), and role-compliance (every wave audited, never-auto-retry honored, RECORD.md written, review lane after code waves). Split mirrors harness-audit: deterministic machine layer core/infra/manager-audit.sh (env-seamed, always exit 0, findings JSON)
    • interpreting skill skills/manager-audit/SKILL.md that judges semantic candidates and writes concrete patch proposals to .agent/plans/<slug>/PROPOSALS.md for one-click user approval. Proposals target conventions/templates/docs only — runtime model-switching stays rejected (docs/model-routing.md). Battery: core/tests/manager-audit-test.sh (28 checks, fixture-injected seams, including review-driven regressions: dangling-flag termination, BSD-grep whitespace sections, non-ASCII wave titles, remediated-FAIL non-flagging, mixed-tier fan-out evidence, unsafe-slug rejection). First real-run smoke (2026-07-17) produced two user-approved fixes: the fan-out lane now honors the documented Explore-at-MID exception, and --since <ISO-ts> scopes the audit window to one run (Step 0 records the run-start timestamp in RESTATEMENT.md for exactly this) — battery now 30 checks.
  • /supervise Step 0 "Intake restatement" — before plan validation, the supervisor now restates the user's chat prompt into a machine-checkable record (skills/supervise/templates/prompt-restatement.md: Original ask verbatim / Interpreted goal / Assumptions / Out of scope / Success criteria / Open questions), persisted to .agent/plans/<slug>/RESTATEMENT.md. Non-full-auto runs surface unresolved Open questions before Wave 1. Completion (Step 5) now offers /manager-audit <slug>, the meta-audit that grades this restatement along with routing waste, relative token spend, and role compliance.
  • model-routing-observer spend signal — each dispatch record in .agent/logs/model-routing.jsonl now carries prompt_chars (always) and total_tokens (best-effort probe of tool_response usage, null when the runtime surfaces none). Pure-observer contract unchanged (silent stdout, exit 0 always). This is the measurement seam for the manager-audit lane (relative dispatch-cost ranking — tiers and token counts only, no price constants, per docs/model-routing.md).

Removed

  • Gemini backend retired from backends.json — second-opinion-review now ships codex-only (fallback: null), and the gemini backend entry is gone. Real-call verification (2026-07-17) showed the path fails by default for individual installs: gemini-cli 0.44–0.46 oauth-personal is deprecated upstream (IneligibleTierError → Antigravity migration), and the API-key path demands paid prepay credits. A shipped fallback that cannot work out of the box misleads worse than no fallback. The dispatcher's fallback mechanism is unchanged and stays stub-tested; to re-enable, add a backend entry and point a role's fallback at it (the removed entry lives in git history).

Fixed

  • backends.json: gemini headless invocation was wrong — gemini -p requires an argument (Not enough arguments following: p on CLI 0.44.x); the prompt rides stdin, so the registry now ships ["gemini", "-p", ""] ("-p is appended to input on stdin" per the CLI's own help). Found by the first real-call smoke of the second-opinion lane — the PATH-stub battery can't catch a vendor's argv contract, which is exactly why the smoke run is part of the lane's rollout checklist.

Added

  • Cross-vendor second-opinion lane (core/infra/backends.json + core/infra/call-worker.sh): a Claude session can now dispatch Codex (primary) or Gemini (fallback) as a review/verification second opinion. The registry is the machine-readable role→backend SSOT; model names never appear in it — adapter profiles own the tier (docs/model-routing.md § Cross-vendor second-opinion lane). The dispatcher refuses without AGENT_WORKER_YES=1 (paid-call gate, env-only because a headless caller can't answer prompts), names a missing CLI (exit 127), preserves fallback reasons in the capture header, and kills hung workers at timeout_s (exit 124). Captures land in .agent/workers/<ts>-<role>.md. verify-completion documents an optional --second-opinion flag (evidence input only — gate logic unchanged). Tests: core/tests/call-worker-test.sh, ten contract paths on PATH-stubbed backends, zero paid calls in CI. (Benchmark input: netwaif/multi-agent-starter's backends.json + call_worker.sh role registry — design borrowed, no code.)

Removed

  • legacy/trim-2026-07-04/ removed from the shipped tree — the last legacy/ payload is gone (preserved on tag archive/legacy-trim-2026-07). The plugin package is the git tree (no exclusion manifest), so these 44K of retired agents/skills rode into every release. The six legacy/-scoped exclusion rules (gitleaks.toml, sanitize-audit.sh, supply-chain-scan.sh, doc-reality.sh, check-hardcoding.py, secret-content-scan.py) stay unchanged: CHANGELOG entries still reference historical legacy/… paths, and doc-reality needs the exclusion to keep ignoring them. Recover anything with git show archive/legacy-trim-2026-07:legacy/trim-2026-07-04/<path>. (Benchmark input: netwaif/multi-agent-starter ships a deliberately minimal generated tree — our tree-is-the-package model makes git removal + archive tag the only diet mechanism.)

Changed

  • Benchmark landscape: 2026-07-16 spot re-check recorded (docs/benchmark/landscape.md): netwaif/multi-agent-starter v3.3.0 (three patterns adopted — cross-vendor lane, install→validate pairing, checksum update mode), inkeep/open-knowledge v0.33.0-beta.5 (sandbox trial, 11-target init pollution measured, conditional-adoption verdict with --scope project --no-skills containment, GPL boundary noted, skill-symlink SSOT logged as a pattern candidate), and the implicit execution-dispatch permission surface closed by the /supervise pre-flight.

  • /supervise gained a dispatch pre-flight for edit waves. Root-cause of "the supervisor can't edit files": a dispatched subagent cannot answer a native permission prompt, and a background dispatch auto-denies any call that would prompt — so an edit wave sent out without a cleared permission surface silently loses its Edit/Write calls. The skill now requires either the plan-approval flag (/tmp/agent-plan-approved, arms plan-scope-allow) or project-layer Edit/Write allow rules before dispatching an edit wave, and falls back to foreground dispatch otherwise. Also made explicit: execution waves must never route to code-reviewer/security-reviewer (read-only toolsets, CI-enforced) — supervisor.py suggestions name the review lane, not the execution lane. Cross-referenced from docs/model-routing.md.

  • setup.sh installs now end in post-install validation. Every install path auto-runs the existing read-only --doctor diagnosis and the script exits non-zero on FAIL rows, so a broken install fails loudly at install time instead of at first use (AGENT_SETUP_NO_DOCTOR=1 skips — test seam / air-gapped bootstrap). Install paths also self-heal lost exec bits on core/hooks/* and adapters/*/adapter.sh before validating, so exec-bit-hostile distribution paths (ZIP download) don't hard-fail a check the script never remediated. Pattern adopted from multi-agent-starter's generate→validate.py PASS/FAIL pairing, reusing our existing doctor instead of a new validator.

  • apply_template() is now idempotent (checksum update mode). A target byte-identical to the fresh render reports up-to-date with no prompt; only a target that actually differs (user-customized, or template changed) asks before overwriting. Re-running setup over an existing install is now a quiet no-op update pass instead of a prompt per file — the equivalent of multi-agent-starter's update mode ("overwrite system files, preserve user data") for our template set.

  • Gate-registry correction: quality-completion RETIRE-CANDIDATE → KEEP-CONDITIONAL (same-day supersede). The retirement investigation refuted its own premise: session-quality-gate.py is also the enforcement layer for the landed P3-1 session.completion_tests feature, making it a conditional-path gate (fires only where a consumer declares completion tests or the default src/ style scan matches) — zero in-window firings is an adoption gap, not dead wiring, and the gate is battery-verified in verify-all. Deleting it would have silently unshipped P3-1. Recorded as a correction row rather than rewriting history; real follow-up is W-6 generalization/adoption. Lesson reinforced: a DEAD flag on a conditional-path gate needs a wiring-vs-adoption diagnosis before any retirement verdict.

  • Gate-registry DEAD review (2026-07-15) + audit follow-up bookkeeping. The four gates the T-2 digest flagged DEAD got their review verdicts recorded and last_reviewed bumped: project-policy KEEP (conditional-path), r4-mutex KEEP (contention gate, active multi-session ops), context-mode KEEP-CONDITIONAL (tied to the plugin's W-10 review), quality-completion RETIRE-CANDIDATE (two audits agree: wired but never fired, sink never created — retirement is a separate reviewed code PR). New backlog row W-10: third-party plugin-injection pollution scan (observed live: a plugin's injected <context_window_protection> block acting as session-wide instructions, independently flagged by two research subagents — the mirror image of the P3-4 self supply-chain scan). E-1 gains a design-input note: llm-council's anonymous peer-review → chairman pattern as an option for the real-LLM judge track (pattern only, no code import). Docs/bookkeeping only, no behavior.

  • docs/benchmark/landscape.md 2026-07-14 spot re-check. Five-lane web re-verification of the 2026-07-08 snapshot: claude-flow renamed ruflo (2026-02-27, npm/CLI unchanged); Aider marked effectively stalled (last release 2025-08); new "Field shift" note — Anthropic Dynamic Workflows GA (Pro plan 2026-07-02) erodes orchestration-first third parties but not this harness's governance position; re-check found no surveyed leader closing the prompt-only-enforcement gap. Survey-only change, no behavior.

Added

  • Evidence-first inventory — kill the ghost-specialist deadlock at its root (rules/policy/evidence-first.md + core/hooks/agent-inventory.py). A gate that demands a specialist with no in-runtime provider deadlocks the session: the gate blocks the work, and the very thing that would unblock it can't be dispatched (observed live — a stale plugin cache required ui-ux-director/fe-architect/… agents that exist only in the legacy/ v0 mirror, not in the active registry). New policy evidence-first.md names the underlying failure — asserting present state from memory instead of a same-turn read — and forbids demanding a provider you haven't confirmed exists. Enforced by a new SessionStart truth pass, agent-inventory.py, which reconciles the active registry (including consumer overrides CI never sees) against the agent *.md providers actually beside it: real (id has a sibling provider), ghost (id has none → quarantined, never demandable), discovered (an unwired provider). A filename is not evidence — a provider is a .md that carries YAML frontmatter, i.e. actually defines an agent. A stray README.md in the registry dir is not dispatchable, so it must never be discovered (--sync would wire it in as a bogus agent), and an id backed by a frontmatter-less .md lands in ghost, not real — the file exists but the runtime cannot dispatch it, and calling that "real" is the deadlock. This is also what makes the inventory strictly stronger than supervisor.py's own is_real_agent() file-exists check rather than a restatement of it, so has_provider()'s AND actually buys something. The verdict is written to .agent/state/agent-inventory.json (gitignored runtime state — no git churn, no drift-guard conflict) and supervisor.py consumes the ghost set as an extra dispatch-time quarantine source (has_provider), fail-open so a session with no inventory behaves exactly as before. Opt-in hybrid auto-correct (--sync / AGENT_REGISTRY_AUTOSYNC=1) additively wires discovered providers into the registry, copying each model: straight from the agent's own frontmatter so the additive write can't introduce the model drift registry-drift.sh check 4 forbids — it only ever adds, never edits or removes. Fail-open is end-to-end: the SessionStart path, supervisor.py's main(), and the manual CLI all exit 0 on a malformed consumer registry — the reconciler must never itself become the thing that breaks a session. AGENTS.md links the rule from "When in doubt." Battery core/tests/agent-inventory-test.sh (8 checks: real/ghost/discovered classification, provider-needs-frontmatter, inventory persistence + ghost-set readback, additive-sync-with-model-copy, fail-open on no registry, fail-open on a malformed one), auto-discovered by verify-all.sh. supervisor-dispatch-test.sh gains two cases: 15 — a ghost that guards a file_globs path, listed ahead of a real guard on the same path. The existing ghost case (7) is a keyword ghost declaring no globs, so it could never reach the file-glob matcher — leaving the deadlock's other form untested, and that form is the incident's own example (a retired edge-fn-dev guarding **/functions/**). It pins both halves of the contract: the ghost is hinted and skipped (continue, not return, so it cannot swallow the guard behind it), and no emitted ask ever names an undispatchable specialist. 16 — inventory quarantine end-to-end: a provider .md is on disk (so is_real_agent() passes) but the reconcile rejected it, and has_provider() must quarantine it anyway. This is the one case where the inventory layer adds power over the file-exists check, so it is the case that proves the layer is load-bearing rather than decorative.

  • model-routing-observer.py — measure the model-tier convention. A 2026-07-11 transcript audit confirmed the call-time model-override convention is not followed: 7/7 subagent dispatches in the audited session inherited the session top model. New PostToolUse (Task/Agent) pure observer classifies every dispatch as override / pinned_specialist / inherit_top into .agent/logs/model-routing.jsonl (analyze: jq -r .verdict … | sort | uniq -c) — measured before enforced, per the gate- registry philosophy. Never blocks, emits nothing, exit 0 always. Battery model-routing-observer-test.sh (15 checks). Companion policy fixes: docs/model-routing.md gains a "Built-in agents (Claude Code)" section (Plan=TOP inherit, Explore=MID default / LOW for bounded lookups — a deliberate exception to fan-out-LOW), and /verify-completion's general-reviewer dispatch now requires an explicit workhorse-tier override.

  • Skill A/B evaluation dataset (H-3, seed). evals/datasets/skill-ab.jsonl is the labeled seed for measuring whether a shipped skill earns its keep: does its description route the right requests (and not the wrong ones), and does running with the skill produce what a no-skill baseline would miss. 35 cases across the 5 shipped skills (spec, supervise, verify-completion, wrap, harness-audit) — 3 assertion cases + 2 trigger-positive + 2 trigger-negative each. Trigger labels are grounded in each skill's own when_to_use (positives) and NOT clauses (negatives), including cross-skill discriminators (e.g. "execute the approved plan" must route to supervise, not spec); every assertion's rationale quotes the skill's shipped description, so the seed is grounded, not invented. Fail-closed floor evals/baseline-skill-ab.json (min_cases, min_assertions_per_skill = 3, the shipped-skill list). New battery core/tests/skill-ab-dataset-test.sh (21 checks) validates only the seed's shape — parse, per-kind required fields, unique slugs, unknown-skill rejection, the ≥3-assertions-per-skill floor, ≥1 trigger-positive and ≥1 trigger-negative per skill, and that every named skill is a real shipped SKILL.md — with 7 malformed RED fixtures proving each guard goes red; it calls no model and is auto-discovered by verify-all.sh. The A/B runner that executes with-skill vs baseline and scores the assertions is a later increment (H-3 본체). See evals/README.md (Skill A/B track) for the n=35 and honest-ceiling disclosure.

  • Failure-mode grader rubric (L-1, doc portion). evals/failure-modes.yaml replaces the autonomous loop's single opaque harness_score scalar with a checklist of 12 named failure modes, each distilled from a real adversarial-review catch in this campaign (caught_in cites the item + PR): silent-drop, vacuous-green, vacuous-parity, glob-scope-miss, bypass-flag, unanchored-skip, infra-as-verdict, lexical-containment, injection-breakout, loose-coercion, stale-ssot, review-false-clean. Every mode carries a detection_signal (the observable a grader looks for) and a grader_check (the boolean question whose safe answer the candidate must satisfy). Naming the holes makes the grader adversarial the way a human review is — "strengthening verification" without naming failure modes does not stop metric-gaming (proxy-hacking measured at 73.8%). The §5.1 correspondence table's val_bpb row is rewritten to describe per-mode mode:<id> PASS|FAIL — reason emission, keeping a rollup harness_score: line on both the GATE-pass and GATE-fail paths so the §5.2 grep consumer and the results.tsv status enum stay intact (GATE-fail emits harness_score: 0 = discard, never an empty grep that would misclassify as crash). New battery core/tests/failure-modes-test.sh (25 checks) validates the file's shape through the same PyYAML parser the grader will use — schema_version, the ≥8-mode floor, unique kebab-case ids, and every required field non-empty — and proves each guard can go red with six malformed RED fixtures (too-few / missing-field / blank-field / duplicate-id / non-kebab-id / unparseable), auto-discovered by verify-all.sh. The grade.sh implementation that consumes the rubric is deferred to a later batch.

  • plan-scope-allow.py — plan-approved auto-allow accelerator. New PreToolUse (Write/Edit/MultiEdit, last in chain) hook: once the user approves a plan this session (the plan-gate.py flag), in-workspace non-risk edits emit permissionDecision: "allow" so the native permission prompt stops firing on every step of approved work. First permission-weakening hook in the harness — emits only allow-or-silence (never deny/ask), fail-open direction is silence, risk areas (spec-gate GUARD_PATTERNS) + .agent/hook-config.yml + .git/ + out-of-workspace (realpath containment) always pass through, env-gated AGENT_PLAN_ALLOW_MODE=on (default off, ships dark), sink .agent/logs/plan-scope-allow.jsonl, registered in docs/gate-registry.md. Battery plan-scope-allow-test.sh (27 checks incl. symlink/.. escapes, case-evasion, sink discipline). Side fix: README/README.ko hook-count drift (17 → live 19) corrected.

  • skills/harness-help/ — router skill (ask-matt pattern, user-invoked): main flow /spec → approval → /supervise → /verify-completion → /wrap, standalone /harness-audit, and what to do when a gate interrupts. Sync rule: any skill add/remove updates the router in the same commit.

  • docs/skill-authoring.md — skill-writing reference distilled from Matt Pocock's writing-great-skills (MIT, attributed): invocation axis, one trigger per branch, information hierarchy, leading words, checkable completion criteria, failure modes. Applied as a surgical pass over the five existing SKILL.md files (trigger dedup + explicit completion criteria).

  • Gate registry + fire-rate digest (T-2). docs/gate-registry.md is the SSOT list of every deny/ask/block gate with the model weakness it assumes and a last_reviewed date (an assumption expires). telemetry-digest.sh --gates cross-references it against the runtime firing logs and reports per gate: DEAD (0 in-window firings), FATIGUE (firings ≥ --fatigue, default 50), STALE (last_reviewed + --stale-days, default 90, is past), and UNINSTRUMENTED (emits a decision but writes no log — reported honestly, not mislabeled DEAD). Test-reproduction records (reproduce_test:true) are excluded from fire-rate so batteries can't inflate it. Still an observer (exit 0 always); telemetry-digest-test.sh gains a synthetic-registry battery covering all four classes plus the reproduce-test exclusion and missing-registry fail-safe.

  • Runtime enforcement of risk_areas.secrets.paths (P1-8). pre-tool-guard.sh now reads project-declared secret paths from hook-config.yml (via a bounded, metacharacter-rejecting hook_config.load_risk_area_secret_paths loader) and denies read/copy/exfil access to them — closing a field that shipped as schema but was never read at runtime. Additive: the built-in secrets/ guards run first and are never weakened; a config value can only add literal paths, never inject a pattern. risk-area-wiring-test.sh proves enforcement + loader safety bounds.

  • Remote-URL credential scan + gitleaks fire drill (W-3). A token baked into a git remote URL lives in .git/config, invisible to every content scanner — core/git-hooks/scan-remote-url.py flags an http(s) remote whose userinfo carries a password or a token-shaped value (no false positives on ssh / clean / bare-username URLs), wired as pre-push step 0 and a /wrap pre-flight. core/infra/gitleaks-fire-test.sh plants a synthetic secret matching the repo's own rule and asserts gitleaks catches it (PASS = gate live, FAIL = misconfigured allowlist, exit 2 = gitleaks absent/SKIP), so a clean result can be trusted. remote-url-scan-test.sh covers both.

  • Tests for plan-gate and tdd-guard (P1-3). Both hooks previously shipped untested. plan-gate-test.sh (7 checks; new AGENT_PLAN_FLAG seam so tests never clobber the live approval flag) and tdd-guard-test.sh (12 checks; isolated mktemp git repos, RGR red/green/no-test verdicts). tdd-guard's misleading comments claiming hook-config override of its risk-area patterns were corrected to match reality (the override is a tracked follow-up, not yet wired).

  • Real-LLM semantic judge (E-1, batch-3) — out-of-CI. evals/judges/llm-judge.py is the layer above the deterministic floor: it catches a cited test that carries a real assertion (so reference-judge.py CONFIRMs it) yet never exercises the claimed change — asserting on an unrelated function, a stale inline copy, a mock, or a tautology. It conforms to the same verifier interface (llm-judge.py --root <root> <claim.json> → shared verdict JSON), reads the cited test_sources and claimed files (bounded, realpath-contained), embeds them as delimited DATA in a prompt, and asks a real model via a subprocess CLI (LLM_JUDGE_CMD, default claude -p; LLM_JUDGE_MODEL / LLM_JUDGE_TIMEOUT). Refute-by-default (unparseable output, missing/mistyped keys, empty test_sources, a root-escaping path, or confidence < 0.6 → REFUTED) is kept distinct from fail-closed on infrastructure (absent CLI, timeout, nonzero exit, empty stdout, or an invalid strict-integer timeout → clear stderr error, nonzero exit, and no verdict on stdout, so a broken backend never masquerades as a confident label). Prompt-injection is contained, not merely hedged: embedded evidence is wrapped in per-call nonce markers (untrusted content cannot forge the closing marker) and any marker-shaped substring in content is defanged before embedding, so a test that embeds a literal closing marker cannot break out of the DATA quarantine. Labeled dataset evals/datasets/llm-judge.jsonl (10 semantic-hard cases, 5 CONFIRMED / 5 REFUTED — the deterministic floor scores 5/10 on it, by construction) with fail-closed floor evals/baseline-llm.json (min_cases 10, enforced by the LOCAL run, not CI). New deterministic battery core/tests/llm-judge-test.sh (34 checks) drives the adapter with a MOCK CLI, so it needs no real model and runs offline — auto-discovered by verify-all.sh. CI is untouched: it must never call a model, and the track runs locally with --repeat 1 (Pass^k's identical-verdict rule would be dishonest for a nondeterministic judge — flakiness is to be observed, not hidden). See evals/README.md (Real-LLM track). Honest ceiling stated: nondeterminism near the threshold, a residual prompt-injection risk from hostile prose that stays within the data block, and model-availability dependence.

  • Teaching gates (T-1). Every deny/ask/block reason emitted by pre-tool-guard.sh, secret-content-scan.py, check-hardcoding.py, and session-quality-gate.py now carries a fixed WHY: tag (which rule fired and why) and a FIX: tag (the concrete allowed alternative), so a blocked agent can self-correct instead of routing around the gate. Machine-enforced: the per-hook batteries assert WHY:/FIX: on every non-allow fixture, and a new core/tests/check-hardcoding-test.sh (14 checks) gives that hook its first dedicated battery.

  • Skill negative-triggers (T-3). All five shipped skills/*/SKILL.md descriptions now include at least one NOT … negative example (when not to fire), enforced as registry-drift.sh check 7 with fixtures (no-NOT → FAIL, NOT present → PASS, no frontmatter → FAIL).

  • Doctor codex-profile check (M-4). setup.sh --doctor gains check 13: WARN when the quick/deep tier-profile files are missing beside the local codex config (CODEX_CONFIG seam; skipped when no codex config exists) — the same "declared template vs actual" observer family as the plugin-cache and command-scan checks. Fixtures cover present/missing/absent-config.

Fixed

  • Guard false-positives (W-7). pre-tool-guard.sh no longer blocks a commit whose message merely mentions a destructive command (git commit -m "fix: guard rm -rf / patterns"): a preamble strips the inert message payload before the destructive guards (1–4) scan — but only provably-inert shapes (single-quoted, $/backtick-free double-quoted, or a quoted-delimiter <<'EOF' heredoc). An unquoted <<EOF body still command-substitutes at shell-eval time, so it stays fully scannable; the secrets guards always scan the full command. A new core/tests/pipefail-idiom-scan.sh gate is the W-7(2) regression floor: it flags unguarded zero-match count pipes (grep … | wc -l, n=$(grep -c …)) in strict-mode runtime scripts, self-checking against a bad/good fixture so it can't rot to always-green; the safe idiom is documented in AGENTS.md.

  • Graceful-wrap traversal guard. supervisor-goal.sh's _emit_graceful_wrap (the budget-limited stub writer) now refuses a traversal-shaped slug the same way write_record_stub does, so a weird plan slug can never place the stub outside its configured dirs (fail-safe skip; regression-tested).

  • Repo-native execution ledger (F-2). supervisor-goal.sh complete now drops .agent/plans/<slug>/RECORD.md — a mechanical execution ledger (waves from the live DB, plus PR / audit-verdict / carried-items slots the supervise skill fills) — so an execution record exists even on runtimes with no global recording layer. Session narrative stays with the global layer; the two never duplicate. Contract: never clobbers a RECORD.md the skill already wrote, and a ledger write failure never blocks completion (fail-safe); AGENT_PLANS_DIR is the test seam. /supervise Step 5 writes the same facts as its closing discipline on non-goal runs. A traversal-shaped slug (path separators or ..) is refused fail-safe — the ledger can never land outside the plans root (a hardening from this change's adversarial review). New battery core/tests/supervisor-goal-record-test.sh (12 checks, isolated in mktemp git repos). Same-PR bugfix the battery exposed: cmd_init's objective="${4:-$slug}" referenced slug on the same local line — bash 3.2 + set -u crashes on every objective-less init <slug> <waves> (latent: the documented 4-arg form never hit it); the declaration is now split.

  • /spec --interview — opt-in deep-interview submode (F-1). For requests fuzzy enough that a wrong guess commits the spec to the wrong shape, the one-shot "ask if ambiguous" brainstorm gains a structured question loop: an unknowns table marking each row decision-changing (Y/N), batched questions over the Y rows only (at most 4 per round, options + recommended default named), re-scoring after each round (answers resolve rows and surface new ones — decision-tree pruning), and two termination conditions (zero open decision-changing unknowns, or 3 rounds — leftovers carry into spec.md under ## Open questions). The Q/A trail lands in ## Interview log, so the spec shows why it has its shape. Opt-in by design: simple requests keep the single pass, and spec-gate.py is untouched — the enforcement boundary neither knows nor cares which submode produced the artifacts.

  • Delegation contract + orchestration guards (O-1). skills/supervise/templates/delegation-contract.md is the per-dispatch contract skeleton: the four elements (goal / output format / tools & scope / boundaries), an explicit **model**: field (execution waves name their tier instead of inheriting the expensive session default), and three sections absorbed from the 2026-07-10 external-harness comparison — Self-contained (subagents inherit no history; the contract carries every path, decision, and constraint), Constraints re-injection (each wave re-states its relevant constraint slice, not whole rulebooks), and Executable acceptance criteria (verify = a command's exit code by default; prose only with a stated reason). Wave shaping travels with it: fan-out cap 3–5 with a worked split-at-the-cap example, write single-threading (one writer per fileset), and verifier isolation (fresh spawn, end-state only). /supervise Step 2 now composes every dispatch from the template; /spec Step 3 defaults → verify: to an executable check. The guardable half is machine-enforced: registry-drift.sh gains check 5 (review/verify agents must carry a read-only toolset — a write-armed or allowlist-less reviewer fails) and check 6 (a tree that ships skills/supervise must ship the template with its **model**: field), each with injected-defect fixtures in registry-drift-test.sh (24 checks — inline and YAML-block tools: forms both parsed; a write tool smuggled into either form fails).

  • Clean-install CI smoke (M-5). New clean-install CI job: installs all three runtimes into a scratch $HOME non-interactively (AGENT_SETUP_YES=1 bash setup.sh --all), then asserts the install itself — all three runtime configs exist and the {{FRAMEWORK_ROOT}} placeholder was actually templated (anti-vacuous: a no-op or mis-templated install must not go green; this assert step, not the doctor, carries the install verification, since the doctor's checks are repo/env-scoped and tolerate an empty home — a gap this change's adversarial review measured live). The doctor must then report 0 fail against the scratch home, and a built-in mutation probe (a hook stripped of its exec bit must turn the doctor red, else the job fails) keeps the job's green load-bearing. Closes the cold-install gap the landscape survey flagged: "install once, use everywhere" was previously verified only by hand.

  • Harness self-audit skill + extracted registry-drift gate (H-2). skills/harness-audit/SKILL.md is an agent-driven, read-only self-audit that sits on top of the machine gates: one verify-all.sh dry-run, a per-check PASS/FAIL/SKIP table, an explicit citation of the P1-1 doc-reality verdict, and for each failure a root-cause + fix + backlog-follow-up. It interprets the gates; it does not reimplement them. Enabling refactor: the CI validate-plugin job's four inline checks (manifest required fields; hooks.json → executable core-hook resolution; agent name: frontmatter; registry↔agent model agreement) are extracted to core/tests/registry-drift.sh (a standalone, cwd-independent gate with a REGISTRY_DRIFT_ROOT test seam) with a non-vacuous fixture battery core/tests/registry-drift-test.sh (11 checks — each drift class injected and asserted caught). The gate is auto-discovered by verify-all.sh, closing the one machine check the unified runner was missing. Same-PR hardening from the 2026-07-10 workflow audit: the skill gains a runtime-layer step (setup.sh --doctor + core/infra/telemetry-digest.sh, with an unmeasured-is-unmeasured reporting rule) and a negative trigger in its description (T-3 applied early); doctor gains check 12 — the phantom-command scan (a runtime commands/*.md invoking a script that does not exist on this machine is a live failure path — reported as a WARN-only observation; AGENT_COMMANDS_DIR test seam, relative refs resolved against the runtime root, unexpanded $VAR refs skipped, and control characters stripped from echoed refs — an escape-sequence display spoofing hardening from this change's security review) with 7 new fixture checks in setup-doctor-test.sh (25 checks) — the rule-ification of a real orphaned-command failure found live in that audit. Backlog bookkeeping lands in the same PR: new §4.11 F-series (F-1 opt-in --interview deep-interview submode for /spec; F-2 repo-native RECORD.md execution ledger), the G-2 global-hygiene decision record, and a matching F-rows series check in doc-reality.sh.

  • doc-reality gate — the harness gates its own doc drift (P1-1). core/tests/doc-reality.sh (+ a 39-case battery) fails CI when a shipped doc contradicts the repo: (A) a referenced in-repo path that does not exist — scanned across every tracked *.md (recursive; nested READMEs and docs/** included), minus the forward-looking plan, the backward-looking CHANGELOG, and legacy/; fenced code blocks (0–3-space fences, CommonMark-tracked) are illustrative examples and stripped, while an unterminated fence is itself a malformed-doc failure; (B) the six backlog-count series and (C) the four artifact counts declared in the plan §7, cross-checked against the live repo. New CI job. Hardened over four adversarial-review rounds; also completed the phantom-ref cleanup it surfaced (core/hooks/README.md, core/hooks/secret-content-scan.py, core/hooks/context-mode-guard.sh, rules/policy/strong-goal-template.md).

  • Eval harness — labeled dataset + Pass^3 + regression gate (E-1, deterministic layer). evals/run-evals.py grades the completion verifier (core/infra/completion-verify.py) against evals/datasets/completion-verify.jsonl (12 labeled CONFIRMED/REFUTED cases): each claim must produce its labeled verdict, the suite runs three times (Pass^3) requiring identical results, and evals/baseline.json gates on a coverage/accuracy regression. New evals CI job; runner battery core/tests/evals-test.sh (28 checks). The LLM-judge semantic layer and the skill A/B dataset are later increments.

  • Eval harness — semantic track, deterministic floor (E-1, batch-2). A judge that catches green-by-construction tests — a cited test that "passes" but asserts nothing real. evals/judges/reference-judge.py consumes a claim's test_sources and classifies each file meaningful iff it holds at least one real, non-constant assertion (line-based bash+python heuristics), emitting the shared verdict schema; it is refute-by-default, reads are bounded, and no path escapes --root. It is deliberately biased to false-REFUTED over false-CONFIRMED (a completion gate must never bless a hollow test). Graded by the existing runner against evals/datasets/semantic-judge.jsonl (17 labeled cases, 8 CONFIRMED / 9 REFUTED) under Pass^3 + evals/baseline-semantic.json regression gate; wired as new steps in the evals CI job; judge battery core/tests/reference-judge-test.sh (52 checks, incl. leak-safety on unsafe ..//absolute/symlink paths and python constant-comparison triviality across decimal/hex/octal/binary/underscore/scientific number forms). This is the deterministic floor only: it catches syntactic triviality (no real / only constant assertions), not semantic triviality (a real-looking assertion that never exercises the changed code path) — that deeper judgment needs a real model and runs via skills/verify-completion or a pluggable --verifier, not in CI.

  • Unified verification runner — one command runs the whole suite (P1-2). core/tests/verify-all.sh fulfills the README "Verification" one-command promise. It discovers the check set dynamically — every core/tests/*.sh except the runner and its own test — so the four gates, all thirteen *-test.sh batteries, and any script added later are picked up with no edit here (the anti-rot property). It then runs the two evals layers (the exact CI invocation) and gitleaks, reporting PASS / FAIL / loud SKIP per check and a final tally; a silently-skipped check reported as pass is exactly the false-green this repo guards against. It is set -u (not -e) so every check runs even when one fails, isolates each in a subshell, and refuses to report success on an empty discovery (zero checks → non-zero exit, not a vacuous green). Battery core/tests/verify-all-test.sh (21 checks) proves completeness against the live dir (no hardcoded list), fail-propagation, all-green, list-equals-run, the empty-set floor, and that an absent gitleaks is counted skipped, never passed. New verify CI job runs the self-test then the full discovered set — the sole CI executor of the ten checks that had no dedicated job, and auto-inclusive of future ones.

Changed

  • README synced to 0.2.5 reality. Version badge/status, skill count and table (spec + verify-completion were missing), the model-tier paragraph rewritten to match docs/model-routing.md (judgment inherits session-top; reviewer pins are the only machine-enforced part; implementation/mechanical dispatch via per-call override — conventions), and the docs index gains model-routing.md + benchmark/landscape.md.
  • Cross-AI parity gate strengthened (P1-6). core/tests/adapter-parity.sh now asserts the three adapters (claude-code / codex / gemini) return the same normalized decision for a logically-identical event — and the same full decision JSON — across all three verbs (deny/allow/ask) and both tool_input shapes, each driven through a hook that actually reads that shape so a mistranslated field flips the decision: the command shape via pre-tool-guard.sh, the file/content shape via check-hardcoding.py (deny on hardcoded content, allow on clean) — plus shell-special (quoted) command and content. The prior version only checked a "deny" substring per adapter independently, so two adapters could disagree and still pass. 24 parity assertions; exit 1 on any divergence. (Kept the filename: cross-ai-parity.sh named in the plan was a phantom P0-1 removed, and the docs already reference adapter-parity.sh.)

Removed

  • legacy/v0-mirror-2026-05-12/ retired from the shipped tree — defence-in-depth for the ghost-specialist deadlock (preserved on tag archive/v0-mirror). The plugin package is the git tree (marketplace.json declares "source": "./" and there is no exclusion manifest — the only things missing from a release are the gitignored ones), so the v0 mirror rode into every release: 194 files, 1.3 MB, including 33 retired agent .md providers (ui-ux-director, fe-architect, edge-fn-dev, …) sitting next to two 64-entry master-registry.json files. That adjacency is exactly the shape is_real_agent() trusts — a registry id is "real" iff a sibling <id>.md exists — so any copy of that tree into an active registry path resurrects the deadlock the rest of this release exists to kill. find_registry() never reaches into legacy/, so this was a latent trap and dead weight rather than a live path (the live defence is agent-inventory.py above); removing it means a retired provider can no longer be resurrected by a stale copy. legacy/trim-2026-07-04/ stays — its 3 agent .md files have no sibling registry, so they cannot satisfy the predicate, and keeping legacy/ alive keeps the five legacy/-scoped exclusion rules (gitleaks.toml, sanitize-audit.sh, supply-chain-scan.sh, doc-reality.sh, check-hardcoding.py) valid and unchanged. Recover anything with git show archive/v0-mirror:legacy/v0-mirror-2026-05-12/<path>.

Security

  • Codex/Gemini adapter synthetic-mode no longer builds canonical JSON by string interpolation. adapters/{codex,gemini}/adapter.sh constructed the event JSON by interpolating --command/--content into a python3 -c literal ('''$TOOL_CMD'''), so a value containing a quote, newline, or ''' broke the literal — mis-parsing the event (a cross-adapter parity break) and, worse, letting a crafted command inject python and force an allow, bypassing the very guard the adapter feeds. The values are now passed via environment variables to a fixed python program (no interpolation), carrying any command/content verbatim. Regression-locked by the new quoted-command parity cases.

[0.2.5] — 2026-07-08

Changed

  • Model policy: judgment stays up, hands get dispatched. The supervise Model policy now names all four judgment classes (planning/design, wave dispatch decisions, gate verdicts & abort/advance, result synthesis) as session-top work, and adds an execution-dispatch row: implementation waves dispatch at the workhorse tier and mechanical work at the low tier via an explicit per-call model override — inline execution at the session's top model is the expensive default this rule exists to prevent. All of this is documented convention (only specialist pins are CI-enforced); O-1's delegation-contract template gains a required model field in its done-condition. docs/model-routing.md adds the matching orchestration- judgment row and reworks the Implementation row.
  • Doctor check 10 now scans every cached plugin, not just this harness. Any <marketplace>/<plugin>/ with more than one cached <version>/ is WARN-listed (a live third-party dual-version cache motivated the generalization — the same stale-cache drift class check 10 was built for). WARN-only, absent cache root still passes. Fixtures: 3 new cases (third-party dual → WARN, multi-plugin single → PASS, stray file at version depth ignored).

Added

  • Floors: long-horizon implementation is not a LOW-tier task. docs/model-routing.md cites an external program-based-verifier benchmark (datacurve-ai/deep-swe, 2026-05 leaderboard, press-reported) showing light-tier models trailing the top tier by ~40+ points on long-horizon SWE work — cost/performance reference data reinforcing that LOW is for bounded mechanical tasks only.
  • docs/benchmark/landscape.md — survey + self-assessment against the most popular agent harnesses on GitHub (2026-07-08 snapshot, API-verified stars): two comparison tables (Claude Code ecosystem / general harnesses), field investments vs field gaps, evidence-linked strengths and weaknesses, explicit non-goals with reversal conditions, and a gap→backlog map. Not a run benchmark — results.md remains the only measured comparison.
  • M-5 (backlog) — clean-install CI smoke: bare checkout → scratch config home → setup.sh --doctor asserts 0 failures. Distribution integrity was a consistent field investment in the survey; the cold-install path is currently only verified by hand.

[0.2.4] — 2026-07-08

Added

  • M-1 — docs/model-routing.md. Canonical cross-runtime model-tier policy: a three-rung ladder (LOW mechanical / MID workhorse / TOP reasoning) with an orthogonal effort dial ("effort before tier-up"), a work-class → tier table for Claude Code / Codex CLI / Gemini CLI, floors (verify-judge ≥ MID, fan-out workers default LOW — worker tier is the dominant cost lever at ~15× multi-agent token spend), and an enforcement map. Explicitly rejected: runtime model-switching hooks, automatic tier escalation, dedicated low-tier agents, price constants in the repo.
  • M-2 — verify-judge tier floor. /verify-completion's Layer 2 semantic judge is documented as never-below-workhorse (sonnet-class): a low-tier refute-by-default judge emits plausible false CONFIRMED verdicts and silently disables the completion gate. Low-tier sessions must pass an explicit model override on the judge dispatch.
  • M-3 — adapter templates carry the tiers. The Codex adapter ships quick.config.toml.template (LOW) / deep.config.toml.template (TOP) as per-profile config files (recent Codex CLI builds reject inline [profiles.*] tables as legacy — verified against a live CLI); the Gemini template ships a workhorse default model with explicit -m escalation. Model IDs are marked as 2026-07 snapshots.

Changed

  • skills/supervise/SKILL.md Model policy: the mechanical-fixes row now names its real mechanism (explicit per-call model override on the Agent dispatch — no low-tier agent is shipped) and the table links to docs/model-routing.md as its cross-runtime generalization.

[0.2.3] — 2026-07-08

Changed

  • I-1 — secret-content-scan matcher consolidation. The 6 non-edit matcher blocks (supabase ×3 tools, firecrawl ×5, WebFetch, Notion ×3, Google Drive ×2, stitch ×2) collapse into one union matcher; the registration inside the Write|Edit|MultiEdit chain stays to preserve chain order. 7→2 blocks, the 19 covered tools verified unchanged, no double-fire (each tool matches exactly one block).

Added

  • I-2 — doctor drift checks (10 & 11). setup.sh --doctor now warns when the plugin install cache holds more than one agent-harness version (a stale cache re-exposing retired agents was a live incident), and reconciles a user-declared global-hook manifest (AGENT_HOOK_MANIFEST, default ~/.claude/LOCAL-LAYER.hooks; AGENT_GLOBAL_SETTINGS for the live file) against the runtime settings in both directions (declared-but-not-live / live-but-undeclared). WARN-only — observers never block; no manifest → check skipped. Manifest lines are trusted substrings authored by the user (an over-broad line makes the check vacuous by choice). Fixtures: setup-doctor-test.sh 12 checks.

[0.2.2] — 2026-07-08

Cumulative since 0.2.0: the 0.2.1 plugin release (2026-07-07, model-routing policy + the P3-1/3/4/5 batch) was cut without sectioning this file, so the entries below span 0.2.1 and 0.2.2. Newly in 0.2.2: P3-2 (/spec + spec-gate), the freedom-vs-enforcement calibration audit, and the T/E + O/L/I backlog series.

Added

  • Freedom-vs-enforcement calibration audit + new backlog series. docs/freedom-enforcement-calibration-2026-07.md tiers every enforcement/ freedom point (deny/ask/block/advisory/observe/aspirational) and grounds keep/promote/measure-first verdicts in 2026 external evidence. New backlog: T-1 teaching gates (WHY/FIX in every gate message), T-2 gate registry + fire-rate + expiry, T-3 negative skill-trigger examples, E-1 public eval harness promotion (Pass^3); O-1 supervise delegation-contract revision (4-part brief, fan-out cap 3–5, single-writer, isolated verify lane), O-2 generic skills/loop (fresh-context / one-task-per-iteration / file+git state), L-1 failure-mode-checklist grader redesign, L-2 grader/tests write-ban + append-only ledger, I-1 secret-scan matcher consolidation, I-2 doctor cache/manifest drift checks (docs/harness-improvement-plan.md §4.8–4.9).
  • P3-1 — completion-gate test verification. core/hooks/session-quality-gate.py (Stop hook) now runs a project's session.completion_tests (declared in .agent/hook-config.yml|json) on the first Stop; any command that exits non-zero, times out, or fails to spawn emits {"decision":"block"} so a session cannot end while the project's own tests fail. A second Stop passes (anti-loop) and AGENT_QUALITY_GATE_BLOCK=0 is advisory; per-command bound is AGENT_COMPLETION_TEST_TIMEOUT (default 120s). New fail-safe/bounded loader hook_config.load_session_config (≤20 cmds, ≤500 chars each; malformed → empty). Trust model = the project's own scripts (docs/hook-config.md). Adversarial-review hardening of the always-exit-0 contract: the timeout env var is parsed at import behind a guard (a typo like 2m/30s degrades to 120 instead of crashing the Stop hook), and completion commands run with start_new_session=True so a process-group teardown idiom (kill 0, trap 'kill 0' EXIT) reaches only the command's own group, not the hook. Test: core/tests/quality-gate-completion-test.sh (21 checks incl. YAML path, advisory, anti-loop, malformed fail-safe, non-numeric timeout, process-group isolation, timeout→block, loader bounds).
  • P3-2 — /spec upstream-planning discipline + spec-gate enforcer. New skill skills/spec/SKILL.md walks brainstorm → .agent/plans/<slug>/spec.md + plan.md → ExitPlanMode approval, borrowing the superpowers methodology as CONTENT while the ENFORCEMENT is a tool boundary. New PreToolUse hook core/hooks/spec-gate.py is the consumer of the plan-approval flag that plan-gate.py already writes (and session-init/close already clear): when no plan is approved this session and a Write/Edit targets substantive impl code (scope covers src/app/pages/lib/ server/components, extension-gated), it acts per AGENT_SPEC_GATE_MODE — off no-op, dryrun (default) advisory-only, block emits permissionDecision: "ask". The approval flag is the dedup (approve once via ExitPlanMode → every edit passes; the two escapes — ExitPlanMode approval and AGENT_SPEC_GATE_MODE=off — are named in the reason). ask not deny, per the reversible-gate escalation principle. Fail-open (any exception → exit 0, empty stdout); mirrors tdd-guard.py. Wired into the Write|Edit|MultiEdit chain (after tdd-guard, before supervisor). Adversarial-review hardening (13-agent refute-by-default; 6 confirmed): SKIP dir tokens are now ANCHORED ((^|/)(types|config|…)/ so src/subtypes/ no longer inherits the types/ exemption — a MAJOR false-allow), the flag path is the shared hardcoded /tmp/agent-plan-approved with NO env override (a per-consumer override would decouple the reader from plan-gate's writer), the default scope was broadened past src/ to real app layouts, and matching is case-insensitive (so src/Pay.TS can't evade on a case-insensitive FS). Test: core/tests/spec-gate-test.sh (42 checks incl. anchored-skip, absolute paths, broadened scope + harness-not-gated, case-insensitive ext, fail-open, and the ask-not-deny + both-escapes assertions).
  • P3-3 — verification-gate-bypass + linter-tamper guards. core/hooks/pre-tool-guard.sh now asks on git commit/push --no-verify (and git commit -n), which skip the repo's own gitleaks+sanitize commit gate, and on Bash edits to linter/formatter/gate configs (eslint, prettier, ruff, flake8, biome, golangci, pre-commit, gitleaks) — the "disable the check instead of fixing the code" anti-pattern. git push -n (dry-run) and normal commits pass; reading a config passes. ask (not deny) per the escalation principle. Adversarial-review hardening: the no-verify guard now also catches git's bundled short-flag forms (-nm, -vn; git parses -nm as -n -m), a -c key=val global option before the subcommand, and inline core.hooksPath= — while a commit whose message merely mentions -n no longer false-asks (the message is stripped before matching); the linter-tamper rule no longer false-asks on a read that redirects elsewhere (cat .eslintrc.json > backup.txt) by requiring the config to be the mutate target. First test for this hook: core/tests/pre-tool-guard-test.sh (28 checks, incl. regression coverage of the existing destructive/secret rules).
  • P3-4 — self supply-chain scan. core/tests/supply-chain-scan.sh statically scans the harness's OWN shipped, auto-loaded instruction files (agents/, skills/, commands/, rules/, templates/, AGENTS.md, AI_BOOTSTRAP.md, CLAUDE.md) for three prose classes — prompt-injection override, unattended observer-loop persistence, no-confirmation coercion — and the auto-fired AI-decision hooks (core/hooks/) for a background-daemon-spawn class (nohup/setsid/disown/crontab), failing on any hit. Patterns are calibrated to zero hits on the clean tree; the no-confirm class is anchored on confirmation/permission/approval so a routing rule like "do not ask for a phantom agent" is not flagged. Explicitly-invoked plumbing with sanctioned one-shot async (autosync post-commit push, agent-session.sh subscribe) is out of scope and documented in rules/policy/security-guards.md. This is the self-integrity analogue of sanitize-audit.sh. Wired as a new CI job (4th). Adversarial-review hardening (13-agent refute-by-default pass) closed real detection gaps: instruction files are now scanned as *.md plus *.template (scaffolding copied verbatim into consumers) and *.json (the agent registry); hooks are scanned as every file (a hook may be extensionless); each prose class is matched against a whitespace-flattened copy so an injection wrapped across soft line breaks can't evade line-oriented grep; and the self-reference exemption is anchored to exact paths (a file merely named security-guards.md elsewhere no longer inherits it). Test: core/tests/supply-chain-scan-test.sh (20 checks: 4-class detection incl. templates, extensionless hooks, wrapped/multi-line, JSON registry, all four daemon tokens + clean-tree pass + false-positive guards for the phantom rule, start_new_session, infra-scope daemons, and the path-anchored exemption).
  • P3-5 — independent completion-claim verifier (eval-harness seed). A builder-validator layer that re-checks a completion claim from a separate context, so "the builder says it's done" is never the last word. core/infra/completion-verify.py is the deterministic core: given a claim (.agent/claims/<slug>.yml|json declaring cited files / tests / assertions) it mechanically checks each file exists (and contains its declared substring), each test exits 0, and each assertion holds, then emits a shared-convention verdict JSON ({verdict, score, dimensions, refutations}). Refute-by-default: a missing file, a failing test, a malformed/empty claim, or an over-cap claim all resolve to REFUTED (exit 1) — never a crash, never a silent pass; exit 0 iff CONFIRMED, so it doubles as a CI/wave GATE. Bounded (≤50 files, ≤20 tests/assertions) and hardened with the same start_new_session=True + guarded-timeout-parse lessons as P3-1. skills/verify-completion/SKILL.md is the semantic layer on top — an independent-context judge that adds the "does the code actually do what the claim says / are the tests meaningful" review scripts can't make, emitting the same schema. New shared spec docs/scoring-convention.md unifies the verdict shape across the verifier, the H-3 skill A/B harness, and the supervisor-goal-audit.sh 25-point scorer. Adversarial-review hardening (11-agent refute-by-default pass): a present-but- non-list section ("files":"x") now REFUTES instead of being silently dropped into a false CONFIRMED; command output is discarded to DEVNULL and the contains read is bounded (5 MB) so a chatty command or a huge file can't OOM the verifier. Test: core/tests/completion-verify-test.sh (25 checks incl. false-claim→refuted, consistent→confirmed, malformed/non-list-section fail-safe, YAML, process-group isolation, over-cap refutation, bare-string file entry, --root default, timeout degrade, partial score).
  • core/infra/telemetry-digest.sh (P1-5) — pillar④ janitor step 1: reads .agent/logs/supervisor.jsonl (path arg, or $AGENT_TELEMETRY_LOG, or <repo-root>/.agent/logs/supervisor.jsonl) and reports action counts, a per-specialist funnel (match → ask → dispatched, with conversion %), top keywords by match count, and a rule-candidate section with three heuristics derived only from what supervisor.py already logs: NO-ACCEPT (a specialist was asked ask-intent+ask-security ≥3 times but never dispatched — specialist-routing.md Lesson 1), GHOST (a specialist logged action=="ghost" — registry references an agent id with no sibling agents/<id>.md), and OVER-GENERAL (a single keyword accounts for >70% of all match records, once total matches ≥3). --window <days> (default 30) scopes records by ts; --json emits a machine-parseable JSON blob instead of the human report. Known limitation, documented in the header: NO-ACCEPT can't fire under AGENT_SUPERVISOR_MODE=observe (it logs observe-intent/observe-security instead, which aren't counted). Dependencies are bash + python3 only — no jq (an earlier draft used jq, which contradicts setup.sh --doctor's own WARN-tier/optional stance on it; JSON parsing runs through an embedded python3 heredoc instead). Always exits 0 (observer, not a gate) — a malformed line, a missing file, or an internal error degrades to a zeroed/"inactive" report rather than failing. Reproduce suite: core/tests/telemetry-digest-test.sh (21 checks: action-count accuracy, funnel notation, all three rule candidates, malformed-line skip counting, --window filtering, missing-log handling, --json output validity, legacy v0.1-record degradation).
  • core/hooks/supervisor.py v0.2 — minimal dispatch-not-advise router (P1-4). Replaces the observation-only v0.1 stub. On UserPromptSubmit it word-boundary matches the prompt against each registry agent's matches.keywords and records a 30-min TTL intent in .agent/state/supervisor-intent.json; the next Write/Edit/MultiEdit then returns permissionDecision: "ask" naming the specialist (once per intent — no repeat nag), and dispatching that specialist via Task/Agent (namespace-agnostic: x:code-reviewer resolves code-reviewer) clears the intent. A separate security matcher asks on matches.file_globs for the tool, independent of intent, once per path. Ghost specialists (a registry id with no sibling <id>.md) never ask — stderr hint + {"action":"ghost"} log only (specialist-routing Lesson 2). AGENT_SUPERVISOR_MODE=observe downgrades every ask to stderr; any exception is fail-open (exit 0, empty stdout). Wired into hooks/hooks.json on all three events. Reproduce suite: core/tests/supervisor-dispatch-test.sh (10 scenarios).
  • session-init now warns (stderr only) when gitleaks or git is missing from PATH — a mini env-doctor surfacing a degraded secret-scan setup at session start. Silent when both are present; never blocks the session or writes stdout.
  • setup.sh --doctor (P1-7) — full environment diagnosis, read-only, zero install side effects. Checks: git present; python3 >= 3.9 (README-declared floor, with version + path); gitleaks present (WARN if not — secret-scan git hook skips); jq present only if a core/hooks/*.sh script actually shells out to it (WARN if used-but-missing); every core/hooks/*.sh/*.py has its executable bit (hook_config.py exempt — a library module imported by secret-content-scan.py, never invoked directly); every adapters/*/adapter.sh is executable; agents/master-registry.json parses and every entry's model matches its sibling agents/<id>.md frontmatter (same drift guard as CI); hooks/hooks.json parses and every referenced hook script exists and is executable; ~/.agent/plans exists (WARN + mkdir hint if not). Prints a [PASS|WARN|FAIL] row per check plus a doctor: N pass, N warn, N fail summary line; exits 1 iff any check FAILs. Reproduce suite: core/tests/setup-doctor-test.sh (clean-repo exit 0, gitleaks WARN under a restricted PATH, exit 1 + named FAIL line when a hook loses its executable bit — exercised against a throwaway mktemp copy, never the real tree).
  • templates/hook-config.yml.template (P1-8 partial) — ships the real, dynamically-loaded python_hooks: schema (core/hooks/hook_config.py) as a commented example, bracketed by LIVE-SCHEMA-EXAMPLE-BEGIN/END markers, clearly labeled as the ONE block in the file actually read at runtime — distinct from the risk_areas:/resources:/hardcoding: blocks above it, which remain declarative-only (docs/customization.md part 2). New drift-guard case in core/tests/hook-config-test.sh (case h) extracts and uncomments that example into a real .agent/hook-config.yml and round-trips it through hook_config.load_extensions(), so template and loader can't silently drift apart. docs/customization.md cross-references the new template section.
  • .github/workflows/ci.yml — CI: gitleaks secret scan + plugin manifest/hook/agent validation + sanitize gate
  • README portfolio polish: badges, Mermaid architecture diagram, agent/skill/hook catalog
  • README.ko.md — Korean mirror of the README (same sections, localized prose)
  • docs/harness-improvement-plan.md — audit scorecard + prioritized backlog + autonomous improvement-loop design (Korean)
  • gitleaks.toml — detect NVIDIA NIM API keys (nvapi- prefix; built-in rules miss it)
  • docs/architecture.md — "Determinism and model-invariance" section: the hooks (gates) are model-invariant and machine-proven so via core/tests/adapter-parity.sh; risk-area denial is a real enforced gate while plan-mode/TDD enforcement is not yet wired (flag is recorded but unconsumed — see P1-4/P1-8); generated content (plans, code, prose) is honestly NOT guaranteed identical across models

Changed

  • docs/harness-improvement-plan.md — added §4.7 P3 series (5 items) from a 2026-07-06 benchmark audit of 8 top personal harnesses (superpowers, ECC, karpathy-skills, gstack, revfactory, hooks-mastery, Chachamaru, showcase): completion-gate test verification (P3-1), upstream spec/plan discipline (P3-2), --no-verify/linter- tamper blocking (P3-3), self-supply-chain scan (P3-4), independent completion-claim verifier (P3-5). Adoptions ranked by demonstrated mechanism value, not stars; catalog-maximalism, prompt-coercion, and unattended instinct-persistence explicitly rejected as design-principle conflicts. §7 backlog count 24 → 29 (P3 pattern P[0-3]).
  • agents/master-registry.json — supervisor keyword matchers hardened to domain anchors (review follow-up, MAJOR-1; specialist-routing Lesson 1). code-reviewer drops bare review/look over for multi-word phrases (code review, review this diff, …); security-reviewer drops bare security/auth for security review/security audit/owasp/… — generic tokens a consumer writes as often as an author (review my plan, the auth flow) no longer false-route a specialist. security-reviewer.matches.tools gains MultiEdit (aligns with the Write|Edit|MultiEdit hook wiring); file_globs (path anchors) unchanged. Default mode stays dispatch — only the match surface narrowed, not the enforcement.
  • core/hooks/README.md — supervisor.py moved from the deferred-roadmap "generic stub" row to the shipped-hooks table as the v0.2 minimal dispatcher; the roadmap row now scopes the remaining deferral to the full 54KB registry-aware orchestrator
  • README rewritten for first-time readers: concept primer table, install-path chooser (plugin vs shell), prerequisites section (incl. previously undocumented python3 dependency), "See it work" example, 4-layer architecture summary, trimmed layout tree
  • Hook count corrected everywhere: 17 executable hooks + 1 shared module (hook_config.py) — previous "~25" claim was stale
  • README.md/README.ko.md's "Why AI-agnostic?" section now cross-links to docs/architecture.md's new "Determinism and model-invariance" section

Fixed

  • plan-gate was wired to UserPromptSubmit in hooks/hooks.json but is a PostToolUse hook (its docstring and logic key off tool_name, a field absent from UserPromptSubmit events). Result: every invocation was a silent no-op and the /tmp/agent-plan-approved flag was never written. Rewired to PostToolUse with matcher ExitPlanMode|Task|Agent, and broadened the plan-class check to accept the Task tool name (subagent dispatch differs by Claude Code version). READMEs' hook tables corrected to match.
  • session-quality-gate wrote its violations log to parents[2] of the hook file — the plugin install cache when installed as a plugin — instead of the user's project. Log destination is now resolved at runtime: stdin event cwd → CLAUDE_PROJECT_DIR → os.getcwd(). Detection and block logic unchanged.
  • session-init crashed at load on Python 3.9 — its pathlib.Path | None return annotation (PEP 604) is evaluated at def-time and raises TypeError before 3.10. Added from __future__ import annotations so annotations are treated as strings; the annotation itself is unchanged. Supported Python floor is 3.9 (now documented in README Prerequisites).
  • Phantom test paths removed from README.md, AGENTS.md, docs/architecture.md, docs/getting-started.md — core/tests/adapter-smoke/*/run.sh, cross-ai-parity.sh, verify-all.sh, bootstrap-test.sh, and a pytest invocation never existed; docs now reference the 4 real test scripts (sanitize-audit, adapter-parity, hook-config-test, post-commit-autosync-test)
  • Documented overwrite behavior corrected: setup.sh has no --force flag — replacements prompt interactively, or set AGENT_SETUP_YES=1
  • README.md infra path corrected: scripts/infra/agent-session.sh → core/infra/agent-session.sh
  • AI_BOOTSTRAP.md Step 5 pledge now names the generic 5 risk areas (per hook-config.yml / rules/policy/security-guards.md) instead of prior-project domain terms; the removed terms were added to the sanitize-audit token list (failure → new rule)
  • core/tests/sanitize-audit.sh now scans git-visible content only (tracked + untracked-unignored via git grep --untracked), mirroring the CI job's excludes — runtime state and gitignored local files no longer cause permanent false FAILs; CI sanitize job additionally runs the full token-set audit as a superset step
  • core/hooks/secret-content-scan.py plan-file comment corrected to the canonical ~/.agent/plans/ path
  • Docs drift sweep: removed phantom hook/file references (memory-explore-verify.py, claude-mem-watch.py, rules/policy/skill-adoption-comparison.md) that described tooling never implemented; standardized risk-area vocabulary to the canonical data / secrets / deploy / payment / domain-output IDs across README, README.ko, docs/customization.md, and docs/concepts/security-guards-generic.md; corrected the security-guard layer count from 5 to 6 (matches hooks.json's "6-layer secret hardening"); canonicalized stale .claude/ path references to the runtime's actual .agent/ and rules//skills/ locations in docs/concepts/multi-session-worktree.md, AI_BOOTSTRAP.md, and docs/concepts/plan-mode.md; and removed the .claude/rules/ scaffold over-claim from docs/architecture.md, README.md, and README.ko.md (setup.sh --project never creates it)
  • Docs drift sweep follow-up: removed the two remaining phantom rules/policy/skill-adoption-comparison.md references (docs/master-registry.md, skills/README.md); removed the phantom classify-prompt.py hook citation from docs/concepts/plan-mode.md (no UserPromptSubmit hook exists beyond agent-session-heartbeat.sh — tier classification is the AI applying the documented heuristics itself, not an automated hook). Rewrote docs/customization.md end to end after discovering its documented hook-config.yml schema doesn't match what any hook actually loads: only core/hooks/hook_config.py (used by secret-content-scan.py) reads a config file dynamically, from .agent/hook-config.yml/.json — a [regex, label]-pair secret_patterns/exempt_paths/credential_key_names schema, optionally nested under python_hooks:. The previously-documented risk_areas:/resources:/ hardcoding: map (from templates/hook-config.yml.template) is not read by any hook at runtime — pre-tool-guard.sh, r4-mutex-check.sh, and check-hardcoding.py each match against patterns hardcoded in the script, not a project's hook-config.yml. The doc's old secret_patterns example ({id, description, regex} objects) was also independently confirmed to silently parse to an empty list under the real loader, which expects [regex, label] pairs — verified by running hook_config._coerce_pattern_list directly against both shapes. README.md/README.ko.md's customization section and other doc mentions of a dynamically-loaded risk_areas:/resources: still describe the same not-yet-implemented mechanism and were out of scope for this sweep — flagged for a follow-up pass.
  • Docs drift sweep, final pass: closed out the flagged follow-up above. README.md/README.ko.md's Customization section no longer claims core/hooks/r4-mutex-check.sh "reads [hook-config.yml] and enforces it" — the risk_areas: block is now described as declarative (a documented policy record), with today's actual enforcement attributed to each hook script's own hardcoded patterns and the one dynamically-loaded mechanism (secret-scan extensions via .agent/hook-config.yml) called out, linking to docs/customization.md for the full real-vs-documented split. AI_BOOTSTRAP.md's Step 5 pledge softened from "definitions live in hook-config.yml" (implies runtime consumption) to "declared in hook-config.yml; enforcement currently lives in the hook scripts." Same fix applied to the last two remaining spots the gate caught: docs/concepts/security-guards-generic.md's "How to extend" section no longer claims "the same pre-tool-guard.sh reads this and enforces it" or shows a fabricated abort_code key — the example is now framed as declarative intent requiring a pre-tool-guard.sh fork to enforce, with a link to docs/customization.md. rules/multi-agent-worktree.md's R4 mutex-resource list dropped the phantom payment-live entry — core/hooks/r4-mutex-check.sh only ever claims production-db, production-deploy, or edge-function-deploy; there is no payment mutex.

Removed

  • (recorded retroactively — the trim shipped before 0.2.0 but was never logged) Shipped agent set reduced 10 → 5 (architect, code-reviewer, security-reviewer, test-engineer, build-error-resolver) and skills 16 → 4 (supervise, tdd, diagnose, wrap); the removed items remain available in legacy/
  • Shipped agent set reduced 5 → 2: architect, test-engineer, and build-error-resolver archived to legacy/trim-2026-07-04/agents/. Basis: 7 weeks of session telemetry showed zero dispatches for these three, and their roles are covered by other tooling. code-reviewer and security-reviewer are retained (they form the benchmarked review pair — see docs/benchmark/results.md). Recoverable via git mv from the archive plus re-adding the entries to agents/master-registry.json.
  • Shipped skill set reduced 4 → 2: tdd and diagnose archived to legacy/trim-2026-07-04/skills/. Basis: 7 weeks of session telemetry showed zero dispatches for either skill. supervise and wrap are retained. The tdd-guard hook is unrelated to the tdd skill and continues to run unchanged. Recoverable via git mv from the archive.
  • codex-skills/ retired to legacy/trim-2026-07-04/codex-skills/. Basis: zero usage recorded in 7 weeks of session telemetry. The Codex CLI adapter (adapters/codex/) is unrelated and remains active; setup.sh no longer offers the ~/.codex/skills symlink install step. See legacy/trim-2026-07-04/ARCHIVE-NOTE.md for the full recovery procedure.

0.2.0 — 2026-06-15

Added

  • Claude Code plugin packaging — .claude-plugin/plugin.json + marketplace.json make the harness installable via /plugin marketplace add joymin5655/Agent → /plugin install agent-harness@agent. One install, every project.
  • hooks/hooks.json — plugin hook wiring (SessionStart / Stop / UserPromptSubmit / PreToolUse / PostToolUse) dispatching through the Claude Code adapter to core/hooks/ via ${CLAUDE_PLUGIN_ROOT}.
  • commands/project-init.md — /project-init slash command to scaffold project-level files.
  • LICENSE — MIT (was TBD).

Changed

  • README now leads with the Claude Code plugin install path; shell setup.sh remains for Codex/Gemini or non-plugin use.

0.1.0 — 2026-05-18

Added

  • Initial AI-agnostic agent framework structure
  • 3-AI adapter layer: Claude Code, Codex CLI, Gemini CLI
  • Canonical hook protocol (docs/hook-protocol.md): stdin JSON event + stdout decision JSON
  • Core hooks (core/hooks/): ~25 portable hooks for security, session coordination, plan-mode, TDD enforcement, drift detection
  • Core infra (core/infra/): multi-session worktree coordination (agent-session.sh), commit/PR automation (auto-ship.sh), session store, supervisor goal mode
  • Core git-hooks (core/git-hooks/): pre-commit (gitleaks + hardcoding scan) + pre-push (gitleaks + secret diff scan)
  • Generic policy rules (rules/): 7 critical + 12 lazy-loaded archive
  • Generic agents (agents/): code-reviewer / architect / build-error-resolver / security-reviewer / performance-optimizer / test-engineer / docs-writer / refactor-cleaner / tdd-guide / copy-humanizer
  • Generic skills (skills/): wrap, supervise, tdd, diagnose, grill-me, grill-with-docs, improve-codebase-architecture, caveman, api-and-interface-design, incremental-implementation, source-driven-development, deprecation-and-migration, design-variant-mockup, hook-reproduce-test, triage-external-draft, weekly-digest
  • Codex-native skills (codex-skills/): code-explorer, code-reviewer, database-reviewer, planner
  • Templates (templates/): generic CLAUDE.md / AGENTS.md / GEMINI.md / RTK.md / karpathy.md / hook-config.yml / project-rules.md / gitleaks.toml
  • setup.sh 4-mode installer: --claude / --codex / --gemini / --project / --hooks-only
  • GitHub Actions workflow templates (github/workflows.template/): secret scan + lint

Changed

  • N/A (first release)

Archived

  • legacy/v0-mirror-2026-05-12/ — original mirror skeleton + domain-specific assets from the prior project version. See legacy/v0-mirror-2026-05-12/ARCHIVE-NOTE.md for migration guide.

Security

  • Base gitleaks.toml with 100+ built-in patterns + extensible per-project allowlist
  • Generic content-scan hook with 7 default patterns covering Python/Node secret-file readers, hardcoded credentials, OpenAI-style sk-... tokens, JWT literals, Bash secret-readers, and exfiltration via find -exec (see core/hooks/secret-content-scan.py for full pattern list)
  • Project-configurable risk-area abort codes via templates/hook-config.yml.template