joymin5655/agent-harness
Portable AI agent harness — curated review/build/test agents, secret-hardening + worktree + plan-gate hooks, and spec/supervise/wrap/verify-completion skills. Install once, use in every project. AI-agnostic core (Claude Code / Codex / Gemini).
Changelog
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
Unreleased
[0.5.13] - 2026-10-02
Added
- Antigravity native hook adapter (W5-3):
adapters/antigravity/adapter.shandadapter.pytranslate agy 1.2.12'sPreToolUse/PostToolUse/Stophook JSON (event name from argv, since agy's stdin carries none) into canonical events and run the same core-hook chains as the Codex template (run_commandasBash,write_to_fileasWrite,replace_file_contentand eachmulti_replace_file_contentchunk asEdit, capped at 100). agy treats a PreToolUse{}as a deny and a hookallowwas not observed to grant more thanask(unmeasured, so never emitted); a pass-through is{"decision":"ask"}, a core-hookaskbecomesforce_ask(plainaskdefers to the user's allow rules), a core deny is a deny, and any hook failure, timeout, bad JSON, exhausted 25s budget or unverified argument shape is a fail-closed deny.PostToolUsealways answers{}.Stopturns a coreblockinto one{"decision":"continue"}guarded by a marker under${AGENT_STATE_DIR:-$HOME/.agent/state}/antigravity-stop/<conversationId>(6h TTL). WithAGENT_ANTIGRAVITY_WORKER=1every matched tool call is denied without running a hook.GEMINI_API_KEY/GOOGLE_API_KEYare scrubbed from the hooks' environment. Review fixes:send_command_input(text typed into a live shell) is now matched and guarded asBash; the Stop chain keeps a reserved time slice forbrain-capture.pyandsession-close.shand reports a gate that overran; a malformedAGENT_ANTIGRAVITY_BUDGET_Sfalls back to 25 instead of crashing the hook;session-quality-gate.pyreads agy'stool_callstranscript shape.core/tests/antigravity-adapter-test.sh(113 checks, includes a drift check against the Codex template) and a newadapter-parity.shantigravity section cover it. - Antigravity plugin install (W5-4):
adapters/antigravity/install-plugin.py,plugin.jsonandhooks.json.template.setup.sh --antigravityinstalls a plugin folder (${AGENT_ANTIGRAVITY_PLUGIN_DIR:-~/.gemini/config/plugins/agent-harness}) whosehooks.jsonpoints at the absolute adapter path. It writes atomically and idempotently, refuses a foreign plugin folder or a framework root containing shell metacharacters, offers--uninstall,--checkand--dry-run, and never reads or writes~/.gemini/config/hooks.jsonor agy'ssettings.json. Setup also prints (never applies) API-key andpermissions.denyguidance.setup.sh --doctorgained an "antigravity native hooks" check (core/tests/antigravity-native-hooks-test.sh, 86 checks). A marker file keeps a folder ours after an uninstall that left a user file behind, so reinstall andsetup.sh --antigravityno longer refuse it, and the setup step warns instead of aborting. Thepermissions.denyguidance uses agy's documented prefix form, not*globs. - Antigravity worker API-key opt-in (W5-2):
ANTIGRAVITY_AUTH=apikeymakesantigravity-worker.shread a Gemini API key from the Keychain (servicegemini-api-key) and exportGEMINI_API_KEYinto agy's environment only, never argv or logs. A missing item orsecuritybinary exits 2; a missing"modelProvider": "gemini"in agy's settings is a stderr warning, and the worker never edits that file. The keyring stays the default. The worker also exportsAGENT_ANTIGRAVITY_WORKER=1for every dispatch. Review fixes: the worker writes a static workspace deny plugin (.agents/plugins/agent-worker-deny/) before agy starts, exports the API key only after that succeeds and exits 2 otherwise, and its sandbox profile denies writes to~/.gemini/config/hooks.json,config/plugins/andsettings.json(core/tests/antigravity-worker-test.sh, 95 checks, includes a realsandbox-execenforcement check).
Changed
- Antigravity docs and registry brought current with agy 1.2.12 (W5-5):
adapters/antigravity/README.mdnow records the 2026-09-29 probe (soft-deny indenied_actionsplus ajetski ... auto-deniedstderr notice, exit 3 on status ERROR), the Gemini CLI individual end date 2026-06-18, the API-key opt-in, the native-hook plugin, and the worker threat-model drift (headless 1.2.12 ranechowith no allow rule).docs/runtime-registry.jsonantigravity:cli_version_measured1.2.12,measured_on2026-09-29,hook_events_wiredPreToolUse/PostToolUse/Stop.docs/hook-protocol.mdgained section 13 and theantigravityai value;docs/cross-runtime-harness-design.mdstates what the plugin enforces and what it does not.
Fixed
- Antigravity lane counted a soft-denied run as success (W5-1). Headless agy
soft-denies a tool call it cannot get approval for: the run continues and exits 0
with a stderr notice.
antigravity-worker.shnow exits 9 on that notice and 10 when the--output-format jsonenvelope is unparseable or itsstatusis notSUCCESS. It captures agy's output outside the sandbox's writable dir, so a prompt-driven write cannot forge the envelope.antigravity-preflight.shreports a soft-deny as exit 8 (lane absent), reads the probe token from.responseonly, and treatsauthentication requiredas an auth failure. Newcore/tests/antigravity-preflight-test.sh. - Antigravity soft-deny detection missed agy 1.2.12 (W5-1b). 1.2.12 reports a
soft-deny as a non-empty
denied_actionsarray in the json envelope plus a stderr notice with new wording, which the 1.1.14 pattern did not match, soantigravity-worker.shreturned exit 0 for a run that was denied. It now exits 9 on a non-emptydenied_actions(single-value stdout only) or a stderr match onpermission check failed|denied permission to|auto-denied|cannot prompt for; an emptydenied_actionsis not a soft-deny, and agy's own nonzero exit codes (including the undocumented 3) pass through unchanged.
[0.5.12] - 2026-10-01
Added
- Supply-chain scan classes 5–7 (adapted from ECC v2.2
pi/core):core/tests/supply-chain-scan.shdelegates tocore/tests/supply-chain-remote.py. Class 5 flags fetch-and-execute (curl … | sh,bash <(curl …),eval "$(curl …)"); class 6 flags unpinned remote runners (npx/npm exec with--yes/--package, bunx, pnpm/yarn dlx, uvx, pipx run); class 7 flags URL hosts in auto-fired hooks andhooks/*.json/.mcp.jsonthat are not incore/tests/supply-chain-allowlist.txt. Threat model documented inrules/policy/security-guards.md(#138). - Impact context (idea from Graft's blast radius):
core/infra/impact-context.pylists dependents outside the diff and the test files the change reaches, from the existing CodeGraph index. Fail-open (10 s budget, 60-line cap, exit 0)./council-reviewadds it to the shared review core;/wrapshows it as an advisory pre-flight step (#138).
[0.5.11] - 2026-10-01
Added
- Codex native hook path (W4-1):
adapters/codex/adapter.py's native mode (run_native) translates Codex's nativePreToolUsestdin (Claude-shaped:hook_event_name,tool_name,tool_input,tool_use_id,session_id,cwd,transcript_path) into canonical events, splittingapply_patch's patch text (carried intool_input.command, same field as Bash) into oneWrite/Editevent per file with absolute paths, and aggregating deny > ask > advisory > allow across files. A canonicalaskand any hook failure (non-zero exit other than2, a timeout, invalid JSON, or a zero-op patch) become a fail-closeddenyJSON; non-PreToolUseevents stay fail-open by design.adapters/codex/hooks.json.template(merged into~/.codex/hooks.jsonby the newadapters/codex/merge-hooks.py, which replaces only Agent-owned entries and leaves other tools' hooks alone) wires SessionStart/UserPromptSubmit/PreToolUse/PostToolUse/Stop/SessionEnd. Hardened after a security review: one 25s check budget per hook across all files of a patch, a 100-file cap, indented patch markers denied as ambiguous,normpathon patch paths,Move todestinations checked with the source file's content (regular files only), a deny for a missing guard, missingpython3, or an unknown decision verb, and every per-file event run onPostToolUse(council review).core/tests/codex-native-hooks-test.sh(41 checks) and a newadapter-parity.shnative section cover it; one livecodex execsession (codex-cli 0.157.0, 2026-09-27) confirmed the Bash andapply_patchdenials end-to-end. - Portable plugin manifest + marketplace (W4-2): root
plugin.json(agent-plugins.org1.0.0,extensions.com.openai.hooks→hooks/codex-hooks.json) and.agents/plugins/marketplace.json, socodex plugin marketplace add joymin5655/Agent+codex plugin add agent-harness@agentinstalls the same hooks and skills — install verified with the codex CLI (skills appear incodex debug prompt-input) on 2026-09-27.core/tests/version-parity.shnow covers the root manifest.
Changed
council-escalation-gate.pyescape 2 now requires a stated reason. A retry on an already-denied council-scale diff passes only when thecode-reviewerdispatch prompt carriescouncil-unavailable: <reason>(≥10 chars after whitespace collapse, not the pasted<placeholder>); the reason is written tosecurity-violations.jsonl. If the diff hash is unavailable, a stated reason alone opens the escape. The deny text used to advertise "re-issue this exact dispatch", and a model was observed retrying reflexively without trying/council-review— a bare identical retry is now denied again, and repeat denials no longer refresh the ledger entry. Tests: 9 new cases incouncil-escalation-gate-test.sh(32 pass).- model-routing-observer records both session ids.
session_idnow prefers the hook event's runtime session UUID (falls back toAGENT_SESSION_ID), and the env id is kept asagent_session_id, so concurrent sessions in one cwd stay distinguishable.manager-audit.sh --session <id>matches either field, keeping existing filters working. setup.sh --codexnow mergeshooks.json.templateinto~/.codex/hooks.jsonbesideconfig.tomlviamerge-hooks.py, symlinks eachskills/<name>into~/.agents/skills, repoints an old~/bin/codex-bashsymlink tolegacy/codex-shell-wrap/, and prints a reminder that Codex only enforces a hook after you trust it with/hooks.setup.sh --doctorgained a "codex native hooks" check: WARN when[features] hooks = falseor nothing is installed, FAIL on invalid JSON or a moved adapter path, WARN when entries are installed but nothing is trusted yet, PASS otherwise (core/tests/setup-doctor-test.sh).docs/runtime-registry.json,docs/ai-adapters.md,docs/cross-runtime-harness-design.md, anddocs/hook-protocol.mdbrought current with the native path: Codexhook_events_wired, Tier A coverage for Bash +apply_patch+ MCP tools, the ask→deny translation now living in the adapter itself, and the fail-open/fail-closed split by adapter in the exit-code table.- Claude Code hook manifests brought current with 2.1.282 (W3-1/W3-2/W3-3/W3-4).
hooks/hooks.jsonandadapters/claude-code/settings.json.template: matchersWrite|Edit|MultiEdit→Write|Edit,Task|Agent→Agent,ExitPlanMode|Task|Agent→ExitPlanMode|Agent(MultiEditis no longer a documented tool;Taskis a legacy alias ofAgent). Wired 6 Claude-only extended events non-canonical to the cross-AI protocol:PostToolUseFailure→circuit-breaker.py,SessionEnd→session-close.sh(timeout: 2),PreModelSwitch/PostModelSwitch→session-tier-observer.py,SubagentStart/SubagentStop→model-routing-observer.py.PermissionRequestdeliberately left unwired (differentdecisionschema; exit 2 not honored), and so areWorktreeCreate/WorktreeRemove(aWorktreeCreatehook replaces git's worktree creation and must print the new path — the observerr4-mutex-check.shwould have broken every Claude worktree; wired during W3, unwired before release) — seedocs/hook-protocol.md§12.secret-content-scan.py's MCP matcher collapsed from an explicit per-tool pipe-list to per-vendormcp__<server>__.*wildcards;rubric-commit-judge.shgained a narrowing"if": "Bash(git commit*)".adapter.shheader now documents the extended events and the no-op-vs-fail-open distinction.docs/hook-protocol.mdgained §12;docs/runtime-registry.jsonclaude-codehook_events_wiredreflects the full wired list. - CI: added
plugin-validatejob runningclaude plugin validate --strict .(best-effort CLI install; explicit::notice::skip if the CLI never lands — never a silent pass, and not a required check).
Deprecated
codex-shell-wrap.shmoved tolegacy/codex-shell-wrap/(W4-3). It remains a fallback only for[features] hooks = falsebuilds, or an adminrequirements.tomlthat allows managed hooks only;adapters/codex/tests/run.sh(T5/T6) still exercises it so the fallback path doesn't rot.
Fixed
session-close.shmissedSessionEndinjson.dumps-style input. The event-namesedrequired:"with no space, so"hook_event_name": "SessionEnd"fell through to the full Stop path (TODO scan, notification, broadcast) inside the 1.5s SessionEnd budget. The extraction now allows whitespace after the colon.circuit-breaker.pyansweredPostToolUseFailurewithhookEventName: "PostToolUse". The advisory now echoes the event it was invoked for.runtime-currency.shread a futuremeasured_on/checked_onas fresh. A date after today is now a FAIL (typo), not a negative age.codex-template-currency-test.shdenylist lagged the registry (gpt-5\.[234]vsdocs/runtime-registry.json'sgpt-5\.[2-6]); synced, registry named as SSOT.- Deferred council findings (runtime-currency-2026-09).
- The
rubric-commit-judgemanifest entry no longer carries"if": "Bash(git commit*)", which missedgit -C <dir> commitandrtk git commit. The hook's internal regex now does all of the filtering. circuit-breaker.pyserializes its shared state file withflockand writes it atomically (tmp + rename). Before this, concurrent sessions lost failure records.setup.sh --bootstrapwith no usable OS package manager (or no Homebrew) now still reaches the PyYAML pip step instead of returning early. It still exits 1.session-tier-observer.pystampsoriginon its session-start record, which the W1 log-origin tag had missed.
- The
[0.5.10] - 2026-09-02
Fixed
session-quality-gate.pylayer 3 fired without end and blamed the wrong session. The unverified-session advisory read the whole dirty work tree while claiming the files "changed this session", and recomputed and reprinted on every Stop:stop_hook_activegates blocking only, and the note is re-injected as model-visible context, so it fed the turn that produced the next Stop. A file left uncommitted by a concurrent session was reported to five unrelated sessions as their own change — 80 firings in 7m22s in one session (median gap 2.4s) and 79 in another, ending only when a verification command happened to land in the sink. The documented exit ("or state explicitly why none applies") was inert: nothing consumed a stated reason, so a session that correctly declined to test someone else's work had none. The diff is now intersected with the files the session actually edited (read offtranscript_path; sidechain kept, since a subagent's edit is still the session's work) and the note is suppressed once recorded for the same (session, changed-file set), with blocking keyed separately so enablingAGENT_VERIFY_OBSERVER_BLOCK=1midway is not swallowed by an earlier advisory. Without a transcript it falls back to the work-tree list and says "attribution unavailable" rather than claiming one. Two ways the intersection could go silently empty are fixed with it:git diffandgit ls-filesdisagree about their base outside the work-tree root (git now runs at the toplevel,diff.relativepinned off), andcore.quotePathC-quotes non-ASCII paths (-zthroughout). The FIFO/S_ISREG/tail-window hardening moved into one shared traversal instead of being copied per predicate. (#122)
Added
top-edit-advisor.pyhook (PostToolUse Write/Edit/MultiEdit). Accumulation-time counterpart tomodel-routing-advisor.py/model-routing-observer.py, which only see Task/Agent dispatches and so cannot see the leak where the TOP model implements directly and never dispatches at all — a 2026-09-02 audit measured 511 such direct edits across 26 sessions, with a single warning proven insufficient. Repeats asystemMessageadvisory every +15 measured main-loop edits (envAGENT_TOP_EDIT_THRESHOLD) instead of warning once and going silent; never caches a not-TOP verdict since the session model can switch mid-session. Advisory only — never blocks. Tests:core/tests/top-edit-advisor-test.sh.- Codex
quick/deeptier profile installation insetup.sh --codex.install_codex()now installs the existingadapters/codex/{quick,deep}.config.toml.templatebeside~/.codex/config.toml, so the tier ladder documented indocs/model-routing.md(codex --profile quick|deep) is wired up by setup instead of requiring a manual copy; doctor check 13's remediation now also points atsetup.sh --codex. - Review-tier ladder (
core/infra/review-tier.sh) +/wrapstep 1d integration. Review dispatches are the second-largest routing cost after implementation, so cadence — not just model tier — is now a lever: every diff gets tier 0 (docs-only or ≤AGENT_REVIEW_SKIP_LINESnon-risk code lines — skip, self-check only), tier 1 (the common case — onecode-reviewerpass at wrap/commit time), or tier 2 (council-scale —/council-review), delegating the tier-2 judgment to the existingcouncil-threshold.shSSOT instead of re-mirroring its risk-area patterns.agents/code-reviewer.mdanddocs/model-routing.mdupdated to the hybrid timing this implies (wrap-time default, immediate only on risk-area paths). Tests:core/tests/review-tier-test.sh. council-escalation-gate.pysame-diff-hash escape visibility upgrade. The loop-safety escape (a council-scale diff already denied once is let through on retry) used to be stderr-only; it now also emits a PreToolUseadditionalContextadvisory (model-routing-advisor.py's emission pattern) so the model/user see that the dispatch skipped review, not just a log line nobody reads. Allow behavior, TTL, and hash binding unchanged.
[0.5.9] - 2026-08-25
Added
- OpenRouter free advisory lane + purpose launcher set + codex global
rules (spec:
.agent/plans/free-lanes-and-launchers/). Newadapters/openrouter/worker-lane bridge to OpenRouter:freeroutes (non-votingadvisor-freerole, sensitive-cwd guard + per-dispatch retention warning, fail-open on 429, free exact-token preflight; pinnvidia/nemotron-3-super-120b-a12b:freein the adapter-owned tiers file, chosen by live probe over the congestedz-ai/glm-5.2:free)./council-reviewgains--with-free(advisory lane, mirrors--with-grok). Newadapters/claude-code/launchers/purpose launcher set (claude-build/claude-quick/claude-researchtier launchers +claude-ox.templateOpenRouter gateway launcher absorbed as repo SSOT, personal blocklist externalized to the shared~/.config/agent-harness/sensitive-pathsfile) — session-start human allocation, the allowed side of the no-runtime-switching policy. Newadapters/codex/AGENTS.global.md.templatedeployed to~/.codex/AGENTS.mdso manual codex sessions carry the portable harness rules (evidence contract, tier discipline, review-before-done).docs/model-routing.mdgraduates the "tier/cost-aware task allocation" follow-up to a designed free-lane allocation section (NIM evaluated and deferred on ToS; Groq documented as strongest future candidate); newdocs/launchers.md. Tests:core/tests/openrouter-worker-test.sh,adapters/claude-code/tests/launcher-test.sh(stubbed, zero paid calls).
Changed
- kiro gateway roster drift absorbed (2.19.1). The kiro-cli 2.19.1 roster
carries no OpenAI models and no
claude-opus-5(live-probed):kiro-openaibackend disabled with a dateddisabled_reason(templates kept for revival);second-opinion-review/second-opinion-verify/advisorfallbackkiro-openai→nullwith rationale (a fallback must not silently change the vote's vendor — same rule asthird-opinion-review);kiro-anthropic-topre-pinned toclaude-sonnet-4.5(1.3x).adapters/kiro/README.mdtier table updated with the roster-recheck rule.
Security
-
Free-lane egress hardening (council + security-reviewer, 2026-08-25). openrouter worker/preflight + ox launcher: sensitive-cwd guard now strips trailing slashes, canonicalizes blocklist entries through
pwd -P(symlink parity with the cwd side), fails CLOSED when no guard file resolves (an explicitly setOPENROUTER_SENSITIVE_PATHS_FILEis honored strictly — no silent template fallback), and announces its guard source / FORCE override on stderr. Prompt egress floor: hard byte cap (OPENROUTER_PROMPT_MAX_BYTES, default 256KiB) + credential-shape refusal (private key / AKIA / sk- / ghp_ / xox-) with loudAGENT_OPENROUTER_UNSAFE_PROMPT=1override. Key-residue windows closed: TERM-first watchdog + TERM/INT/HUP→EXIT trap chain so the key-bearing curl config is scrubbed on timeout, interrupt, and hangup. ox launcher header now documents the isolation tradeoff (deny rules/hooks absent in gateway sessions) and env-token readability. Battery grew to 27 cases incl. a guard-bypass matrix (trailing slash, symlinked entry, fail-closed, /dev/null, egress floor, byte cap). -
Conditional council auto-escalation — a council-scale diff can no longer be signed off by a solo Claude reviewer. New PreToolUse
Task|Agentgatecore/hooks/council-escalation-gate.pydenies a plaincode-reviewerdispatch whencore/infra/council-threshold.shjudges the staged diff council-scale (line/file threshold, or a path in a declared risk area) and points the caller at/council-review --stagedinstead. Every other case is silent: non-dispatch tools, other subagents, small diffs, and any internal error all fail open — a broken gate must not block review, but it says so on stderr and in the audit log rather than failing open silently. Escape hatches, both deliberately narrow:/council-reviewmarks itself active through the gate's own--council-flag set|clearCLI (steps 0.5 and 6) so the council's own internalcode-reviewerdispatch isn't denied by the gate that routed the caller there. The flag lives OUTSIDE the reviewed workspace (keyed by project root under~/.agent/state/council, overrideAGENT_COUNCIL_STATE_DIR), is TTL-bound (AGENT_COUNCIL_ACTIVE_TTL_S, default 300s) and content-bound to the current diff's hash — so a stale flag, a flag for a different diff, and an in-repo file a reviewed diff could plant all fail to open the gate. Grant-side state outside the workspace follows thetrust_tier.pyprecedent.- An identical dispatch re-issued after a denial is let through once, with a
warning, so an agent that genuinely cannot run the council isn't trapped in
a loop. The ledger entries expire (
AGENT_COUNCIL_DENY_TTL_S)./wrapgains an advisory step 1d covering the path the hook can't see — edits made without any Task/Agent dispatch — recommending/council-reviewbefore a solo commit rather than aborting. Tests:core/tests/council-escalation-gate-test.sh(24 checks),core/tests/council-threshold-test.sh(19 checks).
-
/worker-setupskill — guided onboarding for the cross-vendor worker lanes (codex, antigravity, grok, kiro): a read-only status sweep ofcore/infra/backends.json(livejqquery, never a hardcoded lane list), a cost-model/tier-allocation briefing before anything installs, then per-lane install → auth → verify with every real round-trip probe announced first (kiro's is billable). Never a dispatcher itself (core/infra/call-worker.shis) and never a paid probe without explicit user approval. Test:core/tests/worker-setup-skill-test.sh(14 checks: frontmatter, plugin-cache-safe dispatch, vendor coverage, no model IDs, no credential collection). -
council-review dispatch fixed for plugin-cache installs —
skills/council-review/SKILL.mdstep 3 resolvedcore/infra/call-worker.shvia${CLAUDE_PLUGIN_ROOT:-$PWD}instead of a bare cwd-relative path, which silently no-op'd every external lane on a plugin install (docs/claude-plugin-install-lifecycle.md§7 item 1 — now resolved for this one consumer; other consumers named in that item may still carry the gap). A missing dispatcher now names itself and reports every external lane absent rather than guessing a path. Test:core/tests/council-dispatch-path-test.sh(3 checks). -
setup.sh—~/bincreated unconditionally + doctor PATH row — everyinstall_codex/install_gemini/install_grok/install_antigravity/install_kironow calls a sharedensure_home_bin()instead of silently skipping the symlink when~/bindidn't already exist; a NOTE prints once if~/binis not onPATH. New doctor row "worker symlink dir" (PASS when~/binexists and is on PATH; WARN naming exactly which half is missing, with the export one-liner) runs just before the existing generic worker-lane PATH sweep. Vendor-CLI-not-initialized NOTEs added to the gemini/grok/antigravity tier-seeding blocks (config dir missing -> told to run the CLI once, then re-run the install flag). -
install_kiro()+--kiroflag (opt-in, not part of--all/default — kiro is metered/paid, same stance as--grok/--antigravity): seeds every shippedadapters/kiro/*.json.templateinto~/.kiro/agents/(existing user-owned profiles never overwritten), symlinkskiro-preflight, and notes thekiro-cliinstall command +KIRO_API_KEYauth when the CLI isn't on PATH. Always exits 0. Test:setup-doctor-test.sh(r)/(r2) sections — throwaway-HOME smoke (profiles seeded, pre-existing profile preserved, preflight symlink resolves) + the new doctor row's PASS/WARN branches (11 checks). -
antigravity-preflight.shhint bug fixed — the missing-on-PATH message wrongly told users to symlink into~/bin/grok-worker(copy-paste from the grok adapter); now names~/bin/antigravity-workerand mentionsantigravity-preflight. Regression check added tocore/tests/antigravity-worker-test.sh(17th check: no straygrok-workerstring anywhere underadapters/antigravity/). -
codex preflight is now auth-aware —
core/infra/backends.json's codex backend switchedpreflightfrom["codex", "--version"](proves the binary exists, nothing about auth) to["codex", "login", "status"]+preflight_timeout_s: 15. Measured 2026-08-20 on codex 0.147.0: logged out exits 1 and prints "Not logged in"; logged in exits 0 — a local check, no network call, not billable. -
Lane cost-model docs —
docs/model-routing.md§ Cross-vendor lanes gained a "Lane cost models (2026-08-20)" entry (grok/antigravity/codex quota-based, kiro-* metered/paid including its own preflight) as the SSOT/worker-setupcites, plus a note that tier/cost-aware cross-vendor task allocation is a candidate/spec, not designed yet. -
Council lens split — codex and gemini voting lanes get distinct review perspectives (decided 2026-08-20: perspective diversity over duplicate generalists). Each voting external lane's prompt is now a lens preamble + the shared core: codex (
second-opinion-review) reviews implementation correctness (logic, edge cases, error handling, concurrency); gemini/agy (third-opinion-review) reviews architecture & consistency (design/simplification, contract coherence, doc–code drift). A lens states emphasis, not permission — either lane still reports any defect it sees, and cross-lens agreement keeps (strengthens) the ≥2-vendor high-signal rule. Grok's advisory seat stays deliberately unscoped. Framing lives inskills/council-review/SKILL.mdstep 1; role comments incore/infra/backends.jsonrecord the split. -
Antigravity (
agy) review lane restores the seated google reviewer — newadapters/antigravity/(worker + preflight + tiers template + measured posture README), following the grok adapter pattern. The council'sthird-opinion-reviewlane, dead since the gemini CLI's individual OAuth was retired (2026-07), isenabled: trueagain via agy — Google's official successor — authenticating from the OS keyring (no API key; the GEMINI_API_KEY path is upstream-contradicted). Measured 2026-08-19 (agy 1.1.14): default headless mode fails closed on shell exec and created no file in any probe; worker forbids--dangerously-skip-permissionsand runs under a sandbox-exec deny-write/deny-cred-read profile; argv is flags-before--p(flags after are misparsed); tiers differ by model (effort is baked into the ID). Live-verified: preflight exit 0 + one E2E council dispatch returned a real finding withstatus: complete. Tests:core/tests/antigravity-worker-test.sh(16),setup.sh --antigravityinstall path. -
Rate-limit fail-open contract — a vendor quota/rate limit is now a distinct lane condition, not a generic failure:
grok-worker.shclassifies the xAI free-tier "usage limit" ending (measured 2026-08-18: message + generic exit 1) as EX_TEMPFAIL (75),call-worker.shmaps 75 tostatus: rate-limitedin the capture (escaping exit code stays 1), and/council-reviewreports the lane "rate-limited, retry later" and proceeds without it — decided 2026-08-19: the grok lane stays on the free tier and is skipped when exhausted; no upgrade prompt mid-review. Watchdog kills (124/137/143) are excluded from reclassification. Tests: grok battery case (g) — limit→75, plain failure→1, output still streams. -
Grok (xAI) advisor lane —
adapters/grok/worker-lane bridge (grok-worker.shstdin→prompt-file + sandbox-exec deny-write + neutral cwd;grok-preflight.shexact-token probe; tiers template owns the model pin) and registry entry carrying theadvisor-thirdrole only — never a gate vote. Measured (0.2.118, 2026-08-19): the CLI's own flags (--tools '',--permission-mode plan,--deny,--disallowed-tools) do NOT block writes; only the OS sandbox does. The sandbox denies reads of credential stores (~/.ssh/~/.aws/…) and narrows writes to the run's WORK_DIR + the CLI session store minus its own tiers file (argv-persistence hole), forwards signals instead ofexecso the prompt file never outlives the run, and validates$HOMEbefore building the scheme profile (adapters/grok/README.md). -
Gemini worker-lane contract restored —
gemini-worker.shtier bridge +gemini-preflight.shprobe +third-opinion-reviewrole (fallback: nullfor vendor independence). Backend staysenabled: false: re-verified 2026-08-19, a cached oauth credential still gets 401 UNAUTHENTICATED; the registry's re-enable condition is a passing preflight, not a credential file. -
/council-reviewskill — multi-vendor review council: Claudecode-revieweragent + codex + gemini in parallel, grok opt-in (--with-grok, advisory only). Synthesis order: mechanical capture status → dedup → citation verification against the actual files (hallucinated findings dropped and counted per lane) → severity re-rating → source tagging (≥2-vendor findings marked high-signal) → disagreement adjudication. False-council guard: an all-external-lanes-absent run must say so. -
setup.sh --grok+ worker-lane symlink installs for gemini/grok, and a generic doctor check: every enabled backend'scmd[0]/preflight[0]must resolve on PATH. -
Tests:
core/tests/gemini-preflight-test.sh(10 checks),core/tests/grok-worker-test.sh(14 checks: argv/prompt bridge + OS-level sandbox write-block, credential-read-deny, tiers-write-deny, unsafe-HOME regressions) — PATH stubs, zero paid calls. -
Autonomous-loop runner
core/infra/loop-run.sh+skills/loop+skills/harness-loop(P2-1 + P2-4 + O-2). The §5 loop's mechanical enforcement in one script —initcreates the SQLite goal row (attempts modeled assupervisor-goal.shwaves) plus a JSON state file and the active-loop flag;attemptruns the grader (defaultgrade.sh --base/--target, stubbable viaLOOP_RUN_GRADE_CMD) under a background watchdog that TERMs then KILLs at a timeout (600s default), classifies the outcome (keep / discard / crash / timeout), appends exactly one row to.agent/loop/results.tsvvialoop-ledger.sh, and prints aLOOP: continue/LOOP: stop(cap|circuit-breaker)verdict — two consecutiveharness_score: 0attempts trip the circuit-breaker, reaching the cap stops the loop, both remove the active flag. Git reset/keep of the candidate commit stays a skill-level step, not this script's job.skills/loopis the generic fresh-context/one-task/resumable protocol (O-2);skills/harness-loopspecializes it to this harness's own reviewer prompts with §5.2's 9-step procedure verbatim, numbered, human-only-edited.core/tests/loop-run-test.sh(37 checks, hermetic temp-git-repo fixture) andcore/tests/loop-skill-test.sh(13 checks, binds the SKILL prose to the mechanics with RED reorder/removal/token-drop fixtures) drive both.loop-run-test.shneedssqlite3+jq(supervisor-goal.sh's own hard dependency) and exits 2 when they are absent, so verify-all's discovered-check SKIP lane tallies it as skipped with its reason rather than as a pass. -
Autonomous-loop grader
core/tests/grade.sh(P2-2 + L-1 impl). The loop (§5) now has its grader: it runs the GATE floor (sanitize-audit / adapter-parity / hook-config-test / post-commit-autosync / gitleaks) and then evaluates each named failure mode inevals/failure-modes.yamlby RE-RUNNING the battery that encodes it — a mode PASSes iff its guard is still green (the hole stays closed), FAILs iff the candidate re-opened it. Output is the L-1 checklist (mode:<id> PASS|FAIL — reason) plus a rollupharness_score: X.Y= (#PASS) − 0.5×(#untrustworthy), the mandatory last line on every path so an empty grep is always a real crash. It fails closed toharness_score: 0(= discard) on any untrustworthy path — GATE failure, a TARGET-boundary violation (off-target diff), or an unparseable rubric — the operator-chosen baseline verdict. On a clean tree it emitsharness_score: 10.0(10 code-guarded modes PASS;vacuous-parityis GATE-covered — its only guard, adapter-parity, is a GATE battery, and a second checklist path would be unreachable behind the GATE short-circuit, so it is reported N/A on the single GATE path — andreview-false-cleanis N/A), superseding the retired single-scalar 8.0. It is a loop-time tool (re-runs batteries; ~1 min) and is excluded fromverify-allto avoid recursion;grade-test.sh(36 checks) drives it hermetically and gates mode↔guard-map drift, including a ban on GATE batteries in the guard map. The TARGET clean-tree check usesgit status --porcelain, so an UNTRACKED tamper file (which the batteries would execute but a committed-range diff cannot see) also refuses the grade (fail-closed). -
Append-only loop ledger
core/infra/loop-ledger.sh(P2-3). The sanctioned writer for.agent/loop/results.tsv(untracked run state): a 5-column schema (commit / harness_score / duration_s / status / description≤80), a status enum (keep|discard|crash|timeout), numeric validation that rejects rather than coerces a malformed score/duration, and free-text sanitization. It only ever appends — the header is written once. Every successful append notarizes the ledger in a sidecar witness (<ledger>.witness= sha256 + line count) that notarizes a PREFIX: the next append REFUSES a ledger that is missing (delete-recreate), shorter than the witness (truncation), or whose first witnessed lines no longer hash to it (rewritten history), while rows beyond the prefix are append-only extension and are accepted — which self-heals the crash window where an append landed but the process died before the witness update (P2-4 kills runs on timeout, so that window is real). Tamper-evident, honestly bounded (a shell can still delete both files; that two-target act is what loop-write-guard escalates during a loop).loop-ledger-test.sh(30 checks). -
Loop write-ban
core/hooks/loop-write-guard.py(L-2). While an improvement loop is active (envAGENT_LOOP_ACTIVE=1or a flag file), a Write/Edit to the grader/verifier surface (core/tests/,evals/) or a non-append rewrite of the results ledger escalates toask(not deny — the calibration policy reserves deny for secrets) so a human stays on the loop and a run cannot silently rewrite the code that scores it. Outside a loop the guard is fully inert (zero added friction). Containment uses realpath, so a symlink into the guarded dir cannot dodge it. Wired into theWrite|Edit|MultiEditPreToolUse chain;loop-write-guard-test.sh(31 checks) covers the ask/allow matrix, the symlink-escape case, Bash write-path detection, the delete-recreate/witness escalations, and WHY/FIX tags. -
B4 follow-ups: gitleaks allowlist regression + 5-surface drift tripwire + a real SKIP lane for discovered checks.
core/infra/gitleaks-fire-test.shproved the DETECTION side of the secret gate; nothing proved the ALLOW side.core/tests/gitleaks-allowlist-test.shplants eachgitleaks.tomlplaceholder shape (your_*_key,dummy_*,example_*,sk-proj-placeholder…,{{VAR}},<your_x>,$USER_*_JWT) in realistic config/doc contexts and asserts a clean scan. Adversarial review reshaped it three ways: the control is now per shape rather than one aggregate scan (an aggregate passes on a single finding anywhere, which proved only 1 arm of 8); findings are read from gitleaks' JSON report by rule id, never from its exit status (non-zero means "leaks found" OR "gitleaks errored", and conflating them lets a broken control report ok); and the placeholder value is assembled at runtime so no secret-shaped literal is committed — a committed one would make the repo's own secret-scan depend on the very allowlist arm under test, blocking future tightening. It reports an honest ceiling: measured 2026-08-10, only 2 of 10 shapes are load-bearing (sk-proj-…andBearer USER_B_TOKEN); the other 8 are caught by no rule with or without the allowlist, so their clean pass is stated as proving nothing rather than counted as coverage. A floor assertion keeps the load-bearing set non-empty. Separately, the enforcement surface (core/tests/,evals/,loop-write-guard.py,pre-tool-guard.sh,loop-ledger.sh,hooks.json,adapter.sh,gitleaks.toml) is declared five times —grade.sh'ssurface_listpathspec, itsguarded_surface_re,loop-write-guard.py's_guarded_dirs/_guarded_files, and that module'sGUARDED_TOKENS(hoisted to module level here), which is the only copy gating Bash writes: a file present in the other four but missing there escalates on Write/Edit yet stays freely rewritable viased -i/redirect.grade-test.sh's new(g4) SURFACE DRIFTextracts all five LIVE (never hand-mirrored) and asserts they cover one surface,(g4-floor)adds a content floor so a synchronized shrink of every copy cannot pass, and(g4-mutation)drives each comparison through its own shared function with a dropped entry to prove the checks are sensitive rather than vacuously equal.verify-all.shgained a SKIP lane for auto-discovered checks (exit 2 = inapplicable → tallied skipped, reason printed). Without it a battery gated on an optional binary could only signal "cannot run" by exiting 0, which the runner printed asPASSwhile discarding the battery's own SKIP text — and the CIverifyjob has no gitleaks installed, so the new allowlist battery would have reported green in CI having asserted nothing.verify-all-test.shcase (9) pins the lane, including that exit 1 is still a hard FAIL.
Fixed
- Council escape hatch never opened for a session running in a
subdirectory —
council-escalation-gate.pykeyed its state directory on the raw root string, which the--council-flag setCLI (git top-level) and the hook (the PreToolUse event'scwd) spell differently whenever the session's cwd is a subdirectory of the repo, or the checkout is reached through a symlink (macOS/tmpand/varare symlinks — three distinct keys for one repo were measured). The flag was written where the hook never read it, so/council-reviewstep 0.5 was a no-op and the council's own reviewer dispatch was denied. Both paths now canonicalize throughcanonical_root()(git top-level, thenrealpath). Regression: gate test section F13 — verified to fail on the pre-fix code, not just pass on the fixed code. /wrapand/verify-completionshelled intocore/by bare relative path — same plugin-cache defect class already fixed for/council-review:bash core/infra/council-threshold.shandbash core/infra/call-worker.shresolve against$PWD, which on a plugin install is the user's project, not the harness. Both now resolve via${CLAUDE_PLUGIN_ROOT:-$PWD}, and/wrapreports a missing script as SKIPPED rather than reading its nonzero exit as "not council-scale".core/tests/council-dispatch-path-test.shgrew a 4th check sweeping every shipped skill for bare invocations of the three council dispatch scripts. The wider "skills shell into barecore/paths" gap (7 SKILL.md files as of 2026-08-21, mostlycore/tests/helper calls) is NOT fixed here — it stays tracked indocs/claude-plugin-install-lifecycle.md§7 item 1, with a grep to reproduce the current list.
Known limitations (not fixed, tracked here)
- Linux sandboxing for grok/antigravity worker lanes — both workers use
sandbox-exec(macOS-only) for their write-deny/credential-read-deny profile; on Linux they refuse unless the explicit unsandboxed opt-out (GROK_WORKER_ALLOW_UNSANDBOXED=1/ANTIGRAVITY_WORKER_ALLOW_UNSANDBOXED=1) is set./worker-setupwarns about this rather than solving it — a Linux sandbox equivalent is backlog. - Cross-vendor task-allocation routing —
/worker-setupstep 2 briefs on cost models and a general allocation shape (prefer free/subscription lanes for advisory volume, reserve metered kiro lanes for gate votes) but does not implement any routing logic. A candidate/specitem when this is taken up. - Preflight stderr is not redacted at the adapter layer — the skill now
reports only a probe's exit code and first stderr line (the auth-failure
branch is where a vendor CLI is likeliest to echo a key prefix or account
identity), but
antigravity-preflight.sh'semit_capturestill strips control characters only.call-worker.shitself is clean: preflight stdout/stderr go to/dev/nulland never reach.agent/workers/*.md(only the argv does, viaunavailable_reason). Adapter-level redaction of key-shaped tokens is backlog. - No CI gate on remote installers in shipped instruction files — the
/worker-setupinstall step is fenced as user-run with an inspect-before-run note, but nothing mechanically stops a future skill from telling a session to pipe a remote script into a shell. A fifthsupply-chain-scan.shclass overskills/,agents/,commands/,rules/would make this a CI failure instead of a review catch. - Consumer repos do not inherit a
.gitignorefor.agent/workers/— this repo ignores worker captures, butsetup.sh --projectships no gitignore, so a consumer running worker lanes cangit add -Afull reply bodies of a reviewed diff. Surfaced by the 2026-08-20 security review; pre-existing, not introduced here.
[0.5.8] - 2026-08-20
Added
- Grounded completion gate (P1) — a completion CLAIM ("done, tests pass")
is now checkable against evidence, not just self-report. False-success
claims are 45-76% of agent failures in the literature (arXiv 2606.09863);
independent verification cuts that to ~3%, but LLM-judge-only blocking
overtrusts itself (arXiv 2406.07791, 2410.21819) — so this lands as a
deterministic hard gate with the judge kept strictly advisory on top.
core/infra/completion-verify.pygains two opt-in flags, additive to its existing behavior (a regression case pins the no-flags output unchanged):--require-evidencerefutes a claim that changed code but cites zero tests/assertions (newevidencedimension);--diff-base <ref>audits every ADDED line ofgit diff <ref>for a skip/xfail/it.only(/test.skip/describe.only(/broad-mock marker, and flags a changed code file outside the claim's declaredfiles/scope(newtrajectorydimension; lockfiles/build artifacts are exempt via anchored patterns). New consumercore/infra/completion-gate.sh(docs/gate-registry.mdcompletion-gaterow — the first BLOCKING consumer of a completion-verify.py verdict,docs/scoring-convention.md):AGENT_VERIFY_BLOCKING=off|dryrun(default)|block, a liveness canary that fails the gate closed (exit 2) if the verifier cannot REFUTE an obviously-false claim, a JSONL sink written in every reachable mode, and aneeds_semanticexit-0 advisory (never a hard block) when a claim carries no deterministically-checkable evidence at all. Wired into/supervise's wave-audit step via--verify-blocking;/verify-completioncross-references it as the blocking counterpart to its own advisory verdict. New failure modefalse-done-claiminevals/failure-modes.yaml. Tests:core/tests/completion-verify-test.sh(+22 cases, r1-r9),core/tests/completion-gate-test.sh(new, 41 cases).
Fixed (code-review round, same day)
- needs_semantic no longer downgrades a real REFUTED — the gate's "claim
cites nothing checkable, dispatch a semantic pass instead of blocking"
advisory used to trigger off
core_total==0alone, which also fires when--require-evidencelegitimately refutes an empty claim against a real code diff; that refutation was silently getting waved through as advisory even in block mode. Now gated on whether theevidence/trajectorydimensions themselves failed, not just on citation count. - Claim path now falls back to the repo convention —
.agent/claims/<slug>-w<wave-i>.ymlfirst, then.agent/claims/<slug>.yml(skills/verify-completion/SKILL.md,docs/scoring-convention.md); neither existing is reported as a distinct "claim file missing" refutation, not completion-verify.py's generic parse-error text.skills/supervise/SKILL.mdandtemplates/delegation-contract.mdnow instruct the wave's worker to write the claim file on completion. dist/build/.agent/node_modulesscope exemptions are repo-root anchored (^dist/, not(^|/)dist/) —src/dist/real.pyno longer rides the build-artifact exemption.- Liveness canary checks both sub-checks independently (file-existence
AND test-failure refutations both present), not just the top-level verdict
— catches a partial death (e.g. a regressed
_run()that always returns success) that a file-check alone would still report as REFUTED. AGENT_COMPLETION_VERIFY_BINis honored only underAGENT_REPRODUCE_TEST=1— removes the arbitrary-binary-injection surface the seam had on a production block-mode run (security review finding).- Empty/unparseable verifier stdout now logs an explicit named refutation
instead of a bare, unexplained REFUTED; the record also carries
verifier_stderr(truncated) anddiff_base_used(ref ornull) for post-hoc debugging. _TRAJECTORY_SKIP_PATTERNSgained@unittest.skip/skipIf(/skipUnless(.
Known limitations (not fixed, tracked here per code-review request)
- TOCTOU on the sink path:
resolve_sink's confinement check and the later>>append are two separate operations — a symlink swapped in between them could redirect the write. Not exploitable by the gate's own callers (no untrusted input reaches the sink path outside the env override already confinement-checked), but not hardened against a concurrent attacker with write access to the sink's parent directory. Backlogged, not fixed in this round.
[0.5.7] - 2026-08-02
Release-only version bump: ships the already-merged #101 (decision-time
model-routing advisory + telemetry-digest.sh --model), #102 (cross-runtime
harness docs), and #103 (circuit-breaker exit-code fix + verify-observer) to
plugin consumers. The installed-plugin cache is keyed by manifest version, so
without this bump claude plugin update was a no-op and the PreToolUse
model-routing advisor never reached live sessions (measured: dispatches still
inheriting the session top model four days after #101 merged).
Added
- Decision-time model-routing advisory hook (
core/hooks/model-routing-advisor.py). The 2026-07-11 audit measured 7/7 dispatches silently inheriting the session top model with the post-hoc observer already live — observation alone does not change dispatch behavior. This PreToolUse (Task|Agent) hook adds the decision-time counterpart: when a dispatch is about to inherit top (no registry pin, no call-timemodeloverride, and not the intentionally inheritingPlanagent) it injects a one-lineadditionalContextreminder pointing at the tier table. Advisory only — never blocks, never rewrites the model, always exits 0, swallows all errors (same fail-safe contract as the observer), writes no logs. The "no runtime model-switching" rejection indocs/model-routing.md§ What-this-policy-deliberately-does-not-do is refined, not reversed: blocking/auto-switching/classifiers stay rejected; a deterministic decision-time reminder sits inside the line, and the dispatch decision stays with the dispatcher. Registered inhooks/hooks.json, the Claude settings template,docs/gate-registry.md(decision=advise), andcore/hooks/README.md; batterycore/tests/model-routing-advisor-test.sh(18 checks: warn on unpinned, silent on override/pin/Plan, fail-safe on broken stdin, zero observer-log pollution). Doc counts (hooks 25→26, tests 56→57) updated in §7 live-count declarations and READMEs. - Model axis in telemetry (
telemetry-digest.sh --model). New mode (same architecture as--gates) summarizing.agent/logs/model-routing.jsonl: verdict distribution (override / pinned_specialist / inherit_top) and tier-multiplier-weighted relative spend (reuses manager-audit's LOW 0.15 / MID 1 / TOP 3.5 and its env seams). Loud SKIP when the log is absent; always exit 0. Adversarial security-lane hardening (1 MED + 2 LOW, same wave): per-record type validation routes malformed log lines (stringtotal_tokens, non-stringmodel/verdict) into the existingskipped_malformedtally instead of a traceback-then-exit-0 false-green, an outer guard turns any analysis failure into the same loud SKIP as the absent-log path, log-derived strings are whitelist-sanitized ([A-Za-z0-9:._-], 80-char cap) in the text report (JSON keeps raw values for machine consumers), and supervise Step 5 writesrouting: skip (no routing log)— neverclean— when the routing log is absent or empty (manager-audit short-circuits an empty log to lane-clean PASS, so absence must not be transcribed as cleanliness). Battery +16 checks total intelemetry-digest-test.sh(8 model-axis, 3 type-confusion, rendering-fuzz forged-line/raw-ESC probes). - M-9 backlog row — per-model harness re-tuning. Cursor's swarm experiment had to drop a frontier model mid-run (emphasis markers taken literally → runaway loop): external evidence that harness prompts need re-validation per model. M-9 adds a prompt-compat smoke row to the tier-promotion procedure (backlog only, no mechanism).
- M-6 — cost-effective-harness concept doc.
docs/concepts/cost-effective-harnesses.mddistills the 2026-07 happytlog article + ClaudeDevs thread (intelligence-placement patterns orchestrator/advisor/verifier, the advisor checkpoint finding — mid-run re-ranking beats upfront frontier advice, measured anti-correlation — coordination-cost economics with the small-task inversion, prompt-cache worker-reuse), plus a dated 4-guidance audit table mapping each point to harness state. Backlog rows M-6/M-7 (✅, this PR) and M-8 (open — delegation economics telemetry: spawn-vs-reuse ratio, per-wave delegated volume) added todocs/harness-improvement-plan.md§4.10 with the M-series count declaration updated 5→9 (M-8 from the recovered commit, M-9 added in this PR below); article added to §8 references. (Recovered from stranded branchclaude/cost-effective-harness-article, commit7aaa587, 2026-07-16 — authored pre-0.5.x but never merged; landed here unchanged.)
Changed
- M-7 — delegation-economics policy wiring.
docs/model-routing.mdgains an Intelligence placement — the advisor pattern section (three TOP placements by task shape; advisor rule: checkpoints stay TOP and recur mid-run — a single upfront TOP ranking is measured anti-correlated; no new mechanism, stays a documented convention) and two new Floors: a coordination-cost floor (boundary tokens billed ≥2×; below a threshold task size solo TOP is cheaper than orchestration — delegate only when delegated volume dwarfs the handoff) and prompt-cache preservation (route repeat calls to the same worker; fresh spawn re-pays the context write; verifiers always fresh — isolation beats cache). SuperviseSKILL.mdModel policy gains the placement corollaries paragraph and Step 2b's orchestration rules go four→five (Worker reuse (cache));delegation-contract.mdWave shaping gains Handoff must pay for itself and Worker reuse over fresh spawns bullets. Enforcement map labels the new rules honestly as conventions (call-time choices are not statically verifiable).docs/README.mdindex gainsmodel-routing.md(pre-existing omission) and the new concept doc. Docs only, no behavior. (Same recovered commit as M-6 above.) - Cursor swarm evidence folded into the delegation-economics canon.
docs/concepts/cost-effective-harnesses.mdgains a "Planner-output quality is a cost control" section (tier placement dominates the bill — hybrid at ~1/8 of frontier-everywhere at equal final score; ambiguous briefs convert planner savings into worker spend; low-correlation cheap review stacks; per-model re-tuning) with the Cursor source added;docs/model-routing.mdIntelligence-placement gains the orchestrator corollary tying the restatement-quality lane and LE-9 into the cost path (relative numbers only, no price constants). LE-9's evidence column cites the same measurement. - Supervise Step 5 routing visibility. The manager-audit offer is promoted
to an automatic post-run summary of the routing-waste + token-spend lanes,
written to RECORD.md's new
routing:stub field (full 4-lane audit stays an offer; still non-blocking).supervisor-goal.shstub gains the field; battery +1 check.
[0.5.6] - 2026-07-28
Added
- Machine-identity PII guard. After a full-history PII audit (350
commits, gitleaks + targeted scans; zero secrets, two low-sensitivity
history-only artifacts accepted as-is by owner decision), the sanitize
gate now forward-blocks personal machine identifiers:
core/tests/sanitize-audit.shgains a token group for real macOS home paths (placeholder spellings stay legal) and Apple mDNS device hostnames (shaped to never false-positive on.env.local/settings.local.json); the 3 pre-existing placeholder paths in the template/fixtures were normalized in the same commit. Paired RED-mutation cases extendcore/tests/sanitize-audit-test.sh(catch real path, catch hostname, no-false-positive). Rule + git-author noreply-email guidance:rules/public-repo.md§ Machine-identity PII, cross-referenced fromAGENTS.md§2. Adversarial security-lane hardening (4 MED, same wave): hostname pattern covers macOS's real hyphenated ComputerName forms (Mac-mini/Pro/Studio); the PII tokens moved to a case-sensitiveTOKENS_CSgroup (kills latent false positives on lowercase REST-route/users/text and imacros-style.localfilenames);--rangenow also scans each commit's metadata (author/committer name+email, subject, body) — the historical leak class was a hostname-bearing author field that diff lines never show; battery extended to 10 checks (hyphenated RED case, FP guards, metadata RED case). - Community files. Root
CONTRIBUTING.md(verification battery + ground rules),CODE_OF_CONDUCT.md(Contributor Covenant 2.1, contact via issues — no personal email),.github/ISSUE_TEMPLATE/(bug / feature); the PR template moved to the standard.github/path. docs/demo.md— three reproducible gate-catch scenarios (denied secret read, REFUTED false-"done" claim, caught PII plant) with real captured outputs and GIF-recording instructions.docs/launch-checklist.md— maintainer-run distribution guide (directory listings, launch copy drafts); all submissions stay manual.
Changed
- README star-readiness rework (en + ko). Pain-first opening with the install one-liner in the first screen, live CI badge, evidence-cited category comparison table, promoted blind-benchmark result (honest caveat intact), and stale ungated prose counts corrected to live values (agents 2→3, skills 8→9, tests 54→56, hooks measured as 21 wired / 25 scripts). Hero copy and launch-checklist wording keep the enforcement claim honest: secret access and destructive commands hard-deny; the spec gate is disclosed as observation-mode by default with block opt-in.
docs/benchmark/landscape.mdself-row refresh. The document violated its own refresh cadence and still described this repo pre-0.3 ("no eval suite", 4 CI jobs, 2 agents/4 skills); the self-row, strengths and gap→backlog map now match the live repo (E-1/O-1/T-1..3/L-1/M-4/M-5 shipped; O-2/L-2 open), with a dated self-row verification line.- Guard trim (evidence-based, GT series). Executed the
docs/freedom-enforcement-calibration-2026-07.md§3c "instrument, then recalibrate" follow-up against 30 days of gate telemetry (new §3e records the verdicts):check-hardcoding.py: harddeny→ default dryrun withAGENT_HARDCODING_MODE=off|dryrun|block(block restores the old deny); gains a firing sink (.agent/logs/hardcoding.jsonl, schema 2.0.0,reproduce_testexclusion) closing its UNINSTRUMENTED status; the docstring/error claim thathook-config.yml: hardcoding_patterns[]customizes it — never wired — is now honestly labeled as planned (T-4). Battery rewritten to 21 checks (3 modes, unset-env-defaults-to-dryrun contract, sink schema).session-quality-gate.py: the subjective style scan (inline types, hex colors, console.log) is advisory by default — reported and logged, never blocking; theblockdecision is owned solely by the objectivesession.completion_testslayer (P3-1, unweakened). Opt back in withAGENT_QUALITY_STYLE_BLOCK=1. Battery +6 checks (style-only no-block + sink write, opt-in restore, completion-still-blocks).- High-firing safety gates (destructive / secrets / verify-bypass) are explicitly unchanged — firing volume is defensive track record, not creativity suppression.
Fixed
- Adversarial-review hardening (2 lanes, 2026-07-27). code-reviewer
(2 MINOR + 1 NIT) and security-reviewer (2 LOW, verdict APPROVE — both
demoted gates confirmed to carry no security duty;
pre-tool-guard.sh/secret-content-scan.pyverified untouched):- check-hardcoding: unknown
AGENT_HARDCODING_MODEvalues now warn loudly on stderr and degrade to dryrun (a typo likeBlockno longer silently weakens an intended hard gate); sink writes are confined to the repo root or the system temp dir (escaping/empty overrides fall back to the default in-repo sink — mirrors the digest's read-side confinement), and the same writer-side confinement was applied tospec-gate.py. - session-quality-gate: the block reason's FIX remedies now name only the layer that actually blocked (completion vs opt-in style), instead of always listing style fixes that no longer block.
- telemetry-digest: reproduce_test-excluded records are now tallied and
reported as
suppressedinstead of silently dropped, so a lingeringAGENT_REPRODUCE_TEST=1in a real session surfaces as an anomaly rather than invisibly blinding the DEAD/FATIGUE audit. - gate-registry: hardcoding row's decision column reads
deny(opt-in; default=dryrun)so the scannable machine block can't be misread as a live deny.
- check-hardcoding: unknown
- telemetry-digest
--gatesdouble-counting. Two gates sharing one sink AND one guard value (secrets-bashvssecrets-content, bothguard=secretsinsecurity-violations.jsonl) each reported the union (both fired=2179; actual in-window 1353/826).count_sinknow also matches the record'shookfield against the registry hook when present (legacy hook-less records keep guard-only matching). Regression battery section (l), 3 checks.
Added
supervisor-askgate registered (registry blind-spot closure).supervisor.pyhad logged ~440 records / 11 ask-intents in 30 days while absent fromdocs/gate-registry.md(firing but unreviewable). Ask-family records now carryguard/hook/reproduce_teststamps and the registry gains asupervisor-askrow — mode unchanged (measure first; the digest counts from registration forward since older records lack the stamp).- Roadmap v2 (§4.13, growth axes + GT/AP/MC series).
docs/harness-improvement-plan.mdgains a 5-axis growth plan (guard calibration / low-cost-high-power / beginner onboarding + proposal-based adaptation / plugin interop / model currency) that re-sequences the open backlog (P2, LE, T-4, H-4, W-10, O-2) into waves and adds only three new series: GT (this release's executed trim, 5/5 ✅), AP (proposal-based personalization: local usage aggregation →/tuneproposal skill → append-only adopt/rollback ledger → onboarding presets; auto-change stays forbidden), MC (model-currency watch + new-model bench-replay procedure). docs/model-routing.mdCurrency log. Dated re-verification section: 2026-07-27 entry confirms the generation-neutral reviewer alias pins self-update (no edit needed), the codex adapter profiles match the current lineup, and files the above-top client variant as a watch item gated on bench evidence (MC series) per effort-before-tier-up.
[0.5.5] — 2026-07-23
Added
- Persona-review skill + orchestrator agent (citizen/user review lane). A
new
/persona-review <target>skill andpersona-review-orchestratoragent seat a panel of grounded Korean citizen personas in front of UX / copy / content and report how ordinary users would react — a user-perspective lens besidecode-reviewer(correctness) andsecurity-reviewer(vulnerabilities), replacing neither. Ships a stratified persona catalog (skills/persona-review/personas/catalog.json, N=119) subsampled from the publicnvidia/Nemotron-Personas-Koreadataset (CC BY 4.0) via DuckDBhttpfs: age_group × province hard-balanced across a 7×17 grid (one persona per cell — every province gets an even seat, not a population-proportional one) plus occupation_group × education_level soft-balancing, with a 1-2-tagreview_lens(UX/카피/접근성/신뢰/가격민감) per persona. A reproducible builder (skills/persona-review/scripts/build_catalog.py) and a determinism battery (core/tests/persona-catalog-test.sh: schema, CC-BY attribution, stratification sanity, skill/agent wiring) guard it. Personas are synthetic (no real individuals, no name/contact fields). (Converged from two parallel builders: an initial population-proportional HTTP-paging sampler was superseded by the DuckDB full-dataset hard-balanced one for genuinely even province coverage — the earlier approach left small provinces as few as 2 of 120 seats.)
[0.5.4] — 2026-07-21
Truthfulness repair wave. An external cross-AI audit (2026-07-21) found the docs claiming more than the code delivered and the doctor verifying installability but never installed state. This release makes the claims match the code, the two install paths match each other, and the doctor check the machine it runs on — and adds gates so each repaired drift class cannot silently return (version parity, hook-manifest parity, memory-dump pollution).
Added
- Version-parity gate (
core/tests/version-parity.sh+ battery). README (en/ko) badge + status lines,plugin.json,marketplace.json, and the CHANGELOG's latest release heading must all agree on one version; any lag fails the suite (this release repaired a three-file drift: README and marketplace at 0.5.1 vs plugin manifest at 0.5.3). - backends.json v2 lane registry (PR #89). Cross-vendor worker lanes gain
tiers,
enabled/preflight flags, and status frontmatter, replacing the flat single-backend registry. - Doctor real-wiring checks (15–18).
setup.sh --doctornow verifies the install is actually wired, not just installable: codex wiring (brain MCP[mcp_servers.brain]+codex-shell-wrap.shin the live config; WARN when absent, FAIL when a wired path is missing on disk), gemini wiring (same policy for~/.gemini/settings.json— doctor previously had zero Gemini checks; newGEMINI_SETTINGSseam), claude install path (reports whether the plugin or the shell install is live; WARN when both — hooks can run twice — or neither), and brain strict lint over the live store (WARN when the/brain-ingestpromotion gate would be blocked). 12 new fixture cases insetup-doctor-test.sh.
Fixed
- Brain W1 lint measures isolation, not just out-degree.
lint.pyW1 now fires only for notes with no typed edges in either direction and skipsstatus: seedimports (seeds are declared-unconnected; distillations arestatus: growing, where the strict gate keeps its teeth). A freshly seeded store no longer fails--strictwholesale — previously 143 warnings made the/brain-ingestprecondition (--strictat 0) permanently unsatisfiable. - doc-reality prunes
.claude/. A live multi-session worktree under.claude/worktrees/carries another branch's doc snapshots; the gate was scanning them and failing this branch for that branch's paths..claude/is untracked runtime state, pruned like.agent/. - Hook-manifest parity gate (
core/tests/hook-template-parity.sh+ battery).hooks/hooks.json(plugin path) andadapters/claude-code/settings.json.template(shell-install path) had silently diverged by six hooks; the template is now a faithful mirror of the manifest (SSOT:hooks/hooks.json) and the gate fails the suite on any future drift (path prefixes normalized; chain order enforced). Shell installs now also wirespec-gate.py,plan-scope-allow.py,model-routing-observer.py,rubric-commit-judge.sh, the WebFetch/MCPsecret-content-scan.pymatcher, andsupervisor.pyon UserPromptSubmit (replacing a staleplan-gate.py) — behavior change is observation-only since spec-gate/tdd-guard default todryrun. - Plugin session capture.
brain-capture.pyadded to the plugin Stop chain (hooks/hooks.json) — the 0.5.3 brain feature now actually captures sessions on the plugin install path, not just shell installs. - Codex profile templates refreshed to the GPT-5.6 family (PR #88).
quick → light lane, deep → top lane at xhigh effort, ahead of the 2026-07-23
legacy-model sunset; model ids verified against the Codex CLI's own models
cache, with a currency test (
codex-template-currency-test.sh) pinning them. - Memory-pollution guard (
core/tests/memory-pollution-guard.sh+ battery). Fails the suite when an AI-memory plugin's session-context dump (observed injected intoAGENTS.md) is present in any committable file — tracked or untracked-unignored — so a personal session log can never reach a public commit. Wired into the/wrappre-flight gate list.
Docs
- Docs truthfulness. README (en/ko) no longer claims
spec-gate.py/tdd-guard.py"physically block" — both ship in observation mode (off | dryrun | block, defaultdryrun; onlypre-tool-guard.shalways blocks) and the modes are now documented. The "same guardrails no matter which AI" claim is split into decision parity (machine-tested) vs per-runtime event coverage, with an honest Runtime coverage table.docs/architecture.mdno longer claims the plan-approval flag has no consumer — both wired consumers (spec-gate.py,plan-scope-allow.py) are now named. Stale README counts corrected (39→48 test scripts, 7→8 skills —brain-ingestadded to the catalog).
[0.5.3] — 2026-07-19
Consolidation note. The changelog was not rolled from [0.2.5] (2026-07-08) through the 0.5.3 cut, so entries for the 0.3.x–0.5.3 releases accumulated under Unreleased. They are gathered here at 0.5.3 rather than retro-split into per-version sections, because no version tags or per-release
release:commits exist for 0.3.x/0.4.x/0.5.0 to place undated entries reliably (a guess would be fabrication). The inline(vX.Y.Z)/(PR #NN)labels below preserve the finer per-release provenance. Future releases roll Unreleased into a dated section.
Added
- Candidate scoring + project rubric (PR #79).
/spec --score-candidatesscores a field of candidate problems/approaches numerically (candidates × dimensions → top-K) instead of a prose pick, mirroring the--interviewsubmode. A domain-neutral project rubric (templates/rubric.yml.template→.agent/rubric.yml) is scored deterministically bycore/infra/rubric-score.py(shared verdict schema, refute-by-default) and run per commit by the advisorycore/hooks/rubric-commit-judge.shhook — trust-gated to personal-tier repos only (grader_checks are shell commands and.agent/rubric.ymlships with the repo tree, so a foreign clone's rubric is never auto-executed) — and folded intoverify-completion's semantic judge on-demand (the two-layer split that keeps the fresh-spawn judge off the commit path). Follow-up PR #80 hardened the hook's test battery: a.agent/rubric.jsonfallback (no PyYAML dep) makes the trust-gate tests self-standing, plus a foreign-origin collab case proving theowner-decides-alonebranch also blocks auto-execution. sanitize-audit.sh --rangemode (PR #78). The domain-neutrality gate can now scan a commit range (PR/push span), catching taint that a single commit adds and a later commit removes — an add-then-remove sequence the per-commit scan misses. CI wires an explicit--rangeverification step.
Fixed
doc-reality.shno longer scans gitignored.agent/(PR #79). The gate'sfindwalked.agent/plans/, so a/specplan naming to-be-created files false-positived as phantom-path drift..agentis now pruned (matching the gate's own "tracked *.md" intent), with a regression case indoc-reality-test.sh..gitignore—.agent/plans/re-entry gap closed (v0.5.2). The runtime-state block enumerated.agent/locks|logs|state|workers/but not.agent/plans/, so agit add -Acould commit per-run records (RECORD/RESTATEMENT/PROPOSALS) — artifacts that demonstrably carry machine-local absolute paths. Audit note for the record: the tracked tree itself was verified clean (sanitize-audit tokens + manual grep — no personal paths shipped); this closes the door, it does not clean up a leak. If committed specs are ever wanted, un-ignoring.agent/plans/**/spec.mdis the narrower future carve-out.
Docs
- sqlite3 + jq declared as goal-mode prerequisites (v0.5.2).
core/infra/supervisor-goal.shhard-requires both (exit 127) but README Prerequisites,setup.sh --doctor, and the supervise skill never said so — the same undeclared-dependency class the jq/telemetry-digest fix already established as a bug (docs/harness-improvement-plan.mdP1-5 — jq removed there precisely because a hard dep contradicts doctor's WARN tier). Now: README/README.ko Optional entries, doctor check 14 ("goal-mode deps", WARN-tier — goal-mode is optional), and a prerequisite note on the--goal-moderow inskills/supervise/SKILL.md.
Added
docs/concepts/fable-5-prompting.md— frontier-model dispatch guidance (v0.5.2). Distills Anthropic's Fable-5-class prompting guide into 8 harness-mapped rules (effort-before-tier-up backing, anti-wrap-up, evidence-grounded progress claims, boundaries+why, delegation, memory surface, two registers, no reasoning replay). Advisory: cited from the supervise Model policy, the delegation-contract template (evidence-citation- anti-wrap-up lines added), verify-completion (artifacts-not-replayed-
reasoning hard rule), and
docs/model-routing.md. Enforcement lane is backlog (LE-9), not shipped.
- anti-wrap-up lines added), verify-completion (artifacts-not-replayed-
reasoning hard rule), and
- Agent SDK loop cross-check —
docs/loop-engineering-audit-2026-07.md§4 (v0.5.2). Maps the SDK's loop controls (max_turns, budget, effort, permission modes, compaction, subagent isolation, result subtypes, hooks) to harness equivalents. One real gap found: no per-run turn/dispatch cap (token budget only) → LE-8; compaction/scheduling stay intentionally runtime-native. New backlog: LE-8 (dispatch cap), LE-9 (fable-5 prompt audit lane).
Fixed
- v0.5.1 version bump — re-cut of stale 0.5.0 caches. PR #73 changed
shipped content (manager-audit
--since, Explore-MID fan-out exception, Step 0 run-start timestamp) without bumping the version, so installed 0.5.0 caches were cut at the earlier commit and silently missed those fixes. The plugin distribution is the git tree keyed by version — a content change without a bump leaves every existing install stale.
Added
manager-audit --global— slug-less routing sweep. The audit's routing-waste + token-spend lanes previously required a<plan-slug>, so ad-hocExplore/ execution dispatches outside any/superviserun — the common case for TOP-inherit leaks — were never swept.--globaldrops the slug and run scope, skips the slug-scoped lanes (restatement-quality, role-compliance), and reports inherit-top / floor / fan-out leaks across the wholemodel-routing.jsonl. Still after-the-fact, still WARN-only, still exit 0 (no runtime switching). Covered by 8 new cases inmanager-audit-test.sh./manager-audit— meta-audit of the supervisor (v0.5.0). Four lanes answering what the supervisor cannot be trusted to answer about itself:restatement-quality(intake restatement exists, six sections filled, measurable criteria, scope-drift candidates),routing-waste(TOP-inherit leaks, verify/judge MID-floor violations, fan-out not at LOW),token-spend(relative dispatch cost = tokens × tier multiplier — LOW 0.15 / MID 1 / TOP 3.5, midpoints of the docs/model-routing.md relative ranges; still no price constants in the repo), androle-compliance(every wave audited, never-auto-retry honored, RECORD.md written, review lane after code waves). Split mirrors harness-audit: deterministic machine layercore/infra/manager-audit.sh(env-seamed, always exit 0, findings JSON)- interpreting skill
skills/manager-audit/SKILL.mdthat judges semantic candidates and writes concrete patch proposals to.agent/plans/<slug>/PROPOSALS.mdfor one-click user approval. Proposals target conventions/templates/docs only — runtime model-switching stays rejected (docs/model-routing.md). Battery:core/tests/manager-audit-test.sh(28 checks, fixture-injected seams, including review-driven regressions: dangling-flag termination, BSD-grep whitespace sections, non-ASCII wave titles, remediated-FAIL non-flagging, mixed-tier fan-out evidence, unsafe-slug rejection). First real-run smoke (2026-07-17) produced two user-approved fixes: the fan-out lane now honors the documented Explore-at-MID exception, and--since <ISO-ts>scopes the audit window to one run (Step 0 records the run-start timestamp in RESTATEMENT.md for exactly this) — battery now 30 checks.
- interpreting skill
- /supervise Step 0 "Intake restatement" — before plan validation, the
supervisor now restates the user's chat prompt into a machine-checkable
record (
skills/supervise/templates/prompt-restatement.md: Original ask verbatim / Interpreted goal / Assumptions / Out of scope / Success criteria / Open questions), persisted to.agent/plans/<slug>/RESTATEMENT.md. Non-full-auto runs surface unresolved Open questions before Wave 1. Completion (Step 5) now offers/manager-audit <slug>, the meta-audit that grades this restatement along with routing waste, relative token spend, and role compliance. - model-routing-observer spend signal — each dispatch record in
.agent/logs/model-routing.jsonlnow carriesprompt_chars(always) andtotal_tokens(best-effort probe oftool_responseusage,nullwhen the runtime surfaces none). Pure-observer contract unchanged (silent stdout, exit 0 always). This is the measurement seam for the manager-audit lane (relative dispatch-cost ranking — tiers and token counts only, no price constants, per docs/model-routing.md).
Removed
- Gemini backend retired from
backends.json—second-opinion-reviewnow ships codex-only (fallback: null), and thegeminibackend entry is gone. Real-call verification (2026-07-17) showed the path fails by default for individual installs: gemini-cli 0.44–0.46oauth-personalis deprecated upstream (IneligibleTierError→ Antigravity migration), and the API-key path demands paid prepay credits. A shipped fallback that cannot work out of the box misleads worse than no fallback. The dispatcher's fallback mechanism is unchanged and stays stub-tested; to re-enable, add a backend entry and point a role'sfallbackat it (the removed entry lives in git history).
Fixed
backends.json: gemini headless invocation was wrong —gemini -prequires an argument (Not enough arguments following: pon CLI 0.44.x); the prompt rides stdin, so the registry now ships["gemini", "-p", ""]("-p is appended to input on stdin" per the CLI's own help). Found by the first real-call smoke of the second-opinion lane — the PATH-stub battery can't catch a vendor's argv contract, which is exactly why the smoke run is part of the lane's rollout checklist.
Added
- Cross-vendor second-opinion lane (
core/infra/backends.json+core/infra/call-worker.sh): a Claude session can now dispatch Codex (primary) or Gemini (fallback) as a review/verification second opinion. The registry is the machine-readable role→backend SSOT; model names never appear in it — adapter profiles own the tier (docs/model-routing.md§ Cross-vendor second-opinion lane). The dispatcher refuses withoutAGENT_WORKER_YES=1(paid-call gate, env-only because a headless caller can't answer prompts), names a missing CLI (exit 127), preserves fallback reasons in the capture header, and kills hung workers attimeout_s(exit 124). Captures land in.agent/workers/<ts>-<role>.md.verify-completiondocuments an optional--second-opinionflag (evidence input only — gate logic unchanged). Tests:core/tests/call-worker-test.sh, ten contract paths on PATH-stubbed backends, zero paid calls in CI. (Benchmark input: netwaif/multi-agent-starter'sbackends.json+call_worker.shrole registry — design borrowed, no code.)
Removed
legacy/trim-2026-07-04/removed from the shipped tree — the lastlegacy/payload is gone (preserved on tagarchive/legacy-trim-2026-07). The plugin package is the git tree (no exclusion manifest), so these 44K of retired agents/skills rode into every release. The sixlegacy/-scoped exclusion rules (gitleaks.toml,sanitize-audit.sh,supply-chain-scan.sh,doc-reality.sh,check-hardcoding.py,secret-content-scan.py) stay unchanged: CHANGELOG entries still reference historicallegacy/…paths, and doc-reality needs the exclusion to keep ignoring them. Recover anything withgit show archive/legacy-trim-2026-07:legacy/trim-2026-07-04/<path>. (Benchmark input: netwaif/multi-agent-starter ships a deliberately minimal generated tree — our tree-is-the-package model makes git removal + archive tag the only diet mechanism.)
Changed
-
Benchmark landscape: 2026-07-16 spot re-check recorded (
docs/benchmark/landscape.md): netwaif/multi-agent-starter v3.3.0 (three patterns adopted — cross-vendor lane, install→validate pairing, checksum update mode), inkeep/open-knowledge v0.33.0-beta.5 (sandbox trial, 11-target init pollution measured, conditional-adoption verdict with--scope project --no-skillscontainment, GPL boundary noted, skill-symlink SSOT logged as a pattern candidate), and the implicit execution-dispatch permission surface closed by the/supervisepre-flight. -
/supervisegained a dispatch pre-flight for edit waves. Root-cause of "the supervisor can't edit files": a dispatched subagent cannot answer a native permission prompt, and a background dispatch auto-denies any call that would prompt — so an edit wave sent out without a cleared permission surface silently loses its Edit/Write calls. The skill now requires either the plan-approval flag (/tmp/agent-plan-approved, armsplan-scope-allow) or project-layerEdit/Writeallow rules before dispatching an edit wave, and falls back to foreground dispatch otherwise. Also made explicit: execution waves must never route tocode-reviewer/security-reviewer(read-only toolsets, CI-enforced) —supervisor.pysuggestions name the review lane, not the execution lane. Cross-referenced fromdocs/model-routing.md. -
setup.shinstalls now end in post-install validation. Every install path auto-runs the existing read-only--doctordiagnosis and the script exits non-zero on FAIL rows, so a broken install fails loudly at install time instead of at first use (AGENT_SETUP_NO_DOCTOR=1skips — test seam / air-gapped bootstrap). Install paths also self-heal lost exec bits oncore/hooks/*andadapters/*/adapter.shbefore validating, so exec-bit-hostile distribution paths (ZIP download) don't hard-fail a check the script never remediated. Pattern adopted from multi-agent-starter's generate→validate.pyPASS/FAIL pairing, reusing our existing doctor instead of a new validator. -
apply_template()is now idempotent (checksum update mode). A target byte-identical to the fresh render reportsup-to-datewith no prompt; only a target that actually differs (user-customized, or template changed) asks before overwriting. Re-running setup over an existing install is now a quiet no-op update pass instead of a prompt per file — the equivalent of multi-agent-starter's update mode ("overwrite system files, preserve user data") for our template set. -
Gate-registry correction: quality-completion RETIRE-CANDIDATE → KEEP-CONDITIONAL (same-day supersede). The retirement investigation refuted its own premise:
session-quality-gate.pyis also the enforcement layer for the landed P3-1session.completion_testsfeature, making it a conditional-path gate (fires only where a consumer declares completion tests or the defaultsrc/style scan matches) — zero in-window firings is an adoption gap, not dead wiring, and the gate is battery-verified in verify-all. Deleting it would have silently unshipped P3-1. Recorded as a correction row rather than rewriting history; real follow-up is W-6 generalization/adoption. Lesson reinforced: a DEAD flag on a conditional-path gate needs a wiring-vs-adoption diagnosis before any retirement verdict. -
Gate-registry DEAD review (2026-07-15) + audit follow-up bookkeeping. The four gates the T-2 digest flagged DEAD got their review verdicts recorded and
last_reviewedbumped: project-policy KEEP (conditional-path), r4-mutex KEEP (contention gate, active multi-session ops), context-mode KEEP-CONDITIONAL (tied to the plugin's W-10 review), quality-completion RETIRE-CANDIDATE (two audits agree: wired but never fired, sink never created — retirement is a separate reviewed code PR). New backlog row W-10: third-party plugin-injection pollution scan (observed live: a plugin's injected<context_window_protection>block acting as session-wide instructions, independently flagged by two research subagents — the mirror image of the P3-4 self supply-chain scan). E-1 gains a design-input note: llm-council's anonymous peer-review → chairman pattern as an option for the real-LLM judge track (pattern only, no code import). Docs/bookkeeping only, no behavior. -
docs/benchmark/landscape.md2026-07-14 spot re-check. Five-lane web re-verification of the 2026-07-08 snapshot: claude-flow renamed ruflo (2026-02-27, npm/CLI unchanged); Aider marked effectively stalled (last release 2025-08); new "Field shift" note — Anthropic Dynamic Workflows GA (Pro plan 2026-07-02) erodes orchestration-first third parties but not this harness's governance position; re-check found no surveyed leader closing the prompt-only-enforcement gap. Survey-only change, no behavior.
Added
-
Evidence-first inventory — kill the ghost-specialist deadlock at its root (
rules/policy/evidence-first.md+core/hooks/agent-inventory.py). A gate that demands a specialist with no in-runtime provider deadlocks the session: the gate blocks the work, and the very thing that would unblock it can't be dispatched (observed live — a stale plugin cache requiredui-ux-director/fe-architect/… agents that exist only in thelegacy/v0 mirror, not in the active registry). New policyevidence-first.mdnames the underlying failure — asserting present state from memory instead of a same-turn read — and forbids demanding a provider you haven't confirmed exists. Enforced by a new SessionStart truth pass,agent-inventory.py, which reconciles the active registry (including consumer overrides CI never sees) against the agent*.mdproviders actually beside it:real(id has a sibling provider),ghost(id has none → quarantined, never demandable),discovered(an unwired provider). A filename is not evidence — a provider is a.mdthat carries YAML frontmatter, i.e. actually defines an agent. A strayREADME.mdin the registry dir is not dispatchable, so it must never bediscovered(--syncwould wire it in as a bogus agent), and an id backed by a frontmatter-less.mdlands inghost, notreal— the file exists but the runtime cannot dispatch it, and calling that "real" is the deadlock. This is also what makes the inventory strictly stronger thansupervisor.py's ownis_real_agent()file-exists check rather than a restatement of it, sohas_provider()'s AND actually buys something. The verdict is written to.agent/state/agent-inventory.json(gitignored runtime state — no git churn, no drift-guard conflict) andsupervisor.pyconsumes theghostset as an extra dispatch-time quarantine source (has_provider), fail-open so a session with no inventory behaves exactly as before. Opt-in hybrid auto-correct (--sync/AGENT_REGISTRY_AUTOSYNC=1) additively wiresdiscoveredproviders into the registry, copying eachmodel:straight from the agent's own frontmatter so the additive write can't introduce the model driftregistry-drift.shcheck 4 forbids — it only ever adds, never edits or removes. Fail-open is end-to-end: the SessionStart path,supervisor.py'smain(), and the manual CLI all exit 0 on a malformed consumer registry — the reconciler must never itself become the thing that breaks a session.AGENTS.mdlinks the rule from "When in doubt." Batterycore/tests/agent-inventory-test.sh(8 checks: real/ghost/discovered classification, provider-needs-frontmatter, inventory persistence + ghost-set readback, additive-sync-with-model-copy, fail-open on no registry, fail-open on a malformed one), auto-discovered byverify-all.sh.supervisor-dispatch-test.shgains two cases: 15 — a ghost that guards afile_globspath, listed ahead of a real guard on the same path. The existing ghost case (7) is a keyword ghost declaring no globs, so it could never reach the file-glob matcher — leaving the deadlock's other form untested, and that form is the incident's own example (a retirededge-fn-devguarding**/functions/**). It pins both halves of the contract: the ghost is hinted and skipped (continue, notreturn, so it cannot swallow the guard behind it), and no emittedaskever names an undispatchable specialist. 16 — inventory quarantine end-to-end: a provider.mdis on disk (sois_real_agent()passes) but the reconcile rejected it, andhas_provider()must quarantine it anyway. This is the one case where the inventory layer adds power over the file-exists check, so it is the case that proves the layer is load-bearing rather than decorative. -
model-routing-observer.py— measure the model-tier convention. A 2026-07-11 transcript audit confirmed the call-timemodel-override convention is not followed: 7/7 subagent dispatches in the audited session inherited the session top model. New PostToolUse (Task/Agent) pure observer classifies every dispatch asoverride/pinned_specialist/inherit_topinto.agent/logs/model-routing.jsonl(analyze:jq -r .verdict … | sort | uniq -c) — measured before enforced, per the gate- registry philosophy. Never blocks, emits nothing, exit 0 always. Batterymodel-routing-observer-test.sh(15 checks). Companion policy fixes:docs/model-routing.mdgains a "Built-in agents (Claude Code)" section (Plan=TOP inherit, Explore=MID default / LOW for bounded lookups — a deliberate exception to fan-out-LOW), and/verify-completion's general-reviewer dispatch now requires an explicit workhorse-tier override. -
Skill A/B evaluation dataset (H-3, seed).
evals/datasets/skill-ab.jsonlis the labeled seed for measuring whether a shipped skill earns its keep: does itsdescriptionroute the right requests (and not the wrong ones), and does running with the skill produce what a no-skill baseline would miss. 35 cases across the 5 shipped skills (spec,supervise,verify-completion,wrap,harness-audit) — 3assertioncases + 2 trigger-positive + 2 trigger-negative each. Trigger labels are grounded in each skill's ownwhen_to_use(positives) andNOTclauses (negatives), including cross-skill discriminators (e.g. "execute the approved plan" must route tosupervise, notspec); every assertion'srationalequotes the skill's shipped description, so the seed is grounded, not invented. Fail-closed floorevals/baseline-skill-ab.json(min_cases,min_assertions_per_skill= 3, the shipped-skill list). New batterycore/tests/skill-ab-dataset-test.sh(21 checks) validates only the seed's shape — parse, per-kindrequired fields, unique slugs, unknown-skill rejection, the ≥3-assertions-per-skill floor, ≥1 trigger-positive and ≥1 trigger-negative per skill, and that every named skill is a real shippedSKILL.md— with 7 malformed RED fixtures proving each guard goes red; it calls no model and is auto-discovered byverify-all.sh. The A/B runner that executes with-skill vs baseline and scores the assertions is a later increment (H-3 본체). Seeevals/README.md(Skill A/B track) for the n=35 and honest-ceiling disclosure. -
Failure-mode grader rubric (L-1, doc portion).
evals/failure-modes.yamlreplaces the autonomous loop's single opaqueharness_scorescalar with a checklist of 12 named failure modes, each distilled from a real adversarial-review catch in this campaign (caught_incites the item + PR): silent-drop, vacuous-green, vacuous-parity, glob-scope-miss, bypass-flag, unanchored-skip, infra-as-verdict, lexical-containment, injection-breakout, loose-coercion, stale-ssot, review-false-clean. Every mode carries adetection_signal(the observable a grader looks for) and agrader_check(the boolean question whose safe answer the candidate must satisfy). Naming the holes makes the grader adversarial the way a human review is — "strengthening verification" without naming failure modes does not stop metric-gaming (proxy-hacking measured at 73.8%). The §5.1 correspondence table'sval_bpbrow is rewritten to describe per-modemode:<id> PASS|FAIL — reasonemission, keeping a rollupharness_score:line on both the GATE-pass and GATE-fail paths so the §5.2 grep consumer and the results.tsv status enum stay intact (GATE-fail emitsharness_score: 0= discard, never an empty grep that would misclassify as crash). New batterycore/tests/failure-modes-test.sh(25 checks) validates the file's shape through the same PyYAML parser the grader will use — schema_version, the ≥8-mode floor, unique kebab-case ids, and every required field non-empty — and proves each guard can go red with six malformed RED fixtures (too-few / missing-field / blank-field / duplicate-id / non-kebab-id / unparseable), auto-discovered byverify-all.sh. Thegrade.shimplementation that consumes the rubric is deferred to a later batch. -
plan-scope-allow.py— plan-approved auto-allow accelerator. New PreToolUse (Write/Edit/MultiEdit, last in chain) hook: once the user approves a plan this session (theplan-gate.pyflag), in-workspace non-risk edits emitpermissionDecision: "allow"so the native permission prompt stops firing on every step of approved work. First permission-weakening hook in the harness — emits only allow-or-silence (never deny/ask), fail-open direction is silence, risk areas (spec-gateGUARD_PATTERNS) +.agent/hook-config.yml+.git/+ out-of-workspace (realpath containment) always pass through, env-gatedAGENT_PLAN_ALLOW_MODE=on(default off, ships dark), sink.agent/logs/plan-scope-allow.jsonl, registered indocs/gate-registry.md. Batteryplan-scope-allow-test.sh(27 checks incl. symlink/..escapes, case-evasion, sink discipline). Side fix: README/README.ko hook-count drift (17 → live 19) corrected. -
skills/harness-help/— router skill (ask-matt pattern, user-invoked): main flow/spec→ approval →/supervise→/verify-completion→/wrap, standalone/harness-audit, and what to do when a gate interrupts. Sync rule: any skill add/remove updates the router in the same commit. -
docs/skill-authoring.md— skill-writing reference distilled from Matt Pocock'swriting-great-skills(MIT, attributed): invocation axis, one trigger per branch, information hierarchy, leading words, checkable completion criteria, failure modes. Applied as a surgical pass over the five existing SKILL.md files (trigger dedup + explicit completion criteria). -
Gate registry + fire-rate digest (T-2).
docs/gate-registry.mdis the SSOT list of every deny/ask/block gate with the model weakness it assumes and alast_revieweddate (an assumption expires).telemetry-digest.sh --gatescross-references it against the runtime firing logs and reports per gate: DEAD (0 in-window firings), FATIGUE (firings ≥--fatigue, default 50), STALE (last_reviewed+--stale-days, default 90, is past), and UNINSTRUMENTED (emits a decision but writes no log — reported honestly, not mislabeled DEAD). Test-reproduction records (reproduce_test:true) are excluded from fire-rate so batteries can't inflate it. Still an observer (exit 0 always);telemetry-digest-test.shgains a synthetic-registry battery covering all four classes plus the reproduce-test exclusion and missing-registry fail-safe. -
Runtime enforcement of
risk_areas.secrets.paths(P1-8).pre-tool-guard.shnow reads project-declared secret paths fromhook-config.yml(via a bounded, metacharacter-rejectinghook_config.load_risk_area_secret_pathsloader) and denies read/copy/exfil access to them — closing a field that shipped as schema but was never read at runtime. Additive: the built-insecrets/guards run first and are never weakened; a config value can only add literal paths, never inject a pattern.risk-area-wiring-test.shproves enforcement + loader safety bounds. -
Remote-URL credential scan + gitleaks fire drill (W-3). A token baked into a git remote URL lives in
.git/config, invisible to every content scanner —core/git-hooks/scan-remote-url.pyflags an http(s) remote whose userinfo carries a password or a token-shaped value (no false positives on ssh / clean / bare-username URLs), wired as pre-push step 0 and a/wrappre-flight.core/infra/gitleaks-fire-test.shplants a synthetic secret matching the repo's own rule and asserts gitleaks catches it (PASS = gate live, FAIL = misconfigured allowlist, exit 2 = gitleaks absent/SKIP), so a clean result can be trusted.remote-url-scan-test.shcovers both. -
Tests for plan-gate and tdd-guard (P1-3). Both hooks previously shipped untested.
plan-gate-test.sh(7 checks; newAGENT_PLAN_FLAGseam so tests never clobber the live approval flag) andtdd-guard-test.sh(12 checks; isolated mktemp git repos, RGR red/green/no-test verdicts). tdd-guard's misleading comments claiming hook-config override of its risk-area patterns were corrected to match reality (the override is a tracked follow-up, not yet wired). -
Real-LLM semantic judge (E-1, batch-3) — out-of-CI.
evals/judges/llm-judge.pyis the layer above the deterministic floor: it catches a cited test that carries a real assertion (soreference-judge.pyCONFIRMs it) yet never exercises the claimed change — asserting on an unrelated function, a stale inline copy, a mock, or a tautology. It conforms to the same verifier interface (llm-judge.py --root <root> <claim.json>→ shared verdict JSON), reads the citedtest_sourcesand claimedfiles(bounded, realpath-contained), embeds them as delimited DATA in a prompt, and asks a real model via a subprocess CLI (LLM_JUDGE_CMD, defaultclaude -p;LLM_JUDGE_MODEL/LLM_JUDGE_TIMEOUT). Refute-by-default (unparseable output, missing/mistyped keys, emptytest_sources, a root-escaping path, or confidence< 0.6→ REFUTED) is kept distinct from fail-closed on infrastructure (absent CLI, timeout, nonzero exit, empty stdout, or an invalid strict-integer timeout → clear stderr error, nonzero exit, and no verdict on stdout, so a broken backend never masquerades as a confident label). Prompt-injection is contained, not merely hedged: embedded evidence is wrapped in per-call nonce markers (untrusted content cannot forge the closing marker) and any marker-shaped substring in content is defanged before embedding, so a test that embeds a literal closing marker cannot break out of the DATA quarantine. Labeled datasetevals/datasets/llm-judge.jsonl(10 semantic-hard cases, 5 CONFIRMED / 5 REFUTED — the deterministic floor scores 5/10 on it, by construction) with fail-closed floorevals/baseline-llm.json(min_cases 10, enforced by the LOCAL run, not CI). New deterministic batterycore/tests/llm-judge-test.sh(34 checks) drives the adapter with a MOCK CLI, so it needs no real model and runs offline — auto-discovered byverify-all.sh. CI is untouched: it must never call a model, and the track runs locally with--repeat 1(Pass^k's identical-verdict rule would be dishonest for a nondeterministic judge — flakiness is to be observed, not hidden). Seeevals/README.md(Real-LLM track). Honest ceiling stated: nondeterminism near the threshold, a residual prompt-injection risk from hostile prose that stays within the data block, and model-availability dependence. -
Teaching gates (T-1). Every
deny/ask/blockreason emitted bypre-tool-guard.sh,secret-content-scan.py,check-hardcoding.py, andsession-quality-gate.pynow carries a fixedWHY:tag (which rule fired and why) and aFIX:tag (the concrete allowed alternative), so a blocked agent can self-correct instead of routing around the gate. Machine-enforced: the per-hook batteries assertWHY:/FIX:on every non-allow fixture, and a newcore/tests/check-hardcoding-test.sh(14 checks) gives that hook its first dedicated battery. -
Skill negative-triggers (T-3). All five shipped
skills/*/SKILL.mddescriptions now include at least oneNOT …negative example (when not to fire), enforced asregistry-drift.shcheck 7 with fixtures (no-NOT→ FAIL,NOTpresent → PASS, no frontmatter → FAIL). -
Doctor codex-profile check (M-4).
setup.sh --doctorgains check 13: WARN when thequick/deeptier-profile files are missing beside the local codex config (CODEX_CONFIGseam; skipped when no codex config exists) — the same "declared template vs actual" observer family as the plugin-cache and command-scan checks. Fixtures cover present/missing/absent-config.
Fixed
-
Guard false-positives (W-7).
pre-tool-guard.shno longer blocks a commit whose message merely mentions a destructive command (git commit -m "fix: guard rm -rf / patterns"): a preamble strips the inert message payload before the destructive guards (1–4) scan — but only provably-inert shapes (single-quoted,$/backtick-free double-quoted, or a quoted-delimiter<<'EOF'heredoc). An unquoted<<EOFbody still command-substitutes at shell-eval time, so it stays fully scannable; the secrets guards always scan the full command. A newcore/tests/pipefail-idiom-scan.shgate is the W-7(2) regression floor: it flags unguarded zero-match count pipes (grep … | wc -l,n=$(grep -c …)) in strict-mode runtime scripts, self-checking against a bad/good fixture so it can't rot to always-green; the safe idiom is documented in AGENTS.md. -
Graceful-wrap traversal guard.
supervisor-goal.sh's_emit_graceful_wrap(the budget-limited stub writer) now refuses a traversal-shaped slug the same waywrite_record_stubdoes, so a weird plan slug can never place the stub outside its configured dirs (fail-safe skip; regression-tested). -
Repo-native execution ledger (F-2).
supervisor-goal.sh completenow drops.agent/plans/<slug>/RECORD.md— a mechanical execution ledger (waves from the live DB, plus PR / audit-verdict / carried-items slots the supervise skill fills) — so an execution record exists even on runtimes with no global recording layer. Session narrative stays with the global layer; the two never duplicate. Contract: never clobbers a RECORD.md the skill already wrote, and a ledger write failure never blocks completion (fail-safe);AGENT_PLANS_DIRis the test seam./superviseStep 5 writes the same facts as its closing discipline on non-goal runs. A traversal-shaped slug (path separators or..) is refused fail-safe — the ledger can never land outside the plans root (a hardening from this change's adversarial review). New batterycore/tests/supervisor-goal-record-test.sh(12 checks, isolated in mktemp git repos). Same-PR bugfix the battery exposed:cmd_init'sobjective="${4:-$slug}"referencedslugon the samelocalline — bash 3.2 +set -ucrashes on every objective-lessinit <slug> <waves>(latent: the documented 4-arg form never hit it); the declaration is now split. -
/spec --interview— opt-in deep-interview submode (F-1). For requests fuzzy enough that a wrong guess commits the spec to the wrong shape, the one-shot "ask if ambiguous" brainstorm gains a structured question loop: an unknowns table marking each row decision-changing (Y/N), batched questions over the Y rows only (at most 4 per round, options + recommended default named), re-scoring after each round (answers resolve rows and surface new ones — decision-tree pruning), and two termination conditions (zero open decision-changing unknowns, or 3 rounds — leftovers carry intospec.mdunder## Open questions). The Q/A trail lands in## Interview log, so the spec shows why it has its shape. Opt-in by design: simple requests keep the single pass, andspec-gate.pyis untouched — the enforcement boundary neither knows nor cares which submode produced the artifacts. -
Delegation contract + orchestration guards (O-1).
skills/supervise/templates/delegation-contract.mdis the per-dispatch contract skeleton: the four elements (goal / output format / tools & scope / boundaries), an explicit**model**:field (execution waves name their tier instead of inheriting the expensive session default), and three sections absorbed from the 2026-07-10 external-harness comparison — Self-contained (subagents inherit no history; the contract carries every path, decision, and constraint), Constraints re-injection (each wave re-states its relevant constraint slice, not whole rulebooks), and Executable acceptance criteria (verify = a command's exit code by default; prose only with a stated reason). Wave shaping travels with it: fan-out cap 3–5 with a worked split-at-the-cap example, write single-threading (one writer per fileset), and verifier isolation (fresh spawn, end-state only)./superviseStep 2 now composes every dispatch from the template;/specStep 3 defaults→ verify:to an executable check. The guardable half is machine-enforced:registry-drift.shgains check 5 (review/verify agents must carry a read-only toolset — a write-armed or allowlist-less reviewer fails) and check 6 (a tree that shipsskills/supervisemust ship the template with its**model**:field), each with injected-defect fixtures inregistry-drift-test.sh(24 checks — inline and YAML-blocktools:forms both parsed; a write tool smuggled into either form fails). -
Clean-install CI smoke (M-5). New
clean-installCI job: installs all three runtimes into a scratch$HOMEnon-interactively (AGENT_SETUP_YES=1 bash setup.sh --all), then asserts the install itself — all three runtime configs exist and the{{FRAMEWORK_ROOT}}placeholder was actually templated (anti-vacuous: a no-op or mis-templated install must not go green; this assert step, not the doctor, carries the install verification, since the doctor's checks are repo/env-scoped and tolerate an empty home — a gap this change's adversarial review measured live). The doctor must then report 0 fail against the scratch home, and a built-in mutation probe (a hook stripped of its exec bit must turn the doctor red, else the job fails) keeps the job's green load-bearing. Closes the cold-install gap the landscape survey flagged: "install once, use everywhere" was previously verified only by hand. -
Harness self-audit skill + extracted registry-drift gate (H-2).
skills/harness-audit/SKILL.mdis an agent-driven, read-only self-audit that sits on top of the machine gates: oneverify-all.shdry-run, a per-check PASS/FAIL/SKIP table, an explicit citation of the P1-1 doc-reality verdict, and for each failure a root-cause + fix + backlog-follow-up. It interprets the gates; it does not reimplement them. Enabling refactor: the CIvalidate-pluginjob's four inline checks (manifest required fields; hooks.json → executable core-hook resolution; agentname:frontmatter; registry↔agent model agreement) are extracted tocore/tests/registry-drift.sh(a standalone, cwd-independent gate with aREGISTRY_DRIFT_ROOTtest seam) with a non-vacuous fixture batterycore/tests/registry-drift-test.sh(11 checks — each drift class injected and asserted caught). The gate is auto-discovered byverify-all.sh, closing the one machine check the unified runner was missing. Same-PR hardening from the 2026-07-10 workflow audit: the skill gains a runtime-layer step (setup.sh --doctor+core/infra/telemetry-digest.sh, with an unmeasured-is-unmeasured reporting rule) and a negative trigger in its description (T-3 applied early); doctor gains check 12 — the phantom-command scan (a runtimecommands/*.mdinvoking a script that does not exist on this machine is a live failure path — reported as a WARN-only observation;AGENT_COMMANDS_DIRtest seam, relative refs resolved against the runtime root, unexpanded$VARrefs skipped, and control characters stripped from echoed refs — an escape-sequence display spoofing hardening from this change's security review) with 7 new fixture checks insetup-doctor-test.sh(25 checks) — the rule-ification of a real orphaned-command failure found live in that audit. Backlog bookkeeping lands in the same PR: new §4.11 F-series (F-1 opt-in--interviewdeep-interview submode for/spec; F-2 repo-nativeRECORD.mdexecution ledger), the G-2 global-hygiene decision record, and a matchingF-rowsseries check indoc-reality.sh. -
doc-reality gate — the harness gates its own doc drift (P1-1).
core/tests/doc-reality.sh(+ a 39-case battery) fails CI when a shipped doc contradicts the repo: (A) a referenced in-repo path that does not exist — scanned across every tracked*.md(recursive; nested READMEs anddocs/**included), minus the forward-looking plan, the backward-looking CHANGELOG, andlegacy/; fenced code blocks (0–3-space fences, CommonMark-tracked) are illustrative examples and stripped, while an unterminated fence is itself a malformed-doc failure; (B) the six backlog-count series and (C) the four artifact counts declared in the plan §7, cross-checked against the live repo. New CI job. Hardened over four adversarial-review rounds; also completed the phantom-ref cleanup it surfaced (core/hooks/README.md,core/hooks/secret-content-scan.py,core/hooks/context-mode-guard.sh,rules/policy/strong-goal-template.md). -
Eval harness — labeled dataset + Pass^3 + regression gate (E-1, deterministic layer).
evals/run-evals.pygrades the completion verifier (core/infra/completion-verify.py) againstevals/datasets/completion-verify.jsonl(12 labeled CONFIRMED/REFUTED cases): each claim must produce its labeled verdict, the suite runs three times (Pass^3) requiring identical results, andevals/baseline.jsongates on a coverage/accuracy regression. NewevalsCI job; runner batterycore/tests/evals-test.sh(28 checks). The LLM-judge semantic layer and the skill A/B dataset are later increments. -
Eval harness — semantic track, deterministic floor (E-1, batch-2). A judge that catches green-by-construction tests — a cited test that "passes" but asserts nothing real.
evals/judges/reference-judge.pyconsumes a claim'stest_sourcesand classifies each file meaningful iff it holds at least one real, non-constant assertion (line-based bash+python heuristics), emitting the shared verdict schema; it is refute-by-default, reads are bounded, and no path escapes--root. It is deliberately biased to false-REFUTED over false-CONFIRMED (a completion gate must never bless a hollow test). Graded by the existing runner againstevals/datasets/semantic-judge.jsonl(17 labeled cases, 8 CONFIRMED / 9 REFUTED) under Pass^3 +evals/baseline-semantic.jsonregression gate; wired as new steps in theevalsCI job; judge batterycore/tests/reference-judge-test.sh(52 checks, incl. leak-safety on unsafe..//absolute/symlink paths and python constant-comparison triviality across decimal/hex/octal/binary/underscore/scientific number forms). This is the deterministic floor only: it catches syntactic triviality (no real / only constant assertions), not semantic triviality (a real-looking assertion that never exercises the changed code path) — that deeper judgment needs a real model and runs viaskills/verify-completionor a pluggable--verifier, not in CI. -
Unified verification runner — one command runs the whole suite (P1-2).
core/tests/verify-all.shfulfills the README "Verification" one-command promise. It discovers the check set dynamically — everycore/tests/*.shexcept the runner and its own test — so the four gates, all thirteen*-test.shbatteries, and any script added later are picked up with no edit here (the anti-rot property). It then runs the two evals layers (the exact CI invocation) and gitleaks, reporting PASS / FAIL / loud SKIP per check and a final tally; a silently-skipped check reported as pass is exactly the false-green this repo guards against. It isset -u(not-e) so every check runs even when one fails, isolates each in a subshell, and refuses to report success on an empty discovery (zero checks → non-zero exit, not a vacuous green). Batterycore/tests/verify-all-test.sh(21 checks) proves completeness against the live dir (no hardcoded list), fail-propagation, all-green, list-equals-run, the empty-set floor, and that an absent gitleaks is counted skipped, never passed. NewverifyCI job runs the self-test then the full discovered set — the sole CI executor of the ten checks that had no dedicated job, and auto-inclusive of future ones.
Changed
- README synced to 0.2.5 reality. Version badge/status, skill count and
table (spec + verify-completion were missing), the model-tier paragraph
rewritten to match
docs/model-routing.md(judgment inherits session-top; reviewer pins are the only machine-enforced part; implementation/mechanical dispatch via per-call override — conventions), and the docs index gains model-routing.md + benchmark/landscape.md. - Cross-AI parity gate strengthened (P1-6).
core/tests/adapter-parity.shnow asserts the three adapters (claude-code / codex / gemini) return the same normalized decision for a logically-identical event — and the same full decision JSON — across all three verbs (deny/allow/ask) and bothtool_inputshapes, each driven through a hook that actually reads that shape so a mistranslated field flips the decision: the command shape viapre-tool-guard.sh, the file/content shape viacheck-hardcoding.py(deny on hardcoded content, allow on clean) — plus shell-special (quoted) command and content. The prior version only checked a "deny" substring per adapter independently, so two adapters could disagree and still pass. 24 parity assertions; exit 1 on any divergence. (Kept the filename:cross-ai-parity.shnamed in the plan was a phantom P0-1 removed, and the docs already referenceadapter-parity.sh.)
Removed
legacy/v0-mirror-2026-05-12/retired from the shipped tree — defence-in-depth for the ghost-specialist deadlock (preserved on tagarchive/v0-mirror). The plugin package is the git tree (marketplace.jsondeclares"source": "./"and there is no exclusion manifest — the only things missing from a release are the gitignored ones), so the v0 mirror rode into every release: 194 files, 1.3 MB, including 33 retired agent.mdproviders (ui-ux-director,fe-architect,edge-fn-dev, …) sitting next to two 64-entrymaster-registry.jsonfiles. That adjacency is exactly the shapeis_real_agent()trusts — a registry id is "real" iff a sibling<id>.mdexists — so any copy of that tree into an active registry path resurrects the deadlock the rest of this release exists to kill.find_registry()never reaches intolegacy/, so this was a latent trap and dead weight rather than a live path (the live defence isagent-inventory.pyabove); removing it means a retired provider can no longer be resurrected by a stale copy.legacy/trim-2026-07-04/stays — its 3 agent.mdfiles have no sibling registry, so they cannot satisfy the predicate, and keepinglegacy/alive keeps the fivelegacy/-scoped exclusion rules (gitleaks.toml,sanitize-audit.sh,supply-chain-scan.sh,doc-reality.sh,check-hardcoding.py) valid and unchanged. Recover anything withgit show archive/v0-mirror:legacy/v0-mirror-2026-05-12/<path>.
Security
- Codex/Gemini adapter synthetic-mode no longer builds canonical JSON by string
interpolation.
adapters/{codex,gemini}/adapter.shconstructed the event JSON by interpolating--command/--contentinto apython3 -cliteral ('''$TOOL_CMD'''), so a value containing a quote, newline, or'''broke the literal — mis-parsing the event (a cross-adapter parity break) and, worse, letting a crafted command inject python and force anallow, bypassing the very guard the adapter feeds. The values are now passed via environment variables to a fixed python program (no interpolation), carrying any command/content verbatim. Regression-locked by the new quoted-command parity cases.
[0.2.5] — 2026-07-08
Changed
- Model policy: judgment stays up, hands get dispatched. The supervise
Model policy now names all four judgment classes (planning/design, wave
dispatch decisions, gate verdicts & abort/advance, result synthesis) as
session-top work, and adds an execution-dispatch row: implementation waves
dispatch at the workhorse tier and mechanical work at the low tier via an
explicit per-call
modeloverride — inline execution at the session's top model is the expensive default this rule exists to prevent. All of this is documented convention (only specialist pins are CI-enforced); O-1's delegation-contract template gains a requiredmodelfield in its done-condition.docs/model-routing.mdadds the matching orchestration- judgment row and reworks the Implementation row. - Doctor check 10 now scans every cached plugin, not just this harness.
Any
<marketplace>/<plugin>/with more than one cached<version>/is WARN-listed (a live third-party dual-version cache motivated the generalization — the same stale-cache drift class check 10 was built for). WARN-only, absent cache root still passes. Fixtures: 3 new cases (third-party dual → WARN, multi-plugin single → PASS, stray file at version depth ignored).
Added
- Floors: long-horizon implementation is not a LOW-tier task.
docs/model-routing.mdcites an external program-based-verifier benchmark (datacurve-ai/deep-swe, 2026-05 leaderboard, press-reported) showing light-tier models trailing the top tier by ~40+ points on long-horizon SWE work — cost/performance reference data reinforcing that LOW is for bounded mechanical tasks only. docs/benchmark/landscape.md— survey + self-assessment against the most popular agent harnesses on GitHub (2026-07-08 snapshot, API-verified stars): two comparison tables (Claude Code ecosystem / general harnesses), field investments vs field gaps, evidence-linked strengths and weaknesses, explicit non-goals with reversal conditions, and a gap→backlog map. Not a run benchmark —results.mdremains the only measured comparison.- M-5 (backlog) — clean-install CI smoke: bare checkout → scratch config
home →
setup.sh --doctorasserts 0 failures. Distribution integrity was a consistent field investment in the survey; the cold-install path is currently only verified by hand.
[0.2.4] — 2026-07-08
Added
- M-1 —
docs/model-routing.md. Canonical cross-runtime model-tier policy: a three-rung ladder (LOW mechanical / MID workhorse / TOP reasoning) with an orthogonal effort dial ("effort before tier-up"), a work-class → tier table for Claude Code / Codex CLI / Gemini CLI, floors (verify-judge ≥ MID, fan-out workers default LOW — worker tier is the dominant cost lever at ~15× multi-agent token spend), and an enforcement map. Explicitly rejected: runtime model-switching hooks, automatic tier escalation, dedicated low-tier agents, price constants in the repo. - M-2 — verify-judge tier floor.
/verify-completion's Layer 2 semantic judge is documented as never-below-workhorse (sonnet-class): a low-tier refute-by-default judge emits plausible false CONFIRMED verdicts and silently disables the completion gate. Low-tier sessions must pass an explicitmodeloverride on the judge dispatch. - M-3 — adapter templates carry the tiers. The Codex adapter ships
quick.config.toml.template(LOW) /deep.config.toml.template(TOP) as per-profile config files (recent Codex CLI builds reject inline[profiles.*]tables as legacy — verified against a live CLI); the Gemini template ships a workhorse default model with explicit-mescalation. Model IDs are marked as 2026-07 snapshots.
Changed
skills/supervise/SKILL.mdModel policy: the mechanical-fixes row now names its real mechanism (explicit per-callmodeloverride on the Agent dispatch — no low-tier agent is shipped) and the table links todocs/model-routing.mdas its cross-runtime generalization.
[0.2.3] — 2026-07-08
Changed
- I-1 — secret-content-scan matcher consolidation. The 6 non-edit matcher blocks (supabase ×3 tools, firecrawl ×5, WebFetch, Notion ×3, Google Drive ×2, stitch ×2) collapse into one union matcher; the registration inside the Write|Edit|MultiEdit chain stays to preserve chain order. 7→2 blocks, the 19 covered tools verified unchanged, no double-fire (each tool matches exactly one block).
Added
- I-2 — doctor drift checks (10 & 11).
setup.sh --doctornow warns when the plugin install cache holds more than one agent-harness version (a stale cache re-exposing retired agents was a live incident), and reconciles a user-declared global-hook manifest (AGENT_HOOK_MANIFEST, default~/.claude/LOCAL-LAYER.hooks;AGENT_GLOBAL_SETTINGSfor the live file) against the runtime settings in both directions (declared-but-not-live / live-but-undeclared). WARN-only — observers never block; no manifest → check skipped. Manifest lines are trusted substrings authored by the user (an over-broad line makes the check vacuous by choice). Fixtures:setup-doctor-test.sh12 checks.
[0.2.2] — 2026-07-08
Cumulative since 0.2.0: the 0.2.1 plugin release (2026-07-07, model-routing policy + the P3-1/3/4/5 batch) was cut without sectioning this file, so the entries below span 0.2.1 and 0.2.2. Newly in 0.2.2: P3-2 (
/spec+ spec-gate), the freedom-vs-enforcement calibration audit, and the T/E + O/L/I backlog series.
Added
- Freedom-vs-enforcement calibration audit + new backlog series.
docs/freedom-enforcement-calibration-2026-07.mdtiers every enforcement/ freedom point (deny/ask/block/advisory/observe/aspirational) and grounds keep/promote/measure-first verdicts in 2026 external evidence. New backlog: T-1 teaching gates (WHY/FIX in every gate message), T-2 gate registry + fire-rate + expiry, T-3 negative skill-trigger examples, E-1 public eval harness promotion (Pass^3); O-1 supervise delegation-contract revision (4-part brief, fan-out cap 3–5, single-writer, isolated verify lane), O-2 genericskills/loop(fresh-context / one-task-per-iteration / file+git state), L-1 failure-mode-checklist grader redesign, L-2 grader/tests write-ban + append-only ledger, I-1 secret-scan matcher consolidation, I-2 doctor cache/manifest drift checks (docs/harness-improvement-plan.md§4.8–4.9). - P3-1 — completion-gate test verification.
core/hooks/session-quality-gate.py(Stop hook) now runs a project'ssession.completion_tests(declared in.agent/hook-config.yml|json) on the first Stop; any command that exits non-zero, times out, or fails to spawn emits{"decision":"block"}so a session cannot end while the project's own tests fail. A second Stop passes (anti-loop) andAGENT_QUALITY_GATE_BLOCK=0is advisory; per-command bound isAGENT_COMPLETION_TEST_TIMEOUT(default 120s). New fail-safe/bounded loaderhook_config.load_session_config(≤20 cmds, ≤500 chars each; malformed → empty). Trust model = the project's own scripts (docs/hook-config.md). Adversarial-review hardening of the always-exit-0 contract: the timeout env var is parsed at import behind a guard (a typo like2m/30sdegrades to 120 instead of crashing the Stop hook), and completion commands run withstart_new_session=Trueso a process-group teardown idiom (kill 0,trap 'kill 0' EXIT) reaches only the command's own group, not the hook. Test:core/tests/quality-gate-completion-test.sh(21 checks incl. YAML path, advisory, anti-loop, malformed fail-safe, non-numeric timeout, process-group isolation, timeout→block, loader bounds). - P3-2 —
/specupstream-planning discipline + spec-gate enforcer. New skillskills/spec/SKILL.mdwalks brainstorm →.agent/plans/<slug>/spec.md+plan.md→ ExitPlanMode approval, borrowing the superpowers methodology as CONTENT while the ENFORCEMENT is a tool boundary. New PreToolUse hookcore/hooks/spec-gate.pyis the consumer of the plan-approval flag thatplan-gate.pyalready writes (and session-init/close already clear): when no plan is approved this session and a Write/Edit targets substantive impl code (scope coverssrc/app/pages/lib/server/components, extension-gated), it acts perAGENT_SPEC_GATE_MODE—offno-op,dryrun(default) advisory-only,blockemitspermissionDecision: "ask". The approval flag is the dedup (approve once via ExitPlanMode → every edit passes; the two escapes — ExitPlanMode approval andAGENT_SPEC_GATE_MODE=off— are named in the reason).asknotdeny, per the reversible-gate escalation principle. Fail-open (any exception → exit 0, empty stdout); mirrorstdd-guard.py. Wired into theWrite|Edit|MultiEditchain (after tdd-guard, before supervisor). Adversarial-review hardening (13-agent refute-by-default; 6 confirmed): SKIP dir tokens are now ANCHORED ((^|/)(types|config|…)/sosrc/subtypes/no longer inherits thetypes/exemption — a MAJOR false-allow), the flag path is the shared hardcoded/tmp/agent-plan-approvedwith NO env override (a per-consumer override would decouple the reader from plan-gate's writer), the default scope was broadened pastsrc/to real app layouts, and matching is case-insensitive (sosrc/Pay.TScan't evade on a case-insensitive FS). Test:core/tests/spec-gate-test.sh(42 checks incl. anchored-skip, absolute paths, broadened scope + harness-not-gated, case-insensitive ext, fail-open, and the ask-not-deny + both-escapes assertions). - P3-3 — verification-gate-bypass + linter-tamper guards.
core/hooks/pre-tool-guard.shnowasks ongit commit/push --no-verify(andgit commit -n), which skip the repo's own gitleaks+sanitize commit gate, and on Bash edits to linter/formatter/gate configs (eslint, prettier, ruff, flake8, biome, golangci, pre-commit, gitleaks) — the "disable the check instead of fixing the code" anti-pattern.git push -n(dry-run) and normal commits pass; reading a config passes.ask(not deny) per the escalation principle. Adversarial-review hardening: the no-verify guard now also catches git's bundled short-flag forms (-nm,-vn; git parses-nmas-n -m), a-c key=valglobal option before the subcommand, and inlinecore.hooksPath=— while a commit whose message merely mentions-nno longer false-asks (the message is stripped before matching); the linter-tamper rule no longer false-asks on a read that redirects elsewhere (cat .eslintrc.json > backup.txt) by requiring the config to be the mutate target. First test for this hook:core/tests/pre-tool-guard-test.sh(28 checks, incl. regression coverage of the existing destructive/secret rules). - P3-4 — self supply-chain scan.
core/tests/supply-chain-scan.shstatically scans the harness's OWN shipped, auto-loaded instruction files (agents/,skills/,commands/,rules/,templates/,AGENTS.md,AI_BOOTSTRAP.md,CLAUDE.md) for three prose classes — prompt-injection override, unattended observer-loop persistence, no-confirmation coercion — and the auto-fired AI-decision hooks (core/hooks/) for a background-daemon-spawn class (nohup/setsid/disown/crontab), failing on any hit. Patterns are calibrated to zero hits on the clean tree; the no-confirm class is anchored on confirmation/permission/approval so a routing rule like "do not ask for a phantom agent" is not flagged. Explicitly-invoked plumbing with sanctioned one-shot async (autosync post-commit push,agent-session.sh subscribe) is out of scope and documented inrules/policy/security-guards.md. This is the self-integrity analogue ofsanitize-audit.sh. Wired as a new CI job (4th). Adversarial-review hardening (13-agent refute-by-default pass) closed real detection gaps: instruction files are now scanned as*.mdplus*.template(scaffolding copied verbatim into consumers) and*.json(the agent registry); hooks are scanned as every file (a hook may be extensionless); each prose class is matched against a whitespace-flattened copy so an injection wrapped across soft line breaks can't evade line-oriented grep; and the self-reference exemption is anchored to exact paths (a file merely namedsecurity-guards.mdelsewhere no longer inherits it). Test:core/tests/supply-chain-scan-test.sh(20 checks: 4-class detection incl. templates, extensionless hooks, wrapped/multi-line, JSON registry, all four daemon tokens + clean-tree pass + false-positive guards for the phantom rule, start_new_session, infra-scope daemons, and the path-anchored exemption). - P3-5 — independent completion-claim verifier (eval-harness seed). A
builder-validator layer that re-checks a completion claim from a separate
context, so "the builder says it's done" is never the last word.
core/infra/completion-verify.pyis the deterministic core: given a claim (.agent/claims/<slug>.yml|jsondeclaring cited files / tests / assertions) it mechanically checks each file exists (and contains its declared substring), each test exits 0, and each assertion holds, then emits a shared-convention verdict JSON ({verdict, score, dimensions, refutations}). Refute-by-default: a missing file, a failing test, a malformed/empty claim, or an over-cap claim all resolve toREFUTED(exit 1) — never a crash, never a silent pass; exit 0 iffCONFIRMED, so it doubles as a CI/wave GATE. Bounded (≤50 files, ≤20 tests/assertions) and hardened with the samestart_new_session=True+ guarded-timeout-parse lessons as P3-1.skills/verify-completion/SKILL.mdis the semantic layer on top — an independent-context judge that adds the "does the code actually do what the claim says / are the tests meaningful" review scripts can't make, emitting the same schema. New shared specdocs/scoring-convention.mdunifies the verdict shape across the verifier, the H-3 skill A/B harness, and thesupervisor-goal-audit.sh25-point scorer. Adversarial-review hardening (11-agent refute-by-default pass): a present-but- non-list section ("files":"x") now REFUTES instead of being silently dropped into a false CONFIRMED; command output is discarded to DEVNULL and thecontainsread is bounded (5 MB) so a chatty command or a huge file can't OOM the verifier. Test:core/tests/completion-verify-test.sh(25 checks incl. false-claim→refuted, consistent→confirmed, malformed/non-list-section fail-safe, YAML, process-group isolation, over-cap refutation, bare-string file entry, --root default, timeout degrade, partial score). core/infra/telemetry-digest.sh(P1-5) — pillar④ janitor step 1: reads.agent/logs/supervisor.jsonl(path arg, or$AGENT_TELEMETRY_LOG, or<repo-root>/.agent/logs/supervisor.jsonl) and reports action counts, a per-specialist funnel (match → ask → dispatched, with conversion %), top keywords by match count, and a rule-candidate section with three heuristics derived only from whatsupervisor.pyalready logs:NO-ACCEPT(a specialist was askedask-intent+ask-security≥3 times but neverdispatched— specialist-routing.md Lesson 1),GHOST(a specialist loggedaction=="ghost"— registry references an agent id with no siblingagents/<id>.md), andOVER-GENERAL(a single keyword accounts for >70% of allmatchrecords, once total matches ≥3).--window <days>(default 30) scopes records byts;--jsonemits a machine-parseable JSON blob instead of the human report. Known limitation, documented in the header: NO-ACCEPT can't fire underAGENT_SUPERVISOR_MODE=observe(it logsobserve-intent/observe-securityinstead, which aren't counted). Dependencies are bash + python3 only — nojq(an earlier draft usedjq, which contradictssetup.sh --doctor's own WARN-tier/optional stance on it; JSON parsing runs through an embedded python3 heredoc instead). Always exits 0 (observer, not a gate) — a malformed line, a missing file, or an internal error degrades to a zeroed/"inactive" report rather than failing. Reproduce suite:core/tests/telemetry-digest-test.sh(21 checks: action-count accuracy, funnel notation, all three rule candidates, malformed-line skip counting,--windowfiltering, missing-log handling,--jsonoutput validity, legacy v0.1-record degradation).core/hooks/supervisor.pyv0.2 — minimal dispatch-not-advise router (P1-4). Replaces the observation-only v0.1 stub. OnUserPromptSubmitit word-boundary matches the prompt against each registry agent'smatches.keywordsand records a 30-min TTL intent in.agent/state/supervisor-intent.json; the next Write/Edit/MultiEdit then returnspermissionDecision: "ask"naming the specialist (once per intent — no repeat nag), and dispatching that specialist via Task/Agent (namespace-agnostic:x:code-reviewerresolvescode-reviewer) clears the intent. A separate security matcherasks onmatches.file_globsfor the tool, independent of intent, once per path. Ghost specialists (a registry id with no sibling<id>.md) neverask— stderr hint +{"action":"ghost"}log only (specialist-routing Lesson 2).AGENT_SUPERVISOR_MODE=observedowngrades everyaskto stderr; any exception is fail-open (exit 0, empty stdout). Wired intohooks/hooks.jsonon all three events. Reproduce suite:core/tests/supervisor-dispatch-test.sh(10 scenarios).session-initnow warns (stderr only) whengitleaksorgitis missing from PATH — a mini env-doctor surfacing a degraded secret-scan setup at session start. Silent when both are present; never blocks the session or writes stdout.setup.sh --doctor(P1-7) — full environment diagnosis, read-only, zero install side effects. Checks:gitpresent;python3>= 3.9 (README-declared floor, with version + path);gitleakspresent (WARN if not — secret-scan git hook skips);jqpresent only if acore/hooks/*.shscript actually shells out to it (WARN if used-but-missing); everycore/hooks/*.sh/*.pyhas its executable bit (hook_config.pyexempt — a library module imported bysecret-content-scan.py, never invoked directly); everyadapters/*/adapter.shis executable;agents/master-registry.jsonparses and every entry'smodelmatches its siblingagents/<id>.mdfrontmatter (same drift guard as CI);hooks/hooks.jsonparses and every referenced hook script exists and is executable;~/.agent/plansexists (WARN +mkdirhint if not). Prints a[PASS|WARN|FAIL]row per check plus adoctor: N pass, N warn, N failsummary line; exits 1 iff any check FAILs. Reproduce suite:core/tests/setup-doctor-test.sh(clean-repo exit 0, gitleaks WARN under a restricted PATH, exit 1 + named FAIL line when a hook loses its executable bit — exercised against a throwawaymktempcopy, never the real tree).templates/hook-config.yml.template(P1-8 partial) — ships the real, dynamically-loadedpython_hooks:schema (core/hooks/hook_config.py) as a commented example, bracketed byLIVE-SCHEMA-EXAMPLE-BEGIN/ENDmarkers, clearly labeled as the ONE block in the file actually read at runtime — distinct from therisk_areas:/resources:/hardcoding:blocks above it, which remain declarative-only (docs/customization.md part 2). New drift-guard case incore/tests/hook-config-test.sh(case h) extracts and uncomments that example into a real.agent/hook-config.ymland round-trips it throughhook_config.load_extensions(), so template and loader can't silently drift apart.docs/customization.mdcross-references the new template section..github/workflows/ci.yml— CI: gitleaks secret scan + plugin manifest/hook/agent validation + sanitize gate- README portfolio polish: badges, Mermaid architecture diagram, agent/skill/hook catalog
README.ko.md— Korean mirror of the README (same sections, localized prose)docs/harness-improvement-plan.md— audit scorecard + prioritized backlog + autonomous improvement-loop design (Korean)gitleaks.toml— detect NVIDIA NIM API keys (nvapi-prefix; built-in rules miss it)docs/architecture.md— "Determinism and model-invariance" section: the hooks (gates) are model-invariant and machine-proven so viacore/tests/adapter-parity.sh; risk-area denial is a real enforced gate while plan-mode/TDD enforcement is not yet wired (flag is recorded but unconsumed — see P1-4/P1-8); generated content (plans, code, prose) is honestly NOT guaranteed identical across models
Changed
docs/harness-improvement-plan.md— added §4.7 P3 series (5 items) from a 2026-07-06 benchmark audit of 8 top personal harnesses (superpowers, ECC, karpathy-skills, gstack, revfactory, hooks-mastery, Chachamaru, showcase): completion-gate test verification (P3-1), upstream spec/plan discipline (P3-2),--no-verify/linter- tamper blocking (P3-3), self-supply-chain scan (P3-4), independent completion-claim verifier (P3-5). Adoptions ranked by demonstrated mechanism value, not stars; catalog-maximalism, prompt-coercion, and unattended instinct-persistence explicitly rejected as design-principle conflicts. §7 backlog count 24 → 29 (P3 patternP[0-3]).agents/master-registry.json— supervisor keyword matchers hardened to domain anchors (review follow-up, MAJOR-1; specialist-routing Lesson 1). code-reviewer drops barereview/look overfor multi-word phrases (code review,review this diff, …); security-reviewer drops baresecurity/authforsecurity review/security audit/owasp/… — generic tokens a consumer writes as often as an author (review my plan,the auth flow) no longer false-route a specialist.security-reviewer.matches.toolsgainsMultiEdit(aligns with theWrite|Edit|MultiEdithook wiring);file_globs(path anchors) unchanged. Default mode staysdispatch— only the match surface narrowed, not the enforcement.core/hooks/README.md—supervisor.pymoved from the deferred-roadmap "generic stub" row to the shipped-hooks table as the v0.2 minimal dispatcher; the roadmap row now scopes the remaining deferral to the full 54KB registry-aware orchestrator- README rewritten for first-time readers: concept primer table, install-path chooser
(plugin vs shell), prerequisites section (incl. previously undocumented
python3dependency), "See it work" example, 4-layer architecture summary, trimmed layout tree - Hook count corrected everywhere: 17 executable hooks + 1 shared module
(
hook_config.py) — previous "~25" claim was stale - README.md/README.ko.md's "Why AI-agnostic?" section now cross-links to
docs/architecture.md's new "Determinism and model-invariance" section
Fixed
plan-gatewas wired toUserPromptSubmitinhooks/hooks.jsonbut is aPostToolUsehook (its docstring and logic key offtool_name, a field absent fromUserPromptSubmitevents). Result: every invocation was a silent no-op and the/tmp/agent-plan-approvedflag was never written. Rewired toPostToolUsewith matcherExitPlanMode|Task|Agent, and broadened the plan-class check to accept theTasktool name (subagent dispatch differs by Claude Code version). READMEs' hook tables corrected to match.session-quality-gatewrote its violations log toparents[2]of the hook file — the plugin install cache when installed as a plugin — instead of the user's project. Log destination is now resolved at runtime: stdin eventcwd→CLAUDE_PROJECT_DIR→os.getcwd(). Detection and block logic unchanged.session-initcrashed at load on Python 3.9 — itspathlib.Path | Nonereturn annotation (PEP 604) is evaluated at def-time and raisesTypeErrorbefore 3.10. Addedfrom __future__ import annotationsso annotations are treated as strings; the annotation itself is unchanged. Supported Python floor is 3.9 (now documented in README Prerequisites).- Phantom test paths removed from
README.md,AGENTS.md,docs/architecture.md,docs/getting-started.md—core/tests/adapter-smoke/*/run.sh,cross-ai-parity.sh,verify-all.sh,bootstrap-test.sh, and a pytest invocation never existed; docs now reference the 4 real test scripts (sanitize-audit,adapter-parity,hook-config-test,post-commit-autosync-test) - Documented overwrite behavior corrected:
setup.shhas no--forceflag — replacements prompt interactively, or setAGENT_SETUP_YES=1 README.mdinfra path corrected:scripts/infra/agent-session.sh→core/infra/agent-session.shAI_BOOTSTRAP.mdStep 5 pledge now names the generic 5 risk areas (perhook-config.yml/rules/policy/security-guards.md) instead of prior-project domain terms; the removed terms were added to the sanitize-audit token list (failure → new rule)core/tests/sanitize-audit.shnow scans git-visible content only (tracked + untracked-unignored viagit grep --untracked), mirroring the CI job's excludes — runtime state and gitignored local files no longer cause permanent false FAILs; CI sanitize job additionally runs the full token-set audit as a superset stepcore/hooks/secret-content-scan.pyplan-file comment corrected to the canonical~/.agent/plans/path- Docs drift sweep: removed phantom hook/file references (
memory-explore-verify.py,claude-mem-watch.py,rules/policy/skill-adoption-comparison.md) that described tooling never implemented; standardized risk-area vocabulary to the canonicaldata/secrets/deploy/payment/domain-outputIDs across README, README.ko,docs/customization.md, anddocs/concepts/security-guards-generic.md; corrected the security-guard layer count from 5 to 6 (matcheshooks.json's "6-layer secret hardening"); canonicalized stale.claude/path references to the runtime's actual.agent/andrules//skills/locations indocs/concepts/multi-session-worktree.md,AI_BOOTSTRAP.md, anddocs/concepts/plan-mode.md; and removed the.claude/rules/scaffold over-claim fromdocs/architecture.md,README.md, andREADME.ko.md(setup.sh --projectnever creates it) - Docs drift sweep follow-up: removed the two remaining phantom
rules/policy/skill-adoption-comparison.mdreferences (docs/master-registry.md,skills/README.md); removed the phantomclassify-prompt.pyhook citation fromdocs/concepts/plan-mode.md(noUserPromptSubmithook exists beyondagent-session-heartbeat.sh— tier classification is the AI applying the documented heuristics itself, not an automated hook). Rewrotedocs/customization.mdend to end after discovering its documentedhook-config.ymlschema doesn't match what any hook actually loads: onlycore/hooks/hook_config.py(used bysecret-content-scan.py) reads a config file dynamically, from.agent/hook-config.yml/.json— a[regex, label]-pairsecret_patterns/exempt_paths/credential_key_namesschema, optionally nested underpython_hooks:. The previously-documentedrisk_areas:/resources:/hardcoding:map (fromtemplates/hook-config.yml.template) is not read by any hook at runtime —pre-tool-guard.sh,r4-mutex-check.sh, andcheck-hardcoding.pyeach match against patterns hardcoded in the script, not a project'shook-config.yml. The doc's oldsecret_patternsexample ({id, description, regex}objects) was also independently confirmed to silently parse to an empty list under the real loader, which expects[regex, label]pairs — verified by runninghook_config._coerce_pattern_listdirectly against both shapes.README.md/README.ko.md's customization section and other doc mentions of a dynamically-loadedrisk_areas:/resources:still describe the same not-yet-implemented mechanism and were out of scope for this sweep — flagged for a follow-up pass. - Docs drift sweep, final pass: closed out the flagged follow-up above.
README.md/README.ko.md's Customization section no longer claimscore/hooks/r4-mutex-check.sh"reads [hook-config.yml] and enforces it" — therisk_areas:block is now described as declarative (a documented policy record), with today's actual enforcement attributed to each hook script's own hardcoded patterns and the one dynamically-loaded mechanism (secret-scan extensions via.agent/hook-config.yml) called out, linking todocs/customization.mdfor the full real-vs-documented split.AI_BOOTSTRAP.md's Step 5 pledge softened from "definitions live inhook-config.yml" (implies runtime consumption) to "declared inhook-config.yml; enforcement currently lives in the hook scripts." Same fix applied to the last two remaining spots the gate caught:docs/concepts/security-guards-generic.md's "How to extend" section no longer claims "the samepre-tool-guard.shreads this and enforces it" or shows a fabricatedabort_codekey — the example is now framed as declarative intent requiring apre-tool-guard.shfork to enforce, with a link todocs/customization.md.rules/multi-agent-worktree.md's R4 mutex-resource list dropped the phantompayment-liveentry —core/hooks/r4-mutex-check.shonly ever claimsproduction-db,production-deploy, oredge-function-deploy; there is no payment mutex.
Removed
- (recorded retroactively — the trim shipped before 0.2.0 but was never logged)
Shipped agent set reduced 10 → 5 (
architect,code-reviewer,security-reviewer,test-engineer,build-error-resolver) and skills 16 → 4 (supervise,tdd,diagnose,wrap); the removed items remain available inlegacy/ - Shipped agent set reduced 5 → 2:
architect,test-engineer, andbuild-error-resolverarchived tolegacy/trim-2026-07-04/agents/. Basis: 7 weeks of session telemetry showed zero dispatches for these three, and their roles are covered by other tooling.code-reviewerandsecurity-reviewerare retained (they form the benchmarked review pair — seedocs/benchmark/results.md). Recoverable viagit mvfrom the archive plus re-adding the entries toagents/master-registry.json. - Shipped skill set reduced 4 → 2:
tddanddiagnosearchived tolegacy/trim-2026-07-04/skills/. Basis: 7 weeks of session telemetry showed zero dispatches for either skill.superviseandwrapare retained. Thetdd-guardhook is unrelated to thetddskill and continues to run unchanged. Recoverable viagit mvfrom the archive. codex-skills/retired tolegacy/trim-2026-07-04/codex-skills/. Basis: zero usage recorded in 7 weeks of session telemetry. The Codex CLI adapter (adapters/codex/) is unrelated and remains active;setup.shno longer offers the~/.codex/skillssymlink install step. Seelegacy/trim-2026-07-04/ARCHIVE-NOTE.mdfor the full recovery procedure.
0.2.0 — 2026-06-15
Added
- Claude Code plugin packaging —
.claude-plugin/plugin.json+marketplace.jsonmake the harness installable via/plugin marketplace add joymin5655/Agent→/plugin install agent-harness@agent. One install, every project. hooks/hooks.json— plugin hook wiring (SessionStart / Stop / UserPromptSubmit / PreToolUse / PostToolUse) dispatching through the Claude Code adapter tocore/hooks/via${CLAUDE_PLUGIN_ROOT}.commands/project-init.md—/project-initslash command to scaffold project-level files.LICENSE— MIT (was TBD).
Changed
- README now leads with the Claude Code plugin install path; shell
setup.shremains for Codex/Gemini or non-plugin use.
0.1.0 — 2026-05-18
Added
- Initial AI-agnostic agent framework structure
- 3-AI adapter layer: Claude Code, Codex CLI, Gemini CLI
- Canonical hook protocol (
docs/hook-protocol.md): stdin JSON event + stdout decision JSON - Core hooks (
core/hooks/): ~25 portable hooks for security, session coordination, plan-mode, TDD enforcement, drift detection - Core infra (
core/infra/): multi-session worktree coordination (agent-session.sh), commit/PR automation (auto-ship.sh), session store, supervisor goal mode - Core git-hooks (
core/git-hooks/): pre-commit (gitleaks + hardcoding scan) + pre-push (gitleaks + secret diff scan) - Generic policy rules (
rules/): 7 critical + 12 lazy-loaded archive - Generic agents (
agents/): code-reviewer / architect / build-error-resolver / security-reviewer / performance-optimizer / test-engineer / docs-writer / refactor-cleaner / tdd-guide / copy-humanizer - Generic skills (
skills/): wrap, supervise, tdd, diagnose, grill-me, grill-with-docs, improve-codebase-architecture, caveman, api-and-interface-design, incremental-implementation, source-driven-development, deprecation-and-migration, design-variant-mockup, hook-reproduce-test, triage-external-draft, weekly-digest - Codex-native skills (
codex-skills/): code-explorer, code-reviewer, database-reviewer, planner - Templates (
templates/): generic CLAUDE.md / AGENTS.md / GEMINI.md / RTK.md / karpathy.md / hook-config.yml / project-rules.md / gitleaks.toml setup.sh4-mode installer:--claude/--codex/--gemini/--project/--hooks-only- GitHub Actions workflow templates (
github/workflows.template/): secret scan + lint
Changed
- N/A (first release)
Archived
legacy/v0-mirror-2026-05-12/— original mirror skeleton + domain-specific assets from the prior project version. Seelegacy/v0-mirror-2026-05-12/ARCHIVE-NOTE.mdfor migration guide.
Security
- Base
gitleaks.tomlwith 100+ built-in patterns + extensible per-project allowlist - Generic content-scan hook with 7 default patterns covering Python/Node secret-file readers, hardcoded credentials, OpenAI-style
sk-...tokens, JWT literals, Bash secret-readers, and exfiltration viafind -exec(seecore/hooks/secret-content-scan.pyfor full pattern list) - Project-configurable risk-area abort codes via
templates/hook-config.yml.template