Skip to content

0x7067/claude-jev

v0.29.2MIT

Hands the small judgments in a session to TypeSafe's Jev, a System One model that returns typed answers instead of generating text: what kind of request each prompt is, which model tier a subagent needs, whether an edit breaks a rule already written in your CLAUDE.md or AGENTS.md, and which transcript blocks survive compaction. Kept blocks carry over verbatim — no LLM writes a summary.

Changelog

All notable changes to claude-jev. Format follows Keep a Changelog; versions follow the plugin manifest. .claude/skills/release/release.py turns the [Unreleased] section into the next release.

[Unreleased]

  • The AFK prompt router and subagent router are off unless their promptRouter or subagentRouter option is true; both stay registered and exit before calling Jev. Each cost one Jev call per prompt or spawn, AFK drops the subagent router's advice (griffinwork40/agent-afk#2778), and the prompt router does not log on AFK, so neither was measured there. adapters/afk/.claude-plugin/plugin.json declares both options with default: false, and enabled() takes the value to use when an option is unset; the Claude Code hooks still treat unset as on.
  • adapters/afk/README.md: updated Known-gaps and Setup sections against agent-afk 5.286.1 source. Blocking additionalContext on PreToolUse now reaches the model (agent-afk 5.121.0); non-blocking additionalContext is still dropped (#2778). transcript_path is now sent in hook stdin (agent-afk 5.276.14, #2647); null on daemon/chat/web. PreCompact hook event exists and can block compaction but cannot select which blocks to keep (agent-afk 5.10.0). PreToolUse hooks fire inside subagent forks (agent-afk 5.121.0); SessionStart injectContext does not. pluginHookEnv (agent-afk 5.276.21, #2700) is now documented as the supported route for forwarding TYPESAFE_API_KEY / OPENROUTER_API_KEY. CLAUDE_CONFIG_DIR and CLAUDE_PLUGIN_OPTION_* are still not passed to hooks as of 5.286.1 (#2373 open). Issue links added for #2816 (per-hook disable) and #2817 (UserPromptSubmit/Stop REPL-only).
  • The Stop sweep shares one 4.5s budget across rule classification and the Jev judgment call, so a cold-cache loadRules is covered by the same deadline and the whole sweep finishes or fails open with a logged rules-error inside AFK's 5s handler window. Previously loadRules passed no timeoutMs, falling through to the 8s client default, which could silently exhaust the budget before Jev was ever called.
  • An AFK rule check that fails while classifying instruction files logs rules-error. A failure while reading those files, before Jev is asked, stays rules-skip with reason rules-unreadable.
  • Two AFK rules that share an id keep separate probabilities, violation ids, and block ids in the log row.
  • stats.py counts a later rules-skip of a warned file as not rechecked, and takes the score from the next check that asked Jev. AFK rows stay out of Rule outcomes and its flagged-only list. They still join Rule calibration.
  • The AFK rule hook judges patch_apply. Edits made through it used to land unjudged and never reached the Stop sweep. agent-afk hands a plugin hook tool_name: "patch_apply" with the raw { changes, dry_run } input, and tests a /…/ matcher against that name alone, so the matcher now lists it. Each file in the patch is judged in its own Jev call, in parallel, against the rules scoped to that file, and each check writes its own log row. The calls share one 12 s budget inside the hook's 15 s timeout. The patch is atomic, so one file at >= 0.80 blocks the whole call and none of its files are recorded for the Stop sweep. A clean or flagged patch records every file. A dry_run call writes nothing, so it is not judged or recorded.

0.29.2 - 2026-10-02

  • package-lock.json records an integrity hash for the five @earendil-works/* 0.87.1 packages nested under pi-coding-agent. Without a hash, Claude Code's plugin installer skipped them and warned on every install and update.

0.29.1 - 2026-10-02

  • src/rules.ts no longer drops a Bash-written file when the rewrite lands in the same wall-clock second with the same size. The hook diffs write-tree snapshots taken on a copy of the real index; the copy carried a fresh mtime, which cleared git's racily-clean guard, so the new stat was trusted and both snapshots hashed the old content — the snapshot was consumed with files: [], partial: false, and no error. The copy now keeps the source index's mtime, so the scratch index makes the same racy call the real index would. About one write in ten in a fast edit loop hit this; the file's change was never recorded.
  • jev.py's missing-key error names every provider variable when no provider is pinned (set TYPESAFE_API_KEY or OPENROUTER_API_KEY) instead of always naming TYPESAFE_API_KEY; missing_key_message was built for both but ask passed the already-defaulted provider, so the unpinned branch never ran. The AFK client's missing-key error now names only the pinned provider's variable instead of always listing both (claude-jev-afk 0.6.2).
  • src/session-start.ts is deleted. It was copied into src/ by the TypeScript runtime migration but nothing registers a SessionStart hook in the Claude Code plugin — the rule digest is an AFK-only feature and adapters/afk/src/session-start.ts is its implementation.
  • An edit rule that has already blocked its file twice no longer calls the next act-band hit uncertain. The notice says that rule was already raised. A flag-band hit stays an uncertainty notice. A Stop hit held back by the session's two turn blocks, or while the agent is still finishing, says that, and does not claim the rule was cited. On AFK, the Stop sweep says the session already used its two turn blocks (claude-jev-afk 0.6.1), and that text still rides in the next prompt.
  • Two rules that slug to the same id keep separate probabilities. Keys are assigned on the full in-scope list and kept for the escalation call, so a second pass cannot write one rule's score onto the other. A rule the subject filter skips cannot inherit the other rule's score.

0.28.0 - 2026-10-01

  • Pi compaction lives in adapters/pi/. pi install git:github.com/0x7067/claude-jev loads adapters/pi/jev.ts, which handles session_before_compact and /jev. The hook calls selectBlocks in src/compact/strategy.ts. Block shaping, toolCallId pairing, bash and summary role labels, the <read-files> index, and logs under PI_CODING_AGENT_DIR stay in the adapter. Kept text is fit to 14,000 characters. Paths from dropped or truncated calls may add up to 2,000 characters. Pi's own read paths and <modified-files> sit outside that reservation. Claude's compaction budget stays 16,000 characters. A missing key returns nothing, so Pi's own summary runs. 0x7067/pi-jev is unchanged.
  • The rule hook skips files outside the project again, as the README says it does. The TypeScript port's relativeTo swapped an out-of-project path for its absolute form, so isOutside never matched and scratch files under /tmp were judged against the project's rules. In the live log from the TypeScript runtime (2026-09-27 to 2026-10-01), 36 of 194 edit judgments and 4 of 9 edit blocks were on files outside the session's working directory, 33 of those judgments under /tmp. The working directory and the edited file's directory are both resolved through symlinks before they are compared, so an edit or Bash write in a project reached through a symlink such as /tmp still counts as inside it and is recorded by its project-relative path, and a Bash command that deletes the working directory still has its changes recorded.
  • AFK installs the adapter from this repository's marketplace as claude-jev:claude-jev-afk (afk marketplace install 0x7067/claude-jev), which checks out the newest release tag. It needs agent-afk 5.271.1 or later: earlier versions list the plugin as installed but load the root Claude Code plugin's hooks instead of the adapter's (griffinwork40/agent-afk#2456). In a throwaway AFK_HOME on agent-afk 5.282.0, with the marketplace as a real directory, an afk chat turn ran the adapter's SessionStart hook and not the root plugin's. /release no longer pushes a split of adapters/afk to the afk branch, so that branch no longer moves. An install made with --ref afk should run afk plugin remove claude-jev and reinstall from the marketplace. The entry also shows in Claude Code's /plugin list for this marketplace, with a description that says it is AFK only.
  • selectBlocks and fitKept take an optional character budget that defaults to the existing 16,000, and a block may name the earlier block a tool result answers. Claude's rows bridge sets neither, so its budget stays 16,000.
  • /release bumps root plugin.json along with .claude-plugin/plugin.json and .claude-plugin/marketplace.json, and stops when those two plugin versions disagree.
  • A Pi tool result names its tool, and only a read result is held, including after a text preamble and on a turn that also ran another tool. An unlabeled Claude [tool_result] is held when the single tool call it answers is a Read, in any capitalization; a call that ran several tools holds only the result lined up with that Read. The Pi compact log leaves out per-block rows, so /jev can still read the last line. Pairing a kept result does not copy a block the rescue window already kept. The 16,000-character Pi figure is the selected body plus the 2,000-character reservation for paths from dropped or truncated calls. The digest header, ---[jev:…]--- lines, Pi's own read paths, and <modified-files> sit outside it.
  • The summary Pi shows the model includes one <read-files> list and Pi's <modified-files>. The read list merges Pi's own reads with the capped paths of dropped or truncated calls, skips a path that was also modified while that capped list is filled, and does not repeat a path. A [tool_result] with no needs, including a named [tool_result read], walks back past earlier results to the tool call, including when that call starts with other text, so a held Read keeps the path arguments. The next compaction strips those file lists only from the end of the summary. The read and modified lists follow Pi 0.87.1 in adapters/pi/file-ops.ts.
  • The repository root carries an Agent Plugins 1.0.0 plugin.json (name claude-jev, the same version as .claude-plugin/plugin.json, empty extensions). Claude Code still loads .claude-plugin/plugin.json and hooks/.
  • Compaction's checks, thresholds, and selectBlocks live in src/compact/strategy.ts. src/compactor.ts only maps Claude's compact rows on and off stdin. eval/compact.ts imports the strategy module, and its selection cache key covers that module.
  • The AFK adapter's rule hook now blocks edits. It never could on AFK before. AFK dispatches PostToolUse hooks without waiting and only records their decision in its trace (agent-afk src/agent/tools/dispatcher.core-exec.ts), so rules.ts judged every edit and its block never reached the agent. It now runs on PreToolUse for edit_file and write_file, where AFK turns a block into an error result and never runs the tool, and the block message says the edit was not applied. In a live agent-afk 5.259.0 afk chat session, a console.log write under an AGENTS.md ban was blocked at 0.98 about half a second after the call, and the agent rewrote the file without console.log; the same prompt under the old wiring left the violating file in place and the agent reported no errors. An edit the hook blocks is no longer recorded for the Stop sweep, and the matcher no longer lists patch_apply, whose input the hook never read.
  • The rule hook and the Stop sweep stay silent when an event carries no session_id, instead of keying their state to unknown. agent-afk sent command hooks no session id, so every session shared one block counter and one edit log: a rule that had blocked a file twice in any session never blocked it again anywhere, and the sweep judged edits from other sessions and repositories. agent-afk sends the id from griffinwork40/agent-afk#2392.
  • The Stop sweep reports through additionalContext, which AFK prepends to the user's next prompt, because AFK shows a Stop block to the user as a notice and never to the agent. It judges only the edits made since its last completed judgment, keeping them for the next turn when the Jev call fails, reports a violation even when an uncertain match comes back in the same call (the uncertain branch used to return first and drop it), and gives its Jev call 4 s, inside the 5 s AFK allows a Stop handler (claude-jev-afk 0.6.0).
  • The adapter README says which hooks run only in the interactive REPL, that AFK drops every PreToolUse output except a block (so the subagent router's notes and uncertain rule matches reach no one yet), and that afk config env set refuses the key names, so the key goes into afk.env by hand.

0.27.0 - 2026-09-27

  • The AFK adapter installs from GitHub. Its hooks run the TypeScript source under node type stripping, so it needs no build step, and each release pushes a split of adapters/afk to the afk branch. Install with afk plugin install 0x7067/claude-jev claude-jev --ref <afk commit> and update with afk plugin update claude-jev --ref afk. A bare afk plugin update checks out the newest version tag, which is the Claude Code plugin.

0.26.0 - 2026-09-27

  • The subagent router sets a model unless something else already chose one: the call's model, CLAUDE_CODE_SUBAGENT_MODEL, a model: other than inherit in the agent's definition (project .claude/agents/ up from the working directory, ~/.claude/agents/, or an installed plugin's agents/), or a built-in with a fixed model (statusline-setup, claude-code-guide). Explore, Plan, claude, and general-purpose inherit the parent's model, capped at Opus for Explore (code.claude.com/docs/en/sub-agents), so they are routed. An agent whose definition the hook cannot find is left alone. Skipped spawns still get the brief check. On the 279 eval spawns of those inheriting types, the rule gets 145 matches, 45 too cheap, and 89 too dear, against 78, 44, and 157 under the 0.24.0 gate.
  • The AFK adapter's hooks never ran when installed as its README says. hooks/hooks.json pointed at ${CLAUDE_PLUGIN_ROOT}/adapters/afk/dist/, but AFK sets the plugin root to the installed adapters/afk directory, so every command failed with Cannot find module. Commands now use ${CLAUDE_PLUGIN_ROOT}/dist/. The README now says AFK needs "enablePluginHooks": true and agent-afk 5.257.3 or later, and that AFK passes no API key to hooks. The Jev client now reads TYPESAFE_API_KEY or OPENROUTER_API_KEY from AFK's own afk.env when AFK runs the hook (AFK_HOOK_EVENT is set) and the variable is not in the environment. Checked in a live agent-afk 5.257.3 session: the subagent hook called Jev with the key from afk.env.
  • The AFK adapter's subagent hook never fired. hooks/hooks.json matched /Agent/, but agent-afk names the tool agent; the matcher is now /^agent$/. The hook also skips the tier question when the spawn names an agent_type, as the Claude Code router now does for non-general-purpose spawns (claude-jev-afk 0.4.0).
  • The subagent router reads the user's tier text from CLAUDE.md only when it asks the tier question.

0.25.0 - 2026-09-27

  • The subagent router routes every spawn it can answer. It picks the cheapest tier where Jev puts at most a 0.10 chance on the task needing a stronger one, instead of requiring 0.75 confidence in one tier. Four tiers split the probability, so that gate held back 190 of 300 spawns, which then inherited an opus or fable parent. On those 300 spawns, labelled with the model the user named, the new rule gets 160 matches, 51 too-cheap picks, and 89 too-dear, against 92, 50, and 158 before (eval/subagent_eval.py, new). The AFK adapter's advisory tier uses the same rule. Tier text rewritten toward judgment-heavy work did worse on the same 300 (146, 46, 108), so the shipped text stays.
  • The prompt router no longer asks for a model tier or prints a tier systemMessage. In the live log, 197 of its 219 hints told an opus or fable session to switch down to haiku or sonnet. Switching mid-session costs the prompt cache and the context, and subagent routing already covers cheaper work. Jev picked opus 4 times in 659 prompts and fable never. The router log drops tier_hint and model_now, and the Stats row drops its tier line. The AFK adapter's intent bundle drops the question too (claude-jev-afk 0.3.0).
  • A subagent spawn that goes through on its second try with parts still missing, and also gets a routed model, now shows both notes. Before, the routing line replaced the missing-parts warning.

0.24.0 - 2026-09-26

  • The plugin runtime is TypeScript. The prompt router, subagent router, rule hook, and compactor run from src/*.ts under node type stripping, so node >= 22.18 is required and there is no build step; hooks/hooks.json and the /claude-jev compaction bridge call them directly (#20). scripts/compactor.py and scripts/subagent_router.py are gone, scripts/prompt_router.py stays as the routing eval's measurement reference, and eval/compact.ts replaces the Python compaction gate. Stats, jev.py status, and the rules eval stay in Python.

  • Setup guides for using Jev over MCP from OpenCode (#10) and the Vercel AI SDK (#11).

  • The TypeScript client's jev-calls.jsonl rows match the Python ones again. Each row names the hook that made the call instead of hook, carries the error string on failure, and takes its version from the nearest plugin manifest, so an AFK call no longer reports the Claude Code plugin's version. A 200 response with no answers now counts as a failed call instead of a success. The nested-instruction-file walk skips symlinked directories in both scripts/rules.py and the TypeScript rule parser, so a link into a mounted volume cannot stall the rule hook. eval/compact.ts resolves the repository root from eval/ instead of its parent, so the compaction gate can read src/compactor.ts.

  • The compaction gate fails a degraded run rather than reporting a better number. The compactor is fail-open, so a backend chunk that fails yields no answers, its blocks go unscored, and unscored rows are kept, which raises coverage; a wholly failed event was visible only because the planted probe count moved, and a run degraded in part showed up in neither. The gate now wraps the decision backend in a call and failure tally, which needs no change to compactor.ts and so leaves the selection signature and its cache alone, and any failure fails the gate since the floors above it are computed on blocks that were never judged. Distinct errors print with counts, so a 503 from a rate-limited endpoint is distinguishable from a 402 from an exhausted account. The first version of the check rode the floor table with a floor of 0, where 100% failures satisfied val >= floor and a run where every call failed printed ok; the failure test against a nonexistent model caught it, and it is now evaluated as a ceiling. Verified both directions: exit 2 with 1398 named HTTP 400s on the broken run, exit 0 at 0.0% over 1350 calls on a clean one.

  • A truncated block keeps its end. truncateBlock kept a 400-character head and an elision marker, so anything in the tail was lost unconditionally, and that is the shape of the real failure: a user says "also, do not touch the migration files" at the end of a long message, and a head-only cut drops the most recent instruction while keeping the routine output above it. It now keeps a 150-character tail as well, and the marker still reports the true elided count so the re-read pointer does not misstate what reading it back costs. The gate now tallies kept, cut, and dropped per planted kind and reports where the key sits in each cut block, which is what identified this: all five failures sat at 90 to 97% of the way through their block, within the last 25 to 60 characters. Over the full 38 real and 60 synthetic events, buried restatement survival goes 90.9% to 96.6% with cut 5 to 0, re-fetch verbatim coverage 85.4% to 86.1%, verbatim re-reads 649 to 654, for 3476 to 3523 tokens per event. The floor had been failing by one probe at 89.8% and passing by one probe at 90.9%, which is the measurement telling you it cannot distinguish a fix from a coin flip; a truncation-ordering change was tried first, ordering the budget cut by keep score rather than full score, and measured as an exact no-op, so it was reverted rather than kept for sounding right.

  • respan/span-01-lite clears this compaction gate and performs badly against the rules as they are currently written and tuned. Those are narrower than the first version of this entry claimed, and neither is a general statement about the model: the floors and probes were shaped around Jev, so clearing them shows span-01-lite is not disqualified here rather than that it is equivalent, and a planted probe tests literal key retention through selection and truncation, which is not a test of whether a resumed agent obeys a constraint. Measured on the same code, span-01-lite passes all three compaction floors: 81.3% verbatim coverage against Jev's 86.1%, and planted survival of 97.8% for a user constraint and 94.3% for a buried restatement against 97.8% and 96.6%, with every cut eliminated in both. Its coverage and verbatim re-reads still trail, 618 against 654, for 3591 tokens per event against 3523, though that gap is observed on this corpus rather than established: the runs are paired, only 15 of the 56 events that had re-reads are discordant, they come from 19 sessions, and a session-clustered interval on the advantage is wide enough to include zero. One run per model also cannot separate run-to-run model variation from a real difference. An earlier reading of this repo had span-01-lite missing both planted floors at 91.0% and 75.0%, and concluded from that that constraint discrimination was a property of the model rather than the algorithm. That was wrong, and the tail-retention fix above is what proved it: the planted keys sit at 90 to 97 per cent of the way through their blocks, and span-01-lite's lower keep scores put more of its kept blocks through the truncation path, so the same bug cost it more than it cost Jev. Measured before the fix it looked like a model ceiling; after it, the model is within 2.3 points on both planted floors. As the rule decision model and judge over the 303-record corpus it catches 1 of 29 hand-written violations where Jev catches 15 of 28, and blocks 0.0% of real edits at every ACT the sweep covers, 0.60 to 0.95, because its scores for rule violations sit almost entirely below 0.60. That is not the same as no threshold working, and an earlier version of this entry claimed it was. Swept further down, at ACT 0.30 span-01-lite catches 6 of 29 while blocking 33 of 247 real edits and 2 of 24 compliant near-misses, an operating point Jev dominates on all three axes at 0.80; at 0.15 it catches 20 of 29 but blocks 141 of 247 real edits and 12 of 24 near-misses. The defensible summary is that no threshold reproduces Jev's operating point, not that none works: the probabilities are compressed toward the bottom of the scale, and ACT is calibrated to a model that uses the range. Both arms are complete over the same 303 records with 0 errors each: Jev catches 16 of 29 and blocks 8.5% of real edits at ACT 0.80, span-01-lite catches 1 of 29 and blocks none, and its best row is still 1 of 29 at ACT 0.60. Jev was served from OpenRouter rather than api.typesafe.ai here, since the same jev-latest model is served by both under the same System One protocol and the TypeSafe account was out of credits; the answer cache keys on model, state, and questions rather than provider, so switching the billing path cost only the records that had failed.

  • The rule hook stops walking into mounts and the user Library. A record whose cwd is the home directory sent the nested-instruction-file walk across the whole home tree, and it wedged: every worker blocked in os.listdir at near-zero CPU with the answer cache frozen. A mount point is a different project on a different volume, and OrbStack's NFS share under the home directory does not answer at all, so one lstat short of reaching it the walk never returns; same_filesystem compares st_dev without following anything, so the mountpoint is recognised without touching what is mounted behind it. The macOS per-user Library is app containers and File Provider domains rather than source, is enormous, and a backed provider can block a listing for as long as the network takes, which is the same wedge by another route. Both were production hazards and not only eval ones: a session started in the home directory hung the rule hook on the same walk, and the hook's deadline turned that into a silent no-check instead of an error. Loading rules from the home directory now completes and returns 12.

  • The rules corpus loses 28 records and says why. Extraction pins each record to the repo HEAD at extract time, so a rewritten or force-pushed history invalidates every record pinned to it: pi-environment was retired into five fresh commits, and the objects behind its old head are in neither the local repo nor the remote, where git fetch does not recover them. Those 28 records were being judged at the live checkout, which answers a different question under the same name — the rule files as they stand today rather than as they stood when the corpus was built — and did so silently. The run now stops and names the repo, the commit, the cause, and the two honest options, and the corpus is repaired by dropping what can no longer be reproduced: 1541 records kept, the pre-repair file kept beside it. A re-extract repins the current history, and the question does not arise again until that history is rewritten too.

  • Compaction holds a read result instead of voting on it. isReadResult recognizes a tool_result whose call was a Read and keeps it at KEEP_THRESHOLD, because routine file reads are the largest block class in a long session and were out-competing the constraint, decision, and error rows a later turn actually re-reads. Measured over the 38 real compactions, verbatim coverage moved 84.5% to 91.8%, clear of the 69.6-84.5% band the same code produced across reruns, with no event regressing and three improving, at 2761 to 2925 tokens per event.

  • The compaction harness can afford to be run more than once. Summaries were held in memory, so every full rerun re-paid all sixty, and they are now cached to disk by cut hash next to the selection cache that was already there. Generation was spawnSync, which blocks the event loop, so the worker pool serialized behind one summary and a rerun needing fresh ones took over three hours; async spawn lets it pool, and the same sixty summaries went from 2 generated in seven minutes to 49. The generator is pi rather than claude, which also drops a CLI that is weekly rate limited and returned nothing for the whole synthetic arm. --explain attributes each uncovered re-read to its cause, and it is what pointed at this read rule rather than at the keep threshold: the dropped blocks scored 0.04 to 0.16, nowhere a threshold change would have reached.

  • The decision model and the judge are swappable behind one seam. DecisionBackend is a name plus ask(state, questions, timeout) returning typed Answers; production keeps System One as the default, and the eval injects another implementation by parameter. The backend name is folded into the selection signature, which is the part that would otherwise bite — without it a second model's run reads the first model's cached decisions back and measures nothing. Two implementations share one typedAsk over the identical {state, model, questions} to {answers} contract, so the swap is a URL and a model id rather than a second code path: systemone, and decisions for OpenRouter's /api/alpha/decisions, which serves respan/span-01-lite, a System One model behind a different URL. Spec syntax is backend:model, or a bare model id meaning System One, and an unknown backend id raises rather than falling back quietly on both sides. scripts/jev.py takes the same split so eval/rules_eval.py --judge-model can judge with the same model that decides; its answer cache already keyed on model and rules.py passes timeout= by keyword, so the judge model is injected as a default without disturbing either. Every compaction keep question is a noul, which is all Respan accepts. Over the full population, 38 real events and 60 synthetic, with the planted probes at their whole n=89 and n=88, span-01-lite as decision model gives 84.2% verbatim coverage against Jev's 85.9%, but planted user constraint survival falls to 91.0% from 97.8% and buried restatement to 75.0% from 89.8%, missing both floors, and it spends more to get there at 3007 tokens per event against about 2900. Coverage is judge-independent; the discrimination between a user constraint and a routine row is not, and that is what the keep floor and the planted probes exist to protect. Read on the 38 real events alone the picture inverts to 92.3% coverage and only 29 and 28 probes, which is the whole argument for the synthetic arm.

  • adapters/afk/ ships a second, TypeScript implementation of the hooks for the AFK host (#18), with its own plugin manifest (claude-jev-afk 0.2.0) and install path (npm run build, copy into ~/.afk/plugins). It shares the Claude Code plugin's judged surface verbatim — the question bundles, ACT 0.80 / FLAG 0.50, the twice-per-session block downgrade, fail-open on everything — and adds one hook the Claude Code plugin lacks: a SessionStart rule digest that reads the instruction files and injects a structured summary with no API call. It lacks what AFK cannot express, and adapters/afk/README.md says so: no compaction (AFK hooks expose no transcript), advisory-only subagent routing (no updatedInput, so a bad brief is annotated, never denied), prompt routing on the prompt alone (no previous turn), no ast-grep enrichment, and no shell-write tracking (hooks.json matches only edit_file|write_file|patch_apply). README.md's install section claimed "Same hooks, TypeScript" — false in both directions, now corrected to state the delta. The recorded plan: this adapter's shared core becomes the single implementation, brought up to Claude Code's enforcement level, before scripts/ is deleted.

  • control-jev pins CLAUDE_CONFIG_DIR=$VERIFY_HOME/.claude for every child process. jev.config_dir() prefers that variable over ~/.claude, so a home isolated by HOME alone still leaked to whatever config directory the calling shell exported — which this repo's own instructions tell an agent to export for hand-fed hook events. Every log, cache, and rule-count assertion in the verification feature map silently described the wrong directory; the harness fix makes them true, and a re-drive with a decoy exported confirmed writes stay in the verify home.

  • README.md's rule-hook table gains a v0.23.0 column beside v0.21.0, from re-scoring the same 247 judged real edits at current source on cached answers: 23 blocked (9.3%) against 22 (8.9%), 9 flagged only against 8, hand-written violations 17/29, compliant near-misses 0/24, median 8 rule questions per edit — unchanged. Latency is published as not measured: a cached re-score reports a 0.01s median against the 0.80s a live run costs, so the cell would measure the cache rather than the hook.

  • The feature map now carries a live proof of the settings pane and the function-hook bridge, run against the installed copy at 0.23.0. Toggling Rule checks writes pluginConfigs["claude-jev@claude-jev"].options.rules and the effect was measured as an A/B on one identical Write: zero "kind":"rules" log rows with it false, two with it true. Toggling Compaction off makes /compact log nothing from the plugin, because register.ts returns next(e) before the $.ui.log that reports a fallback — so plugin silence is the off signal and built-in summary runs is a failure, a distinction that would otherwise let a verifier read a broken bridge as a disabled toggle. Also recorded: claude plugin update claude-jev@claude-jev brings the installed copy level with main (it was four versions behind here, which is why --plugin-dir was needed at all), and the @inline vs @claude-jev module-load ids differ by load path.

  • eval/rules_eval.py run takes JEV_RULES_CACHE to point the answer cache at another path. --out never bypassed it — it moves the prediction file while eval/observed/rules_cache.jsonl is still read — so a re-scored run reported latency median 0.01s, which is the speed of a dict lookup, not of the hook. Judged live with the override (326 calls, 0 errors, pinned ast-grep present) the same 247 edits return the same 23 blocks and 9 flags at a 0.40s median, and that is the figure now in the README.md table against v0.21.0's 0.80s. The live pass costs ~40s at four workers; the cached one is free.

  • eval/sweep.py and eval/planted.py run again. Both inserted only their own directory before import compactor, so python3 eval/sweep.py --help died with ModuleNotFoundError: No module named 'compactor' while sweep.py's own docstring prescribes python3 eval/sweep.py --synth 60 and README.md publishes survival figures attributed to planted.py. They now insert ../scripts like compare.py, replay.py, and rules_eval.py do. The pair only worked when reached through compare.py, which puts scripts/ on the path before importing planted — so the compaction gate passed while the diagnostics that read a change were dead, and python3 -m compileall -q scripts eval could not catch it because it compiles without importing.

  • The feature map under .cursor/skills/verify-claude-jev/ was reconciled against source and one live run per feature. New features/comparators.md covers the pinned ast-grep: which and its path / cached / none states, the degraded judgment with no binary, and the detached fetch. Corrected: stats.py bails on an empty router log alone, --log redirects only that log, --days 0 means no cutoff, and report reads cached predictions so re-reading a published number costs nothing while run bills. Softened two claims the repo cannot substantiate — the settings.json path behind a pane toggle and two debug-log strings that are Claude Code's own output, not the plugin's.

0.23.0 - 2026-09-25

  • The scorer takes shell writes from Claude Code's own diff, not a regex. A Bash result's toolUseResult carries bashEditDiff, the file-state diff Claude Code took around the command, naming the git-visible project files it changed — the same notion of a write the live hook in rules.py already used, so the scorer and the hook now agree. observed.walk_session joins each Bash call to its result and hands summarize the changed paths; summarize treats that verdict as authoritative and lets it override the pattern in both directions, so a command that only wrote /tmp or a gitignored path no longer makes the turn an edit. Claude Code 2.1.274 is where the field first appears in the transcript corpus, so BASH_DIFF_VERSION gates it: below that, a missing field means "unknown" rather than "wrote nothing", and the patterns still decide.
  • Measured against that diff as ground truth — 6,143 real Bash calls across the 145 sessions that record it — the command-text patterns score precision 0.25 and recall 0.84. Three quarters of the writes they claimed were a scratch file outside the project, a gitignored path, or a quoted string that happens to contain >, such as curl -w '%{redirect_url}'. The misses were writes the text cannot name at all: docker compose run … pnpm add, which edits package.json through a bind mount.
  • The fallback patterns also lost two misses of their own, the same pair compact-adviser fixed in its PRs #52 and #53: sed --in-place, the GNU long form, which takes its backup suffix only attached with = and so is never ambiguous the way bare -i is; and a write, ops, or read command anchored after then, else, do, or {, or at the start of a line in a multi-line command, which the old (^|[;&|]\s*) position could not reach. case $x in a) mkdir -p out ;; esac still misses — a ) prefix bought nothing on the corpus and would match subshells. Over the ground-truth set these move recall 0.843 to 0.846; they matter as correctness of the fallback, not as a measurement change.
  • Labels moved on 262 of 4,857 rows (5.4%): fix 857 to 699, lookup 1634 to 1707, feature 793 to 840, ops 239 to 282, chat 1284 to 1297. Bash-changed files now count toward n_files, so a turn that wrote three files through the shell reaches the substantiality band the way a turn that wrote three through Edit already did. Re-scored over the whole current corpus with the same cached predictions (eval/replay.py run --variant v7_no_unclear, 2,449 of 2,464 rows cached, 15 new calls, 0 errors), holding the row set fixed to isolate the scorer: v7 accuracy 32.9% to 34.0% and lift +7.6 to +8.6, hinted-only accuracy 49.9% to 48.4% and lift +7.2 to +5.8, harmful hints 8 either way. Against the 120 hand labels in eval/audit_labels.json it is a wash — 52.5% to 51.7%, six rows changed, two toward the human label and three away.
  • README.md's router table is re-run on the current corpus rather than carried over: 2,464 prompts instead of 1,613, and every figure now reproduces from the commands printed beside it. The published coarse-taxonomy pair did not — 58.8% and 64.7% only come back at --floor 0.5, off the shipped 0.75 floor the rest of the table uses — so it is replaced by v5_three_way at the shipped floor, 65.0% against a constant 63.6%.

0.22.0 - 2026-09-25

  • Compaction asks a fifth rerunnable check: output a rerun would print again — a listing, a passing check's log, build or install output, warnings a rebuild repeats. A rerunnable answer at KEEP_THRESHOLD silences error's verbatim claim, so a regenerable dump keeps as a HEAD_CHARS head with the re-run pointer instead of KEEP_CHARS whole; an exact error no command would repeat, and any user constraint, still hold verbatim. The four keep reasons are unchanged, so pure noise still drops and its ref stays in the row list. Replaces rule-based scrubbing of verbose output with the same judgment Jev already makes per block. Gate behind it (eval/compare.py compact --synth 60 on this source, 60 events, live re-judgment): re-fetch verbatim coverage 76.8% (floor 70%, n=760 reads — inside the standing 76-82% band), planted user constraint 100% (n=87), buried restatement 98.8% (n=82).

  • jev.ask accepts a deadline (monotonic timestamp) alongside timeout, and retries once when a network failure returns in under a second — a refused connection costs milliseconds, a timeout costs the budget.

  • The rules hook now works inside its 10s hooks.json budget instead of hoping two sequential Jev calls fit: main stamps a 9s deadline, every ask gets the time left, and the escalation pass is skipped when under 3s remain — the first-pass verdict still lands instead of the hook being killed with no output.

  • Rule classification runs its per-chunk requests in parallel and caches by chunk content hash instead of whole-file hash. A large instruction file no longer serializes into a hook timeout, a killed classification keeps the chunks that landed, and editing one paragraph re-judges only its chunk.

  • Compaction chunks ask with a 4s timeout instead of the 8s default, so one slow request can't stretch a ~1s compaction; a timed-out chunk's blocks are kept unscored, as before.

  • Compaction asks four checks per block instead of five: the artifact check ("could this be re-fetched?") fed no decision — verdicts reads only the keep and verbatim scores — so it cost a fifth of every compaction request for a diagnostic nobody consumed. The /compact <text> directive moved from a suffix on every question to one line in the shared state header, where session_context already names it.

  • Blocks just older than the 150-row judgment window are no longer dropped unseen: the 150 before them get the constraint check alone, and any that score a user requirement are kept verbatim ahead of the window's rows. A constraint stated early in a long session now reaches Jev instead of falling off the window silently.

  • Ruff formats and lints the Python sources. ruff.toml selects pycodestyle errors (E4, E7, E9), Pyflakes (F), and import sorting (I), at a line length of 100, and skips eval/observed and eval/authored. It is tool config, not a dependency: hooks still import only the standard library. ruff format --check and ruff check run in release preflight and in GitHub Actions, next to compileall and the two check_no_*.py scripts. A machine with only python3 can still claim those stdlib checks.

  • eval/data/ is now eval/observed/ and eval/private/ is now eval/authored/ — the names say what they hold: machine-extracted transcripts and predictions vs hand-written cases. Both stay gitignored.

0.21.0 - 2026-09-24

  • The rules eval now measures the AskUserQuestion answers the way the hook sends them. Its corpora carried only the typed request, so no rerun could see the v0.20.0 answers feature: extract records each edit's answers since its request, and run composes the request through the same request_with_answers the hook calls; the rules decision log rows gained a user_answers count, so live outcomes can split by whether the user answered questions. Extraction collects answers unconditionally, like the hook: an AskUserQuestion's result line normally carries no tool name, so the substring gate this replaces saw 5 answers-bearing edits where the corpus has 140 of 1,569 — enough that a normal 250-sample now measures the feature (~22 records). Records without answers compose to the old request text, so their cache entries still hit. The hook itself gains a fix: a transcript that could not be opened made last_user_prompt return a 2-tuple, crashing the handler before any rule was asked. scripts/check_no_stubs.py joins the Verify contract: .ask = assignments are banned everywhere except the cached_ask wiring shapes in cmd_run, so a probe cannot fake the client and present the result as verification. Live evidence through the real client: a313bcf9#932 ("rspec check in CI failed") reached the judge with its picked option and none acted, its re-judgment hitting at 0.0s the cache key only its answers-carrying state could produce; a synthetic result line with no tool name rides in the request exactly as the hook composes it. Records whose directory stopped being a git checkout drop their sha and join the live-checkout bucket the report already counts, instead of killing the run. The rerun on the new corpus: 300/303 judged, 0 errors; 22/247 real-edit blocks (8.9%), 17 of them symphony's own code-comments-are-banned-in rule (0.82–0.91) on accepted edits from before this change — none of the blocked rows carry answers; 17/29 violations blocked, 0/24 compliant; 22/247 edits carried AskUserQuestion answers, the line this change adds.

0.20.0 - 2026-09-24

  • The rules hook counts the user's AskUserQuestion answers as part of the request. Jev saw only typed prompts, so an approval given by picking an option never reached it. In one session that blocked three approved edits: two to an axe spec (0.83, 0.85) and one to feature_flag_spec.rb (0.91), each under "do not weaken tests", after the user had picked the option that said the test expectation would change. The request now lists each answer since the latest prompt, as the question, the picked label and that option's description, capped at MAX_ANSWER_CHARS. Subagent results are skipped. The rules eval has not been rerun on this change.

0.19.2 - 2026-09-24

  • The rules hook records what a Bash command changes. A PreToolUse hook snapshots the git working tree, untracked files included, in a scratch index. The matching PostToolUse diffs it and records each changed file as a hunk, so the Stop check judges shell writes like edits. A turn whose snapshot fails, such as outside a git repo, still tells Jev its diff is partial. Snapshots took under 0.7s on the five slowest repos checked.
  • A failed Bash command's writes are recorded too, through a PostToolUseFailure hook. A turn with no recorded hunks but a partial diff still gets the Stop check, and so does a Bash write to a git-ignored path, which marks the diff partial.

0.19.1 - 2026-09-24

  • The rules Stop check judges only the edits made since the latest user prompt, and skips a turn with none. It had pooled every edit since the session began and judged them against the newest prompt. In session ccb4b920, that blocked twice on approved .zprofile/.zshenv edits from an earlier turn (0.91, 0.92) and spent the session's block budget. Hunks are keyed by the transcript uuid of the prompt they were made under. The request and the repair message now name this turn's files. See docs/stop-hook-false-positive.md.

  • When a turn with recorded edits also ran Bash, the Stop request tells Jev the diff is partial, since shell writes never reach the recorded hunks.

  • The rules hook reads and writes its session state under a file lock. Twenty parallel writers kept 20 of 20 hunks; without the lock they kept 5. It logs a swallowed exception as a rules-error row, which stats never scores, so a missing decision row can be explained.

  • python3 eval/rules_eval.py turns splits live Stop-hook checks by whether the turn made an Edit/Write of its own. It is offline and free. On the live log before the fix: 216 checks. 82 judged the turn's own edits (0 blocks, 30 flags). 107 came after Bash-only turns (2 blocks, 23 flags) and 27 after turns with no changes (0 blocks, 6 flags). Both blocks were the ccb4b920 false positives. Under the turn scoping, those 134 checks and 29 flags no longer run.

  • Add a bounded rule-prompt comparison tool and document a reviewed live false positive, evidence limits, and candidate questions. Shipped rule prompts and thresholds are unchanged.

  • Stats breaks down failed calls by HTTP status or timeout for each caller, shows recent call health and last failure/success timestamps, and explains that rule outcomes are edit heuristics rather than verified repairs.

  • scripts/stats.py lists ~/.claude/projects once instead of globbing it for every session. On this machine the report went from 9.1s to 0.9s cold and from 2.15s to 0.59s warm, with byte-identical output on a frozen copy of the logs.

  • The Status row in /claude-jev shows the status the pane already loaded as soon as it is pressed, then refreshes it from jev.py status.

0.19.0 - 2026-09-23

  • The prompt router no longer shows an ops hint. Live, 18 of 60 were right; on the 120 hand-labeled prompts humans agreed with 6 of 18. With feature already silent, replaying 1,613 prompts through the shipped rule goes from 31.1% accuracy and −0.9 lift to 50.2% and +18.4, at 13.8% coverage. The cleaned live log goes from 49.5% (111 hints) to 72.5% (51 hints). An intent is silent by having no entry in GUIDANCE; SILENT_INTENTS is gone. eval/variants.py gains v9_hinted_only, which scores v7 answers only where the hook shows a hint.

  • The feature silence (commit 0f67b21) cited 67 predictions against 6 observed; most of those predictions were compaction requests. It stays silent on current evidence: 8 of 28 right in the cleaned live log.

  • scripts/stats.py scores only prompt-router entries; rule checks and subagent decisions were counted as suppressed prompts (1,040 reported, 351 real). It prints how many entries were Claude Code's own requests, reports the confidence floor as held-back count and hit rate (35 held back, 15 of them right, against 72.5% for the hints the router now shows) instead of counting chat predictions that never show, and calls a block with no transcript unscorable instead of unknown.

  • eval/replay.py report and compare default --floor to prompt_router.MIN_CONFIDENCE (0.75) instead of 0.55, so the documented command reproduces the README row.

  • scripts/check_no_comments.py scans only files git tracks or would track. It had been failing on the gitignored eval/data/at/ checkouts.

  • AGENTS.md: run hand-fed hook events with CLAUDE_CONFIG_DIR set to a temp directory. Test events with session ids like t1 had landed in the live log, and were the 15 "unknown" rule outcomes.

  • scripts/stats.py no longer scores Claude Code's own pre-compact request ("Your task is to create a detailed summary…") as a user prompt. Older plugin versions logged 73 of them with a hint, and the session that followed used no tools, so the report counted them as wrong feature hints against an inflated always-chat baseline. On the same log (1,316 decisions), the report moved from 72/195 agreed (36.9%) against 48.7% for always chat, to 58/122 (47.5%) against 26.2% for always lookup. The weak class is now ops: 60 hints, 20 observed.

  • The prefix lives in SYNTHETIC in scripts/observed.py, so the router, stats, and eval/replay.py share one filter; COMPACT_PROMPT in scripts/prompt_router.py is gone. eval/replay.py report --variant v7_no_unclear is unchanged (1,613 scored, 33.7% accuracy, +4.4 lift): none of its predictions were compaction prompts. Derived labels still agree with eval/audit_labels.json on 63/120.

0.18.0 - 2026-09-23

  • OpenRouter as a second Jev provider. Set OPENROUTER_API_KEY, or TYPESAFE_API_KEY to an OpenRouter key (sk-or-...), and every call goes to https://openrouter.ai/api/v1/systemone, which takes the same request, model IDs, and answer shape. TYPESAFE_API_KEY is read first. A TypeSafe key still calls api.typesafe.ai. The Provider row in /claude-jev (provider in /config: auto, typesafe, openrouter) pins one provider and reads only its variable, so OpenRouter works with both variables set. Live through OpenRouter: a noul answered 0.99, and the prompt router answered in 472 ms. Each jev-calls.jsonl line records its provider.
  • /claude-jev settings pane (needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1): save an API key for all sessions, see which provider it calls, turn the prompt router, subagent router, rule checks, and compaction on or off, and see the last Jev call. The on/off rows also appear in /config. A key in the environment wins over a saved key. jev.py status prints what the pane shows. The pane's Stats row shows the scripts/stats.py report.
  • Removed the /claude-jev:jev and /claude-jev:stats skills. No session outside this repository ever called the first, and it could not see a key saved in the pane. Stats moved into the pane; python3 scripts/stats.py still prints the full report.
  • scripts/compactor.py changed only its docstring; hooks/register.ts now passes the saved key and provider to it. Compaction gate, run through OpenRouter: re-fetch verbatim coverage 79.8% (floor 70%, n=678); planted user constraint survival 100.0% (floor 95%, n=72); planted buried restatement survival 97.1% (floor 90%, n=69).

0.17.0 - 2026-09-22

  • Subagent briefs are checked before the spawn: a brief that changes files but omits paths, acceptance criteria, a verification command, or a commit policy is denied once per session with the missing parts listed, then goes through with a systemMessage. Read-only briefs are exempt. Live, a thin rename brief scored 0.06–0.16 on all four parts and a complete one 0.97+.
  • Subagent tier criteria rewritten with sharper boundaries and a fourth fable option for adversarial review and cross-system debugging with conflicting evidence, never implementation. On seven hand-written briefs the live-vs-eval bug moved from opus 0.44 to 0.77, above the routing gate. The prompt router's advisory tier hint recognizes fable model names. Tier text can be overridden by a ## Delegating to sub-agents section with - tier: text bullets in the user's global CLAUDE.md.
  • Every file under the user's config directory (logs, caches, ast-grep binary, global CLAUDE.md) resolves through CLAUDE_CONFIG_DIR when set, else ~/.claude.
  • scripts/compactor.py changed only its log path. Compaction gate: re-fetch verbatim coverage 76.5% (floor 70%, n=727); planted user constraint survival 100.0% (floor 95%, n=72); planted buried restatement survival 98.6% (floor 90%, n=69).

0.16.1 - 2026-09-22

  • The release skill and script moved out of the plugin into .claude/skills/release; /claude-jev:release no longer exists for plugin users.

0.16.0 - 2026-09-22

  • Compaction judges each block with five concrete checks instead of two aggregate questions. The checks: user constraint, decision with reason, exact error, open work, re-fetchable output; keep and verbatim scores derive from them in code. A constraint planted mid-session survives 100% (was 77%); a restatement buried in a later reply survives 98% (was 35%). Scores in the 0.35–0.65 band fall from 54% to 21%.
  • eval/compare.py compact gates both compaction goals at once: re-fetch verbatim coverage (floor 70%) and planted-constraint survival (floors 95% and 90%), exiting 2 below either. eval/sweep.py and eval/planted.py added as diagnostics.
  • Rule hook escalates the uncertain band with one focused second call and adds ast-grep comparators so the fact outside the hunk reaches the judgment.
  • Rules are structured with a local relevance gate and violation-framed questions; every Jev call and every compaction row is logged for /claude-jev:stats.
  • Code comments are banned in scripts/, eval/, and hooks/, enforced by scripts/check_no_comments.py.
  • Docs: docs/prompt-craft.md records measured effects of question wording; Pstack verify skill under .cursor/skills/verify-claude-jev.
  • New /claude-jev:release skill and scripts/release.py.

0.12.0 - 2026-09-21

  • Compaction through the experimental session.compact function hook: Jev's kept rows replace the built-in summary.