0x7067/claude-jev
Hands the small judgments in a session to TypeSafe's Jev, a System One model that returns typed answers instead of generating text: what kind of request each prompt is, which model tier a subagent needs, whether an edit breaks a rule already written in your CLAUDE.md or AGENTS.md, and which transcript blocks survive compaction. Kept blocks carry over verbatim — no LLM writes a summary.
Changelog
All notable changes to claude-jev. Format follows Keep a Changelog; versions follow the plugin manifest. .claude/skills/release/release.py turns the [Unreleased] section into the next release.
[Unreleased]
- The AFK prompt router and subagent router are off unless their
promptRouterorsubagentRouteroption istrue; both stay registered and exit before calling Jev. Each cost one Jev call per prompt or spawn, AFK drops the subagent router's advice (griffinwork40/agent-afk#2778), and the prompt router does not log on AFK, so neither was measured there.adapters/afk/.claude-plugin/plugin.jsondeclares both options withdefault: false, andenabled()takes the value to use when an option is unset; the Claude Code hooks still treat unset as on. adapters/afk/README.md: updated Known-gaps and Setup sections against agent-afk 5.286.1 source. BlockingadditionalContextonPreToolUsenow reaches the model (agent-afk 5.121.0); non-blockingadditionalContextis still dropped (#2778).transcript_pathis now sent in hook stdin (agent-afk 5.276.14, #2647); null on daemon/chat/web.PreCompacthook event exists and can block compaction but cannot select which blocks to keep (agent-afk 5.10.0).PreToolUsehooks fire inside subagent forks (agent-afk 5.121.0);SessionStartinjectContext does not.pluginHookEnv(agent-afk 5.276.21, #2700) is now documented as the supported route for forwardingTYPESAFE_API_KEY/OPENROUTER_API_KEY.CLAUDE_CONFIG_DIRandCLAUDE_PLUGIN_OPTION_*are still not passed to hooks as of 5.286.1 (#2373 open). Issue links added for #2816 (per-hook disable) and #2817 (UserPromptSubmit/Stop REPL-only).- The Stop sweep shares one 4.5s budget across rule classification and the Jev judgment call, so a cold-cache
loadRulesis covered by the same deadline and the whole sweep finishes or fails open with a loggedrules-errorinside AFK's 5s handler window. PreviouslyloadRulespassed notimeoutMs, falling through to the 8s client default, which could silently exhaust the budget before Jev was ever called. - An AFK rule check that fails while classifying instruction files logs
rules-error. A failure while reading those files, before Jev is asked, staysrules-skipwith reasonrules-unreadable. - Two AFK rules that share an id keep separate probabilities, violation ids, and block ids in the log row.
stats.pycounts a laterrules-skipof a warned file as not rechecked, and takes the score from the next check that asked Jev. AFK rows stay out of Rule outcomes and its flagged-only list. They still join Rule calibration.- The AFK rule hook judges
patch_apply. Edits made through it used to land unjudged and never reached the Stop sweep. agent-afk hands a plugin hooktool_name: "patch_apply"with the raw{ changes, dry_run }input, and tests a/…/matcher against that name alone, so the matcher now lists it. Each file in the patch is judged in its own Jev call, in parallel, against the rules scoped to that file, and each check writes its own log row. The calls share one 12 s budget inside the hook's 15 s timeout. The patch is atomic, so one file at >= 0.80 blocks the whole call and none of its files are recorded for the Stop sweep. A clean or flagged patch records every file. Adry_runcall writes nothing, so it is not judged or recorded.
0.29.2 - 2026-10-02
package-lock.jsonrecords anintegrityhash for the five@earendil-works/*0.87.1 packages nested underpi-coding-agent. Without a hash, Claude Code's plugin installer skipped them and warned on every install and update.
0.29.1 - 2026-10-02
src/rules.tsno longer drops a Bash-written file when the rewrite lands in the same wall-clock second with the same size. The hook diffswrite-treesnapshots taken on a copy of the real index; the copy carried a fresh mtime, which cleared git's racily-clean guard, so the new stat was trusted and both snapshots hashed the old content — the snapshot was consumed withfiles: [],partial: false, and no error. The copy now keeps the source index's mtime, so the scratch index makes the same racy call the real index would. About one write in ten in a fast edit loop hit this; the file's change was never recorded.jev.py's missing-key error names every provider variable when no provider is pinned (set TYPESAFE_API_KEY or OPENROUTER_API_KEY) instead of always namingTYPESAFE_API_KEY;missing_key_messagewas built for both butaskpassed the already-defaulted provider, so the unpinned branch never ran. The AFK client's missing-key error now names only the pinned provider's variable instead of always listing both (claude-jev-afk0.6.2).src/session-start.tsis deleted. It was copied intosrc/by the TypeScript runtime migration but nothing registers a SessionStart hook in the Claude Code plugin — the rule digest is an AFK-only feature andadapters/afk/src/session-start.tsis its implementation.- An edit rule that has already blocked its file twice no longer calls the next act-band hit uncertain. The notice says that rule was already raised. A flag-band hit stays an uncertainty notice. A Stop hit held back by the session's two turn blocks, or while the agent is still finishing, says that, and does not claim the rule was cited. On AFK, the Stop sweep says the session already used its two turn blocks (
claude-jev-afk0.6.1), and that text still rides in the next prompt. - Two rules that slug to the same id keep separate probabilities. Keys are assigned on the full in-scope list and kept for the escalation call, so a second pass cannot write one rule's score onto the other. A rule the subject filter skips cannot inherit the other rule's score.
0.28.0 - 2026-10-01
- Pi compaction lives in
adapters/pi/.pi install git:github.com/0x7067/claude-jevloadsadapters/pi/jev.ts, which handlessession_before_compactand/jev. The hook callsselectBlocksinsrc/compact/strategy.ts. Block shaping,toolCallIdpairing, bash and summary role labels, the<read-files>index, and logs underPI_CODING_AGENT_DIRstay in the adapter. Kept text is fit to 14,000 characters. Paths from dropped or truncated calls may add up to 2,000 characters. Pi's own read paths and<modified-files>sit outside that reservation. Claude's compaction budget stays 16,000 characters. A missing key returns nothing, so Pi's own summary runs.0x7067/pi-jevis unchanged. - The rule hook skips files outside the project again, as the README says it does. The TypeScript port's
relativeToswapped an out-of-project path for its absolute form, soisOutsidenever matched and scratch files under/tmpwere judged against the project's rules. In the live log from the TypeScript runtime (2026-09-27 to 2026-10-01), 36 of 194 edit judgments and 4 of 9 edit blocks were on files outside the session's working directory, 33 of those judgments under/tmp. The working directory and the edited file's directory are both resolved through symlinks before they are compared, so an edit or Bash write in a project reached through a symlink such as/tmpstill counts as inside it and is recorded by its project-relative path, and a Bash command that deletes the working directory still has its changes recorded. - AFK installs the adapter from this repository's marketplace as
claude-jev:claude-jev-afk(afk marketplace install 0x7067/claude-jev), which checks out the newest release tag. It needs agent-afk 5.271.1 or later: earlier versions list the plugin as installed but load the root Claude Code plugin's hooks instead of the adapter's (griffinwork40/agent-afk#2456). In a throwawayAFK_HOMEon agent-afk 5.282.0, with the marketplace as a real directory, anafk chatturn ran the adapter'sSessionStarthook and not the root plugin's./releaseno longer pushes a split ofadapters/afkto theafkbranch, so that branch no longer moves. An install made with--ref afkshould runafk plugin remove claude-jevand reinstall from the marketplace. The entry also shows in Claude Code's/pluginlist for this marketplace, with a description that says it is AFK only. selectBlocksandfitKepttake an optional character budget that defaults to the existing 16,000, and a block may name the earlier block a tool result answers. Claude's rows bridge sets neither, so its budget stays 16,000./releasebumps rootplugin.jsonalong with.claude-plugin/plugin.jsonand.claude-plugin/marketplace.json, and stops when those two plugin versions disagree.- A Pi tool result names its tool, and only a
readresult is held, including after a text preamble and on a turn that also ran another tool. An unlabeled Claude[tool_result]is held when the single tool call it answers is a Read, in any capitalization; a call that ran several tools holds only the result lined up with that Read. The Pi compact log leaves out per-block rows, so/jevcan still read the last line. Pairing a kept result does not copy a block the rescue window already kept. The 16,000-character Pi figure is the selected body plus the 2,000-character reservation for paths from dropped or truncated calls. The digest header,---[jev:…]---lines, Pi's own read paths, and<modified-files>sit outside it. - The summary Pi shows the model includes one
<read-files>list and Pi's<modified-files>. The read list merges Pi's own reads with the capped paths of dropped or truncated calls, skips a path that was also modified while that capped list is filled, and does not repeat a path. A[tool_result]with noneeds, including a named[tool_result read], walks back past earlier results to the tool call, including when that call starts with other text, so a held Read keeps the path arguments. The next compaction strips those file lists only from the end of the summary. The read and modified lists follow Pi 0.87.1 inadapters/pi/file-ops.ts. - The repository root carries an Agent Plugins 1.0.0
plugin.json(nameclaude-jev, the same version as.claude-plugin/plugin.json, emptyextensions). Claude Code still loads.claude-plugin/plugin.jsonandhooks/. - Compaction's checks, thresholds, and
selectBlockslive insrc/compact/strategy.ts.src/compactor.tsonly maps Claude's compact rows on and off stdin.eval/compact.tsimports the strategy module, and its selection cache key covers that module. - The AFK adapter's rule hook now blocks edits. It never could on AFK before. AFK dispatches
PostToolUsehooks without waiting and only records their decision in its trace (agent-afksrc/agent/tools/dispatcher.core-exec.ts), sorules.tsjudged every edit and its block never reached the agent. It now runs onPreToolUseforedit_fileandwrite_file, where AFK turns a block into an error result and never runs the tool, and the block message says the edit was not applied. In a live agent-afk 5.259.0afk chatsession, aconsole.logwrite under anAGENTS.mdban was blocked at 0.98 about half a second after the call, and the agent rewrote the file withoutconsole.log; the same prompt under the old wiring left the violating file in place and the agent reported no errors. An edit the hook blocks is no longer recorded for the Stop sweep, and the matcher no longer listspatch_apply, whose input the hook never read. - The rule hook and the Stop sweep stay silent when an event carries no
session_id, instead of keying their state tounknown. agent-afk sent command hooks no session id, so every session shared one block counter and one edit log: a rule that had blocked a file twice in any session never blocked it again anywhere, and the sweep judged edits from other sessions and repositories. agent-afk sends the id from griffinwork40/agent-afk#2392. - The Stop sweep reports through
additionalContext, which AFK prepends to the user's next prompt, because AFK shows a Stop block to the user as a notice and never to the agent. It judges only the edits made since its last completed judgment, keeping them for the next turn when the Jev call fails, reports a violation even when an uncertain match comes back in the same call (the uncertain branch used to return first and drop it), and gives its Jev call 4 s, inside the 5 s AFK allows a Stop handler (claude-jev-afk0.6.0). - The adapter README says which hooks run only in the interactive REPL, that AFK drops every
PreToolUseoutput except a block (so the subagent router's notes and uncertain rule matches reach no one yet), and thatafk config env setrefuses the key names, so the key goes intoafk.envby hand.
0.27.0 - 2026-09-27
- The AFK adapter installs from GitHub. Its hooks run the TypeScript source under node type stripping, so it needs no build step, and each release pushes a split of
adapters/afkto theafkbranch. Install withafk plugin install 0x7067/claude-jev claude-jev --ref <afk commit>and update withafk plugin update claude-jev --ref afk. A bareafk plugin updatechecks out the newest version tag, which is the Claude Code plugin.
0.26.0 - 2026-09-27
- The subagent router sets a model unless something else already chose one: the call's
model,CLAUDE_CODE_SUBAGENT_MODEL, amodel:other thaninheritin the agent's definition (project.claude/agents/up from the working directory,~/.claude/agents/, or an installed plugin'sagents/), or a built-in with a fixed model (statusline-setup,claude-code-guide).Explore,Plan,claude, andgeneral-purposeinherit the parent's model, capped at Opus forExplore(code.claude.com/docs/en/sub-agents), so they are routed. An agent whose definition the hook cannot find is left alone. Skipped spawns still get the brief check. On the 279 eval spawns of those inheriting types, the rule gets 145 matches, 45 too cheap, and 89 too dear, against 78, 44, and 157 under the 0.24.0 gate. - The AFK adapter's hooks never ran when installed as its README says.
hooks/hooks.jsonpointed at${CLAUDE_PLUGIN_ROOT}/adapters/afk/dist/, but AFK sets the plugin root to the installedadapters/afkdirectory, so every command failed withCannot find module. Commands now use${CLAUDE_PLUGIN_ROOT}/dist/. The README now says AFK needs"enablePluginHooks": trueand agent-afk 5.257.3 or later, and that AFK passes no API key to hooks. The Jev client now readsTYPESAFE_API_KEYorOPENROUTER_API_KEYfrom AFK's ownafk.envwhen AFK runs the hook (AFK_HOOK_EVENTis set) and the variable is not in the environment. Checked in a live agent-afk 5.257.3 session: the subagent hook called Jev with the key fromafk.env. - The AFK adapter's subagent hook never fired.
hooks/hooks.jsonmatched/Agent/, but agent-afk names the toolagent; the matcher is now/^agent$/. The hook also skips the tier question when the spawn names anagent_type, as the Claude Code router now does for non-general-purpose spawns (claude-jev-afk0.4.0). - The subagent router reads the user's tier text from
CLAUDE.mdonly when it asks the tier question.
0.25.0 - 2026-09-27
- The subagent router routes every spawn it can answer. It picks the cheapest tier where Jev puts at most a 0.10 chance on the task needing a stronger one, instead of requiring 0.75 confidence in one tier. Four tiers split the probability, so that gate held back 190 of 300 spawns, which then inherited an opus or fable parent. On those 300 spawns, labelled with the model the user named, the new rule gets 160 matches, 51 too-cheap picks, and 89 too-dear, against 92, 50, and 158 before (
eval/subagent_eval.py, new). The AFK adapter's advisory tier uses the same rule. Tier text rewritten toward judgment-heavy work did worse on the same 300 (146, 46, 108), so the shipped text stays. - The prompt router no longer asks for a model tier or prints a tier
systemMessage. In the live log, 197 of its 219 hints told an opus or fable session to switch down to haiku or sonnet. Switching mid-session costs the prompt cache and the context, and subagent routing already covers cheaper work. Jev pickedopus4 times in 659 prompts andfablenever. The router log dropstier_hintandmodel_now, and the Stats row drops its tier line. The AFK adapter's intent bundle drops the question too (claude-jev-afk0.3.0). - A subagent spawn that goes through on its second try with parts still missing, and also gets a routed model, now shows both notes. Before, the routing line replaced the missing-parts warning.
0.24.0 - 2026-09-26
-
The plugin runtime is TypeScript. The prompt router, subagent router, rule hook, and compactor run from
src/*.tsunder node type stripping, so node >= 22.18 is required and there is no build step;hooks/hooks.jsonand the/claude-jevcompaction bridge call them directly (#20).scripts/compactor.pyandscripts/subagent_router.pyare gone,scripts/prompt_router.pystays as the routing eval's measurement reference, andeval/compact.tsreplaces the Python compaction gate. Stats,jev.py status, and the rules eval stay in Python. -
Setup guides for using Jev over MCP from OpenCode (#10) and the Vercel AI SDK (#11).
-
The TypeScript client's
jev-calls.jsonlrows match the Python ones again. Each row names the hook that made the call instead ofhook, carries the error string on failure, and takes its version from the nearest plugin manifest, so an AFK call no longer reports the Claude Code plugin's version. A 200 response with no answers now counts as a failed call instead of a success. The nested-instruction-file walk skips symlinked directories in bothscripts/rules.pyand the TypeScript rule parser, so a link into a mounted volume cannot stall the rule hook.eval/compact.tsresolves the repository root fromeval/instead of its parent, so the compaction gate can readsrc/compactor.ts. -
The compaction gate fails a degraded run rather than reporting a better number. The compactor is fail-open, so a backend chunk that fails yields no answers, its blocks go unscored, and unscored rows are kept, which raises coverage; a wholly failed event was visible only because the planted probe count moved, and a run degraded in part showed up in neither. The gate now wraps the decision backend in a call and failure tally, which needs no change to
compactor.tsand so leaves the selection signature and its cache alone, and any failure fails the gate since the floors above it are computed on blocks that were never judged. Distinct errors print with counts, so a 503 from a rate-limited endpoint is distinguishable from a 402 from an exhausted account. The first version of the check rode the floor table with a floor of 0, where 100% failures satisfiedval >= floorand a run where every call failed printed ok; the failure test against a nonexistent model caught it, and it is now evaluated as a ceiling. Verified both directions: exit 2 with 1398 named HTTP 400s on the broken run, exit 0 at 0.0% over 1350 calls on a clean one. -
A truncated block keeps its end.
truncateBlockkept a 400-character head and an elision marker, so anything in the tail was lost unconditionally, and that is the shape of the real failure: a user says "also, do not touch the migration files" at the end of a long message, and a head-only cut drops the most recent instruction while keeping the routine output above it. It now keeps a 150-character tail as well, and the marker still reports the true elided count so the re-read pointer does not misstate what reading it back costs. The gate now tallies kept, cut, and dropped per planted kind and reports where the key sits in each cut block, which is what identified this: all five failures sat at 90 to 97% of the way through their block, within the last 25 to 60 characters. Over the full 38 real and 60 synthetic events, buried restatement survival goes 90.9% to 96.6% with cut 5 to 0, re-fetch verbatim coverage 85.4% to 86.1%, verbatim re-reads 649 to 654, for 3476 to 3523 tokens per event. The floor had been failing by one probe at 89.8% and passing by one probe at 90.9%, which is the measurement telling you it cannot distinguish a fix from a coin flip; a truncation-ordering change was tried first, ordering the budget cut by keep score rather than full score, and measured as an exact no-op, so it was reverted rather than kept for sounding right. -
respan/span-01-liteclears this compaction gate and performs badly against the rules as they are currently written and tuned. Those are narrower than the first version of this entry claimed, and neither is a general statement about the model: the floors and probes were shaped around Jev, so clearing them shows span-01-lite is not disqualified here rather than that it is equivalent, and a planted probe tests literal key retention through selection and truncation, which is not a test of whether a resumed agent obeys a constraint. Measured on the same code, span-01-lite passes all three compaction floors: 81.3% verbatim coverage against Jev's 86.1%, and planted survival of 97.8% for a user constraint and 94.3% for a buried restatement against 97.8% and 96.6%, with every cut eliminated in both. Its coverage and verbatim re-reads still trail, 618 against 654, for 3591 tokens per event against 3523, though that gap is observed on this corpus rather than established: the runs are paired, only 15 of the 56 events that had re-reads are discordant, they come from 19 sessions, and a session-clustered interval on the advantage is wide enough to include zero. One run per model also cannot separate run-to-run model variation from a real difference. An earlier reading of this repo had span-01-lite missing both planted floors at 91.0% and 75.0%, and concluded from that that constraint discrimination was a property of the model rather than the algorithm. That was wrong, and the tail-retention fix above is what proved it: the planted keys sit at 90 to 97 per cent of the way through their blocks, and span-01-lite's lower keep scores put more of its kept blocks through the truncation path, so the same bug cost it more than it cost Jev. Measured before the fix it looked like a model ceiling; after it, the model is within 2.3 points on both planted floors. As the rule decision model and judge over the 303-record corpus it catches 1 of 29 hand-written violations where Jev catches 15 of 28, and blocks 0.0% of real edits at everyACTthe sweep covers, 0.60 to 0.95, because its scores for rule violations sit almost entirely below 0.60. That is not the same as no threshold working, and an earlier version of this entry claimed it was. Swept further down, atACT0.30 span-01-lite catches 6 of 29 while blocking 33 of 247 real edits and 2 of 24 compliant near-misses, an operating point Jev dominates on all three axes at 0.80; at 0.15 it catches 20 of 29 but blocks 141 of 247 real edits and 12 of 24 near-misses. The defensible summary is that no threshold reproduces Jev's operating point, not that none works: the probabilities are compressed toward the bottom of the scale, andACTis calibrated to a model that uses the range. Both arms are complete over the same 303 records with 0 errors each: Jev catches 16 of 29 and blocks 8.5% of real edits atACT0.80, span-01-lite catches 1 of 29 and blocks none, and its best row is still 1 of 29 atACT0.60. Jev was served from OpenRouter rather thanapi.typesafe.aihere, since the samejev-latestmodel is served by both under the same System One protocol and the TypeSafe account was out of credits; the answer cache keys on model, state, and questions rather than provider, so switching the billing path cost only the records that had failed. -
The rule hook stops walking into mounts and the user
Library. A record whose cwd is the home directory sent the nested-instruction-file walk across the whole home tree, and it wedged: every worker blocked inos.listdirat near-zero CPU with the answer cache frozen. A mount point is a different project on a different volume, and OrbStack's NFS share under the home directory does not answer at all, so onelstatshort of reaching it the walk never returns;same_filesystemcomparesst_devwithout following anything, so the mountpoint is recognised without touching what is mounted behind it. The macOS per-userLibraryis app containers and File Provider domains rather than source, is enormous, and a backed provider can block a listing for as long as the network takes, which is the same wedge by another route. Both were production hazards and not only eval ones: a session started in the home directory hung the rule hook on the same walk, and the hook's deadline turned that into a silent no-check instead of an error. Loading rules from the home directory now completes and returns 12. -
The rules corpus loses 28 records and says why. Extraction pins each record to the repo HEAD at extract time, so a rewritten or force-pushed history invalidates every record pinned to it:
pi-environmentwas retired into five fresh commits, and the objects behind its old head are in neither the local repo nor the remote, wheregit fetchdoes not recover them. Those 28 records were being judged at the live checkout, which answers a different question under the same name — the rule files as they stand today rather than as they stood when the corpus was built — and did so silently. The run now stops and names the repo, the commit, the cause, and the two honest options, and the corpus is repaired by dropping what can no longer be reproduced: 1541 records kept, the pre-repair file kept beside it. A re-extract repins the current history, and the question does not arise again until that history is rewritten too. -
Compaction holds a read result instead of voting on it.
isReadResultrecognizes atool_resultwhose call was aReadand keeps it atKEEP_THRESHOLD, because routine file reads are the largest block class in a long session and were out-competing the constraint, decision, and error rows a later turn actually re-reads. Measured over the 38 real compactions, verbatim coverage moved 84.5% to 91.8%, clear of the 69.6-84.5% band the same code produced across reruns, with no event regressing and three improving, at 2761 to 2925 tokens per event. -
The compaction harness can afford to be run more than once. Summaries were held in memory, so every full rerun re-paid all sixty, and they are now cached to disk by cut hash next to the selection cache that was already there. Generation was
spawnSync, which blocks the event loop, so the worker pool serialized behind one summary and a rerun needing fresh ones took over three hours; asyncspawnlets it pool, and the same sixty summaries went from 2 generated in seven minutes to 49. The generator ispirather thanclaude, which also drops a CLI that is weekly rate limited and returned nothing for the whole synthetic arm.--explainattributes each uncovered re-read to its cause, and it is what pointed at this read rule rather than at the keep threshold: the dropped blocks scored 0.04 to 0.16, nowhere a threshold change would have reached. -
The decision model and the judge are swappable behind one seam.
DecisionBackendis a name plusask(state, questions, timeout)returning typedAnswers; production keeps System One as the default, and the eval injects another implementation by parameter. The backend name is folded into the selection signature, which is the part that would otherwise bite — without it a second model's run reads the first model's cached decisions back and measures nothing. Two implementations share onetypedAskover the identical{state, model, questions}to{answers}contract, so the swap is a URL and a model id rather than a second code path:systemone, anddecisionsfor OpenRouter's/api/alpha/decisions, which servesrespan/span-01-lite, a System One model behind a different URL. Spec syntax isbackend:model, or a bare model id meaning System One, and an unknown backend id raises rather than falling back quietly on both sides.scripts/jev.pytakes the same split soeval/rules_eval.py --judge-modelcan judge with the same model that decides; its answer cache already keyed on model andrules.pypassestimeout=by keyword, so the judge model is injected as a default without disturbing either. Every compaction keep question is anoul, which is all Respan accepts. Over the full population, 38 real events and 60 synthetic, with the planted probes at their whole n=89 and n=88, span-01-lite as decision model gives 84.2% verbatim coverage against Jev's 85.9%, but planted user constraint survival falls to 91.0% from 97.8% and buried restatement to 75.0% from 89.8%, missing both floors, and it spends more to get there at 3007 tokens per event against about 2900. Coverage is judge-independent; the discrimination between a user constraint and a routine row is not, and that is what the keep floor and the planted probes exist to protect. Read on the 38 real events alone the picture inverts to 92.3% coverage and only 29 and 28 probes, which is the whole argument for the synthetic arm. -
adapters/afk/ships a second, TypeScript implementation of the hooks for the AFK host (#18), with its own plugin manifest (claude-jev-afk0.2.0) and install path (npm run build, copy into~/.afk/plugins). It shares the Claude Code plugin's judged surface verbatim — the question bundles,ACT0.80 /FLAG0.50, the twice-per-session block downgrade, fail-open on everything — and adds one hook the Claude Code plugin lacks: a SessionStart rule digest that reads the instruction files and injects a structured summary with no API call. It lacks what AFK cannot express, andadapters/afk/README.mdsays so: no compaction (AFK hooks expose no transcript), advisory-only subagent routing (noupdatedInput, so a bad brief is annotated, never denied), prompt routing on the prompt alone (no previous turn), no ast-grep enrichment, and no shell-write tracking (hooks.jsonmatches onlyedit_file|write_file|patch_apply).README.md's install section claimed "Same hooks, TypeScript" — false in both directions, now corrected to state the delta. The recorded plan: this adapter's shared core becomes the single implementation, brought up to Claude Code's enforcement level, beforescripts/is deleted. -
control-jevpinsCLAUDE_CONFIG_DIR=$VERIFY_HOME/.claudefor every child process.jev.config_dir()prefers that variable over~/.claude, so a home isolated byHOMEalone still leaked to whatever config directory the calling shell exported — which this repo's own instructions tell an agent to export for hand-fed hook events. Every log, cache, and rule-count assertion in the verification feature map silently described the wrong directory; the harness fix makes them true, and a re-drive with a decoy exported confirmed writes stay in the verify home. -
README.md's rule-hook table gains a v0.23.0 column beside v0.21.0, from re-scoring the same 247 judged real edits at current source on cached answers: 23 blocked (9.3%) against 22 (8.9%), 9 flagged only against 8, hand-written violations 17/29, compliant near-misses 0/24, median 8 rule questions per edit — unchanged. Latency is published asnot measured: a cached re-score reports a 0.01s median against the 0.80s a live run costs, so the cell would measure the cache rather than the hook. -
The feature map now carries a live proof of the settings pane and the function-hook bridge, run against the installed copy at 0.23.0. Toggling
Rule checkswritespluginConfigs["claude-jev@claude-jev"].options.rulesand the effect was measured as an A/B on one identicalWrite: zero"kind":"rules"log rows with it false, two with it true. TogglingCompactionoff makes/compactlog nothing from the plugin, becauseregister.tsreturnsnext(e)before the$.ui.logthat reports a fallback — so plugin silence is the off signal andbuilt-in summary runsis a failure, a distinction that would otherwise let a verifier read a broken bridge as a disabled toggle. Also recorded:claude plugin update claude-jev@claude-jevbrings the installed copy level withmain(it was four versions behind here, which is why--plugin-dirwas needed at all), and the@inlinevs@claude-jevmodule-load ids differ by load path. -
eval/rules_eval.py runtakesJEV_RULES_CACHEto point the answer cache at another path.--outnever bypassed it — it moves the prediction file whileeval/observed/rules_cache.jsonlis still read — so a re-scored run reportedlatency median 0.01s, which is the speed of a dict lookup, not of the hook. Judged live with the override (326 calls, 0 errors, pinned ast-grep present) the same 247 edits return the same 23 blocks and 9 flags at a 0.40s median, and that is the figure now in theREADME.mdtable against v0.21.0's 0.80s. The live pass costs ~40s at four workers; the cached one is free. -
eval/sweep.pyandeval/planted.pyrun again. Both inserted only their own directory beforeimport compactor, sopython3 eval/sweep.py --helpdied withModuleNotFoundError: No module named 'compactor'whilesweep.py's own docstring prescribespython3 eval/sweep.py --synth 60andREADME.mdpublishes survival figures attributed toplanted.py. They now insert../scriptslikecompare.py,replay.py, andrules_eval.pydo. The pair only worked when reached throughcompare.py, which putsscripts/on the path before importingplanted— so the compaction gate passed while the diagnostics that read a change were dead, andpython3 -m compileall -q scripts evalcould not catch it because it compiles without importing. -
The feature map under
.cursor/skills/verify-claude-jev/was reconciled against source and one live run per feature. Newfeatures/comparators.mdcovers the pinned ast-grep:whichand itspath/cached/nonestates, the degraded judgment with no binary, and the detached fetch. Corrected:stats.pybails on an empty router log alone,--logredirects only that log,--days 0means no cutoff, andreportreads cached predictions so re-reading a published number costs nothing whilerunbills. Softened two claims the repo cannot substantiate — thesettings.jsonpath behind a pane toggle and two debug-log strings that are Claude Code's own output, not the plugin's.
0.23.0 - 2026-09-25
- The scorer takes shell writes from Claude Code's own diff, not a regex. A Bash result's
toolUseResultcarriesbashEditDiff, the file-state diff Claude Code took around the command, naming the git-visible project files it changed — the same notion of a write the live hook inrules.pyalready used, so the scorer and the hook now agree.observed.walk_sessionjoins each Bash call to its result and handssummarizethe changed paths;summarizetreats that verdict as authoritative and lets it override the pattern in both directions, so a command that only wrote/tmpor a gitignored path no longer makes the turn an edit. Claude Code 2.1.274 is where the field first appears in the transcript corpus, soBASH_DIFF_VERSIONgates it: below that, a missing field means "unknown" rather than "wrote nothing", and the patterns still decide. - Measured against that diff as ground truth — 6,143 real Bash calls across the 145 sessions that record it — the command-text patterns score precision 0.25 and recall 0.84. Three quarters of the writes they claimed were a scratch file outside the project, a gitignored path, or a quoted string that happens to contain
>, such ascurl -w '%{redirect_url}'. The misses were writes the text cannot name at all:docker compose run … pnpm add, which editspackage.jsonthrough a bind mount. - The fallback patterns also lost two misses of their own, the same pair compact-adviser fixed in its PRs #52 and #53:
sed --in-place, the GNU long form, which takes its backup suffix only attached with=and so is never ambiguous the way bare-iis; and a write, ops, or read command anchored afterthen,else,do, or{, or at the start of a line in a multi-line command, which the old(^|[;&|]\s*)position could not reach.case $x in a) mkdir -p out ;; esacstill misses — a)prefix bought nothing on the corpus and would match subshells. Over the ground-truth set these move recall 0.843 to 0.846; they matter as correctness of the fallback, not as a measurement change. - Labels moved on 262 of 4,857 rows (5.4%):
fix857 to 699,lookup1634 to 1707,feature793 to 840,ops239 to 282,chat1284 to 1297. Bash-changed files now count towardn_files, so a turn that wrote three files through the shell reaches the substantiality band the way a turn that wrote three throughEditalready did. Re-scored over the whole current corpus with the same cached predictions (eval/replay.py run --variant v7_no_unclear, 2,449 of 2,464 rows cached, 15 new calls, 0 errors), holding the row set fixed to isolate the scorer: v7 accuracy 32.9% to 34.0% and lift +7.6 to +8.6, hinted-only accuracy 49.9% to 48.4% and lift +7.2 to +5.8, harmful hints 8 either way. Against the 120 hand labels ineval/audit_labels.jsonit is a wash — 52.5% to 51.7%, six rows changed, two toward the human label and three away. README.md's router table is re-run on the current corpus rather than carried over: 2,464 prompts instead of 1,613, and every figure now reproduces from the commands printed beside it. The published coarse-taxonomy pair did not — 58.8% and 64.7% only come back at--floor 0.5, off the shipped 0.75 floor the rest of the table uses — so it is replaced byv5_three_wayat the shipped floor, 65.0% against a constant 63.6%.
0.22.0 - 2026-09-25
-
Compaction asks a fifth
rerunnablecheck: output a rerun would print again — a listing, a passing check's log, build or install output, warnings a rebuild repeats. A rerunnable answer atKEEP_THRESHOLDsilenceserror's verbatim claim, so a regenerable dump keeps as aHEAD_CHARShead with the re-run pointer instead ofKEEP_CHARSwhole; an exact error no command would repeat, and any user constraint, still hold verbatim. The four keep reasons are unchanged, so pure noise still drops and its ref stays in the row list. Replaces rule-based scrubbing of verbose output with the same judgment Jev already makes per block. Gate behind it (eval/compare.py compact --synth 60on this source, 60 events, live re-judgment): re-fetch verbatim coverage 76.8% (floor 70%, n=760 reads — inside the standing 76-82% band), planted user constraint 100% (n=87), buried restatement 98.8% (n=82). -
jev.askaccepts adeadline(monotonic timestamp) alongsidetimeout, and retries once when a network failure returns in under a second — a refused connection costs milliseconds, a timeout costs the budget. -
The rules hook now works inside its 10s
hooks.jsonbudget instead of hoping two sequential Jev calls fit:mainstamps a 9s deadline, everyaskgets the time left, and the escalation pass is skipped when under 3s remain — the first-pass verdict still lands instead of the hook being killed with no output. -
Rule classification runs its per-chunk requests in parallel and caches by chunk content hash instead of whole-file hash. A large instruction file no longer serializes into a hook timeout, a killed classification keeps the chunks that landed, and editing one paragraph re-judges only its chunk.
-
Compaction chunks ask with a 4s timeout instead of the 8s default, so one slow request can't stretch a ~1s compaction; a timed-out chunk's blocks are kept unscored, as before.
-
Compaction asks four checks per block instead of five: the
artifactcheck ("could this be re-fetched?") fed no decision —verdictsreads only the keep and verbatim scores — so it cost a fifth of every compaction request for a diagnostic nobody consumed. The/compact <text>directive moved from a suffix on every question to one line in the shared state header, wheresession_contextalready names it. -
Blocks just older than the 150-row judgment window are no longer dropped unseen: the 150 before them get the constraint check alone, and any that score a user requirement are kept verbatim ahead of the window's rows. A constraint stated early in a long session now reaches Jev instead of falling off the window silently.
-
Ruff formats and lints the Python sources.
ruff.tomlselects pycodestyle errors (E4,E7,E9), Pyflakes (F), and import sorting (I), at a line length of 100, and skipseval/observedandeval/authored. It is tool config, not a dependency: hooks still import only the standard library.ruff format --checkandruff checkrun in release preflight and in GitHub Actions, next tocompilealland the twocheck_no_*.pyscripts. A machine with onlypython3can still claim those stdlib checks. -
eval/data/is noweval/observed/andeval/private/is noweval/authored/— the names say what they hold: machine-extracted transcripts and predictions vs hand-written cases. Both stay gitignored.
0.21.0 - 2026-09-24
- The rules eval now measures the AskUserQuestion answers the way the hook sends them. Its corpora carried only the typed request, so no rerun could see the v0.20.0 answers feature:
extractrecords each edit's answers since its request, andruncomposes the request through the samerequest_with_answersthe hook calls; the rules decision log rows gained auser_answerscount, so live outcomes can split by whether the user answered questions. Extraction collects answers unconditionally, like the hook: an AskUserQuestion's result line normally carries no tool name, so the substring gate this replaces saw 5 answers-bearing edits where the corpus has 140 of 1,569 — enough that a normal 250-sample now measures the feature (~22 records). Records without answers compose to the old request text, so their cache entries still hit. The hook itself gains a fix: a transcript that could not be opened madelast_user_promptreturn a 2-tuple, crashing the handler before any rule was asked.scripts/check_no_stubs.pyjoins the Verify contract:.ask =assignments are banned everywhere except thecached_askwiring shapes incmd_run, so a probe cannot fake the client and present the result as verification. Live evidence through the real client: a313bcf9#932 ("rspec check in CI failed") reached the judge with its picked option and none acted, its re-judgment hitting at 0.0s the cache key only its answers-carrying state could produce; a synthetic result line with no tool name rides in the request exactly as the hook composes it. Records whose directory stopped being a git checkout drop their sha and join the live-checkout bucket the report already counts, instead of killing the run. The rerun on the new corpus: 300/303 judged, 0 errors; 22/247 real-edit blocks (8.9%), 17 of them symphony's owncode-comments-are-banned-inrule (0.82–0.91) on accepted edits from before this change — none of the blocked rows carry answers; 17/29 violations blocked, 0/24 compliant; 22/247 edits carried AskUserQuestion answers, the line this change adds.
0.20.0 - 2026-09-24
- The rules hook counts the user's
AskUserQuestionanswers as part of the request. Jev saw only typed prompts, so an approval given by picking an option never reached it. In one session that blocked three approved edits: two to an axe spec (0.83, 0.85) and one tofeature_flag_spec.rb(0.91), each under "do not weaken tests", after the user had picked the option that said the test expectation would change. The request now lists each answer since the latest prompt, as the question, the picked label and that option's description, capped atMAX_ANSWER_CHARS. Subagent results are skipped. The rules eval has not been rerun on this change.
0.19.2 - 2026-09-24
- The rules hook records what a Bash command changes. A
PreToolUsehook snapshots the git working tree, untracked files included, in a scratch index. The matchingPostToolUsediffs it and records each changed file as a hunk, so theStopcheck judges shell writes like edits. A turn whose snapshot fails, such as outside a git repo, still tells Jev its diff is partial. Snapshots took under 0.7s on the five slowest repos checked. - A failed Bash command's writes are recorded too, through a
PostToolUseFailurehook. A turn with no recorded hunks but a partial diff still gets theStopcheck, and so does a Bash write to a git-ignored path, which marks the diff partial.
0.19.1 - 2026-09-24
-
The rules
Stopcheck judges only the edits made since the latest user prompt, and skips a turn with none. It had pooled every edit since the session began and judged them against the newest prompt. In sessionccb4b920, that blocked twice on approved.zprofile/.zshenvedits from an earlier turn (0.91, 0.92) and spent the session's block budget. Hunks are keyed by the transcript uuid of the prompt they were made under. The request and the repair message now name this turn's files. Seedocs/stop-hook-false-positive.md. -
When a turn with recorded edits also ran Bash, the
Stoprequest tells Jev the diff is partial, since shell writes never reach the recorded hunks. -
The rules hook reads and writes its session state under a file lock. Twenty parallel writers kept 20 of 20 hunks; without the lock they kept 5. It logs a swallowed exception as a
rules-errorrow, which stats never scores, so a missing decision row can be explained. -
python3 eval/rules_eval.py turnssplits live Stop-hook checks by whether the turn made an Edit/Write of its own. It is offline and free. On the live log before the fix: 216 checks. 82 judged the turn's own edits (0 blocks, 30 flags). 107 came after Bash-only turns (2 blocks, 23 flags) and 27 after turns with no changes (0 blocks, 6 flags). Both blocks were theccb4b920false positives. Under the turn scoping, those 134 checks and 29 flags no longer run. -
Add a bounded rule-prompt comparison tool and document a reviewed live false positive, evidence limits, and candidate questions. Shipped rule prompts and thresholds are unchanged.
-
Stats breaks down failed calls by HTTP status or timeout for each caller, shows recent call health and last failure/success timestamps, and explains that rule outcomes are edit heuristics rather than verified repairs.
-
scripts/stats.pylists~/.claude/projectsonce instead of globbing it for every session. On this machine the report went from 9.1s to 0.9s cold and from 2.15s to 0.59s warm, with byte-identical output on a frozen copy of the logs. -
The Status row in
/claude-jevshows the status the pane already loaded as soon as it is pressed, then refreshes it fromjev.py status.
0.19.0 - 2026-09-23
-
The prompt router no longer shows an
opshint. Live, 18 of 60 were right; on the 120 hand-labeled prompts humans agreed with 6 of 18. Withfeaturealready silent, replaying 1,613 prompts through the shipped rule goes from 31.1% accuracy and −0.9 lift to 50.2% and +18.4, at 13.8% coverage. The cleaned live log goes from 49.5% (111 hints) to 72.5% (51 hints). An intent is silent by having no entry inGUIDANCE;SILENT_INTENTSis gone.eval/variants.pygainsv9_hinted_only, which scores v7 answers only where the hook shows a hint. -
The
featuresilence (commit 0f67b21) cited 67 predictions against 6 observed; most of those predictions were compaction requests. It stays silent on current evidence: 8 of 28 right in the cleaned live log. -
scripts/stats.pyscores only prompt-router entries; rule checks and subagent decisions were counted as suppressed prompts (1,040 reported, 351 real). It prints how many entries were Claude Code's own requests, reports the confidence floor as held-back count and hit rate (35 held back, 15 of them right, against 72.5% for the hints the router now shows) instead of countingchatpredictions that never show, and calls a block with no transcriptunscorableinstead ofunknown. -
eval/replay.py reportandcomparedefault--floortoprompt_router.MIN_CONFIDENCE(0.75) instead of 0.55, so the documented command reproduces the README row. -
scripts/check_no_comments.pyscans only files git tracks or would track. It had been failing on the gitignoredeval/data/at/checkouts. -
AGENTS.md: run hand-fed hook events withCLAUDE_CONFIG_DIRset to a temp directory. Test events with session ids liket1had landed in the live log, and were the 15 "unknown" rule outcomes. -
scripts/stats.pyno longer scores Claude Code's own pre-compact request ("Your task is to create a detailed summary…") as a user prompt. Older plugin versions logged 73 of them with a hint, and the session that followed used no tools, so the report counted them as wrongfeaturehints against an inflated always-chatbaseline. On the same log (1,316 decisions), the report moved from 72/195 agreed (36.9%) against 48.7% for alwayschat, to 58/122 (47.5%) against 26.2% for alwayslookup. The weak class is nowops: 60 hints, 20 observed. -
The prefix lives in
SYNTHETICinscripts/observed.py, so the router, stats, andeval/replay.pyshare one filter;COMPACT_PROMPTinscripts/prompt_router.pyis gone.eval/replay.py report --variant v7_no_unclearis unchanged (1,613 scored, 33.7% accuracy, +4.4 lift): none of its predictions were compaction prompts. Derived labels still agree witheval/audit_labels.jsonon 63/120.
0.18.0 - 2026-09-23
- OpenRouter as a second Jev provider. Set
OPENROUTER_API_KEY, orTYPESAFE_API_KEYto an OpenRouter key (sk-or-...), and every call goes tohttps://openrouter.ai/api/v1/systemone, which takes the same request, model IDs, and answer shape.TYPESAFE_API_KEYis read first. A TypeSafe key still callsapi.typesafe.ai. The Provider row in/claude-jev(providerin/config:auto,typesafe,openrouter) pins one provider and reads only its variable, so OpenRouter works with both variables set. Live through OpenRouter: anoulanswered 0.99, and the prompt router answered in 472 ms. Eachjev-calls.jsonlline records itsprovider. /claude-jevsettings pane (needsCLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1): save an API key for all sessions, see which provider it calls, turn the prompt router, subagent router, rule checks, and compaction on or off, and see the last Jev call. The on/off rows also appear in/config. A key in the environment wins over a saved key.jev.py statusprints what the pane shows. The pane's Stats row shows thescripts/stats.pyreport.- Removed the
/claude-jev:jevand/claude-jev:statsskills. No session outside this repository ever called the first, and it could not see a key saved in the pane. Stats moved into the pane;python3 scripts/stats.pystill prints the full report. scripts/compactor.pychanged only its docstring;hooks/register.tsnow passes the saved key and provider to it. Compaction gate, run through OpenRouter: re-fetch verbatim coverage 79.8% (floor 70%, n=678); planted user constraint survival 100.0% (floor 95%, n=72); planted buried restatement survival 97.1% (floor 90%, n=69).
0.17.0 - 2026-09-22
- Subagent briefs are checked before the spawn: a brief that changes files but omits paths, acceptance criteria, a verification command, or a commit policy is denied once per session with the missing parts listed, then goes through with a
systemMessage. Read-only briefs are exempt. Live, a thin rename brief scored 0.06–0.16 on all four parts and a complete one 0.97+. - Subagent tier criteria rewritten with sharper boundaries and a fourth
fableoption for adversarial review and cross-system debugging with conflicting evidence, never implementation. On seven hand-written briefs the live-vs-eval bug moved from opus 0.44 to 0.77, above the routing gate. The prompt router's advisory tier hint recognizesfablemodel names. Tier text can be overridden by a## Delegating to sub-agentssection with- tier: textbullets in the user's globalCLAUDE.md. - Every file under the user's config directory (logs, caches, ast-grep binary, global
CLAUDE.md) resolves throughCLAUDE_CONFIG_DIRwhen set, else~/.claude. scripts/compactor.pychanged only its log path. Compaction gate: re-fetch verbatim coverage 76.5% (floor 70%, n=727); planted user constraint survival 100.0% (floor 95%, n=72); planted buried restatement survival 98.6% (floor 90%, n=69).
0.16.1 - 2026-09-22
- The release skill and script moved out of the plugin into
.claude/skills/release;/claude-jev:releaseno longer exists for plugin users.
0.16.0 - 2026-09-22
- Compaction judges each block with five concrete checks instead of two aggregate questions. The checks: user constraint, decision with reason, exact error, open work, re-fetchable output; keep and verbatim scores derive from them in code. A constraint planted mid-session survives 100% (was 77%); a restatement buried in a later reply survives 98% (was 35%). Scores in the 0.35–0.65 band fall from 54% to 21%.
eval/compare.py compactgates both compaction goals at once: re-fetch verbatim coverage (floor 70%) and planted-constraint survival (floors 95% and 90%), exiting 2 below either.eval/sweep.pyandeval/planted.pyadded as diagnostics.- Rule hook escalates the uncertain band with one focused second call and adds ast-grep comparators so the fact outside the hunk reaches the judgment.
- Rules are structured with a local relevance gate and violation-framed questions; every Jev call and every compaction row is logged for
/claude-jev:stats. - Code comments are banned in
scripts/,eval/, andhooks/, enforced byscripts/check_no_comments.py. - Docs:
docs/prompt-craft.mdrecords measured effects of question wording; Pstack verify skill under.cursor/skills/verify-claude-jev. - New
/claude-jev:releaseskill andscripts/release.py.
0.12.0 - 2026-09-21
- Compaction through the experimental
session.compactfunction hook: Jev's kept rows replace the built-in summary.