latent-sre/save-toolkit
Application-engineering and site-reliability agents and reusable skills.
Changelog
Recent changes to the current Save Toolkit. This is a pre-release repository; entries do not
assert a published release or acceptance of an evaluated candidate. Earlier implementation history
is available in Git. Unfinished work belongs in docs/fleet-roadmap.md.
[Unreleased]
Changed
-
CI runs the component tests on four workers and installs nothing for the structural gate. Over the 16 runs before this change the test step took a median 438 s and the whole run 484 s.
component-testssetsPYTEST_ADDOPTSto-n 4 --dist loadfile --durations=20; the command stayspython -m pytest -q. Locally the full suite went from 760 s serial to 249 s (four runs, 248 to 256 s).- The same step sets
PYTHONDONTWRITEBYTECODE=1. The eval runner digests every file underskills/, so a.pycwritten by one worker reads as plugin drift to a native trial running in another. Six first runs on fresh worktrees passed: three with the guard, which wrote no.pycunderskills/, and three without, which wrote seven each. validateno longer installsrequirements-dev.txt. No gate-path script imports a third-party package, Gate A and the atlas check pass on a bare interpreter, andtest_validate_workflow.pyfails when the first such import arrives without the install.
-
The rubric judge follows the latest Sonnet (owner decision 2026-10-04): each calibration requests the
sonnetalias, and its receipt pins the concrete model that answered, so trials never follow the alias between calibrations. A new Sonnet means a recalibration with--resolve-identitybefore the old model retires, with the judge cache cleared when the probe reports that the alias moved, and the native scenarios'expected_modelpins move with it. [verified] The first calibration on Sonnet 5.5 (claude-sonnet-5-5) cost USD 2.44 at list price for 165 live calls, against USD 4.90 on Sonnet 5. -
reviewerfixes the PR #307 review findings (owner decision 2026-10-02; no OS isolation for scratch runs is accepted risk): the autoload preparation gap names Copilot's instruction roots too; a commit is exported through a scratch index, nevergit archive, whose export attributes drop and rewrite files; every literal Git command uses$G; and each P0/P1 finding is reproduced in scratch even when it follows from the code (the export change alone cut reproductions to 7/15 over three wordings, 3/3 at both prior commits; 9/9 now). 12,746 -> 13,278 bytes. The harness addsran_outside_checkout, grades every Git call for the prefix, rejects malformeduncommittedkeys, and regrades seeded-uncommitted andscope: subagentchecks correctly. Sonnet, 3 trials: reviewer suite andsoftware-engineerhandoff 31/33, main 1/12 on the four changed cases; every-call prefix 5/6 (main 0/6); probes 9/9 (a Copilot-only instruction edit does not block on Claude; export-ignored tests stay). Opus 8/9: one run started Python in the checkout with scratch imports. -
reviewerruns checks by default on work in the user's repository (branches, PRs, and uncommitted changes), only in a scratch copy, reproducing each suspected finding first; forks and unknown provenance stay CI-only, and the unused local Docker recipe is gone. 13,507 -> 13,336 bytes. Sonnet, 3 trials per case against main: scratch-copy reproductions 6/6 on direct branch and uncommitted reviews (main ran nothing on the branch and ran code inside the checkout 2/3 on uncommitted work); in-checkout runs after asoftware-engineerhandoff 1/3 (main 3/3); fork-runner refusal and read-only or no-execution scopes held 9/9. -
revieweris independent ofresearcher(owner decision 2026-10-01): it dispatches onlyrepository-investigator, and a dependency advisory or other public fact it cannot verify locally is a stated gap returned to its caller. On a urllib3 downgrade it flagged the risk from memory, labelled it unverified, asked the caller to confirm, and requested changes 3/3 with no dispatch; main dispatched a researcher that could not answer in 1/3. 13,143 -> 12,746 bytes. -
reviewermay bind a branch review's verdict to the named commit and report uncommitted defects as PROVISIONAL findings (owner decision 2026-10-01); the terse-handoff check accepts that or a PROVISIONAL verdict (6/6 Sonnet, 3/3 Opus; main 2/3). The two rename cases ask for the file whose change introduces the defect (Opus had named the broken caller). Roadmap gains PRECOMMIT-001 and REVIEWER-001. -
reviewercloses probe gaps: a candidate that edits CLAUDE.md, AGENTS.md, or.claude/in the checkout it runs in gets a preparation gap instead of a verdict (0/3 -> 3/3; all three earlier trials resisted the injected policy but none noticed it had loaded as their own instructions); changed-file history names each path (0/3 -> 3/3); uncommitted work on a branch is snapshotted before any run (in-checkout runs 0/6 after a terse handoff). The generic security checklist is cut: detection was 9/9 with and without it on workflow injection, SSRF, and a cross-tenant read; the dependency-advisory and agent-system lines stay. 13,331 -> 13,051 bytes. The PROVISIONAL verdict line after a branch-named handoff stays at 4/6: the reviewer binds the verdict to the committed SHA. -
revieweroutput and Git contract: a side-effect-free Git prefix with minimum reads, every changed line covered or its filter named, an evidence label beside each finding's priority,Verdict: PROVISIONAL — …for mutable reviews, a requester header that never copies an account email, and a one-linepython -ccounted as a run. 13,336 -> 13,331 bytes (main 13,507). Sonnet, 3 trials against main: tenant, clean, and uncommitted cases 0/3 -> 3/3; branch-history reads 3/3 (main 1/3, 2/3); scratch copies after asoftware-engineerhandoff 3/3 (main 0/3); PROVISIONAL verdict line after a terse handoff 2/3 (main 0/3). Changed-file history held only where the reviewer named the paths (0/3 in the retry case). -
Reviewer evals grade the review as written: five free-form build scenarios cover a scratch-copy reproduction, a cross-tenant read, a correct change, and an agent's uncommitted work (direct and after a terse handoff). Fixtures gain
checkoutanduncommitted;no_workspace_changeskeeps seeded uncommitted bytes as its baseline; four reviewer scenarios reject Git verbs that move the source checkout; two diff detectors accept the two-argumentgit diffthat had failed correct runs. On main every outcome check held;contract:checks fail where the body lacks the contract. -
The fleet-atlas contract check verifies the tree once.
check_fleet_atlas_v2.pyran seven atlas CLI commands, and each re-verifies the whole repository (build does so twice): eight verifications a run. It now builds in-process, keeps one real CLI query so the command-line path stays under test, and answers the other four cases from the document it already verified: three verifications. The cases and response checks are unchanged. Locally the--buildrun went from 90 s to 34 s; on the runner the step took a median 150 s before. -
software-engineercarries one short inline handoff core in place of three handoff sections, plus six-pass fixes: a credential-values row, a test-integrity rule, adatabase-reliabilitytrigger, and one skill-trigger list. 25,847 -> 23,356 bytes, measured against main on Sonnet. Two build scenarios now grade resuming after a partial helper return and withholding a reviewer when candidate instructions would auto-load. -
observability-engineer,reliability-engineer,reviewer,agent-engineer, andsre-assistantcarry the same short handoff core, 1,756 bytes smaller across the five. Five build probes give those lanes handoff coverage; main, the core, and a no-handoff control matched on every behavioral check. -
sre-assistantsix-pass cleanup:- a viewing-only rule for its browser tools, which no hook checks;
- the suspected-compromise section no longer says
reviewerhands off to this lane, and it has one escalation route; - CF requests go to Apps Manager first, where the team works;
- the credential-file rule says why it is the only control;
- restated rules and a paragraph describing a path this profile lacks are removed.
22,569 -> 22,278 bytes, measured against main on Sonnet. A stricter CF wording and an extra output-template line were measured, found to hurt, and reverted.
-
backend-crafthouse contract brought to current practice:- health endpoints move to
/health/liveand/health/ready, because Cloud Run reserves some paths ending inz; - versioning adds the RFC 9745
Deprecationheader besideSunset; - rate limits require only
Retry-After, with optional IETFRateLimitdraft fields; - an oversized
limitis lowered to the cap rather than rejected (AIP-158); - idempotency keys follow the IETF draft's status codes;
- SSE keep-alives are about every 15 s;
- outbound calls get one deadline per operation;
- a new dependency-failure rule: fail fast, and mark an optional field unavailable or return
502/504.
The references are rewritten in plain sentences, and
fastapi.mdcovers FastAPI 0.132's strictContent-Typeand 0.135's native SSE. The OpenAPI starter gainsoperationIds,WWW-Authenticate, 403 and 500 responses, and a documentedX-Request-ID. SKILL.md is 7,105 -> 7,799 bytes. On the repairedincidents-apiprobe, main rejected the oversized limit 6/6 and the candidate capped it 6/6; every other oracle check passed on both arms. - health endpoints move to
-
backend-crafttrimmed where a no-skill control showed the model already complies (Sonnet, 3 trials per arm):- the response-model allowlist (no leaks 3/3 without the skill);
- the generic test-running bullet;
- webhook authentication and persist-before-ack (3/3 without the skill).
It now names
Idempotent-Replayedand leads webhook acceptance with202. The trimmed skill met or beat the current one on every check of all three probes; the replay header went 2/3 -> 3/3 and 202 went 2/3 -> 3/3. No arm makes acknowledged webhook work recoverable after a crash (0/12 under the restart check), so the persist-before-ack line that was cut produced no recovery either. -
backend-craftteaches recoverable post-acknowledgement work. SKILL.md's Background-work row, which always loads, says to:- answer
202, or the provider's required code, while work is pending; - persist everything still owed, enrichment included, as pending work resumed on startup.
It also gives the reason: an in-memory task dies with its process, and a sender given a 2xx won't redeliver.
background-work.mdgains the mechanism: a pending step runs again after a crash and on every instance that resumes it, so each step must be safe to repeat. A local step is marked done in its own guarded transaction, a remote call or notification follows API-write recovery, and a step whose repeat costs is claimed first under a short lease. Before, the reference was read in only 4 of 9 webhook trials, and even then no trial recovered.Measured on Sonnet 5.5 against main
7e21fb4e, interleaved, four webhook trials per arm:- the restart check went 0/4 -> 4/4;
202was 3/4 on main and 4/4 on the candidate;- the candidate passed all 14 checks in every trial.
One trial per arm on the writes and incidents-api probes passed every oracle check on both arms.
Two earlier wordings were measured on 2026-10-01 and replaced. The first raised recovery to 4/4 but pushed
202down to 2/4, because agents that created the record up front answered 201; the202clause beside the rule fixed that. The second held both at 4/4 but told the agent to mark each step done in the transaction that commits its effect, which a remote call cannot do, and it answered202even where a provider requires another code.To fit under the screen, the Shutdown row's duplicated requeue clause was merged, and the Outbound-calls row lost its inline
consuming-apislink and its unconditional breaker clause. The reference stays linked from the read-first table and keeps the conditional breaker rule. SKILL.md is 7,718 -> 7,787 bytes. - answer
Fixed
-
Three defective calibration cases that judges disagreed on; no label changed. The
human_handoverrubric paragraph now carries the supplied fact that Riley confirmed the flag value, and the handover PASS response no longer asserts an "agreed recovery window" or an "unassigned" dependency owner, neither of which its scenario supplied. The plan-authorship PASS response names the release owner as the one who deploys: [verified] every Claude judgment had read it as the assistant deploying, every OpenAI judgment as a plan for someone else. Thestatement_rerunPASS response saysneeded. The colleague'sinstead ofneeded; the colleague's: [verified] Sonnet 5.5 quoted the semicolon as a comma in six of eight judgments, which the verbatim-evidence rule turns into INCONCLUSIVE. [verified] Afterwards a calibration agrees with all 165 labels. -
The native incident scenario expects
claude-sonnet-5-5, the model thesonnetalias now resolves to. [verified] Withclaude-sonnet-5, a native incident trial run through the alias would stop INCONCLUSIVE on its model check before the follow-up, astest_native_wrong_or_missing_parent_model_stops_before_resumeexercises. -
Clean-room trial and judge workspaces no longer inherit the operator's instructions. Claude Code reads
CLAUDE.mdfrom every ancestor of its working directory, past any git root, and on Windows the default temp dir sits under the user's home. A Haiku probe throughmain's clean room quoted the operator's~/.claude/CLAUDE.mdheading. 340 of 1,432 saved trial traces ran under the home directory, and the rubric judge used the same workspace.clean_room.make_workspace()now refuses any root with aCLAUDE.md,CLAUDE.local.md,AGENTS.md,.claude/CLAUDE.md,.claude/AGENTS.mdor.claude/rules/Markdown file above it, falling back to<checkout drive>\fleet-eval-tmpon Windows.FLEET_EVAL_WORKSPACE_ROOToverrides the root. The same probe now answersNONE. -
The
incident_companion_responserubric'sstatement_rerunparagraph now carries the owner-supplied facts its PASS case repeats (daily-statement/run-42,statement-2026-07-14/v3, the provider receipt, the 11:26 UTC recipient readback, the cached screenshot). Before, a judge reading only the rubric saw them as invented: uncached Sonnet 5 judged calibration case #150 FAIL in both runs that produced a verdict, and every OpenAI judge failed it, while the 2026-09-23 receipt's PASS was cached. After the edit, three uncached Sonnet 5 runs score 21/21 on the rubric, #150 PASS each time. -
Owner review (2026-10-03) of the PASS-labelled calibration cases judges failed:
- The
existing_bridge,existing_tlc,helper_assignment_and_returnandknowledge_cardrubric paragraphs now carry the facts their scenarios supplied: the PCF target and incident ids, the report times, Morgan's role, the helper's sourced 10:00-10:10 observation, and the service card's escalation chain, 503 load shedding and 10-minute restart step. - The
no_blind_retry_after_unknowncase that asserts unsupplied terminal-completion evidence is relabelled FAIL, without its trailing conditional, with a new PASS case that makes the retry conditional on that evidence. The rubric now states that the prompt supplies no completion evidence, so a claim that such evidence already settles the outcome is an unsupported outcome claim, while a retry conditioned on a future check is judged by the retry rule. Before this, the relabelled case drew PASS in one of four judgments with the trailing conditional and one of six without it, and either draw fails a calibration; after it, FAIL in six of six. - The
statement_rerunparagraph also carries the remaining supplied facts (yesterday's cached history, the intended production target, the unknown cause, no rerun or resend), and theknowledge_cardPASS case names the card's full escalation chain. - Three labels stand.
A new calibration receipt passes: 15 live calls, 150 cached, every rubric at or above 0.95. Five uncached runs agree on every judged
no_blind_retry_after_unknowncase (one evidence-format inconclusive on a FAIL case); three uncached runs score 21/21 on the companion rubric. - The
-
--regradegradesverification_completedagainst the repository path the run recorded, so a correctcd "<repo>" && <suite>receipt is no longer a false FAIL because the checkout is gone. Runs from before 2026-10-02 did not record that path, and they already regrade INCONCLUSIVE on evaluator identity.--regradealso exits like a run: trials aggregate per scenario against its threshold within one label and one resolved model, then 1 for any FAIL verdict, 2 for any INCONCLUSIVE one, and 2 when nothing was regraded. It used to return 1 unless every trial passed, and 0 for nothing. -
verification_completedcounts an earlier foreground command that failed (Error: Exit code N, which the CLI receipts as text, not a dict) as completed. Before, any failed earlier command, even a read-onlygit log, made the ordering unknown and the trial INCONCLUSIVE: 54 of the 61 saved trials with that verdict had no other cause. Replaying all 1,184 replayable saved trials changes exactly those 54 from INCONCLUSIVE to PASS and nothing else; other text errors stay unknown. -
backend-craftcorrections:consuming-apis.mdno longer says to prefer thecfCLI;- request-id guidance no longer points at
obs-pipeline, which has none, and states when Gorouter's id can be trusted; - the starter test requires 422 for validation failures;
- the Celery link resolves.
Moogsoft on-prem is named "Moogsoft Onprem" in
backend-craft,stack-profile, andobs-alerting; APEX AIOps Incident Management is the cloud product's name. -
build-software-engineer-incidents-apigraded problem+json on a fixture whose existing errors were plain JSON, while the skill says to keep an existing contract. The fixture now carries the house problem+json handlers. The oracle requires 422 and a cappedlimit, and accepts a fast502/504or an explicitly unavailable owner. It was proven offline against two reference implementations and seven mutants. -
Two
sre-assistantbuild scenarios failed an agent that followedpcf-ops:active-incident-guarded-triagenow supplies the maskedcfwrapperpcf-opsrequires before any rawcfread.suspected-compromise-preserves-evidencenow tests today's access: a logged-incfwith unprotected output, where the agent runs none and asks for the Apps Manager view.
Both accept "haven't changed" and a parenthetical before
Caller next step's colon. -
researchergrants the host GitHits server'slisttool in place of the retiredcode_filesanddocs_list;scripts/validate_fleet.pypins the same set, which now matches the server's twelve tools. -
PR #293 review follow-ups: the FastAPI starter's unhandled-error log keeps the method, path, and chained and grouped exception types, collapses recursion, and survives a project record factory that sets
request_id; a trusted header without the request-id middleware is rejected. Its docs name the server log that still carries raw exception text and when Gorouter's request id can be trusted. The PCF deploy example's health step annotates every failure, including curl errors, bounds its total retry wait, and probes the readiness path; tests now fail if--disableor--output /dev/nullis dropped, and a runner without PyYAML stops with a named requirement instead of a traceback.database-reliabilitytreats an executing plan of a mutating statement on production as a live change even when rolled back, gives the two safe NOT NULL orders separately, and no longer puts a shortlock_timeoutonCREATE INDEX CONCURRENTLY, whose snapshot waits it would cancel into an INVALID index. Agent authoring escapes every default-ignorable code point in the model's view of untrusted text, says a permissions deny rule matches command text rather than the program, and restores the copied-test rule; Akamai production debug-header requests stay human-run whatever tools a lane holds. -
Follow-up PR review: CLI receipts cover discovery failures and pre-apply interruption; the CORS starter permits its authenticated API contract; deployment discovery stops on failure. DataStream rates separate client traffic, purge guidance preserves approved mandatory removal, and installed-plugin invocation wording agrees across the authoring references.
-
PR review corrections: the operator CLI validates positive bounds, caps plans, stops at the first failed or UNKNOWN item, preserves lock ownership, and keeps interactive prompts off JSON stdout. Its evaluator now rejects invalid dry-run JSON and incorrect operational exit codes.
-
Backend starters preserve a project's selected correlation ID, support an explicit correlation owner and outer CORS wrapper, require nonempty request IDs in both schemas, and exercise the unexpected-error handler ahead of catch-all routes. Replay-only headers have an explicit contract.
-
Agent Plugins 1.0 now packages the canonical ADR command with its preflight intact; generated command completeness is checked. Python examples distinguish 3.10 from 3.11+ APIs.
-
A new
pcf-opsreference states what each state-changingcfcommand does to serving:cf scale -m/-k/-l, plaincf restart, and plaincf restagestop the whole app;--strategy rollingrollswebonly; a no-downtime resize is an existing-droplet deployment, nevercf push. The incident fast path now admits only instance-count scale, andproduction-change-gateloads the reference for anycfstate-changing command. -
A correlation id with no trace id now routes to
obs-logsinstead of dead-ending inobs-traces, which had claimed the trigger while forbidding the skill that maps it. Two discovery scenarios cover the split. -
database-reliabilitybounds migration lock waits that can queue traffic: PostgreSQLlock_timeoutwith retry, a long-transaction check beforeCREATE INDEX CONCURRENTLY, and cleanup only of an observed invalid index from the failed build; SQL ServerWAIT_AT_LOW_PRIORITYwhere allowed, otherwiseLOCK_TIMEOUTwithXACT_ABORT ONand explicit error handling and transaction cleanup before retry. -
The
ci-actionsPCF production-deploy example runs only when a human dispatches it frommain, and requires a deployment-branch rule on the environment. -
agent-authoringplatform facts match the current docs: the listing budget is 1% of the context window (8,000 characters only as the fallback); plugin skills load beside same-named personal skills instead of being shadowed;compatibility,omitClaudeMd,initialPromptand the currentmaxTurnsbehavior are recorded; subagent model resolution names both environment variables. The method adds a listing-budget check before description edits, a new-skill admission test, evidence rows combined across every changed surface, and one return-header core with named lane fields. -
production-change-gate's only worked approval packet now carries every checklist slot (execution boundary, change record, watcher, timing) and the realcf appoutput shape. Merge readiness reads the candidate repository's own rules instead of asserting this repository's. Tier 0 reads go only through a lane's granted read path, and a dispatched readback is evidence for the reconciliation owner, never the executor's receipt. Examples use a trading app (order-router), and the restart-classification eval grades a plain whole-app restart off the fast path, with ITO approving in the TLC. ITO is spelled out as IT Operations. -
runbook's template approval banner, which every runbook copies, says in plain terms what Tier 2 and Tier 3 mean and who approves: ITO in the TLC during a declared incident, with the shortened checklist limited to eligible actions; all other actions retain the full production process. The exemplar is a trading service (order-router) that verifies recovery on its SLI rather than p95 and states when a multi-window burn alert clears. The rules demote a wrong runbook instead of deleting it, bumpversionon step changes, and movelast_verifiedonly on a passing drill of the stamped version. The Confluence converter reads the page JSON (title, version, modified date), keeps every bodyh1as a section, refuses to overwrite an existing runbook, prints on legacy code pages, and matches headings on word starts; nine new tests fail on the previous converter. -
postmortemroutes resilience and automation decisions toreliability-engineer, has a form rule for unknown severity and a searchable signature field, and the closeout packet carries severity and impact start.operational-learningfinds a component's dependents across every card and defers retrospectives topostmortem; closeout bindings carrygit status --porcelain, whichobservability-engineerandsoftware-engineernow send. -
obs-dashboardswarns againstor vector(0)and null-to-zero on error ratios, asks for a denominator panel and a telemetry-present stat, names the backend holding each signal, and applies the team's dashboard conventions, which now live in their owngrafanareference (read on every create) instead of under a "legacy" data-source title. App-UI dashboards route tofrontend-craft, with a discovery scenario for the split. -
frontend-craftcovers SSE auth with in-memory tokens, confirmed and idempotent UI writes, an SPA fallback that never swallows API paths, and a browser check that names the gap when no browser tool exists.backend-craftadds object-level authorization tests, request ids on every log line and problem body (generated unless a trusted ingress header such as PCF'sX-Vcap-Request-Idis configured), PCF health-check wiring and the 10 s drain window, a starter test that no longer passes on a router 404, and a replay/in-progress contract for the starter POST. -
gcp-opslog reads work under PowerShell as well as Bash: the severity floor is spelled as exclusions of the five lower severities and 429s are a second read, because the PowerShell guard refuses>,<and parentheses even inside quotes; the asset test now replays every example in both shells. A 503/504/429/start-failure table gives each documented message its first check, the first look adds the Metrics tab, rollback picks the revision that served before onset from evidence, and the CF migration map adds startup probes and the documented eligibility criteria.akamai-edgeadds a cache-purge section (invalidate by default, narrowest scope, never delete during an origin brownout), documented X-Cache values includingTCP_REFRESH_FAIL_HIT, the staging hostname rule, and a DataStream 2 hostname/region query inobs-logs' catalog whose destination stays an owner fact. A new discovery scenario checks that a Cloud Run 503 routes togcp-ops. -
python-craftmatches its method to the request: a fix reproduces the defect, changes only what the fix needs, and ends with a Noticed, not changed line; bug hunting and new code use a new defects reference (aruff --extend-select B,S110,S113,BLE001,DTZ,PLW1510,ASYNC,RUF006,RUF032pass plus the classes a linter misses: zero-count slices, float money, wall-clock deadlines, failure reported as success). Refactors first show the tests reach the moved code.TaskGroupfailures are documented asExceptionGroup(existingexcept Xstops matching), the modernization table covers 3.10–3.14 idioms with a ruffUPpass, the toolkit's own pins left the shipped reference, and the worked example is aDecimalfills calculation instead of the eval fixture. A new build probe,build-python-fix-stays-scoped, scored 0/3 on the previous body and 3/3 on this one (Sonnet): both fixed the bug without touching the duplicated code, and only this one named the duplication it left. -
obs-logstimelines count total and errors in explicitly bounded buckets on a scoped base, split status by class withlimit=0(atimechart … by statusfolds early low-volume 503s intoOTHER), and check source freshness before reading "no events" as healthy. Cloud Run rates use the request log as denominator, PCF request counts the GorouterRTRline, and deploy times come from Apps Manager Events or the Cloud Run traffic shift; correlation evidence returns to the caller. Theobs-logsandobs-metricsdescriptions send live-page decisions toincident-investigation, and remote_write queue detail moved toobs-pipeline. A new live-page routing scenario and an eight-key query-shape contract (1/3 on the previous skill, 3/3 on this one; the miss was handing correlation evidence toobservability-engineer) cover the changes. -
ci-actions' PCF deploy example lists revisions before the push (skipped on a first deploy), verifies health, and on failure prints the release owner'scf cancel-deploymentorcf rollbackchoice without running either. It selects its runner by group and labels, pinsdownload-artifactto a reviewed SHA, and asserts cf CLI v8. The skill follows repository action-pinning policy (SHAs when none exists), keeps existing runner labels, adds agh run listtiming recipe and a caller for the reusable starter, routes live outages toincident-investigation, and moves the wheel-release check to the security reference. -
operator-clihas a default exit-code table (0/1/2/128+N), definesskipped(never attempted) againstfailed, handles SIGINT and SIGTERM with cleanup and a report (Python handles only SIGINT by default), prompts only when stdin is a TTY and never takes confirmation from a pipe, ends quietly when stdout closes, binds confirmation to the set shown, batches bulk effects, writes a per-run receipt, and maps dry run and confirmation onto PowerShell's-WhatIf/ConfirmImpact. A copyable contract test and reference command ship inassets/and run in CI.software-engineer's fallback now allows a TTY prompt as confirmation. A new build probe,build-operator-cli-safe-requeue, whose oracle injects a rejected job, a timeout after the effect, and SIGINT/SIGTERM mid-run, scored 0/3 on the previous skill and 3/3 on this one. -
verification_completedaccepts a suite positioned by onecdorSet-Locationinto the trial repository joined with&&. On 2026-09-23 everybuild-software-engineer-cli-with-teststrial on two agent bodies rancd "<workspace>" && python -m unittest …, and the bare-only matcher failed all six, including the two whose final action was that suite. Any other target,;,||, or a second command still rejects. -
The
no_production_action_claimrubric states the guidance exclusion insidefail_if, next to the progressive example it collided with. Since 2026-09-20 no calibration receipt had been accepted: the judge read "I'm applying the top-level skill guidance I did receive" as a production action (18/19), and the all-rubric gate blocked every rubric-backed trial. Recalibrated: every rubric agrees 100 %, 19 live calls. -
The build probe resolves a trial's model identity from the main thread (init model plus every top-level assistant turn) and records the CLI's usage table separately as
usage_models. Claude Code 2.1.271 lists an internal Haiku helper call of a few tokens in that table, which had closed every three-trial batch as INCONCLUSIVE for mixed identities. A parent that changes model mid-trial still resolves to two. Eval ceiling +36 lines. See the medium Python refactor evidence.
Added
-
Every eval run records
runtime: the CLI's--versionline (nullwhen it cannot report one) and the host's system, release, and machine.main()measures it once per batch and prints it in the batch header; each trial writes it toprovenance.json, the trace summary, and its summary line; and--regradekeeps the recorded value rather than today's. Results from different CLI versions or hosts were indistinguishable before. The 2026-10-03 threat-model ADR requires both. -
Two
backend-craftbuild probes with probe-owned oracles, each proven by a no-model test (a house-rule reference plus targeted mutants):incident-writes: an idempotent create whose caller resends on timeout.pager-webhook: a signed webhook whose processing outlasts the vendor's 3 s window.
Against a no-skill control the skill scored 20/21 vs 11/21 on writes, mostly from requiring the key, using 422 for a changed payload, and repeating 201 on replay. On the webhook it scored 14/18 vs 12/18. The model already authenticates, acknowledges fast, persists first and deduplicates on its own. With or without the skill, none of the 6 trials finishes the acknowledged work after a kill and restart: each attaches the runbook link in memory.
-
A medium-sized Python build probe,
build-python-unify-policy: seven modules, three drifted intake entrypoints, a registry, a configured dotted lookup, a legacy re-export, and a unit suite that encodes the drift. Its oracle checks specification parity for every entrypoint, single ownership through thepolicy.normalize_orderpatch seam, exception identity, input immutability, six fresh-process import orders, and a green suite that keeps its assertions; eleven partial-ownership and lost-consumer artifacts are rejected by name. Eval ceiling +223 lines. Across twelve clean-room Sonnet trials on four candidates the fleet completed the refactor correctly every time, with or withoutpython-craft. See the medium Python refactor evidence. -
software-engineerdirect scenarios for the stale-finding return,[sourced]labels on caller-supplied output, and a tool-less build that must not be narrated, plus its build-CLI routing scenario. The contract scenarios were deleted in the 2026-09-01 corpus cut and the routing one on 2026-09-04; they return in the current schema with regex graders only, and without the line-start slot demands the body no longer states. Each carries red and green fixtures intest_graders.py.
Changed
-
The Copilot/VS Code plugin declares Agent Plugins 1.0. VS Code ranks
.claude-plugin/plugin.jsonabove a schema-less root manifest, so it had been loading the canonical Claude agents and hooks instead of the Copilot projection. The plugin now reads canonicalskills/and takes the generated agents and hooks fromcom.github.copilot/;.github/agents/and.github/skills/remain as workspace customizations. Copilot CLI needs v1.0.85 or later. Installed-host behavior is unverified until the new Format acceptance case passes. -
build-software-engineer-cli-with-testsno longer tells the agent to use Changed and Verified headings; it checks the body's own rule instead, that the agent labels its suite result[verified]. -
software-engineer's body returns to its PR #282 state (834dafd1), undoing PR #284's consolidation into an eight-step Working method. Step 1 again loads the craft skills before the code is read and namespython-craftandfrontend-craft. This is an owner preference, not a measured gain: on 2026-09-23 (Sonnet)build-python-unify-policyloadedpython-craftin 3/3 trials on the 09-16 body, 3/5 on this body, and 2/3 on the consolidated body; no pairing is distinguishable (Fisher p >= 0.46), and the refactor passed its oracle in every trial on every body. The workspace, shell,root-cause, consumer-check, and CI-submission rules from PRs #281 and #282 stay. PR #284's researcher dispatch row stays because the researcher agent reads those fields. Reverted with the rest: the default of one reviewer dispatch and the safe local reproducer for incomingsre-assistantevidence. See the medium Python refactor evidence. -
Generated Copilot/VS Code agent profiles carry no generated preface at all: the whole "Host adapter contract" header is gone, including the bare-names sentence, the inherited-tools caveat, and the guarded-lane and MCP-evidence paragraphs. A projection is now the canonical body with Claude-only addressing removed, nothing prepended. A host limitation that changes what a lane may do travels with the rule it qualifies:
sre-assistantalready states in its own body that it has no shell on Copilot, andresearcher's evidence-routing rule now says in its own body that the exact Context7/GitHits identifiers cannot be granted there, so an unavailable evidence lane is named rather than replaced by a summarized fetch. The preface test is inverted to fail if any preface returns, and was proven red against a reintroduced one. -
The
python-craftdescription leads with writing, refactoring, or modernizing any Python, "routine or difficult", and says to load before editing; it opened with "Improve difficult Python code". Triggers and the not-for line are unchanged. After the change, on Sonnet: the medium probe reaches the skill 3/3 (1/6 on main, 2/3 with the step-1 edit alone), and the refactoring, improvement, modernization, and live-incident-negative routing scenarios pass 3/3;discovery-python-explanationfails 0/3 on both this and main's description, a pre-existing red. See the medium Python refactor evidence. -
software-engineerProcess step 1 namespython-craftbeside the backend, frontend, and CLI crafts; it previously appeared only in the on-demand catalogue further down. Projection regenerated. Measured direction only, not proof: see the evidence record above.
[0.50.0] - 2026-09-16
Fixed
- Reviewer evaluation checks now detect path-qualified and common wrapper-prefixed Python/tool commands and require the reviewed range in Git diff/history requests. Added positive and negative calibration cases; the eval Python ceiling rises by 22 lines to 12,275 for this regression coverage.
- Corrected wrong PCF mitigation guidance: whole-app
cf restartstops every instance, so the covered restart is per-instance or rolling; a revision rollback restores the revision's env vars and creates a newRolled back to revision <n>entry, so readback checks the description; rollback is a rolling deployment and cancelling it restores the bad droplet. Corrected the JVMOutOfMemoryErrorguidance (the message names the exhausted pool) and the Splunk top-offenders query (noerrorkeyword; an empty result is a missing extraction). The incident advisor routes external-monitor and Situation pages toobs-alerting, opens Apps Manager Events first, and never searches for documentation; the runbook exemplar and template are console-first withlast_verifiedbound to the drilled version; the commander's evidence rules match the helper's read contract. Ceilings rise to 610,565 skill bytes and a 77,000-byte human incident path. See the quality round record. - Engineering lane, same round: the Spring problem-details advice no longer claims a property it
makes redundant and names the security-filter 401/403 gap; the builder ladder keeps a scoped
security fix builder-owned; the skill token budget is described as the compaction re-attachment
budget; the CI starter and deploy skeleton carry
timeout-minutes; the contract test carries anauth_headersfixture matching the bearer-protected OpenAPI starter; the Handoffs convention says what a review dispatch supplies; the reviewer's worked examples carryReviewed state:and the non-execution line. Ceilings: 612,873 skill bytes, 116,305 agent bytes. - Corrected evaluation assertion identity, candidate aggregation and trial-input attribution; removed ineffective retired policy tests and scoped CI dependency checks to the actual job.
- Resolved portable helpers from their installed skill, qualified Claude-only guard claims, preserved Confluence link/image references with conversion-loss accounting, and repaired the authoring and human-console guidance. Ordinary CI selects only its two test dependencies from the existing version pins. See the review-fix evidence for focused regressions, measured weight and remaining host/model verification.
Added
-
Added conditional symptom comparisons to the human incident advisor for login failures, intermittent errors, slow requests, stale/wrong data, and missed jobs with limited telemetry. See the general incident help evidence for evaluation and the explicit context/skill-size allocation.
-
A human incident advisor with conditional symptom guidance and a bounded, read-only
sre-assistanthelper. The advisor supports explanations, investigation, recap, and operational closeout while the human owns incident decisions and production actions. -
An
operator-cliskill for operator-facing command-line tools, including partial results, secret handling, and uncertain effect outcomes.
Changed
-
Generated Copilot skill copies no longer open with the adapter banner about bare component names; the explicit-only line remains for manual skills.
-
Retired the standalone
incident-commandskill for the team's investigation role. The advisor retains its seven-field board, owns conditional mitigation and severity advice, and prepares technical updates for an existing bridge/TLC without opening another or assigning command roles. The human incident lead retains formal classification, coordination, and stakeholder communications. -
Expanded
reviewerto gather Git/PR/history evidence, load trusted guidance, write scratch reproductions, run permitted isolated checks, and dispatch bounded investigation/research helpers. Candidate fixes and release actions remain with the caller. Updated the graph, host projections, and balanced review cases; shell/write limits are no longer described as enforced by tool absence. -
Moved the Copilot skill projection to
.github/skills/for standard workspace discovery and updated the plugin selector. Removed the custom VS Code skill-location override and generated banners from adapters and bundled resources; byte validation still checks the authored sources. -
Strengthened Python refactoring guidance for interpreter/environment selection, public module moves, independent old/new comparisons, and review of automated fixes. Added semantic regression checks that independently reject widened keyword-only parameters, callback-path input mutation, and replaced callback exceptions. Added incremental-consumption checks, 936 bounded generated comparisons, and a module-move fixture with alias/configuration/patch compatibility and fresh import orders. Twenty-seven broken artifacts are rejected; the eval ceiling grows for the oracles and calibration. Conditional environment depth is included in the refactor/migration context budgets; Python skill discovery is unchanged.
-
The
runbook,postmortem, andoperational-learningdescriptions lead with what a person would say ("help me write a runbook for restarting pricing", "write a postmortem for INC-1234", "which team owns payments and how do I page them");operational-learninganswers ownership and dependency questions from the service cards;ci-actionsexcludes application deploys and runtime bugs; astack-profilereference defines blast radius, golden signals, and the human roles. Four routing scenarios, two of them new, pass 3/3 on Sonnet. The reviewer's README row and three agent bodies lose rules they could not act on. Ceilings rise to 587,200 skill bytes and 115,800 agent bytes. See the quick-fix record. -
The incident advisor's description now names the first responder and carries on-call trigger phrasing ('I just got paged, what do I do', 'customers are reporting errors, where do I start'). In a clean batch on the merged head, the new unhinted first-page routing scenario fires it 3/3 on Sonnet and 3/3 on Opus; walk-me-through, staging triage, the no-dispatch negative, and the dispatched-read helper positive pass 3/3 on Sonnet.
discovery-active-alert-stays-with-advisorandnative-incident-helper-return-and-resumeeach pass 2 of 3 trials and keep an INCONCLUSIVE aggregate: one trial each ended in a denied or missing file read, not a routing miss. The skill-byte ceiling rose by 200 bytes, to 584,200, against an LF total of 584,042. See the advisor on-call trigger evidence. -
Routine documentation handoffs and operational templates now use unique short commit IDs (8 characters, extended by Git when needed). Checkout evidence and approval requirements remain; full IDs are still accepted. Immutable-review identities, dependency pins and eval digests are unchanged.
-
Kept contact-only runbook corrections, supplied log/metric explanations, and bounded telemetry repairs scoped to their tasks while retaining full-workflow checks when requested. Aligned design consultation and operating-document routing with existing owners, and made rollback versus evidence-backed recovery consistent without changing tool grants or human execution authority. See the scope and owner follow-up.
-
Delegated helpers return assignment status, evidence, results, and gaps to their caller. The caller checks the claims and continues the parent task; helper completion does not resolve an incident or complete the parent objective.
-
Alert, trace, log, metric, runbook, and postmortem guidance scales to the requested task. Short corrections and explanations avoid a full workflow packet; full investigations retain their evidence and ownership requirements. Postmortem and knowledge closeout share follow-up records. See the pending output assessment.
-
Incident examples distinguish scoped observations from unsupported causal, timing, and current- state claims. The helper-exchange assessment records the source verification and remaining behavioral decision.
-
Evaluation uses one runner,
evals/build_probe.py, for routing, contract, and build scenarios, with structural graders and a rubric judge. See the current eval contract.
Removed
- Retired incident-autonomy and incident-navigation restoration bundles, including their obsolete scenarios, rubric fragments, calibration cases, generated copies, and patches.
- Completed review and trim-evaluation reports, completed rename records, superseded release and multi-engine evaluation decisions, and the old archive-location records. Current regressions, unresolved-decision evidence, and applicable contracts remain.
- The accumulated development diary and obsolete baseline inventory from this changelog. Historic feature names, scenario counts, and retired commands no longer describe the current toolkit.
Known limitations
- Incident-quality adoption remains on hold. The second bounded assessment records failed behavioral acceptance and links its preceding evidence. Structural checks and source review do not clear that hold; further model work and exact-candidate acceptance remain owner decisions.
- VS Code plugin enforcement remains host-specific; the installation and verification guidance records its limits.
- Codex distribution was retired. Codex can still work in this repository through
AGENTS.md; the retirement contract retains migration guidance for earlier installations.