Save Toolkit
A Claude Code plugin that helps a human SRE do their job on this team's stack: PCF (through Apps
Manager), Splunk, Wavefront and PCF App Metrics, Grafana, Akamai, with Cloud Run migration guidance.
The GCP landing runtime remains undecided in stack-profile. It fits together in
three layers. You own the work: the incident, the change, the runbook. An advisor thinks with
you: incident-investigation asks what to check next, says what each result means, and tells you
when to mitigate. Agents are your helpers, dispatched by you or the invoking workflow for bounded jobs:
sre-assistant performs a read-only lookup or investigation, observability-engineer tunes an alert, scribe writes the runbook
afterward. The skills serve you and the agents alike: the same logs skill hands you a paste-ready
Splunk search and hands the sre-assistant agent the method to build one, which is why a PCF check is always
the Apps Manager view with the cf command beside it. A human executes every production action,
with one narrow exception: an invoked observability-engineer may apply scoped Grafana dashboard,
folder, and alert-rule writes and temporary silences under its
change-authority rule.
Pre-release (0.50.0). Installs track
mainand may change without notice. The repository has no supported immutable release channel.
Install (Claude Code)
claude plugin marketplace add latent-sre/save-toolkit
claude plugin install save-toolkit@latent-sre
Then just describe the problem
Routing is by description; there are no commands to memorize.
- "walk me through INC-4132, checkout is 502-ing since 14:02 UTC" →
incident-investigationsits beside you: what to check next in Apps Manager, Splunk, or Wavefront, what each result would mean, when to mitigate, who to page, and a board so nothing learned is lost. - "why is orders 502-ing in prod?" →
incident-investigationagain; you own it, it advises. - "check cf events and recent logs for orders since 14:00 UTC" → the
sre-assistantagent gathers the requested evidence and limits, then returns to its caller. - "investigate the checkout latency dashboard while I check the database" →
sre-assistantinterprets available evidence, follows related leads within the assignment, and returns findings and recommendations. Missing browser/API access stays explicit; a prompt cannot supply it. - "this alert is too noisy" →
observability-engineerwithobs-alerting. - "write a runbook for the checkout deploy" → the
scribeagent with therunbookskill. - "is PR #42 ready to merge?" →
production-change-gate's merge-readiness checklist, including disposition of any known blocking findings; independent exact-SHA review is reserved for production deployments.
To select the advisor explicitly, start with /save-toolkit:incident-investigation INC-4132 and
describe the symptom. This is a fallback when automatic selection misses, not a requirement for
ordinary use. A direct request for specific counts or logs stays a bounded helper job; asking the
advisor to use a helper keeps the original incident question with the advisor and you.
Before first use: the canonical stack-profile skill bundle declares
this team's stack and pending migration decisions. Every platform-touching skill
routes through it — if that is not your stack, update its entrypoint and matching references first
or the fleet will confidently recommend someone else's tools.
Two entry points are manual only, so type them: /save-toolkit:adr (ADR scaffold) and
/save-toolkit:pcf-deploy (PCF/TAS deploy plan — model invocation is disabled, so nothing
will offer it).
The fleet
| Agent | Lane | Routing |
|---|---|---|
sre-assistant | Bounded read-only lookup or investigation (guarded Bash/PowerShell on Claude), dispatched by a human or invoking workflow | Returns findings and recommendations; the caller retains the broader investigation; delegates only sanitized public fact checks to researcher |
observability-engineer | Observability and dispatched Grafana changes (unguarded Bash; applies scoped dashboards, alert rules, and silences) | Delegates docs to scribe and sanitized public lookups to researcher; active diagnosis stays with the responder using incident-investigation (bounded investigative work to sre-assistant), and automation goes to software-engineer |
scribe | Write evidence-bound runbooks, resolved-incident postmortems, and approved service/application/alert knowledge | Local document writer with no shell, web, external MCP, or delegation authority |
software-engineer | Build, fix, refactor, and test code or operations tooling | Routes requested or risk-triggered review to reviewer, operational docs to scribe, and sanitized public lookups to researcher |
repository-investigator | Local-only answers about private, current, or uncommitted checkout behavior | Cites file:line; no shell, write, web, external MCP, skill, or delegation |
reliability-engineer | Service reliability assessment, resilience design, and toil reduction | Local evidence and design documents; bounded facts from repository-investigator, protected observations from sre-assistant, and sanitized research from researcher; implementation stays with its existing owner |
researcher | Extended public research or bounded external-source help for another agent | Returns cited evidence from official docs, upstream code, packages, and advisories; quick lookups stay with a caller that already has evidence or approved retrieval tools |
reviewer (for maintainers and builders) | Independent investigation, isolated verification, and correctness/security review | Gathers Git/PR evidence and uses focused evidence helpers; owns the verdict, while its caller dispatches fixes to software-engineer |
agent-engineer (for maintainers) | The fleet's prompts, agents, skills, descriptions, evals, bounded prompt/eval loops, roster/delegation graphs, and portable executable workflow-graph designs | Delegates only sanitized public lookups to researcher; the caller separately dispatches helper code to software-engineer and injection-surface review to reviewer |
The skills, by area (each skills/<name>/SKILL.md carries its own description and triggers):
- Incident and operations —
incident-investigation,root-cause,postmortem,runbook,operational-learning,service-lifecycle(audit, onboard, and retire modes) - Observability —
obs-logs,obs-metrics,obs-traces,obs-dashboards,obs-alerting,obs-pipeline,grafana(Grafana UI/API operations and implementation) - Platform —
stack-profile,pcf-ops,pcf-deploy,gcp-ops,akamai-edge - Change gates —
production-change-gate - Reliability improvement —
resilience-analysis,toil-reduction; existing readiness coverage stays inservice-lifecycle - Engineering craft —
backend-craft,python-craft,frontend-craft,operator-cli,ci-actions,database-reliability,eng-ladder - For maintainers: the fleet itself and the graphs it designs —
agent-authoring
The roster's tool postures, enforcement model, and design disciplines are in AGENTS.md; the repository layout and its consequences are under Start here.
How it works
Save Toolkit is a host-loaded control layer, not a background orchestrator. The host supplies the
model and executes tools; the plugin supplies the specialist roles, reusable methods, routing, and
authority boundaries. Claude Code reads the canonical agents/ and skills/
directly; GitHub Copilot/VS Code receives a committed projection from one deterministic generator,
never edited by hand.
agents/ + skills/ (canonical)
|-- Claude Code reads them directly
`-- generator -> .github/agents/ + .github/skills/ -> VS Code/Copilot
- An agent owns a lane with a distinct prompt, tool posture, and return contract.
- A skill adds a method or checklist without changing the current owner.
- Model delegation uses the host's subagent tool to give a bounded task to a named child, which
returns its result to the caller. The calling agent checks that result against the assignment,
combines the evidence, and continues the original task within its authority; the human need not
relay reports between helpers. Canonical Claude
Agent(save-toolkit:target, ...)grants become VS Code'sagenttool plus the parent'sagents:allowlist. - A VS Code handoff is a separate human-selected ownership transition from
handoffs:. It keeps relevant conversation context but does not grant approval or model-delegation authority. - Production-facing or materially irreversible effects remain human decisions. An invoked
observability-engineermay apply requested Grafana dashboard/folder changes, individual Grafana-managed alert-rule create/update or pause/resume, and temporary silence create/update/expire under its complete change-authority rule. Thegrafanaskill covers operational reads and alert/silence procedures and dashboard implementation;obs-dashboardsowns dashboard design. A handoff alone does not authorize a write. - The team's own inventories are not in this repository. The log-index, metrics, PCF-foundation,
and GCP-project references under
skills/obs-logs,skills/obs-metrics,skills/pcf-ops, andskills/gcp-opsship as<app>/<index>placeholders, and service cards, alert cards, and runbooks are read fromoperations/,runbooks/, andpostmortems/under the team's knowledge repository root. Until those are filled in, "where are the dashboards, logs, and runbooks for this service" has no answer here by design; the skills say so rather than guess.
Host guarantees and limits
A field present in an agent file proves what the plugin requested, not what every host enforces. Treat these as build-bound evidence, and rerun the linked probe after host upgrades.
| Host surface | Contract shipped | Current evidence boundary |
|---|---|---|
| Claude Code | Canonical agents and skills load directly; tool absence is the primary role boundary, with the plugin-level read-only Bash guard for sre-assistant | Claude has the richest enforceable contract, but Agent(target) is enforced only on the main thread; subagent-depth restrictions remain documentary. See AGENTS.md |
VS Code 1.135.0 (08d4889f) | Generated agents, skills, model-call agents:, and human-selected handoffs: | [verified] On 2026-08-30, plugin registration, 8 agents, 33 skills, the separate ADR prompt, and a synthetic allowed child passed. A forbidden child still ran, the real software-engineer -> reviewer call was inconclusive, and the separate hook canary was not run. The live transcript was removed in the 2026-09-02 retention pass; recover it with git show e77fc672^:docs/reviews/evidence/host-002/2026-08-30-vscode-plugin-delegation-transcript.md |
| VS Code 1.137.0 | Generated agents, skills, model-call agents:, and human-selected handoffs: | [verified] On 2026-09-10 the maintainer ran the acceptance cases against an installed 0.40.0 candidate and reported all passing, including the forbidden-child case that ran anyway on 1.135.0. Owner-reported from a live session; no transcript was filed, so this row is the record. The separate agent-scoped hook canary remains unrun and hooks/copilot-hooks.json still ships empty |
First installed VS Code build proven to contain d679b159 | Upstream adds prepare/invoke rejection outside agents: and forwards each child's own list | [sourced] The upstream change is merged; [unverified] the installed plugin path until the RELEASE-001 acceptance run passes on that exact build (procedure removed 2026-09-02; recover it with git show e77fc672^:docs/probes/host-002-vscode-agent-delegation.md) |
Other hosts
For read-only observability without MCP, see Windows/macOS command access. Claude's candidate supports a small native command set and the bundled Grafana read/query helper. Selected native VS Code and Playwright browser interactions require a protected read-only session. The standard Copilot profile remains without terminal tools; its command preview must pass the installed-host canary before adoption.
VS Code / Copilot Chat (beta plugin): confirm chat.plugins.enabled is on, run
Chat: Install Plugin From Source, and enter https://github.com/latent-sre/save-toolkit.
VS Code clones the repository and loads it as an Agent Plugins 1.0 plugin, which the root
plugin.json declares: canonical skills/ plus the generated Copilot agents and
hooks in com.github.copilot/. Alternatively, install the same marketplace through GitHub Copilot
CLI v1.0.85 or later (earlier versions do not discover com.github.copilot/agents/); VS Code
automatically discovers Copilot CLI-installed plugins:
copilot plugin marketplace add latent-sre/save-toolkit
copilot plugin install save-toolkit@latent-sre
For an unpublished local branch, use an isolated VS Code profile and register the branch worktree
with chat.pluginLocations instead:
{
"chat.pluginLocations": {
"/absolute/path/to/save-toolkit": true
}
}
Open a neutral test workspace for that plugin check; opening this repository itself also discovers
.github/agents/ and .github/skills/ as workspace customizations and can hide duplicate-install
mistakes. Opening the repository without installing the plugin discovers those standard
directories directly; no custom skill-location setting is needed.
Use the maintained VS Code plugin acceptance procedure for
discovery, installed helpers, delegation, return/resume, and disable/uninstall checks. Its release
section names the immutable-artifact and rollback evidence still required for supported use.
Codex: the fleet is not distributed to Codex. Codex working in this repository picks up the
root AGENTS.md automatically, which is all it needs
(ADR).
Validate
One structural gate (on Windows use python or py -3, never python3 — the Microsoft Store
stub):
python scripts/gate_a.py # the whole structural gate
python scripts/generate_platform_adapters.py --write # after any canonical edit
python scripts/test_platform_adapters.py # Copilot projection + plugin contract
claude plugin validate . --strict # Claude platform contract
Gate A proves the fleet is well-formed, never that it is correct — the change-shaped checks in
CONTRIBUTING.md's verification table are separate. Active behavioral and routing
evals live in evals/README.md. Accepted fleet failures become focused
regressions and ordinary PR evidence.
Browse the fleet atlas
The local browser dashboard provides searchable guidance, an entity registry, relationship views, and cited source excerpts from a verified atlas snapshot. It opens as a standalone HTML file without a server or external assets.
Contribute
Start with AGENTS.md (the fleet guide and conditional rule map, loaded into every
session) and CONTRIBUTING.md (authoring, verification, and promotion policy).
Live and deferred work is tracked solely in docs/fleet-roadmap.md.