Skip to content

latent-sre/save-toolkit

v0.50.0MIT

Application-engineering and site-reliability agents and reusable skills.

Save Toolkit

A Claude Code plugin that helps a human SRE do their job on this team's stack: PCF (through Apps Manager), Splunk, Wavefront and PCF App Metrics, Grafana, Akamai, with Cloud Run migration guidance. The GCP landing runtime remains undecided in stack-profile. It fits together in three layers. You own the work: the incident, the change, the runbook. An advisor thinks with you: incident-investigation asks what to check next, says what each result means, and tells you when to mitigate. Agents are your helpers, dispatched by you or the invoking workflow for bounded jobs: sre-assistant performs a read-only lookup or investigation, observability-engineer tunes an alert, scribe writes the runbook afterward. The skills serve you and the agents alike: the same logs skill hands you a paste-ready Splunk search and hands the sre-assistant agent the method to build one, which is why a PCF check is always the Apps Manager view with the cf command beside it. A human executes every production action, with one narrow exception: an invoked observability-engineer may apply scoped Grafana dashboard, folder, and alert-rule writes and temporary silences under its change-authority rule.

Pre-release (0.50.0). Installs track main and may change without notice. The repository has no supported immutable release channel.

Install (Claude Code)

claude plugin marketplace add latent-sre/save-toolkit
claude plugin install save-toolkit@latent-sre

Then just describe the problem

Routing is by description; there are no commands to memorize.

  • "walk me through INC-4132, checkout is 502-ing since 14:02 UTC" → incident-investigation sits beside you: what to check next in Apps Manager, Splunk, or Wavefront, what each result would mean, when to mitigate, who to page, and a board so nothing learned is lost.
  • "why is orders 502-ing in prod?" → incident-investigation again; you own it, it advises.
  • "check cf events and recent logs for orders since 14:00 UTC" → the sre-assistant agent gathers the requested evidence and limits, then returns to its caller.
  • "investigate the checkout latency dashboard while I check the database" → sre-assistant interprets available evidence, follows related leads within the assignment, and returns findings and recommendations. Missing browser/API access stays explicit; a prompt cannot supply it.
  • "this alert is too noisy" → observability-engineer with obs-alerting.
  • "write a runbook for the checkout deploy" → the scribe agent with the runbook skill.
  • "is PR #42 ready to merge?" → production-change-gate's merge-readiness checklist, including disposition of any known blocking findings; independent exact-SHA review is reserved for production deployments.

To select the advisor explicitly, start with /save-toolkit:incident-investigation INC-4132 and describe the symptom. This is a fallback when automatic selection misses, not a requirement for ordinary use. A direct request for specific counts or logs stays a bounded helper job; asking the advisor to use a helper keeps the original incident question with the advisor and you.

Before first use: the canonical stack-profile skill bundle declares this team's stack and pending migration decisions. Every platform-touching skill routes through it — if that is not your stack, update its entrypoint and matching references first or the fleet will confidently recommend someone else's tools.

Two entry points are manual only, so type them: /save-toolkit:adr (ADR scaffold) and /save-toolkit:pcf-deploy (PCF/TAS deploy plan — model invocation is disabled, so nothing will offer it).

The fleet

AgentLaneRouting
sre-assistantBounded read-only lookup or investigation (guarded Bash/PowerShell on Claude), dispatched by a human or invoking workflowReturns findings and recommendations; the caller retains the broader investigation; delegates only sanitized public fact checks to researcher
observability-engineerObservability and dispatched Grafana changes (unguarded Bash; applies scoped dashboards, alert rules, and silences)Delegates docs to scribe and sanitized public lookups to researcher; active diagnosis stays with the responder using incident-investigation (bounded investigative work to sre-assistant), and automation goes to software-engineer
scribeWrite evidence-bound runbooks, resolved-incident postmortems, and approved service/application/alert knowledgeLocal document writer with no shell, web, external MCP, or delegation authority
software-engineerBuild, fix, refactor, and test code or operations toolingRoutes requested or risk-triggered review to reviewer, operational docs to scribe, and sanitized public lookups to researcher
repository-investigatorLocal-only answers about private, current, or uncommitted checkout behaviorCites file:line; no shell, write, web, external MCP, skill, or delegation
reliability-engineerService reliability assessment, resilience design, and toil reductionLocal evidence and design documents; bounded facts from repository-investigator, protected observations from sre-assistant, and sanitized research from researcher; implementation stays with its existing owner
researcherExtended public research or bounded external-source help for another agentReturns cited evidence from official docs, upstream code, packages, and advisories; quick lookups stay with a caller that already has evidence or approved retrieval tools
reviewer (for maintainers and builders)Independent investigation, isolated verification, and correctness/security reviewGathers Git/PR evidence and uses focused evidence helpers; owns the verdict, while its caller dispatches fixes to software-engineer
agent-engineer (for maintainers)The fleet's prompts, agents, skills, descriptions, evals, bounded prompt/eval loops, roster/delegation graphs, and portable executable workflow-graph designsDelegates only sanitized public lookups to researcher; the caller separately dispatches helper code to software-engineer and injection-surface review to reviewer

The skills, by area (each skills/<name>/SKILL.md carries its own description and triggers):

  • Incident and operations — incident-investigation, root-cause, postmortem, runbook, operational-learning, service-lifecycle (audit, onboard, and retire modes)
  • Observability — obs-logs, obs-metrics, obs-traces, obs-dashboards, obs-alerting, obs-pipeline, grafana (Grafana UI/API operations and implementation)
  • Platform — stack-profile, pcf-ops, pcf-deploy, gcp-ops, akamai-edge
  • Change gates — production-change-gate
  • Reliability improvement — resilience-analysis, toil-reduction; existing readiness coverage stays in service-lifecycle
  • Engineering craft — backend-craft, python-craft, frontend-craft, operator-cli, ci-actions, database-reliability, eng-ladder
  • For maintainers: the fleet itself and the graphs it designs — agent-authoring

The roster's tool postures, enforcement model, and design disciplines are in AGENTS.md; the repository layout and its consequences are under Start here.

How it works

Save Toolkit is a host-loaded control layer, not a background orchestrator. The host supplies the model and executes tools; the plugin supplies the specialist roles, reusable methods, routing, and authority boundaries. Claude Code reads the canonical agents/ and skills/ directly; GitHub Copilot/VS Code receives a committed projection from one deterministic generator, never edited by hand.

agents/ + skills/ (canonical)
  |-- Claude Code reads them directly
  `-- generator -> .github/agents/ + .github/skills/ -> VS Code/Copilot
  • An agent owns a lane with a distinct prompt, tool posture, and return contract.
  • A skill adds a method or checklist without changing the current owner.
  • Model delegation uses the host's subagent tool to give a bounded task to a named child, which returns its result to the caller. The calling agent checks that result against the assignment, combines the evidence, and continues the original task within its authority; the human need not relay reports between helpers. Canonical Claude Agent(save-toolkit:target, ...) grants become VS Code's agent tool plus the parent's agents: allowlist.
  • A VS Code handoff is a separate human-selected ownership transition from handoffs:. It keeps relevant conversation context but does not grant approval or model-delegation authority.
  • Production-facing or materially irreversible effects remain human decisions. An invoked observability-engineer may apply requested Grafana dashboard/folder changes, individual Grafana-managed alert-rule create/update or pause/resume, and temporary silence create/update/expire under its complete change-authority rule. The grafana skill covers operational reads and alert/silence procedures and dashboard implementation; obs-dashboards owns dashboard design. A handoff alone does not authorize a write.
  • The team's own inventories are not in this repository. The log-index, metrics, PCF-foundation, and GCP-project references under skills/obs-logs, skills/obs-metrics, skills/pcf-ops, and skills/gcp-ops ship as <app>/<index> placeholders, and service cards, alert cards, and runbooks are read from operations/, runbooks/, and postmortems/ under the team's knowledge repository root. Until those are filled in, "where are the dashboards, logs, and runbooks for this service" has no answer here by design; the skills say so rather than guess.

Host guarantees and limits

A field present in an agent file proves what the plugin requested, not what every host enforces. Treat these as build-bound evidence, and rerun the linked probe after host upgrades.

Host surfaceContract shippedCurrent evidence boundary
Claude CodeCanonical agents and skills load directly; tool absence is the primary role boundary, with the plugin-level read-only Bash guard for sre-assistantClaude has the richest enforceable contract, but Agent(target) is enforced only on the main thread; subagent-depth restrictions remain documentary. See AGENTS.md
VS Code 1.135.0 (08d4889f)Generated agents, skills, model-call agents:, and human-selected handoffs:[verified] On 2026-08-30, plugin registration, 8 agents, 33 skills, the separate ADR prompt, and a synthetic allowed child passed. A forbidden child still ran, the real software-engineer -> reviewer call was inconclusive, and the separate hook canary was not run. The live transcript was removed in the 2026-09-02 retention pass; recover it with git show e77fc672^:docs/reviews/evidence/host-002/2026-08-30-vscode-plugin-delegation-transcript.md
VS Code 1.137.0Generated agents, skills, model-call agents:, and human-selected handoffs:[verified] On 2026-09-10 the maintainer ran the acceptance cases against an installed 0.40.0 candidate and reported all passing, including the forbidden-child case that ran anyway on 1.135.0. Owner-reported from a live session; no transcript was filed, so this row is the record. The separate agent-scoped hook canary remains unrun and hooks/copilot-hooks.json still ships empty
First installed VS Code build proven to contain d679b159Upstream adds prepare/invoke rejection outside agents: and forwards each child's own list[sourced] The upstream change is merged; [unverified] the installed plugin path until the RELEASE-001 acceptance run passes on that exact build (procedure removed 2026-09-02; recover it with git show e77fc672^:docs/probes/host-002-vscode-agent-delegation.md)

Other hosts

For read-only observability without MCP, see Windows/macOS command access. Claude's candidate supports a small native command set and the bundled Grafana read/query helper. Selected native VS Code and Playwright browser interactions require a protected read-only session. The standard Copilot profile remains without terminal tools; its command preview must pass the installed-host canary before adoption.

VS Code / Copilot Chat (beta plugin): confirm chat.plugins.enabled is on, run Chat: Install Plugin From Source, and enter https://github.com/latent-sre/save-toolkit. VS Code clones the repository and loads it as an Agent Plugins 1.0 plugin, which the root plugin.json declares: canonical skills/ plus the generated Copilot agents and hooks in com.github.copilot/. Alternatively, install the same marketplace through GitHub Copilot CLI v1.0.85 or later (earlier versions do not discover com.github.copilot/agents/); VS Code automatically discovers Copilot CLI-installed plugins:

copilot plugin marketplace add latent-sre/save-toolkit
copilot plugin install save-toolkit@latent-sre

For an unpublished local branch, use an isolated VS Code profile and register the branch worktree with chat.pluginLocations instead:

{
  "chat.pluginLocations": {
    "/absolute/path/to/save-toolkit": true
  }
}

Open a neutral test workspace for that plugin check; opening this repository itself also discovers .github/agents/ and .github/skills/ as workspace customizations and can hide duplicate-install mistakes. Opening the repository without installing the plugin discovers those standard directories directly; no custom skill-location setting is needed. Use the maintained VS Code plugin acceptance procedure for discovery, installed helpers, delegation, return/resume, and disable/uninstall checks. Its release section names the immutable-artifact and rollback evidence still required for supported use.

Codex: the fleet is not distributed to Codex. Codex working in this repository picks up the root AGENTS.md automatically, which is all it needs (ADR).

Validate

One structural gate (on Windows use python or py -3, never python3 — the Microsoft Store stub):

python scripts/gate_a.py                                # the whole structural gate
python scripts/generate_platform_adapters.py --write    # after any canonical edit
python scripts/test_platform_adapters.py                 # Copilot projection + plugin contract
claude plugin validate . --strict                       # Claude platform contract

Gate A proves the fleet is well-formed, never that it is correct — the change-shaped checks in CONTRIBUTING.md's verification table are separate. Active behavioral and routing evals live in evals/README.md. Accepted fleet failures become focused regressions and ordinary PR evidence.

Browse the fleet atlas

The local browser dashboard provides searchable guidance, an entity registry, relationship views, and cited source excerpts from a verified atlas snapshot. It opens as a standalone HTML file without a server or external assets.

Contribute

Start with AGENTS.md (the fleet guide and conditional rule map, loaded into every session) and CONTRIBUTING.md (authoring, verification, and promotion policy). Live and deferred work is tracked solely in docs/fleet-roadmap.md.