observability-scale
Diagnose slowness, plan scaling, and run production operations. Covers query and endpoint performance, caching layers, read replicas, queues, SLOs and error budgets, logging/metrics/tracing, alerting, incidents, and postmortems. Use this skill whenever the user mentions slow, performance, scaling, bottleneck, timeout, caching, logs, monitoring, alerts, SLO, uptime, incident, outage, postmortem, or asks "can this handle X users", even if they only say "the app is slow" or "we're going viral tomorrow". Not for initial builds (see stack skills), architecture choices (see system-design), or deploy pipelines (see devops-delivery).
Pinned to revision 9bf36daa2e81, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/observability-scale/SKILL.md
- skills/observability-scale/evals/evals.json
- skills/observability-scale/references/incident-response.md
- skills/observability-scale/references/scaling-ladders.md
- skills/observability-scale/templates/postmortem.md
- skills/observability-scale/templates/slo.md
Every link opens the file at its source, pinned to the revision this page describes.