three-dev-failure-modes
Investigates production quality problems in an LLM feature that three.dev already records. Triages the AI Judge's failure modes, reads the flagged conversations, finds the root cause, and proposes a prompt, model, or code fix. Use when the user asks why an LLM feature is failing or misbehaving in production, what its top failure modes, issues, or quality problems are, whether a problem is growing or started after a deploy, or wants to see bad conversations or count or segment requests. Also when the user arrives with a three.dev conversation id, to continue an investigation started in the three.dev chat. Needs the three.dev MCP server. Not for setting up three.dev or routing calls through the proxy; that is the three-dev-setup skill. Not for testing a change on recorded traffic or reading experiment results; that is the three-dev-experiments skill. Not for defining or reporting quality metrics; that is the three-dev-quality-metrics-setup skill.
Pinned to revision d19757064e92, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/three-dev-failure-modes/SKILL.md
- skills/three-dev-failure-modes/references/filters.md
- skills/three-dev-failure-modes/references/report-template.md
Every link opens the file at its source, pinned to the revision this page describes.