error-analysis
Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes. Use to learn what to measure before writing evals. Not for CI failures.
- Version
- 1.0.0
- License
- MIT
- Compatibility
- Claude Code 2.1.277+. Needs Langfuse credentials in env or an exported traces JSONL.
Pinned to revision ee47e33b17bd, so it is the text this page describes rather than whatever the author pushed since.
Pre-approved tools experimental
Experimental field. Support varies between clients, so this list is what the author declared, not what your client will enforce.
- AskUserQuestion
- Bash
- Read
- Write
- Edit
- Grep
- Glob
- Agent
- TaskCreate
- TaskUpdate
- TaskList
- TaskGet
- TaskStop
- WebFetch
- WebSearch
Files
- skills/error-analysis/SKILL.md
- skills/error-analysis/references/judge-alignment.md
- skills/error-analysis/references/langfuse-traces.md
- skills/error-analysis/references/method.md
- skills/error-analysis/test-cases.json
Every link opens the file at its source, pinned to the revision this page describes.