ml-stack-evaluation
Evaluate ML models or systems when a user asks for benchmarks, metrics, slices, calibration, robustness, LLM evaluation, smoke tests, pass/fail thresholds, or go/no-go decisions; preserve revisions and uncertainty.
Pinned to revision 6829bde67fb6, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/ml-stack-evaluation/SKILL.md
- skills/ml-stack-evaluation/agents/openai.yaml
- skills/ml-stack-evaluation/references/benchmark-matrix.md
- skills/ml-stack-evaluation/references/evaluation-report.md
- skills/ml-stack-evaluation/references/metrics-and-slices.md
- skills/ml-stack-evaluation/references/robustness.md
- skills/ml-stack-evaluation/references/smoke-to-scale.md
Every link opens the file at its source, pinned to the revision this page describes.