benchmark-curator
Design, curate, version, and quality-control benchmark or evaluation corpora for Agent Skills and quality workflows, including task taxonomies, discovery/forced/negative controls, adversarial and regression cases, difficulty strata, holdout isolation, provenance, duplication, contamination risk, coverage balance, and immutable benchmark hashes. Use when the user asks to build or maintain an eval dataset, golden set, benchmark suite, regression corpus, holdout, challenge set, or representative test cases. Do not use to execute the model experiment or claim lift (use skill-evaluator), to design the grading rubric (use rubric-designer), to edit the candidate skill (use skill-creator), or to treat leaked/known cases as a clean holdout.
Pinned to revision 6626bb65beb0, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/benchmark-curator/SKILL.md
- skills/benchmark-curator/CHANGELOG.md
- skills/benchmark-curator/LICENSE
- skills/benchmark-curator/README.md
- skills/benchmark-curator/VERSION
- skills/benchmark-curator/agents/openai.yaml
- skills/benchmark-curator/assets/icon.svg
- skills/benchmark-curator/evals/cases.json
- skills/benchmark-curator/evals/real-host.json
- skills/benchmark-curator/references/benchmark-model.md
- skills/benchmark-curator/references/coverage-and-balance.md
- skills/benchmark-curator/references/evaluation.md
- skills/benchmark-curator/references/holdout-and-contamination.md
- skills/benchmark-curator/references/leakage-detection.md
- skills/benchmark-curator/references/output-contract.md
- skills/benchmark-curator/references/untrusted-input.md
- skills/benchmark-curator/scripts/kernel.py
- skills/benchmark-curator/scripts/run_evals.py
Every link opens the file at its source, pinned to the revision this page describes.