skill-evaluator
Design and evaluate Agent Skill experiments that measure whether a skill improves model behavior, discovery, task success, reliability, cost, or latency relative to a no-skill/prior-version baseline, including host/model comparisons, judge agreement, quality-cost Pareto trade-offs, and runtime drift under frozen measurement identity. Use when the user asks to benchmark, A/B test, evaluate, compare, prove, regress-test, or measure a skill across supported harnesses, or to prepare executable real-host suites when execution is unavailable. Do not use for static package/routing audits (use skill-auditor), to create/edit a skill (use skill-creator), to fabricate real-host results, or to claim universal superiority from one configuration.
Pinned to revision 6626bb65beb0, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/skill-evaluator/SKILL.md
- skills/skill-evaluator/CHANGELOG.md
- skills/skill-evaluator/LICENSE
- skills/skill-evaluator/README.md
- skills/skill-evaluator/VERSION
- skills/skill-evaluator/agents/openai.yaml
- skills/skill-evaluator/assets/icon.svg
- skills/skill-evaluator/evals/cases.json
- skills/skill-evaluator/evals/real-host.json
- skills/skill-evaluator/references/evaluation.md
- skills/skill-evaluator/references/experiment-design.md
- skills/skill-evaluator/references/external-results.md
- skills/skill-evaluator/references/output-contract.md
- skills/skill-evaluator/references/paired-statistics-and-stability.md
- skills/skill-evaluator/references/promotion-policy.md
- skills/skill-evaluator/references/real-host-harness.md
- skills/skill-evaluator/references/runtime-observability.md
- skills/skill-evaluator/references/statistics.md
- skills/skill-evaluator/references/untrusted-input.md
- skills/skill-evaluator/scripts/kernel.py
- skills/skill-evaluator/scripts/run_evals.py
Every link opens the file at its source, pinned to the revision this page describes.