graph-agents-cli-eval
This skill should be used when the user wants to "run an evaluation", "evaluate my agent", "write an eval dataset", "add an eval case", "analyze eval failures", "compare eval results", "set a quality threshold", "why did eval exit 1", or "upload evals to LangSmith". Covers the enforceable eval gate (one rule, case statuses, exit codes), the dataset schema, deterministic expect checks, judge and quality metrics, the local-versus-disconnected distinction, and eval submit. Applies to any graph-agents-cli project. Do NOT use for agent code (graph-agents-cli-langgraph-code), deployment (graph-agents-cli-deploy), or scaffolding (graph-agents-cli-scaffold).
- Version
- 0.3.1
Pinned to revision 540e2a351795, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/graph-agents-cli-eval/SKILL.md
- skills/graph-agents-cli-eval/references/dataset_schema.md
- skills/graph-agents-cli-eval/references/metrics-guide.md
Every link opens the file at its source, pinned to the revision this page describes.