autoeval
Run Plumloom Autoeval evaluations and release gates for AI agents and LLM features from the terminal. Use when the user asks to evaluate or test an agent, prompt, model or RAG change, run evals or an eval suite, check whether a change is safe to release, gate a release, or read evaluation results, scores or a PASS / FAIL / INCONCLUSIVE verdict. Reports the verdict Autoeval returns and never edits evals or thresholds to change it.
- License
- Apache-2.0
- Compatibility
- Requires the autoeval CLI (npm package @plumloom/cli, Node.js 22.13 or newer), AUTOEVAL_API_BASE_URL, a Plumloom CLI key, and network access to the Autoeval API.
Pinned to revision 2239a1dcdce1, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/autoeval/SKILL.md
- skills/autoeval/references/commands.md
- skills/autoeval/references/verdicts.md
Every link opens the file at its source, pinned to the revision this page describes.