run-eval
Run a single (pointwise) or pairwise Mizan evaluation of a response/asset against a metric template and explain the verdict, driving the mizan CLI over -o json and reasoning over the parsed result. Use when a user asks to evaluate, score, grade, or A/B-compare a text response or media asset against a metric, rubric, or guidance, or to understand why an asset passed/failed an eval.
- Version
- 0.1.0
- License
- Apache-2.0
- Compatibility
- Requires the `mizan` CLI on PATH (go install github.com/ghchinoy/mizan/cmd/mizan@latest). Live evals call Vertex AI and need Google Application Default Credentials (ADC) plus a configured project/location; this skill never takes or stores credentials — it relies on the user's existing ADC exactly as the CLI does.
Pinned to revision 4618d919fcec, so it is the text this page describes rather than whatever the author pushed since.
Files
Every link opens the file at its source, pinned to the revision this page describes.