run-eval-set
Run a curated multi-concern Mizan eval-set (an EvalSet manifest) against a shared asset/response, interpret the weighted scorecard — per-member verdicts, the aggregate, the overall PASS/FAIL, and the gate — and honor the gate exit code for CI, driving the mizan CLI over -o json. Use when a user wants to evaluate an asset against a whole suite of metrics at once, run a set of evals, get a single weighted verdict across several concerns, or wire an eval-set into a CI gate.
- Version
- 0.1.0
- License
- Apache-2.0
- Compatibility
- Requires the `mizan` CLI on PATH (go install github.com/ghchinoy/mizan/cmd/mizan@latest). Live evals call Vertex AI and need Google Application Default Credentials (ADC) plus a configured project/location; this skill never takes or stores credentials — it relies on the user's existing ADC exactly as the CLI does.
Pinned to revision 4618d919fcec, so it is the text this page describes rather than whatever the author pushed since.
Files
Every link opens the file at its source, pinned to the revision this page describes.