Skip to content

ghchinoy/mizan-eval

v0.1.0Apache-2.0

Run Mizan evaluations from an agent: score a single response or compare two responses against a metric template via the mizan CLI, reasoning over the -o json result. Use when you need to evaluate an asset/response against guidance and explain the verdict.

run-eval-set

Run a curated multi-concern Mizan eval-set (an EvalSet manifest) against a shared asset/response, interpret the weighted scorecard — per-member verdicts, the aggregate, the overall PASS/FAIL, and the gate — and honor the gate exit code for CI, driving the mizan CLI over -o json. Use when a user wants to evaluate an asset against a whole suite of metrics at once, run a set of evals, get a single weighted verdict across several concerns, or wire an eval-set into a CI gate.

Version
0.1.0
License
Apache-2.0
Compatibility
Requires the `mizan` CLI on PATH (go install github.com/ghchinoy/mizan/cmd/mizan@latest). Live evals call Vertex AI and need Google Application Default Credentials (ADC) plus a configured project/location; this skill never takes or stores credentials — it relies on the user's existing ADC exactly as the CLI does.
Read SKILL.md at the source

Pinned to revision 4618d919fcec, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.