ghchinoy/mizan-eval
Run Mizan evaluations from an agent: score a single response or compare two responses against a metric template via the mizan CLI, reasoning over the -o json result. Use when you need to evaluate an asset/response against guidance and explain the verdict.
Run a single (pointwise) or pairwise Mizan evaluation of a response/asset against a metric template and explain the verdict, driving the mizan CLI over -o json and reasoning over the parsed result. Use when a user asks to evaluate, score, grade, or A/B-compare a text response or media asset against a metric, rubric, or guidance, or to understand why an asset passed/failed an eval.
Run a curated multi-concern Mizan eval-set (an EvalSet manifest) against a shared asset/response, interpret the weighted scorecard — per-member verdicts, the aggregate, the overall PASS/FAIL, and the gate — and honor the gate exit code for CI, driving the mizan CLI over -o json. Use when a user wants to evaluate an asset against a whole suite of metrics at once, run a set of evals, get a single weighted verdict across several concerns, or wire an eval-set into a CI gate.