naturali-score-an-agent-change
Measure whether a naturali.ai agent change is good enough to ship with an eval (dataset, llm_judge scorer, pass threshold) run before and after, compared per scorer. Use when asked to test or regression-test an agent, check a prompt change before shipping, compare two eval runs or agent versions, get a pass/fail verdict, read why items failed, or when creating an eval answers 400 VALIDATION_FAILED or a run reports pass_rate null.
- License
- Apache-2.0
Pinned to revision e9daa03587de, so it is the text this page describes rather than whatever the author pushed since.
Files
Every link opens the file at its source, pinned to the revision this page describes.