orq-build-evaluator
Create validated LLM-as-a-Judge evaluators following best practices — binary Pass/Fail judges with TPR/TNR validation for measuring specific failure modes. Use when you need to automate quality checks, build guardrails, or measure a specific failure mode identified during trace analysis. Do NOT use when failures are fixable with prompt changes (use orq-optimize-prompt) or when failure modes are unknown (use orq-analyze-trace-failures first).
Pinned to revision 9634e1d956e4, so it is the text this page describes rather than whatever the author pushed since.
Pre-approved tools experimental
Experimental field. Support varies between clients, so this list is what the author declared, not what your client will enforce.
- Read
- Write
- Edit
- Grep
- Glob
- WebFetch
- Task
- AskUserQuestion
- mcp__orq-workspace__get_llm_eval
- mcp__orq-workspace__get_python_eval
- mcp__orq-workspace__search_entities
- mcp__orq-workspace__search_docs
Files
- skills/orq-build-evaluator/SKILL.md
- skills/orq-build-evaluator/resources/data-split-guide.md
- skills/orq-build-evaluator/resources/judge-prompt-template.md
- skills/orq-build-evaluator/resources/validation-checklist.md
Every link opens the file at its source, pinned to the revision this page describes.