evaluatorq
Write and run evaluatorq evaluation scripts (Python or TypeScript) for a single agent or deployment — custom scorers, built-in evaluators, and dataset-driven evaluation. For CLI workflows, use the companion skills: orq-red-team for eq redteam adversarial testing and orq-simulate-agent for eq sim multi-turn user simulation. Do NOT use when comparing multiple agents head-to-head (use orq-compare-agents) or when running orq.ai-native experiments only (use orq-run-experiment).
Pinned to revision 9634e1d956e4, so it is the text this page describes rather than whatever the author pushed since.
Pre-approved tools experimental
Experimental field. Support varies between clients, so this list is what the author declared, not what your client will enforce.
- Bash(eq:*)
- Bash(pip:*)
- Bash(python:*)
- Bash(npx:*)
- Read
- Write
- Edit
- Grep
- Glob
- WebFetch
- Task
- AskUserQuestion
- mcp__orq-workspace__search_entities
Files
Every link opens the file at its source, pinned to the revision this page describes.