orq-run-experiment
Create and run orq.ai experiments — compare configurations against datasets using evaluators, analyze results, and generate prioritized action plans. Use when evaluating LLM agents, deployments, conversations, or RAG pipelines end-to-end. Do NOT use without a dataset and evaluators. Do NOT use for cross-framework comparisons with external agents (use orq-compare-agents).
Pinned to revision 9634e1d956e4, so it is the text this page describes rather than whatever the author pushed since.
Pre-approved tools experimental
Experimental field. Support varies between clients, so this list is what the author declared, not what your client will enforce.
- Read
- Write
- Edit
- Grep
- Glob
- Task
- AskUserQuestion
- mcp__orq-workspace__list_experiment_runs
- mcp__orq-workspace__get_experiment_run
- mcp__orq-workspace__search_entities
- mcp__orq-workspace__list_datapoints
- mcp__orq-workspace__search_docs
Files
- skills/orq-run-experiment/SKILL.md
- skills/orq-run-experiment/resources/agent-evaluation.md
- skills/orq-run-experiment/resources/anti-patterns.md
- skills/orq-run-experiment/resources/api-reference.md
- skills/orq-run-experiment/resources/conversation-evaluation.md
- skills/orq-run-experiment/resources/rag-evaluation.md
Every link opens the file at its source, pinned to the revision this page describes.