Skip to content

orq-ai/orq

v2.5.1MIT

Agent skills for building, deploying, evaluating, and monitoring LLM pipelines on the orq.ai platform.

orq-build-evaluator

Create validated LLM-as-a-Judge evaluators following best practices — binary Pass/Fail judges with TPR/TNR validation for measuring specific failure modes. Use when you need to automate quality checks, build guardrails, or measure a specific failure mode identified during trace analysis. Do NOT use when failures are fixable with prompt changes (use orq-optimize-prompt) or when failure modes are unknown (use orq-analyze-trace-failures first).

Read SKILL.md at the source

Pinned to revision 9634e1d956e4, so it is the text this page describes rather than whatever the author pushed since.

Pre-approved tools experimental

Experimental field. Support varies between clients, so this list is what the author declared, not what your client will enforce.

  • Read
  • Write
  • Edit
  • Grep
  • Glob
  • WebFetch
  • Task
  • AskUserQuestion
  • mcp__orq-workspace__get_llm_eval
  • mcp__orq-workspace__get_python_eval
  • mcp__orq-workspace__search_entities
  • mcp__orq-workspace__search_docs

Files

Every link opens the file at its source, pinned to the revision this page describes.