three-dev-experiments
Runs an offline experiment in three.dev to test a prompt, model, or reasoning change on recorded production traffic, and reads the results. Use when the user has a change in hand and wants to know whether it is better, cheaper, or faster before shipping, wants to compare two models or providers on their own traffic, asks whether a cheaper or smaller model holds the same quality, wants to A/B test a prompt, or asks what an existing offline experiment showed. Also routes live experiments, which run on real users and start in the three.dev app. Needs the three.dev MCP server. Not for finding out why a feature is failing in production; that is the three-dev-failure-modes skill. Not for wiring calls through the proxy; that is the three-dev-setup skill. Not for defining or reporting the quality metric a live experiment measures; that is the three-dev-quality-metrics-setup skill.
Pinned to revision d19757064e92, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/three-dev-experiments/SKILL.md
- skills/three-dev-experiments/references/results.md
- skills/three-dev-experiments/references/worked-calls.md
Every link opens the file at its source, pinned to the revision this page describes.