rehearsal
Test and improve AI agents with Rehearsal. Builds a sandboxed practice world from a real application (git repository plus pinned images), runs an agent through jobs where a simulated customer changes their mind and writes time out, scores each attempt against the application's own database, explains failures, and proposes a better agent version. Use when the user wants to evaluate, benchmark, stress-test, compare or improve an AI agent that works inside a business application (help desk, team chat, git hosting, CMS, shop, CRM), wants to test their own agent (LangGraph, OpenAI Agents SDK, custom code) against realistic tasks, or mentions Rehearsal, practice worlds, rehearsals, episodes or held-out evaluation.