openai-on-aws/bedrock-bench
Plan agent experiments, benchmark real tasks, compare saved results, diagnose failures, audit graders, and explore charts and tool replays.
Audit benchmark graders with correct solutions, alternate valid answers, malformed artifacts, and deliberate semantic mutants. Use to test false positives, false negatives, or AWS invariants before trusting benchmark results.
Run AWS CDK repair, Terminal-Bench, SWE-bench, AWS-Bench, or starter agent tasks from chat; compare success, latency, token usage, and cost per successful task across Bedrock, OpenAI, OpenRouter, Codex, and OpenCode.
Compare compatible saved Bedrock Bench experiments using task-paired effects, task-cluster uncertainty, complete cost accounting, and chart-ready exports. Use for baseline versus candidate analysis of recorded results.
Turn a benchmark question into a bounded, reproducible Bedrock Bench experiment plan with explicit conditions, runnable configs, task counts, and budget assumptions before execution.
Inventory and triage failures across saved Bedrock Bench runs, with grounded status categories, missing attempts, cost coverage, and evidence citations. Use for recurring failure patterns or failure heatmaps; use inspect-results for one attempt and compare-experiments for comparable performance differences.
Browse the Bedrock Bench run library, replay recorded tools, compare compatible runs, and explain outcomes using saved reports and per-attempt evidence.
Create charts and diagrams from saved Bedrock Bench runs, including cost versus success, timing, token usage, compatible run history, and recorded tool sequences.