Skip to content

openai-on-aws/bedrock-bench

v0.3.0MIT-0 AND CC-BY-SA-4.0

Plan agent experiments, benchmark real tasks, compare saved results, diagnose failures, audit graders, and explore charts and tool replays.

audit-benchmark

Audit benchmark graders with correct solutions, alternate valid answers, malformed artifacts, and deliberate semantic mutants. Use to test false positives, false negatives, or AWS invariants before trusting benchmark results.

benchmark-agent-tasks

Run AWS CDK repair, Terminal-Bench, SWE-bench, AWS-Bench, or starter agent tasks from chat; compare success, latency, token usage, and cost per successful task across Bedrock, OpenAI, OpenRouter, Codex, and OpenCode.

compare-experiments

Compare compatible saved Bedrock Bench experiments using task-paired effects, task-cluster uncertainty, complete cost accounting, and chart-ready exports. Use for baseline versus candidate analysis of recorded results.

design-experiment

Turn a benchmark question into a bounded, reproducible Bedrock Bench experiment plan with explicit conditions, runnable configs, task counts, and budget assumptions before execution.

diagnose-failures

Inventory and triage failures across saved Bedrock Bench runs, with grounded status categories, missing attempts, cost coverage, and evidence citations. Use for recurring failure patterns or failure heatmaps; use inspect-results for one attempt and compare-experiments for comparable performance differences.

inspect-results

Browse the Bedrock Bench run library, replay recorded tools, compare compatible runs, and explain outcomes using saved reports and per-attempt evidence.

visualize-results

Create charts and diagrams from saved Bedrock Bench runs, including cost versus success, timing, token usage, compatible run history, and recorded tool sequences.