Bedrock Bench
Compare AI agents on real tasks and see how often they succeed, how long they take, and what each successful task costs. Run AWS CDK repairs, terminal tasks, and software-engineering benchmarks from a Codex conversation.
Install the plugin once, then select Bedrock Bench in chat and describe what you want to test.
Install
You need the Codex app, a Codex CLI with plugin support, and Python 3.12+. Docker is needed for the CDK, Terminal-Bench, SWE-bench, and AWS-Bench suites. You can try the offline demo without Docker or AWS credentials.
Run these commands in a terminal:
codex plugin marketplace add openai-on-aws/benchmarks-openai \
--ref main \
--sparse .agents/plugins --sparse plugins/bedrock-bench
codex plugin add bedrock-bench@openai-on-aws
Open the project folder where you want to keep your benchmark plans and results. Start a fresh Codex chat and choose Bedrock Bench from the plugin picker. The installed plugin works without cloning this repository.
Update an existing installation
codex plugin marketplace upgrade openai-on-aws
codex plugin add bedrock-bench@openai-on-aws
If your marketplace still follows the old preview branch, remove that
registration and repeat the installation above to follow main:
codex plugin marketplace remove openai-on-aws
Check the installed version with codex plugin list --marketplace openai-on-aws.
Get your first report
With Bedrock Bench selected, ask:
Run the offline demo and explain the results.
You will get a report for six synthetic attempts, including one intentional failure. This shows how success, timing, and cost appear in a report. It makes no model calls and needs no credentials.
Ask for the interactive view:
Open the results explorer. Show the failed attempt and explain its evidence.
Choose a target or task to compare outcomes, timing, and cost. Inspect an attempt's grader result, token usage, settings, and saved files. Supported Codex inline views let you continue with Explain in chat; the standalone HTML report provides a copyable prompt.
To browse your history:
Show my run library. Open the latest live run and replay its tool activity.
Search by model, task, region, or run ID; compare matching runs; then select a model to scrub through its own recorded tools. When a Bedrock Bench skill presents a result, it suggests follow-up prompts. Try Plot cost vs success, Diagram this attempt, or ask for your own graph. Replays currently support saved Harbor/Codex sessions.
Then check a real coding task:
Prepare and run the AWS CDK reference smoke test. Explain what passed.
Codex sets up the benchmark tools and runs the checks in Docker. The broken app should fail and the reference repair should pass. This downloads dependencies and images, but makes no model calls and deploys no AWS resources.
Compare models on Bedrock
For live runs, you need access to the selected Bedrock models. Container agents
also need AWS_BEARER_TOKEN_BEDROCK in the environment that starts the run;
your Codex desktop login alone does not provide it. See
provider setup
for authentication options. Keep credentials out of chat and experiment files.
Start by asking for a plan:
Plan the CDK repair benchmark for GPT-6 Astra, Sol, and Luna on Amazon Bedrock. Use Codex as the runner, low reasoning, and one attempt per model. Show the model IDs, regions, limits, and pricing assumptions. Do not run it yet.
The models being benchmarked are listed in the plan; they are independent of the model you selected for the current chat. The prepared GPT-6 experiments provide smoke and full-suite configurations you can review and reuse.
When you are happy with the plan:
Run that experiment and compare success, time, and cost per successful task.
This step makes paid model calls. Start with the small smoke selection before expanding to a full suite. Current runs have timeouts but no enforced dollar cap.
Choose a benchmark
| Benchmark | What it tests | Available tasks |
|---|---|---|
| Starter | Invoice reconciliation, deployment ordering, and inference-log triage | 3 |
| AWS CDK repair | Fix an SQS → Lambda → DynamoDB app and its message handler | 1 |
| Terminal-Bench 2.0 | Terminal work and repository tasks | 89 |
| SWE-bench Verified | Fix real software issues | 500 |
| AWS-Bench | Complete scenarios in a provisioned AWS testing environment | 9 |
Ask for a specific task or say what you want to measure. Codex starts repository benchmarks with one selected task unless you request a larger set. AWS-Bench requires its own testing environment and model-based grader; review setup and cleanup before using it.
You can also compare Codex with OpenCode, attach an AWS skill to one target, or use the starter's fixed native agent loop to compare model/provider behavior. The suite guide covers those options and which runners each suite supports.
Understand the results
Each run saves an interactive REPORT.html, readable REPORT.md, structured
results, and per-attempt evidence in your project's bench-results directory.
The HTML works offline; its evidence links need the original run files.
Ask Codex to explore a saved run or compare compatible runs:
Compare my latest live run with the previous run using the same task protocol. Show the cost coverage and explain any new failures.
Cost per successful task includes spend on failed attempts. Check cost coverage and whether the figure comes from a rate card or runner estimate. Missing cost remains unknown. Reported agent inference spend excludes AWS infrastructure, subscriptions, and verifier/judge costs.
Reference checks have passed for the default CDK, Terminal-Bench, and SWE-bench tasks. Your first live smoke run still needs to establish model access, agent behavior, and usable accounting. Small smoke runs do not establish a general model ranking.
Explore a research question
The plugin can help you plan an experiment, analyze saved results, and check whether the benchmark is measuring what you intended.
| Ask in chat | Skill |
|---|---|
| “Plan a comparison with and without an AWS skill. Show attempts and budget assumptions.” | Design an experiment |
| “Compare these two targets on the same tasks. Show the differences and uncertainty.” | Compare experiments |
| “Group the failures in these runs and show the evidence for each pattern.” | Diagnose failures |
| “Check that our grader accepts correct solutions and rejects broken ones.” | Audit a benchmark |
These skills include reusable commands, examples, and interpretation guides. Ask for a task heatmap, a paired comparison chart, or a failure breakdown after the analysis. Plans make no model calls; local grader audits are separate from live model measurements. The research command reference also includes an offline walkthrough.
Go further
- Full Codex walkthrough
- CLI and configuration reference
- Cost accounting reference
- Agent operating instructions
- Results inspection skill
- Charts and diagrams skill
- Contributor guidance
- Upstream task sources and licenses
Code: MIT-0. Documentation: CC-BY-SA-4.0.