benchmark-agent-tasks
Run AWS CDK repair, Terminal-Bench, SWE-bench, AWS-Bench, or starter agent tasks from chat; compare success, latency, token usage, and cost per successful task across Bedrock, OpenAI, OpenRouter, Codex, and OpenCode.
Pinned to revision 3ab12aa7c05c, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/benchmark-agent-tasks/SKILL.md
- skills/benchmark-agent-tasks/references/accounting.md
- skills/benchmark-agent-tasks/references/cli.md
- skills/benchmark-agent-tasks/references/follow-ups.md
- skills/benchmark-agent-tasks/references/suites.md
Every link opens the file at its source, pinned to the revision this page describes.