Skip to content

openai-on-aws/bedrock-bench

v0.3.0MIT-0 AND CC-BY-SA-4.0

Plan agent experiments, benchmark real tasks, compare saved results, diagnose failures, audit graders, and explore charts and tool replays.

diagnose-failures

Inventory and triage failures across saved Bedrock Bench runs, with grounded status categories, missing attempts, cost coverage, and evidence citations. Use for recurring failure patterns or failure heatmaps; use inspect-results for one attempt and compare-experiments for comparable performance differences.

Read SKILL.md at the source

Pinned to revision 3ab12aa7c05c, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.