Done is a claim
English · Türkçe
Your agent says “done”. What would prove it?
An evidence-first development workflow for coding agents: understand → plan → implement → diagnose → review → verify → deliver. 13 connected skills, 18 incident-backed rules, and optional tools that keep the evidence attached to the work it tested. Small tasks take a shorter path.
Get started · Follow the workflow · Choose a skill · Try the demo
Watch a false green
Try the local example with Python 3.10+:
git clone https://github.com/Grit-77/done-is-a-claim.git
cd done-is-a-claim
python examples/false_green.py
Synthetic demo: the verifier fails with exit 1; the summarizer displays its output and exits 0. The receipt follows the verifier and says FAIL. The demo itself exits 0 because it exposed the mismatch—not because the verifier passed.
Runnable source · Static proof · Recording and limits
Recorded output excerpts, paced for reading. This is a synthetic demonstration, not a production log or an agent-performance benchmark.
Get started
Give your coding agent this request:
Install https://github.com/Grit-77/done-is-a-claim. Then explain simply what changed, what it helps with, and how I start using it.
For the installing agent: follow INSTALL.md. Identify the current host and installation scope, preserve existing work, verify the result, and finish with a short explanation in the user's language.
Manual installation
Install all skills as a native plugin:
Claude Code
claude plugin marketplace add Grit-77/done-is-a-claim --scope local
claude plugin install done-is-a-claim@done-is-a-claim-marketplace --scope local
Codex
codex plugin marketplace add Grit-77/done-is-a-claim
codex plugin add done-is-a-claim@done-is-a-claim-marketplace
Start a fresh session. In Claude Code, invoke
/done-is-a-claim:using-done-is-a-claim; in Codex, select
using-done-is-a-claim through /skills or the $ skill picker.
Plugin scope, updates and validation limits →
Prefer project-local copies or the rules alone? Use the clone from the demo above:
| You use | Add to your project |
|---|---|
Codex or an agent that reads AGENTS.md | Copy AGENTS.md to the project root. |
| Claude Code | Copy AGENTS.md and CLAUDE.md, which imports it with @AGENTS.md. |
| Existing instructions | Merge the relevant rules into your file. Keep your project-specific commands and constraints. |
You can start with that one rules file. Native plugin installation does not add these rules to the adopting project's instructions. Installation options on macOS, Linux or Windows →
Want the workflow in your project without overwriting existing files? From this clone,
replace ../my-project with your existing project path and preview the copy:
python tools/install_skills.py --project ../my-project --agent codex --profile workflow --dry-run
Then omit --dry-run to install. For Claude Code, use --agent claude-code.
Existing skill destinations are refused; your project instructions stay untouched.
Use --list to see profiles and individual skills. Explicit --update can replace
unmodified installer-owned copies while retaining backups; local edits are refused.
Try this first request in your own project:
Use using-done-is-a-claim for this task. Identify observable acceptance, take a proportionate path through implementation and review, and report the actual result, tested inputs and remaining gaps. Continue within my existing authorization.
This checks whether the agent understood the request. It does not enforce compliance.
One workflow, short paths
Start with using-done-is-a-claim. It routes substantial work through planning and execution, sends unexpected results into diagnosis, and closes review findings with fixes and fresh checks. A clear, small change goes straight from inspection to implementation, review and verification. Workers are optional; an unavailable reviewer is reported as a self-review limit.
See the flow and a worked task →
Want to test your installation on real work? From this repository's clone, create a small broken CSV project, let your installed skill handle it, then check the result independently:
python tools/native_smoke.py create --project ../claim-smoke
python tools/native_smoke.py verify --project ../claim-smoke --json
The initial independent check deliberately fails even though the starter's narrow tests pass. The tool creates and checks the fixture; you run the agent through your normal host and account. Real session procedure and observed limits →
Check the delivered artifact
A successful copy can also deliver the wrong artifact. The delivery demo contrasts a stale copy, the correct copy, and identical bytes that still fail to parse:
python examples/wrong_delivery.py
It uses temporary local fixtures. Exit zero means the traps were demonstrated, not that a real delivery succeeded.
Get a real receipt
Capture a command's own exit, full output and Git state with the optional local tool:
python tools/receipt.py -- python tools/run_tests.py
It writes a fresh receipt.json and command.log under .local/receipts/.
Use --input to fingerprint selected files before and after the run. A zero exit
records command success; it does not declare the user's task complete.
Without selected inputs, the receipt remains incomplete for freshness rechecks.
Usage, input scope and limits →
Reusing a saved receipt after an edit or handoff? Compare it with the current project using the exact receipt path printed earlier:
python tools/check_receipt.py path/to/receipt.json --project .
What this rechecks—and what a match cannot prove →
What changes in the report
| A claim | Evidence that can support it |
|---|---|
| “The tests passed.” | The command, collected cases, actual result, its own exit status and the tested tree. |
| “That failure was already on main.” | Comparable branch and clean-base runs, with matching failure signatures. |
| “The worker finished.” | The worker's actual diff, checked inputs and acceptance result. Delivery is recorded separately. |
| “The screenshot was saved.” | An image that decodes and has been visually inspected. |
| “Ready to publish.” | The reviewed artifact still matches the artifact being published. |
Use the completion receipt, acceptance brief and handoff as small, reusable formats. A missing check belongs in the report.
Practical Bash and PowerShell evidence recipes →
Choose a skill
| When this happens | Use this skill | It helps you decide |
|---|---|---|
| You want one entry point for the task | using-done-is-a-claim | What is the next useful step, given the task, authority and available evidence? |
| The change has dependencies or substantial scope | planning-changes | Which bounded tasks, owners and observations will reach the requested outcome? |
| A plan has ready, authorized work | executing-plans | What can proceed now, and what evidence closes each task? |
| Behavior contradicts expectations | debugging-with-evidence | Which experiment distinguishes the plausible causes? |
| A change is ready for review | reviewing-changes | Does the actual diff meet the request, and are findings resolved and rechecked? |
| The task's “pass” condition is vague | acceptance-design | Does this check cross the boundary the user actually cares about? |
| A number looks convincing | reading-measurements | What does this output establish, and what does it leave unknown? |
| A test fails on your branch | whose-red | Is there comparable evidence for a regression, an existing failure or an unresolved cause? |
| A subagent reports success | collecting-worker-results | Is there relevant work, was that work tested, and was it delivered? |
| Work changed after a check | evidence-freshness | Does the receipt still describe the inputs you are about to act on? |
| A README or launch post makes a claim | public-claims | Can the reader trace it, and are the verification limits stated? |
| You inherit unfinished work | resuming-work | Which checkpoint, artifacts and next action still apply in this checkout? |
| A send or publication reports success | checking-delivery | Does the actual destination contain the reviewed result, and can its consumer use it? |
Each skill is a standalone SKILL.md. No Grit service, account or CLI is needed.
The field rules
The complete wording lives in AGENTS.md. Each original rule has a failure story in INCIDENTS.md.
| Moment | Rules |
|---|---|
| Before you start | 01 Define acceptance before the work. 02 Check the paths. |
| Before you say “done” | 03 Re-run on the current work. 04 Check the wider suite. 05 Capture the command's own exit code. 06 No tests is no pass. 07 Unknown is not pass. 08 Read every required check. |
| When you report | 09 Read the artifact. 10 Open what was written. 11 Look at it yourself. 12 Verify citations. 13 State the base and re-check it before publication. |
| When you test and fix | 14 Test behavior, not agreement with a constant. 15 Fix the defect class. 16 Test the detector on a known-good case. 17 Register cleanup before the work. |
| Always preserve the work | 18 Never use a destructive command to answer a question. |
Where this came from
These rules grew out of running Claude Code and Codex on Grit's own repository. The original internal records report 3,489 task acceptance re-runs between 14–29 September 2026, with 2,282 passing the first independent re-run. A separate internal count reports 736 tasks whose own acceptance passed while the full suite broke on main.
These are author-reported historical observations, not a public benchmark. The raw internal logs are not included. Some unsuccessful re-runs were environment failures; these figures do not establish an agent-error rate or the effectiveness of this toolkit. Claims and limits →
The additional workflows distill Grit's operational runbooks. They are identified as guidance, not newly measured incidents. We also studied how related projects organize installation, skills and verification. Sources and design decisions →
Try to fool it
An agent can repeat a rule and still make the wrong decision. The scenario pack puts the instructions under pressure: a passing pipe, mismatched controls, a worker with no relevant changes, stale evidence and other traps. Participant prompts are separate from evaluator rubrics. Run the prompts in fresh sessions and keep the responses.
These are manual behavioral evaluations, not published success-rate claims.
For the repository itself:
python tools/run_tests.py
python tools/check_repository.py
python tools/check_package.py
The test runner rejects empty or entirely skipped suites. The checks exercise the tools and validate local document targets, skill metadata and imports. CI runs on Windows and Linux. Passing these checks does not prove that an agent follows the instructions.
Rules are not a gate
This repository supplies instructions, small local tools, examples and evaluation material. It does not intercept an agent's tools, block a merge or enforce a deployment policy. Use your project's actual test and release gates for enforcement.
Bring the failure that taught you
A useful contribution starts with what the agent claimed, what was true, and the evidence that revealed the difference. Propose a rule or read the contribution guide.
Tried it in your own project? Share a concrete use report: which piece you used, what happened, and what remains uncertain.
Apache-2.0 · Made by Grit, Ankara.