Skip to content

grit-77/done-is-a-claim

v1.5.0Apache-2.0

Evidence-first workflows for acceptance, verification, reviews, handoffs and delivery.

Done is a claim

English · Türkçe

Replay stamp · Static cover

Your agent says “done”. What would prove it?

An evidence-first development workflow for coding agents: understand → plan → implement → diagnose → review → verify → deliver. 13 connected skills, 18 incident-backed rules, and optional tools that keep the evidence attached to the work it tested. Small tasks take a shorter path.

Get started · Follow the workflow · Choose a skill · Try the demo

Watch a false green

Try the local example with Python 3.10+:

git clone https://github.com/Grit-77/done-is-a-claim.git
cd done-is-a-claim
python examples/false_green.py

Synthetic demo: the verifier fails with exit 1; the summarizer displays its output and exits 0. The receipt follows the verifier and says FAIL. The demo itself exits 0 because it exposed the mismatch—not because the verifier passed.

Runnable source · Static proof · Recording and limits

Recorded output excerpts, paced for reading. This is a synthetic demonstration, not a production log or an agent-performance benchmark.

Get started

Give your coding agent this request:

Install https://github.com/Grit-77/done-is-a-claim. Then explain simply what changed, what it helps with, and how I start using it.

For the installing agent: follow INSTALL.md. Identify the current host and installation scope, preserve existing work, verify the result, and finish with a short explanation in the user's language.

Manual installation

Install all skills as a native plugin:

Claude Code

claude plugin marketplace add Grit-77/done-is-a-claim --scope local
claude plugin install done-is-a-claim@done-is-a-claim-marketplace --scope local

Codex

codex plugin marketplace add Grit-77/done-is-a-claim
codex plugin add done-is-a-claim@done-is-a-claim-marketplace

Start a fresh session. In Claude Code, invoke /done-is-a-claim:using-done-is-a-claim; in Codex, select using-done-is-a-claim through /skills or the $ skill picker. Plugin scope, updates and validation limits →

Prefer project-local copies or the rules alone? Use the clone from the demo above:

You useAdd to your project
Codex or an agent that reads AGENTS.mdCopy AGENTS.md to the project root.
Claude CodeCopy AGENTS.md and CLAUDE.md, which imports it with @AGENTS.md.
Existing instructionsMerge the relevant rules into your file. Keep your project-specific commands and constraints.

You can start with that one rules file. Native plugin installation does not add these rules to the adopting project's instructions. Installation options on macOS, Linux or Windows →

Want the workflow in your project without overwriting existing files? From this clone, replace ../my-project with your existing project path and preview the copy:

python tools/install_skills.py --project ../my-project --agent codex --profile workflow --dry-run

Then omit --dry-run to install. For Claude Code, use --agent claude-code. Existing skill destinations are refused; your project instructions stay untouched. Use --list to see profiles and individual skills. Explicit --update can replace unmodified installer-owned copies while retaining backups; local edits are refused.

Try this first request in your own project:

Use using-done-is-a-claim for this task. Identify observable acceptance, take a proportionate path through implementation and review, and report the actual result, tested inputs and remaining gaps. Continue within my existing authorization.

This checks whether the agent understood the request. It does not enforce compliance.

One workflow, short paths

Start with using-done-is-a-claim. It routes substantial work through planning and execution, sends unexpected results into diagnosis, and closes review findings with fixes and fresh checks. A clear, small change goes straight from inspection to implementation, review and verification. Workers are optional; an unavailable reviewer is reported as a self-review limit.

See the flow and a worked task →

Want to test your installation on real work? From this repository's clone, create a small broken CSV project, let your installed skill handle it, then check the result independently:

python tools/native_smoke.py create --project ../claim-smoke
python tools/native_smoke.py verify --project ../claim-smoke --json

The initial independent check deliberately fails even though the starter's narrow tests pass. The tool creates and checks the fixture; you run the agent through your normal host and account. Real session procedure and observed limits →

Check the delivered artifact

A successful copy can also deliver the wrong artifact. The delivery demo contrasts a stale copy, the correct copy, and identical bytes that still fail to parse:

python examples/wrong_delivery.py

It uses temporary local fixtures. Exit zero means the traps were demonstrated, not that a real delivery succeeded.

Get a real receipt

Capture a command's own exit, full output and Git state with the optional local tool:

python tools/receipt.py -- python tools/run_tests.py

It writes a fresh receipt.json and command.log under .local/receipts/. Use --input to fingerprint selected files before and after the run. A zero exit records command success; it does not declare the user's task complete. Without selected inputs, the receipt remains incomplete for freshness rechecks. Usage, input scope and limits →

Reusing a saved receipt after an edit or handoff? Compare it with the current project using the exact receipt path printed earlier:

python tools/check_receipt.py path/to/receipt.json --project .

What this rechecks—and what a match cannot prove →

What changes in the report

A claimEvidence that can support it
“The tests passed.”The command, collected cases, actual result, its own exit status and the tested tree.
“That failure was already on main.”Comparable branch and clean-base runs, with matching failure signatures.
“The worker finished.”The worker's actual diff, checked inputs and acceptance result. Delivery is recorded separately.
“The screenshot was saved.”An image that decodes and has been visually inspected.
“Ready to publish.”The reviewed artifact still matches the artifact being published.

Use the completion receipt, acceptance brief and handoff as small, reusable formats. A missing check belongs in the report.

Practical Bash and PowerShell evidence recipes →

Choose a skill

When this happensUse this skillIt helps you decide
You want one entry point for the taskusing-done-is-a-claimWhat is the next useful step, given the task, authority and available evidence?
The change has dependencies or substantial scopeplanning-changesWhich bounded tasks, owners and observations will reach the requested outcome?
A plan has ready, authorized workexecuting-plansWhat can proceed now, and what evidence closes each task?
Behavior contradicts expectationsdebugging-with-evidenceWhich experiment distinguishes the plausible causes?
A change is ready for reviewreviewing-changesDoes the actual diff meet the request, and are findings resolved and rechecked?
The task's “pass” condition is vagueacceptance-designDoes this check cross the boundary the user actually cares about?
A number looks convincingreading-measurementsWhat does this output establish, and what does it leave unknown?
A test fails on your branchwhose-redIs there comparable evidence for a regression, an existing failure or an unresolved cause?
A subagent reports successcollecting-worker-resultsIs there relevant work, was that work tested, and was it delivered?
Work changed after a checkevidence-freshnessDoes the receipt still describe the inputs you are about to act on?
A README or launch post makes a claimpublic-claimsCan the reader trace it, and are the verification limits stated?
You inherit unfinished workresuming-workWhich checkpoint, artifacts and next action still apply in this checkout?
A send or publication reports successchecking-deliveryDoes the actual destination contain the reviewed result, and can its consumer use it?

Each skill is a standalone SKILL.md. No Grit service, account or CLI is needed.

The field rules

The complete wording lives in AGENTS.md. Each original rule has a failure story in INCIDENTS.md.

MomentRules
Before you start01 Define acceptance before the work. 02 Check the paths.
Before you say “done”03 Re-run on the current work. 04 Check the wider suite. 05 Capture the command's own exit code. 06 No tests is no pass. 07 Unknown is not pass. 08 Read every required check.
When you report09 Read the artifact. 10 Open what was written. 11 Look at it yourself. 12 Verify citations. 13 State the base and re-check it before publication.
When you test and fix14 Test behavior, not agreement with a constant. 15 Fix the defect class. 16 Test the detector on a known-good case. 17 Register cleanup before the work.
Always preserve the work18 Never use a destructive command to answer a question.

Where this came from

These rules grew out of running Claude Code and Codex on Grit's own repository. The original internal records report 3,489 task acceptance re-runs between 14–29 September 2026, with 2,282 passing the first independent re-run. A separate internal count reports 736 tasks whose own acceptance passed while the full suite broke on main.

These are author-reported historical observations, not a public benchmark. The raw internal logs are not included. Some unsuccessful re-runs were environment failures; these figures do not establish an agent-error rate or the effectiveness of this toolkit. Claims and limits →

The additional workflows distill Grit's operational runbooks. They are identified as guidance, not newly measured incidents. We also studied how related projects organize installation, skills and verification. Sources and design decisions →

Try to fool it

An agent can repeat a rule and still make the wrong decision. The scenario pack puts the instructions under pressure: a passing pipe, mismatched controls, a worker with no relevant changes, stale evidence and other traps. Participant prompts are separate from evaluator rubrics. Run the prompts in fresh sessions and keep the responses.

These are manual behavioral evaluations, not published success-rate claims.

For the repository itself:

python tools/run_tests.py
python tools/check_repository.py
python tools/check_package.py

The test runner rejects empty or entirely skipped suites. The checks exercise the tools and validate local document targets, skill metadata and imports. CI runs on Windows and Linux. Passing these checks does not prove that an agent follows the instructions.

Rules are not a gate

This repository supplies instructions, small local tools, examples and evaluation material. It does not intercept an agent's tools, block a merge or enforce a deployment policy. Use your project's actual test and release gates for enforcement.

Bring the failure that taught you

A useful contribution starts with what the agent claimed, what was true, and the evidence that revealed the difference. Propose a rule or read the contribution guide.

Tried it in your own project? Share a concrete use report: which piece you used, what happened, and what remains uncertain.

Apache-2.0 · Made by Grit, Ankara.