The problem
Agentic coding rarely fails in one step. It fails slowly.
Every change is locally reasonable. The new caller gets its own branch. An unclear state gets a fallback, just in case. Code that feels off-limits gets a wrapper. The old path stays, for compatibility. A review comment gets one more special case. Each diff passes review. Six months later, nobody can say which of the three ways the system saves a setting is the real one.
The agent is not careless. It is timid. It treats every existing line as a constraint, every decision as something to hedge, and every disagreement between two parts of the system as something to reconcile downstream.
Fred Brooks called the missing property conceptual integrity: a system that reflects one coherent set of design ideas. This skill asks the agent to protect it.
Before / after
You ask for file export on OpenHarmony. Your agent opens the export code, finds a switch with one case per platform, and adds another.
// src/scripts/file-export.js, 366 lines
switch (hostPlatform()) {
case 'ios': return shareBlobWithIosRuntime(payload, fileName, { fallbackName });
case 'android': return saveBlobWithAndroidPublicDownloadRuntime(payload, fileName, { fallbackName });
case 'windows': case 'macos': case 'linux': break;
case 'ohos': throw new Error('File export is not implemented for OpenHarmony yet');
}
Behind that switch: 9 staging folders, 7 ways to move bytes, and 4 ideas of
"done" (mode: 'ios-native-share' is one of them). Another case adds to the
pile.
With conceptual integrity, the question becomes: why does export code know about platforms at all? Saving a file is the host's job. So the host does it.
// every export, every platform
const path = await staging.stageBlob(payload, { kind: 'export', preferredName: fileName });
return invoke('deliver_staged_file', { path, fileName }); // { delivered: true | false }
OpenHarmony reuses Android's code and never gets a case. file-export.js went
from 366 lines to 22. The whole change added 1.2k lines and deleted 6.5k.
Real code, from TauriTavern, the project this skill was extracted from.
How it works
1. Facts first read the project's rules; trace the real flow end to end
2. Contracts the hard boundary is what a named consumer depends on;
everything else is changeable implementation
3. Converge shared semantics in one place; fix disagreements at the source
4. Decide recommend one; no flags, fallbacks, or kept-alive paths
to postpone the call
5. Finish migrate every caller, delete the old path, update the docs
6. Tests no test by default; guard behavior, never implementation
7. Proportion simplest complete design; cover 95% with 60% of the complexity
8. Done? one fewer path to learn, and the next change is easier
Taste cuts both ways. The skill refuses the fifth special case, and it also refuses the universal framework nobody needs. Real differences stay explicit. A shared abstraction needs a shared need. A larger change is fine when it removes paths; a rewrite that only moves them is not.
Better with your project's rules
The skill is the method. Your AGENTS.md, CONTRIBUTING.md, and architecture
docs are the facts: which contracts are real, where the boundaries sit, which
trade-offs the project already made. The skill reads them first and defers to
them where they are specific.
A few lines in AGENTS.md go a long way:
## Architecture
- Contracts: the `/api/*` routes, the extension events in docs/Events.md,
the on-disk chat format. Everything else may be refactored.
- Boundaries: docs/Architecture.md.
- Use the conceptual-integrity skill for plans, refactors, new capabilities,
and reviews.
Review mode
Ask for a review ("review this diff for conceptual integrity") and the agent reports one line per finding:
src/api/settings.ts:42: compensation watcher re-syncs the panel after API writes. → route API writes through setSetting().
src/llm/tokens.ts:88: split-path token counting branches per provider in 3 files. → Provider.countTokens().
src/llm/client.ts:15: hedge useLegacyBuilder flag, no caller sets it. → delete it.
net: -3 paths
Tags: split-path, second-truth, compensation, half-move,
special-case, extra-layer, defensive-layer, hedge, test-pile.
Definitions live in
the skill.
Install
Claude Code
/plugin marketplace add Darkatse/conceptual-integrity
/plugin install conceptual-integrity@conceptual-integrity
Codex
codex plugin marketplace add Darkatse/conceptual-integrity
codex plugin add conceptual-integrity@conceptual-integrity
Start a new session afterwards. Inside Codex CLI, /plugins opens the same
plugin browser if you prefer to install from there.
Other agents with skill support
Copy skills/conceptual-integrity/ into your
agent's skills directory: .claude/skills/ for Claude Code without the
plugin, .agents/skills/ for Codex and other hosts that follow the Agent
Skills layout. Commit it to the project so every contributor's agent gets it.
Rule-only agents
For agents that only read a rules file (Cursor rules, Windsurf, Cline, Copilot
instructions), paste the body of
SKILL.md into that file.
FAQ
Does it work with ponytail? Yes. Ponytail shrinks what the agent builds in one change; this keeps the number of ways the system does one thing at one. When they disagree, say a 3-line local patch against a 40-line move that deletes two paths, this skill judges by the cost of the whole system and asks before changing behavior or ownership.
Will my agent start refactoring everything? No. It converges within the semantic scope of the task. When convergence reaches further, it proposes the move in one line and asks, rather than silently expanding the change or silently leaving the split in place.
Will my agent write fewer tests? Yes, on purpose. A test has to guard observable behavior, a real invariant, or a bug that actually happened. A test that pins implementation turns it into a contract, and the next refactor has to fight it. Deleting a test that guards nothing is progress.
Is it always on? No. The agent loads it for planning, refactoring, new capabilities, integrations, and reviews. A typo fix needs no architecture.
Are there benchmarks? Not yet. The case above is one codebase, not a measurement. Maintainability is a property of a sequence of changes, not of one diff. The honest test is a series of tickets on one repository, scored on what ticket N+1 costs and how many parallel paths remain at the end. If you build that harness, please open an issue.