Skip to content
v0.3.0MIT

Skill that makes your coding agent stop and ask instead of silently resolving spec ambiguity, picking up workarounds, or shipping 'close enough'. Accuracy over autonomy.

Urd

Your AI agent hit a gap in the spec, so it guessed, and shipped. Urd makes it stop and ask instead.

The moment the agent moves from a plan into code and the plan runs out, Urd makes it consult the source of truth or ask you — instead of inventing an answer and burying the decision in the diff.

/plugin marketplace add krzysztofdudek/UrdSkill
/plugin install urd@urd-marketplace

Run both, then /reload-plugins to activate it in this session (or restart Claude Code). No config, no API key. The skill activates on its own the moment the agent moves from a plan, design doc, or spec into implementation.

MIT licensed · single markdown file · works with any agent that reads skills · part of the Yggdrasil family · full skill body


See it

Step 4 of the plan: implement the retry logic. The plan says, verbatim:

"On 5xx response, retry up to 3 times with exponential backoff."

Mid-implementation the agent hits a 429 (rate-limited) — not a 5xx, but retryable. The spec is silent on it. The default agent picks something and moves on. Urd makes it stop:

Question — retry logic, step 4. The plan covers 5xx, but I've hit 429 (rate-limited): not 5xx, yet retryable. The spec doesn't say how to treat it. Three options:

  1. Treat 429 like 5xx — retry with backoff.
  2. Honor the Retry-After header if present, otherwise back off.
  3. Leave 429 to the caller — don't retry it here.

I'd go with (2). Proceed?

Note what it did: it didn't resolve the ambiguity silently. Ten seconds on a decision now, instead of finding it in the diff an hour later. That's the whole skill — at the moment of doubt, ask, don't guess.


When it fires

The skill activates the moment the agent moves from a plan/spec into implementation. Once active, it pushes toward asking rather than guessing whenever:

  • the spec is silent on a case the agent hit
  • the spec contradicts itself
  • the agent is about to add a fallback, exception, or TODO the spec didn't sanction
  • a test would have to be weakened or skipped to make the code pass
  • a framework or type-system constraint is forcing a workaround
  • the agent catches itself thinking "this is probably fine" without certainty

Well-written plans trigger few stops. Sparse plans trigger many. That's the point.


How it asks

When Urd asks, it gives you enough to decide in seconds — never an open-ended "what should I do?":

It tells youExample
Where it is in the plan"Step 4, retry logic"
What the spec says, verbatim"On 5xx, retry 3× with backoff"
What it found that doesn't fit"hit a 429, not covered"
The options, with trade-offsthe numbered list above
Its recommendation"I'd go with (2)"

It also distinguishes verified from believed: it states as fact only what it has checked against the source. An unverified inference — a cause, a mechanism, how an API behaves — it labels as a guess or verifies before you rely on it. Challenge a claim and it re-reads the source and corrects itself, instead of defending an answer it never confirmed.


Install

Claude Code plugin (recommended)

Two slash commands. The first registers this repo as a marketplace; the second installs the plugin from it.

/plugin marketplace add krzysztofdudek/UrdSkill
/plugin install urd@urd-marketplace

Then run /reload-plugins to activate it in the current session (or restart Claude Code). No config, no API key.

To upgrade later, refresh the marketplace and reinstall:

/plugin marketplace update urd-marketplace
/plugin install urd@urd-marketplace

GitHub Copilot CLI plugin

The same repo is also a GitHub Copilot CLI marketplace. Register it, then install the plugin:

copilot plugin marketplace add krzysztofdudek/UrdSkill
copilot plugin install urd@urd-marketplace

To upgrade later: copilot plugin update urd. The same skill body powers both Claude Code and Copilot — nothing changes in how it behaves.

Codex CLI plugin

Codex reads the same skill. Register this repo as a marketplace, then install:

codex plugin marketplace add krzysztofdudek/UrdSkill
codex plugin install urd@urd-marketplace

To upgrade later: codex plugin marketplace upgrade urd-marketplace. Or drop the single file into ~/.agents/skills/urd/SKILL.md (user-level) or .agents/skills/urd/SKILL.md (project-level).

Cursor plugin

Cursor auto-discovers the skill from the plugin manifest at the repo root. Install it locally:

git clone https://github.com/krzysztofdudek/UrdSkill.git
ln -s "$(pwd)/UrdSkill" ~/.cursor/plugins/local/urd

Then reload Cursor (Developer: Reload Window). Or drop the single file into ~/.cursor/skills/urd/SKILL.md (user-level) or .cursor/skills/urd/SKILL.md (project-level).

Single-file drop-in (any agent)

The whole skill is one frontmatter-tagged markdown file: skills/urd/SKILL.md. Copy it into your agent's skill directory.

  • Claude Code, user-level: ~/.claude/skills/urd/SKILL.md
  • Claude Code, project-level: .claude/skills/urd/SKILL.md in your repo
  • Other agents: wherever your tool reads markdown skills

Nothing else in this repo affects behavior — all of it lives in that one file.


What it doesn't claim

Urd is a skill that biases the agent toward asking — not a runtime guarantee it will always stop. It does not eliminate bugs, and it does not fix a vague spec for you: it surfaces the gaps so you can. If what you want is a fast pass with deviations resolved silently, this is deliberately the wrong tool — it trades clarification rounds for fewer wrong outcomes delivered confidently.


FAQ

Won't this make the agent constantly stop?

Only when the spec doesn't cover the case. The skill carries an explicit ask-vs-proceed table — if the spec answers the question, it proceeds; if it doesn't, it asks. Most well-written plans don't trigger many stops. Sparse plans do — and that's the signal you wanted.

Does it conflict with TDD, debugging, or verification skills?

No — it composes. It governs the attitude moving from spec to code, not the process. TDD still says write the test first; Urd says ask when the test you'd write isn't covered by the spec.

What if I'd rather the agent just ship something?

Then don't install it. It's deliberately the opposite of "ship something." A correct outcome after a few clarification rounds beats a wrong outcome delivered without questions — but if you want the fast pass, this is the wrong tool.

Why "Urd"?

Urð is the Norn who keeps the well at the root of Yggdrasil — the source of truth consulted before acting. The skill plays the same role for your codebase: the plan is the source of truth, not the agent's judgment. It's part of the Yggdrasil family.


The Yggdrasil family

Jarl is the loop. Yggdrasil is the law. Grain is the survey. Horde plans the mission onto the law before anyone writes, lands every change through a gate no agent can argue with, and turns what the mission learned into law — on Jarl's loop, with Grain in the architect's hands. Those four are the core, and they ship under one version number, Jarl on it from 6.1.0: one set of tools built and tested against each other. Yggdrasil, Grain and Jarl each work alone; Horde is the one built on the other three. Where a repository has a check, the check decides what lands, in a Horde mission and in a Jarl loop alike: a fresh reviewer can only stop a change, never make a failing check pass, and its word is recorded as testimony. A Jarl loop in a repository with no check lands on testimony alone, and says so. Law that stays inside one component is raised freely by the agent working it. Law or decisions that reach a whole type of code are shared vocabulary: the agent proposes them and they run as advice at once, and the client — the one person the whole system answers to — admits them in one batch when the work closes. Only the client lowers or vetoes law. The core's shared machine contracts are registered on one page.

Start where it hurts; there is no ladder to climb first.

Where it hurtsStart with
More issues than one agent can hold in its headJarl
The agent keeps breaking what was agreedYggdrasil
Nobody knows what was agreedGrain, which earns its keep the moment you are about to write law
A task too big for one head to plan up front, in a repository that already has lawHorde
CoreWhat it holds
JarlThe loop. Everything seen becomes an issue, each issue gets one worker in its own worktree, nothing closes without evidence and a review, and no code closes without a fresh reviewer's word. Rulings keep their history, and a ruling about a whole type of code goes to the client in one batch when the loop closes.
YggdrasilThe law. The architecture graph, the rules over it and the log of why, checked before the agent moves on and re-proved in CI without a key. A rule that reaches a whole type runs as advice until the client ratifies it.
GrainThe survey. Mines a repository's own code and history into a first graph — components, dependencies, and the rules the code already keeps, each with the count of places that break it today; Yggdrasil accepts it with one command. It measures and never blocks.
HordeThe mission on the law. A one-shot architect plans the whole mission onto the graph once, measuring with Grain; a worker per ticket in its own worktree; every change lands through a nine-item gate; what the mission learned becomes law. Its record is a Jarl loop. The client orders the mission and is the only one who can lower or veto a rule.

Three add-ons attach to the agent rather than to the graph; each works alone, depends on nothing in the family and keeps its own version. Horde doesn't assume any of them is installed — it carries its own minimum discipline in each role's law — but uses Ratatoskr, Urd and Researcher when they are, one sentence per row below.

Add-onStageWhat it makes the agent proveIn Horde's loop
Ratatoskrrequest → intentKeeps the agent talking to you in plain words, not code, so you can follow what it's doing.Keeps the client's plain-language registry open at both ends of a mission.
Urd (this one)intent → codeWhen the spec is ambiguous, it consults the source of truth and asks, it doesn't guess.The stop a worker hits before it guesses.
Researchercode → measured resultPoint it at a metric and it runs experiments, hypotheses kept and discarded.Runs the retrospective's measurement.

License

MIT © Krzysztof Dudek