Skip to content
v6.1.0MIT

Run a mission too big for one agent: turns your coding agent into a director that raises a horde — workers in worktrees, one ticket each, a one-shot review of every change that can send work back but never approve it, an architect subagent with a veto over the graph, legislate that writes each territory's law down, and a retrospective that closes the mission out.

Horde

A mission is too big for one agent's context, so you either watch it lose track of its own earlier decisions, or you split it up yourself and babysit every piece. Horde does the splitting for you: it turns your coding agent into a director, and the director raises a horde.

ci

/plugin marketplace add krzysztofdudek/Horde
/plugin install horde@horde-marketplace

Run both, then /reload-plugins to activate it in this session (or restart Claude Code). Requires Node.js on your PATH (any recent version), since the skill's tools are plain ES modules with zero dependencies, Yggdrasil and Grain; see Requirements below before you invoke it. Invoke it by handing over a mission: /horde <mission>, or your own words for it ("let's run this as a horde").

MIT licensed · Node scripts, zero dependencies · needs Yggdrasil and Grain, and creates the graph if your repository has none · part of the Yggdrasil family · full skill body


Requirements

Three things, and only three.

Node.js on your PATH (any recent version). The skill's tools are plain ES modules with zero dependencies.

Yggdrasil 6.1.0 or newer. Horde works on an architecture graph: the map it cuts the work by, the rules every ticket is held to, and the verdict that says a change is safe to merge all come from it. Install it once:

npm i -g @chrisdudek/yg

Horde reads the graph only through Yggdrasil's versioned documents (yg-node/1, yg-context/1, yg-impact/1, and others), and reads a refusal by its code (yg-error/1), never by its wording. A release that doesn't answer with the exact document Horde expects — older than 6.1.0, or one that has since changed shape — is refused by name, naming what it saw and which release to install, never read around.

Grain 6.1.0 or newer. The architect measures with it. It cuts the mission along the places your repository's own history changes together, and each cut comes with a score against random cuts. The rules a part of the mission writes down start from the rules Grain reads out of your code. Your report says what the mission did to its part of the code, before and after. Grain only advises: every number comes with what it was counted out of, and none of them stops anything. Install it (clone it, or /plugin install grain@grain-marketplace after /plugin marketplace add krzysztofdudek/Grain) and tell Horde where it is, once:

horde.mjs init <mission> --base <branch> --grain "node /path/to/Grain/plugins/grain/bin/grain.mjs"

With grain on your PATH there is nothing to name. Horde checks the version before it starts a mission and stops, saying what to install, when Grain is missing or older than 6.1.0.

If your repository already has a graph, Horde reads it. If it doesn't, horde init makes one for you before anything else happens, read out of your own code by Grain: the components you actually have and the rules you already follow. It tells you up front how much of the code you have today those rules would refuse. Without Yggdrasil or Grain, Horde stops and says so.

Jarl comes inside Horde. A mission's record — its tickets, your rulings, the questions it asks you — is a Jarl loop, written through the Jarl record Horde carries with it, so there is nothing more to install. With the Jarl plugin installed as well, jarl resume --root .horde/hordes/<mission> shows the mission live, the same view any Jarl loop gives.

The plugin starts an MCP server named horde by itself, so an agent calls every Horde command as a tool (horde_tick, horde_tk_log, …) with nothing to configure; the same commands stay on the command line for a session without MCP.

That's it.

On Windows, Horde runs the same way, with three things to know. The commands you configure — gates.*, notify, runner.spawn — are sh command lines everywhere, and on Windows they run under the sh.exe that ships with Git for Windows (set HORDE_SH to use another); a machine without one is told so rather than handed cmd.exe, where the same line would mean something else. Write a Windows path in such a command with forward slashes (C:/work/log.txt), or inside double quotes, since sh reads a bare backslash as an escape. A CLI npm installed as a .cmd shim (yg, grain) is started as node on its script, so a ygCommand of plain yg works; any other batch file is refused, because Windows cannot start one without cmd.exe reading its arguments as shell text — name the program it runs instead. And a worker's process started by --runner external is told apart from a later process given the same pid only where ps exists, so on Windows a reused pid reads as the worker still alive until its own landed or stopped line ends it.


See it

The mission: migrate the app's permission system from roles to a policy engine, twelve modules touched. Alone, an agent either tries to hold all twelve modules in its head and drifts by module nine, or works through them one at a time and forgets what it agreed with itself three files back.

You and Horde write the charter together first: the goal, what's out of scope, and the evidence that proves it's done (tests, scenarios, nothing vaguer than that). Then it cuts the mission into territories — sets of whole components, sized so one agent can hold one and still have room to work — and spawns one consultant per territory, all at once. Each reads its own area and nothing else, and writes the tickets and the law it thinks the work has earned. Nothing any of them writes is dispatched yet: an architect rules the whole plan once, before anything starts — what's missing, what's buildable, what's circular, what's grown too big to be one piece.

Workers pick up tickets, each in its own worktree, each against a locked spec. When a worker is done, a one-shot review reads that ticket's change once. It cannot approve anything — it names what is wrong or says nothing — and a serious finding sends the ticket back to a worker before anything else runs; whatever it wrote, the change still has to pass the checklist. Nothing merges by hand: a nine-item checklist runs fresh on the branch itself, and the moment every item is green it makes the merge commit and moves on — new tests proven to fail without the change, the architecture rules read and satisfied, the diff kept inside what the ticket declared. After a wave, a legislate pass reads what that area's own work was refused for and writes the pattern down as a rule, so the next ticket in that area gets it enforced rather than repeated by hand.

You never read a diff. What comes to you: a worker that ran out of spec and stopped rather than guess, a request to weaken a rule, a change to the mission card itself. Everything else, the horde rules on itself and writes down why, so the next session picks it up cold from the files, not from your memory of the conversation. At the very end, a retrospective reads everything nobody read twice — every refusal, every note a worker left the next one — and sorts it into what the law could have said, what's worth one line in a component's own history, and what no rule will ever capture; that last list is yours to read, because it's the one thing the horde cannot learn on its own.

Every line the horde ever merged has a custody chain: point at a file and a line and it tells you the commit that introduced it, the ticket that commit belongs to, who wrote it and what the merge checklist proved before it was let through, what evidence that ticket was supposed to prove and whether it did, and what the graph currently says about the rules standing over that code. A closed mission is searched too; a line from before the horde ever touched the repository is reported as exactly that.


When it fires

Reach for it when the plan itself doesn't fit in one agent's context with room to work, the actual cutting rule the skill uses for sizing a node. A cross-cutting refactor, a migration, anything where "just have the agent do it" means either a context that overflows or a human babysitting every step.

Say /horde <mission>, or hand it over in your own words. A session that resumes an existing horde boots straight into status: what's alive, what's waiting on your ruling, what happens next.


How it thinks

Three planes, two loops:

INTENT   charter, evidence catalogue, asks           you and the director
META     the graph: nodes, ports, contracts, rules   the architect, consultants, legislate
CODE     worktrees, branches, tests, scenarios       workers, reviews

Three functions carry the work, plus a one-shot review of the whole plan, a one-shot review of each ticket's change, and the one client in the loop:

FunctionModelDecidesNever
Directoryour own sessionwhat to ask the client, the charter, build decisions between ticketsreads worker output, merges, dispatches tickets by hand
Consultationthe territory's own class, one per territorywhat a territory's own tickets and law should bethe boundary between territories
Workthe cheapest model that will pass the merge checklistimplementation detail, one ticket at a timecontracts, decisions, other branches
Legislationthe territory's own class, once per territory per wavewhich pattern the code has already earned as a rulelowering a rule
Architect (one-shot)Opus, no node of its ownapproves or vetoes graph changes, rules the whole plan onceimplementation
Review (one-shot)the ticket's own class, once per ticket, before it landsnothing — it names what is wrong or stays silent; a serious finding sends the ticket back before the merge checklist runsapproving anything: the checklist runs whatever it wrote
Clientyouthe mission, every answered ask, whether it shipsreviewing every diff

Every change belongs to exactly one ticket. Liveness is judged by files and branches, never by silence. And it never pushes: starting a mission is your consent to local commits on the horde's own branches (and, through your "go", to paid reviewer runs); nothing else, and never a push. The pull request, and the push, stay yours.


Install

Claude Code plugin (recommended)

Two slash commands. The first registers this repo as a marketplace; the second installs the plugin from it.

/plugin marketplace add krzysztofdudek/Horde
/plugin install horde@horde-marketplace

Then run /reload-plugins to activate it in the current session (or restart Claude Code). Requires Node.js on PATH and Yggdrasil, see Requirements. No API key.

To upgrade later, refresh the marketplace and reinstall:

/plugin marketplace update horde-marketplace
/plugin install horde@horde-marketplace

GitHub Copilot CLI plugin

The same repo is also a GitHub Copilot CLI marketplace. Register it, then install the plugin:

copilot plugin marketplace add krzysztofdudek/Horde
copilot plugin install horde@horde-marketplace

To upgrade later: copilot plugin update horde.

Codex CLI plugin

Register this repo as a marketplace, then install:

codex plugin marketplace add krzysztofdudek/Horde
codex plugin install horde@horde-marketplace

To upgrade later: codex plugin marketplace upgrade horde-marketplace. Or drop the whole skills/horde/ directory into ~/.agents/skills/horde/ (user level) or .agents/skills/horde/ (project level).

Cursor plugin

Cursor auto-discovers the skill from the plugin manifest at the repo root. Install it locally:

git clone https://github.com/krzysztofdudek/Horde.git
ln -s "$(pwd)/Horde" ~/.cursor/plugins/local/horde

Then reload Cursor (Developer: Reload Window). Or drop skills/horde/ into ~/.cursor/skills/horde/ (user level) or .cursor/skills/horde/ (project level).

Manual drop-in (any agent that reads markdown skills)

Copy the whole skills/horde/ directory (SKILL.md, reference/, templates/, scripts/) into your agent's skill directory, keeping the structure intact, since SKILL.md reads its own reference/ and templates/ by relative path.

  • Claude Code, project level: .claude/skills/horde/ in your repo
  • Other agents: wherever your tool reads directory-shaped skills

Nothing else in this repo affects behavior, all of it lives in that one directory.

The skill's own files, and your graph

Dropping a skill into a repository adds several dozen files to it, and a repository under Yggdrasil counts every file: the new ones land outside every node's mapping and show up as uncovered, which drops the coverage figure without anything having gone wrong. Two ways to settle it — map the skill's directory to a node of its own, if you want the graph to hold your tooling to rules as well, or exclude it in yg-config.yaml if you don't. Either is a deliberate answer; leaving it is a number that quietly reads worse than the repository deserves.

The promises rule package

This repository also publishes one set of Yggdrasil rules, promises: an evidence layer for a repository that has none. Horde never installs it for you; it offers it in one sentence when it finds nothing in your repository that proves anything. Installing it takes Yggdrasil 6.1.0 or later:

yg pack add https://github.com/krzysztofdudek/Horde#promises          # the newest published version
yg pack add https://github.com/krzysztofdudek/Horde#promises@6.1.0    # this version, pinned

A version is published as the tag pack/promises@<version>. What is on main after the last such tag is not installed.


What it doesn't claim

It's not free. Every spawned agent is a real run on your account, at whatever model class its ticket carries (Haiku, Sonnet, or Opus for the hard nodes and the rulings). The horde does not track or cap what a mission spends — watch your own account the way you would for any other agent work.

The default runner is your own session: your turn calls the loop, reads what it says to dispatch, and spawns the workers. The loop can also be driven from outside any agent — a cron job or a script of your own, pointed at how your tools start an agent — and then that is what starts each worker instead; nothing about the loop itself changes either way, only who starts what it hands out. The skill installs the same way on Codex, Cursor, and Copilot, and the discipline travels with it (evidence over reports, ask rather than guess, nothing merges without a green checklist), but whether those hosts spawn agents the way this one leans on hasn't been checked at all. Try it there and watch whether the spawning holds before trusting it with something you can't easily undo.

It's not a substitute for reading the result. You get the final branch and the evidence catalogue; whether the mission actually did what you meant is still your call.


FAQ

Won't this spawn agents endlessly and burn my budget?

Dispatch is bounded by config.parallelism and by the queue itself: a wave only hands out what is ready, and the mission stops handing out new work once the queue is empty or a ticket needs your answer. The horde does not track spend on its own, so watch your account the way you would for any other agent work.

Do I have to use Yggdrasil?

Yes, and you don't have to set it up first. Horde works on an architecture graph — that is where the components come from, where the rules over each one come from, and what says a change is safe to merge. On a repository that already has one, Horde reads it and never edits it behind your back. On a repository that doesn't, horde init creates the graph, and Grain reads that first graph out of your own code rather than handing you a blank one. What it will not do is invent a second, weaker map of its own and pretend that is the same thing.

What if I'd rather just have one agent do the whole thing?

Then you may not need Horde at all. Start where it hurts. If the pain is more issues than one agent can hold in its head, Jarl is the loop for that: an issue per finding, a worker per issue in its own worktree, evidence and a fresh reviewer's word before each merge, and, where your repository has a check, that check deciding what lands. Horde is for a task too big for one head to plan up front: hand it any mission and it runs the same charter-first loop whether that mission turns out to need one worker or twelve, on that same Jarl loop, with an architect who plans the whole mission onto the law first and a gate no agent can argue with. A mission whose charter, contracts and code fit in one session with room to work moves through Horde as a single ticket, with no ceremony beyond the charter itself.

Why "Horde"?

Not a Norse name, unlike Ratatoskr and Urd. It says what it does: raise many cheap hands under one will, the same plain-word naming Researcher already uses in this family.


The Yggdrasil family

Jarl is the loop. Yggdrasil is the law. Grain is the survey. Horde plans the mission onto the law before anyone writes, lands every change through a gate no agent can argue with, and turns what the mission learned into law — on Jarl's loop, with Grain in the architect's hands. Those four are the core, and they ship under one version number, Jarl on it from 6.1.0: one set of tools built and tested against each other. Yggdrasil, Grain and Jarl each work alone; Horde is the one built on the other three. Where a repository has a check, the check decides what lands, in a Horde mission and in a Jarl loop alike: a fresh reviewer can only stop a change, never make a failing check pass, and its word is recorded as testimony. A Jarl loop in a repository with no check lands on testimony alone, and says so. Law that stays inside one component is raised freely by the agent working it. Law or decisions that reach a whole type of code are shared vocabulary: the agent proposes them and they run as advice at once, and the client — the one person the whole system answers to — admits them in one batch when the work closes. Only the client lowers or vetoes law. The core's shared machine contracts are registered on one page.

Start where it hurts; there is no ladder to climb first.

Where it hurtsStart with
More issues than one agent can hold in its headJarl
The agent keeps breaking what was agreedYggdrasil
Nobody knows what was agreedGrain, which earns its keep the moment you are about to write law
A task too big for one head to plan up front, in a repository that already has lawHorde
CoreWhat it holds
JarlThe loop. Everything seen becomes an issue, each issue gets one worker in its own worktree, nothing closes without evidence and a review, and no code closes without a fresh reviewer's word. Rulings keep their history, and a ruling about a whole type of code goes to the client in one batch when the loop closes.
YggdrasilThe law. The architecture graph, the rules over it and the log of why, checked before the agent moves on and re-proved in CI without a key. A rule that reaches a whole type runs as advice until the client ratifies it.
GrainThe survey. Mines a repository's own code and history into a first graph — components, dependencies, and the rules the code already keeps, each with the count of places that break it today; Yggdrasil accepts it with one command. It measures and never blocks.
Horde (this one)The mission on the law. A one-shot architect plans the whole mission onto the graph once, measuring with Grain; a worker per ticket in its own worktree; every change lands through a nine-item gate; what the mission learned becomes law. Its record is a Jarl loop. The client orders the mission and is the only one who can lower or veto a rule.

Three add-ons attach to the agent rather than to the graph; each works alone, depends on nothing in the family and keeps its own version. Horde doesn't assume any of them is installed — it carries its own minimum discipline in each role's law — but uses Ratatoskr, Urd and Researcher when they are, one sentence per row below.

Add-onStageWhat it makes the agent proveIn Horde's loop
Ratatoskrrequest → intentKeeps the agent talking to you in plain words, not code, so you can follow what it's doing.Keeps the client's plain-language registry open at both ends of a mission.
Urdintent → codeWhen the spec is ambiguous, it consults the source of truth and asks, it doesn't guess.The stop a worker hits before it guesses.
Researchercode → measured resultPoint it at a metric and it runs experiments, hypotheses kept and discarded.Runs the retrospective's measurement.

Acknowledgements

The law each role carries — tests that can fail, finding the cause before the fix, evidence before the claim, findings with a severity, framing before anything runs — is modelled on obra/superpowers (MIT), by Jesse Vincent, and on its method of testing a document by the behaviour of the agent that reads it. The wording here is Horde's own, and the rules travel as role law enforced at the merge rather than as skills of their own.

License

MIT © Krzysztof Dudek