Research Orchestrator
A lightweight, hypothesis-driven research workflow for one or many AI agent sessions.
It deliberately uses only four shared Markdown files while adding a scored hypothesis queue, one-time cross-host verification, adaptive workers, and CPU/GPU/Other resource routing.
Why use it?
- Resume long-running research across fresh sessions.
- Let multiple agents explore independently without sharing unfinished plans.
- Keep
plan.mdsmall by removing completed work. - Never queue the same idea twice: every completed experiment leaves a discovery — positive, negative, or inconclusive — and every new item is checked against discoveries and other agents' queues.
- Turn experiment results into reusable shared discoveries.
- Have a different host check each finding once — Codex's discovery is verified by Claude Code, and vice versa — instead of every session re-reviewing it.
- Preserve promising but incomplete findings with
HOLD. - Re-open discoveries as stronger hypotheses with explicit evidence and improvements.
- Scale from two workers to the machine's practical full load.
Core workflow
flowchart LR
D[DISCOVERIES<br/>shared evidence<br/>cross-checked once by another host]
H[New hypotheses<br/>Sources + Evidence + Improvement]
P[PLAN<br/>scored live queue]
W[Adaptive workers<br/>CPU / GPU / Other]
E[Experiments]
O[HANDOFF<br/>history + artifacts + blockers + next state]
D -->|0..N follow-ups| H
H --> P --> W --> E --> D
P --> O
E --> O
D --> O
A discovery can generate zero, one, or many new hypotheses. One hypothesis can also combine evidence from multiple discoveries.
The four files
flowchart TD
A[AGENTS.md<br/>rules, scoring, review, routing]
P[PLAN.md<br/>active unfinished work only]
D[DISCOVERIES.md<br/>shared reusable findings]
H[HANDOFF.md<br/>operational ledger]
A --> P
P -->|completed result| D
P -->|execution history| H
D -->|new evidence / review| P
| File | Purpose |
|---|---|
agents.md | Stable rules for reading, editing, scoring, review, and resource routing. |
plan.md | Only active unfinished work. Each agent owns its own section and works from a scored queue. |
discoveries.md | What is known: claims, numbers, interpretation, and verification. Every completed experiment leaves one entry; each is cross-checked once by a different host. |
handoff.md | What happened and where to resume: who, when, which host, artifact paths, project-wide values such as the current best, and each slot's resume point. Points to discoveries instead of repeating them. |
Completed items leave plan.md. Each one leaves a discovery (whatever the outcome) and a handoff event that points to it, and its follow-up hypotheses return to plan.md with fresh scores.
Templates and a worked example
Each file is created from a template. The worked example shows the same four files in the middle of a real project: two active slots (A on Claude Code, B on Codex) and one released slot (C), a VERIFIED discovery, a discovery under review, a negative result logged with zero follow-ups and the reason, a same-host take-over, and the handoff log that produced them.
| File | Template | Worked example |
|---|---|---|
agents.md | AGENTS.md.template | agents.md |
plan.md | PLAN.md.template | plan.md |
discoveries.md | DISCOVERIES.md.template | discoveries.md |
handoff.md | HANDOFF.md.template | handoff.md |
Install
The repository packages the same skill for Claude Code, Codex, and Google Antigravity.
Repository:
https://github.com/TaeyanG4/research-orchestrator-skill
Claude Code
Plugin marketplace install
Inside Claude Code:
/plugin marketplace add TaeyanG4/research-orchestrator-skill
/plugin install research-orchestrator@research-orchestrator
Start a new Claude Code session after the first install.
Project-only skill install
Clone or copy:
skills/research-orchestrator-skill/
into:
<project>/.claude/skills/research-orchestrator-skill/
Codex
Plugin marketplace install
From a terminal:
codex plugin marketplace add TaeyanG4/research-orchestrator-skill
codex
Inside Codex:
/plugins
Select the Research Orchestrator marketplace and install research-orchestrator. Start a new chat before first use.
Direct skill install
Repo-scoped:
<project>/.codex/skills/research-orchestrator-skill/
User-scoped:
~/.codex/skills/research-orchestrator-skill/
Copy the repository's skills/research-orchestrator-skill/ directory into one of those locations.
Google Antigravity
Clone the repository:
git clone https://github.com/TaeyanG4/research-orchestrator-skill.git
Then copy skills/research-orchestrator-skill/ to one of these locations.
Project/workspace scope:
<project>/.agents/skills/research-orchestrator-skill/
Global Antigravity scope:
~/.gemini/config/skills/research-orchestrator-skill/
Antigravity CLI legacy/global location:
~/.gemini/antigravity-cli/skills/research-orchestrator-skill/
Use /skills in Antigravity CLI to confirm discovery.
Quick start
Initialize the four project files:
python <installed-skill>/scripts/init_research_orchestrator.py . -n "My Project"
Start with multiple independent agents (2 → A, B):
python <installed-skill>/scripts/init_research_orchestrator.py . -n "My Project" --agents 2
--agents also accepts an explicit list such as A,B,C. The initializer rejects duplicate or non-standard names, creates missing files only, and never overwrites existing project files.
Then run python <installed-skill>/scripts/check_project.py . at every session start and before closing (see Consistency check below).
Agent names
An agent name is a work slot, not the tool or model running it. Every host — Claude Code, Codex, Antigravity — uses the same names:
A, B, C, ... Z, AA, AB, ... ZZ
- Single-agent work always uses
A; each additional concurrent session takes the next unused letter, continuing with two-letter names afterZ. Released slots are reused first, so new letters appear only when every existing slot is taken at once. - Never name an agent
Claude,Codex,GPT,Gemini, or any other host/model name. - Any host may resume any slot. Which host owns a slot is recorded in
handoff.md, not in the name:
### Agent: A
- Current host: Claude Code
- Current thread: H-A-04
...
### Agent: B
- Current host: Codex
- Current thread: H-B-02
...
A slot is free when Current host reads unassigned or released. A new session takes the first free slot in order — skipping one that still holds plan items left by a different host — sets Current host to its own host, and sets it back to released when it closes — so liveness is read from the file, never guessed.
Each completed-log event also records its Host, so the history shows which host did each step even after a slot changes hands.
IDs embed the owning agent so concurrent agents never collide:
| Object | Format | Examples |
|---|---|---|
| Plan item | H-<Agent>-<NN> | H-A-01, H-B-07 |
| Discovery | D-<Agent>-<NNN> | D-A-001, D-C-014 |
Follow-ups and take-over
- Plan follow-ups from every discovery. Each time an agent records a discovery — a new finding, a negative result, or a cross-check verdict — it decides what to test next: zero, one, or many new plan items, each passing the duplicate check. It also rescores or removes its own items the discovery affects. Zero is a valid answer, but it is logged with a reason:
New plan items: none — <reason>. - When a queue runs out, the agent looks for work in this order: re-read its own section → claim a
PENDINGcross-check from another host → take over an item from a same-host slot → cross-host take-over by judgment (one item at most) → derive new hypotheses → note it and release the slot. - Take-over stays within one host by default. A host is the platform, not the model: two Claude Code sessions on different models are the same host. If
A's queue is empty andC(same host) still has queued items,Amoves the highest-priority one into its own section under its own next ID —H-C-02becomes### H-A-04 — … (from H-C-02)— and logs the move. It never takes the item in the owner'sCurrent threador one under anActive computeclaim. - Cross-host take-over by judgment. Items queued by a different host normally stay with that host. As an exception, an agent may take over one of them when everything closer is exhausted, the source slot is released (never an active session on another host), the item has
Priority≥ 15 and clearly beats any new hypothesis, and the agent logs a one-line reason:took over H-C-01 from C (cross-host: Codex → Claude Code; reason: …) as H-A-12. After finishing it, the agent starts the order again. The user can change the floor, forbid this, or approve specific items.
Reading rule
Each agent reads:
agents.md
→ all discoveries.md
→ shared handoff log + its own active handoff
→ only its own detailed PLAN section
Agents do not read another active agent's detailed PLAN just to coordinate work. The only cross-section peeks are a dispatcher reading task metadata, the duplicate check reading item headings and Hypothesis lines, a session reading Current host and Current thread lines to find a free slot or a take-over source, and the one item a session takes over.
Standard PLAN format
Keep only active unfinished items.
### H-A-07 — Separate duplicate leakage from group leakage
- Sources: D-B-014, D-A-021
- Hypothesis: exact duplicates explain most apparent group leakage
- Evidence: D-B-014 weakens after deduplication; D-A-021 identifies repeated rows
- Improvement: isolate exact duplicates before constructing candidate groups
- Impact: 3
- Information: 3
- Confidence: 2
- Unblock: 3
- Diversity: 2
- Cost: 1
- Priority: 21
- Resource: CPU
- Parallel: YES
- Other: NONE
- Next test: compare random CV and GroupKFold after exact-duplicate removal
Priority formula
Score every factor from 0 to 3.
Priority = 2*Impact + 2*Information + Confidence + Unblock + Diversity + (3-Cost)
Run the highest score first unless blocked or explicitly overridden. Break ties by lower Cost, then higher Information.
When a discovery returns to PLAN, these fields are mandatory:
Sources— which discoveries motivated it.Evidence— why it deserves another test.Improvement— what is materially different or stronger than the prior attempt.
Do not simply rerun an old idea under a new task ID.
Duplicate check before adding an item
- Search
discoveries.md— every completed experiment is there, positive, negative (Finding: <claim> does not hold under <conditions>), or inconclusive. SkipVERIFIEDclaims unless you have a realImprovement; skipCHALLENGEDones until resolved. - Scan the other agents' plan sections, reading only item headings and
Hypothesislines. If it is already queued, do not add it. - Re-read
plan.mdright before writing; if the same item appeared meanwhile, keep the earlier one.
Standard DISCOVERIES format
## D-B-014 — Random CV may leak groups
- Source: B
- Host: Codex
- Cross-check: HOLD
- Finding: duplicated groups cross random folds
- Evidence: e014_group_check.py; random CV 0.9162 vs group CV 0.9027
- Implication: current validation may be optimistic
- Reviews:
- Claude Code (A): HOLD — plausible, but exact duplicates must be separated first
Verification is per host, not per agent. A discovery made on Codex is checked once by Claude Code (or another different host), and vice versa. Sessions on the same host share the same blind spots, so they do not re-review each other — ten Claude Code sessions never review the same finding ten times.
flowchart LR
N[New discovery<br/>Host: Codex] --> P[Cross-check: PENDING]
P -->|a Claude Code session claims it| R[REVIEWING Claude Code]
R -->|CLOSED| V[VERIFIED]
R -->|HOLD| H[HOLD]
R -->|CHALLENGED| C[CHALLENGED]
H -->|source adds evidence| P
C -->|source revises| P
Cross-check states:
PENDING— no different host has reviewed it yet.REVIEWING <Host> (<Agent>)— a different-host session claimed the review, so nobody else duplicates it.VERIFIED— a different host reviewed it and recordedCLOSED.HOLD— plausible, but specific evidence or an improvement is needed first.CHALLENGED— a material contradiction, flaw, or missing assumption was found.
Rules:
- The source never reviews its own discovery, and same-host sessions do not review each other.
- One cross-host review is enough; add another only for high-impact or disputed findings.
HOLDandCHALLENGEDreviews must state what would make the discovery acceptable. When the source revises it,Cross-checkreturns toPENDING.VERIFIEDdiscoveries may be used freely. Building on an unverified one must be stated in the plan item'sEvidence;CHALLENGEDones are not used until resolved.- With only one host available, a different slot on the same host may cross-check, starting its reason with
same host —.
Standard HANDOFF format
### 2026-10-05 21:10 — A — H-A-07
- Host: Claude Code
- Action: cross-checked D-B-014; removed exact duplicates and rebuilt group candidates
- Result: duplicates explain most of the gap; see D-A-003 and the review on D-B-014
- Artifacts: experiments/e027_dedup_groups.py; outputs/e027.csv
- Discovery updates: D-B-014 (review), D-A-003
- Review verdict: D-B-014 HOLD
- Resource: CPU
- Other executor: none
- New plan items: H-A-08, H-A-09
When handoff.md becomes hard to scan, archive older completed entries under docs/ and leave a short summary/link in the root handoff.
An event records what happened and points to knowledge; it never restates it. Result is one line that names the discovery, Artifacts lists paths only, and there is no next-step line — each slot's Next action in its active section is the only one.
What goes where
| Information | Home | Elsewhere |
|---|---|---|
| Claim, numbers, interpretation | discoveries.md | handoff Result names the discovery ID |
| Verification state and reasons | discoveries.md (Cross-check, Reviews) | handoff Review verdict names the ID and verdict only |
| Who, when, which host, what was done | handoff.md event | — |
| Files produced or changed | handoff.md Artifacts | discovery Evidence cites what reproduces the claim |
| Current best and other project-wide values | handoff.md Shared state, citing a discovery | never in plan.md |
| What to test next | plan.md Next test | handoff names the plan item ID |
Consistency check
Run the bundled checker at session start, after a take-over, and before closing:
python <installed-skill>/scripts/check_project.py .
It reports mismatches between the files: an active agent without its plan or handoff section, a plan item that handoff names as queued but that is missing from plan.md, a finished or retired item still in the queue, champion-style values written into plan.md, a stale or uncited Current best, and cross-check states that do not match their reviews. Each problem is tagged with the agent that owns it; agents fix their own and list others under Open consistency issues.
Adaptive workers and compute routing
Start with two workers (A, B) when parallelism is useful. Add workers only while independent high-value work and actual resource headroom remain.
Each PLAN item declares:
Resource: CPU | GPU | EITHER
Parallel: YES | NO
Other: NONE | <external executor>
Example:
Other: Kaggle
Routing order:
- Idle GPU → highest-priority compatible GPU/EITHER item.
- Idle CPU → highest-priority compatible CPU/EITHER item.
- One local resource busy, the other idle → fill the idle one with worthwhile independent work.
- Both local resources saturated → eligible work may overflow to
Otherwhen available and authorized. - Never run low-value work merely to keep hardware busy.
- Scale down when RAM pressure, I/O contention, duplicated work, or lower throughput appears.
Worker count is not a goal. Useful throughput is the goal.
Claims and waiting for busy compute
- Claim before you launch.
Active computein handoff Shared state lists who holds which resource:GPU — A (H-A-04, since 2026-10-05 18:20). Before a heavy job, check both the claims and real usage (nvidia-smi, task manager); add the claim and launch in one step, and remove it in the step that records the job's end. When two sessions both see an idle GPU, the claim is what stops them from launching together. - Busy resource → wait, but keep working. Do not launch alongside a job unless the item is
Parallel: YESand measured free memory and load clearly fit. Meanwhile: run an item that fits an idle resource (or a permittedOtherexecutor) → do work that needs no heavy compute (cross-checks, preparing and smoke-testing the waiting experiment, analysis and follow-up planning, rescoring) → only if nothing is left, recordBlocker: waiting for GPU (held by A …)and re-check at an interval that matches the running job. - Stale claims. A claim held by a released slot, or for an item no longer in
plan.md, is stale; the checker reports it. A session never releases its slot while its own job is still running.
Repository layout
research-orchestrator-skill/
├── README.md
├── README.ko.md
├── README.zh-CN.md
├── README.ja.md
├── LICENSE
├── .gitignore
├── plugin.json
├── .agents/plugins/marketplace.json
├── .claude-plugin/
│ ├── plugin.json
│ └── marketplace.json
├── .codex-plugin/plugin.json
├── assets/readme/
│ ├── hero.svg
│ ├── cross-host-check.svg
│ └── take-over.svg
├── examples/cv-leakage-study/
│ ├── agents.md
│ ├── plan.md
│ ├── discoveries.md
│ └── handoff.md
├── scripts/validate_release.py
└── skills/
└── research-orchestrator-skill/
├── SKILL.md
├── agents/openai.yaml
├── scripts/
│ ├── init_research_orchestrator.py
│ └── check_project.py
├── templates/
│ ├── AGENTS.md.template
│ ├── PLAN.md.template
│ ├── DISCOVERIES.md.template
│ └── HANDOFF.md.template
└── assets/icon.svg
Design principles
- Minimal shared state — four coordination documents, no per-agent folder hierarchy.
- Independent exploration — unfinished agent plans remain separated.
- Shared evidence — completed findings flow through discoveries.
- Cross-host verification — each discovery is checked once by a different host, not by every session.
- Live queue only — completed work does not accumulate in PLAN.
- Evidence-backed retries — returning discoveries state evidence and improvements.
- Adaptive concurrency — worker count follows useful work and compute headroom.
- Host-agnostic agents —
A,B,C, ... are slots any host can resume; hosts are recorded in handoff. - Safe shared edits — re-read before patching shared files.
Validation
Run the built-in consistency check before publishing changes:
python scripts/validate_release.py
A release should pass all of these checks:
- Plugin and marketplace manifests parse as valid JSON and share one version.
- No legacy skill name remains.
- Skill frontmatter contains only
nameanddescription. - PLAN, DISCOVERIES, and HANDOFF examples in every README, SKILL.md, the templates, and the worked example use the exact field order.
- Discoveries record their
Hostand a validCross-checkstate; reviews use<Host> (<Agent>)withCLOSED,HOLD, orCHALLENGED, and never come from the source host (unless markedsame host —). - Resource values are
CPU/GPU/EITHERin PLAN andCPU/GPU/Other/nonein HANDOFF. - Real handoff events list their discovery updates and new plan items or say
none — <reason>; concretePriorityvalues match the formula. - The worked example and a freshly initialized project pass
check_project.py. - Agent names are
A-ZorAA-ZZ; IDs followH-<Agent>-NNandD-<Agent>-NNN. - Every README has the language switcher, its relative links resolve, and translations keep the same images and code-block structure.
- The worked example's
agents.mdmatches what the initializer generates today. - Initializer writes LF files, rejects duplicate or non-standard agent names, and never overwrites existing project files.
License
MIT.
Host documentation
- Claude Code plugins: https://docs.claude.com/en/docs/claude-code/plugins
- OpenAI plugin packaging and marketplaces: https://developers.openai.com/plugins/build/plugins
- Codex skills: https://developers.openai.com/blog/eval-skills
- Google Antigravity Skills: https://codelabs.developers.google.com/getting-started-with-antigravity-skills