Skip to content

taeyang4/research-orchestrator

v4.2.0MIT

Hypothesis-driven research orchestration with scored plans, independent discovery review, adaptive workers, resource-aware execution, and durable handoffs.

Research Orchestrator

A lightweight, hypothesis-driven research workflow for one or many AI agent sessions.

It deliberately uses only four shared Markdown files while adding a scored hypothesis queue, one-time cross-host verification, adaptive workers, and CPU/GPU/Other resource routing.

Why use it?

  • Resume long-running research across fresh sessions.
  • Let multiple agents explore independently without sharing unfinished plans.
  • Keep plan.md small by removing completed work.
  • Never queue the same idea twice: every completed experiment leaves a discovery — positive, negative, or inconclusive — and every new item is checked against discoveries and other agents' queues.
  • Turn experiment results into reusable shared discoveries.
  • Have a different host check each finding once — Codex's discovery is verified by Claude Code, and vice versa — instead of every session re-reviewing it.
  • Preserve promising but incomplete findings with HOLD.
  • Re-open discoveries as stronger hypotheses with explicit evidence and improvements.
  • Scale from two workers to the machine's practical full load.

Core workflow

flowchart LR
    D[DISCOVERIES<br/>shared evidence<br/>cross-checked once by another host]
    H[New hypotheses<br/>Sources + Evidence + Improvement]
    P[PLAN<br/>scored live queue]
    W[Adaptive workers<br/>CPU / GPU / Other]
    E[Experiments]
    O[HANDOFF<br/>history + artifacts + blockers + next state]

    D -->|0..N follow-ups| H
    H --> P --> W --> E --> D
    P --> O
    E --> O
    D --> O

A discovery can generate zero, one, or many new hypotheses. One hypothesis can also combine evidence from multiple discoveries.

The four files

flowchart TD
    A[AGENTS.md<br/>rules, scoring, review, routing]
    P[PLAN.md<br/>active unfinished work only]
    D[DISCOVERIES.md<br/>shared reusable findings]
    H[HANDOFF.md<br/>operational ledger]

    A --> P
    P -->|completed result| D
    P -->|execution history| H
    D -->|new evidence / review| P
FilePurpose
agents.mdStable rules for reading, editing, scoring, review, and resource routing.
plan.mdOnly active unfinished work. Each agent owns its own section and works from a scored queue.
discoveries.mdWhat is known: claims, numbers, interpretation, and verification. Every completed experiment leaves one entry; each is cross-checked once by a different host.
handoff.mdWhat happened and where to resume: who, when, which host, artifact paths, project-wide values such as the current best, and each slot's resume point. Points to discoveries instead of repeating them.

Completed items leave plan.md. Each one leaves a discovery (whatever the outcome) and a handoff event that points to it, and its follow-up hypotheses return to plan.md with fresh scores.

Templates and a worked example

Each file is created from a template. The worked example shows the same four files in the middle of a real project: two active slots (A on Claude Code, B on Codex) and one released slot (C), a VERIFIED discovery, a discovery under review, a negative result logged with zero follow-ups and the reason, a same-host take-over, and the handoff log that produced them.

FileTemplateWorked example
agents.mdAGENTS.md.templateagents.md
plan.mdPLAN.md.templateplan.md
discoveries.mdDISCOVERIES.md.templatediscoveries.md
handoff.mdHANDOFF.md.templatehandoff.md

Install

The repository packages the same skill for Claude Code, Codex, and Google Antigravity.

Repository:

https://github.com/TaeyanG4/research-orchestrator-skill

Claude Code

Plugin marketplace install

Inside Claude Code:

/plugin marketplace add TaeyanG4/research-orchestrator-skill
/plugin install research-orchestrator@research-orchestrator

Start a new Claude Code session after the first install.

Project-only skill install

Clone or copy:

skills/research-orchestrator-skill/

into:

<project>/.claude/skills/research-orchestrator-skill/

Codex

Plugin marketplace install

From a terminal:

codex plugin marketplace add TaeyanG4/research-orchestrator-skill
codex

Inside Codex:

/plugins

Select the Research Orchestrator marketplace and install research-orchestrator. Start a new chat before first use.

Direct skill install

Repo-scoped:

<project>/.codex/skills/research-orchestrator-skill/

User-scoped:

~/.codex/skills/research-orchestrator-skill/

Copy the repository's skills/research-orchestrator-skill/ directory into one of those locations.

Google Antigravity

Clone the repository:

git clone https://github.com/TaeyanG4/research-orchestrator-skill.git

Then copy skills/research-orchestrator-skill/ to one of these locations.

Project/workspace scope:

<project>/.agents/skills/research-orchestrator-skill/

Global Antigravity scope:

~/.gemini/config/skills/research-orchestrator-skill/

Antigravity CLI legacy/global location:

~/.gemini/antigravity-cli/skills/research-orchestrator-skill/

Use /skills in Antigravity CLI to confirm discovery.


Quick start

Initialize the four project files:

python <installed-skill>/scripts/init_research_orchestrator.py . -n "My Project"

Start with multiple independent agents (2 → A, B):

python <installed-skill>/scripts/init_research_orchestrator.py . -n "My Project" --agents 2

--agents also accepts an explicit list such as A,B,C. The initializer rejects duplicate or non-standard names, creates missing files only, and never overwrites existing project files.

Then run python <installed-skill>/scripts/check_project.py . at every session start and before closing (see Consistency check below).

Agent names

An agent name is a work slot, not the tool or model running it. Every host — Claude Code, Codex, Antigravity — uses the same names:

A, B, C, ... Z, AA, AB, ... ZZ
  • Single-agent work always uses A; each additional concurrent session takes the next unused letter, continuing with two-letter names after Z. Released slots are reused first, so new letters appear only when every existing slot is taken at once.
  • Never name an agent Claude, Codex, GPT, Gemini, or any other host/model name.
  • Any host may resume any slot. Which host owns a slot is recorded in handoff.md, not in the name:
### Agent: A
- Current host: Claude Code
- Current thread: H-A-04
...

### Agent: B
- Current host: Codex
- Current thread: H-B-02
...

A slot is free when Current host reads unassigned or released. A new session takes the first free slot in order — skipping one that still holds plan items left by a different host — sets Current host to its own host, and sets it back to released when it closes — so liveness is read from the file, never guessed.

Each completed-log event also records its Host, so the history shows which host did each step even after a slot changes hands.

IDs embed the owning agent so concurrent agents never collide:

ObjectFormatExamples
Plan itemH-<Agent>-<NN>H-A-01, H-B-07
DiscoveryD-<Agent>-<NNN>D-A-001, D-C-014

Follow-ups and take-over

  • Plan follow-ups from every discovery. Each time an agent records a discovery — a new finding, a negative result, or a cross-check verdict — it decides what to test next: zero, one, or many new plan items, each passing the duplicate check. It also rescores or removes its own items the discovery affects. Zero is a valid answer, but it is logged with a reason: New plan items: none — <reason>.
  • When a queue runs out, the agent looks for work in this order: re-read its own section → claim a PENDING cross-check from another host → take over an item from a same-host slot → cross-host take-over by judgment (one item at most) → derive new hypotheses → note it and release the slot.
  • Take-over stays within one host by default. A host is the platform, not the model: two Claude Code sessions on different models are the same host. If A's queue is empty and C (same host) still has queued items, A moves the highest-priority one into its own section under its own next ID — H-C-02 becomes ### H-A-04 — … (from H-C-02) — and logs the move. It never takes the item in the owner's Current thread or one under an Active compute claim.
  • Cross-host take-over by judgment. Items queued by a different host normally stay with that host. As an exception, an agent may take over one of them when everything closer is exhausted, the source slot is released (never an active session on another host), the item has Priority ≥ 15 and clearly beats any new hypothesis, and the agent logs a one-line reason: took over H-C-01 from C (cross-host: Codex → Claude Code; reason: …) as H-A-12. After finishing it, the agent starts the order again. The user can change the floor, forbid this, or approve specific items.

Reading rule

Each agent reads:

agents.md
→ all discoveries.md
→ shared handoff log + its own active handoff
→ only its own detailed PLAN section

Agents do not read another active agent's detailed PLAN just to coordinate work. The only cross-section peeks are a dispatcher reading task metadata, the duplicate check reading item headings and Hypothesis lines, a session reading Current host and Current thread lines to find a free slot or a take-over source, and the one item a session takes over.


Standard PLAN format

Keep only active unfinished items.

### H-A-07 — Separate duplicate leakage from group leakage
- Sources: D-B-014, D-A-021
- Hypothesis: exact duplicates explain most apparent group leakage
- Evidence: D-B-014 weakens after deduplication; D-A-021 identifies repeated rows
- Improvement: isolate exact duplicates before constructing candidate groups
- Impact: 3
- Information: 3
- Confidence: 2
- Unblock: 3
- Diversity: 2
- Cost: 1
- Priority: 21
- Resource: CPU
- Parallel: YES
- Other: NONE
- Next test: compare random CV and GroupKFold after exact-duplicate removal

Priority formula

Score every factor from 0 to 3.

Priority = 2*Impact + 2*Information + Confidence + Unblock + Diversity + (3-Cost)

Run the highest score first unless blocked or explicitly overridden. Break ties by lower Cost, then higher Information.

When a discovery returns to PLAN, these fields are mandatory:

  • Sources — which discoveries motivated it.
  • Evidence — why it deserves another test.
  • Improvement — what is materially different or stronger than the prior attempt.

Do not simply rerun an old idea under a new task ID.

Duplicate check before adding an item

  1. Search discoveries.md — every completed experiment is there, positive, negative (Finding: <claim> does not hold under <conditions>), or inconclusive. Skip VERIFIED claims unless you have a real Improvement; skip CHALLENGED ones until resolved.
  2. Scan the other agents' plan sections, reading only item headings and Hypothesis lines. If it is already queued, do not add it.
  3. Re-read plan.md right before writing; if the same item appeared meanwhile, keep the earlier one.

Standard DISCOVERIES format

## D-B-014 — Random CV may leak groups
- Source: B
- Host: Codex
- Cross-check: HOLD
- Finding: duplicated groups cross random folds
- Evidence: e014_group_check.py; random CV 0.9162 vs group CV 0.9027
- Implication: current validation may be optimistic
- Reviews:
  - Claude Code (A): HOLD — plausible, but exact duplicates must be separated first

Verification is per host, not per agent. A discovery made on Codex is checked once by Claude Code (or another different host), and vice versa. Sessions on the same host share the same blind spots, so they do not re-review each other — ten Claude Code sessions never review the same finding ten times.

flowchart LR
    N[New discovery<br/>Host: Codex] --> P[Cross-check: PENDING]
    P -->|a Claude Code session claims it| R[REVIEWING Claude Code]
    R -->|CLOSED| V[VERIFIED]
    R -->|HOLD| H[HOLD]
    R -->|CHALLENGED| C[CHALLENGED]
    H -->|source adds evidence| P
    C -->|source revises| P

Cross-check states:

  • PENDING — no different host has reviewed it yet.
  • REVIEWING <Host> (<Agent>) — a different-host session claimed the review, so nobody else duplicates it.
  • VERIFIED — a different host reviewed it and recorded CLOSED.
  • HOLD — plausible, but specific evidence or an improvement is needed first.
  • CHALLENGED — a material contradiction, flaw, or missing assumption was found.

Rules:

  • The source never reviews its own discovery, and same-host sessions do not review each other.
  • One cross-host review is enough; add another only for high-impact or disputed findings.
  • HOLD and CHALLENGED reviews must state what would make the discovery acceptable. When the source revises it, Cross-check returns to PENDING.
  • VERIFIED discoveries may be used freely. Building on an unverified one must be stated in the plan item's Evidence; CHALLENGED ones are not used until resolved.
  • With only one host available, a different slot on the same host may cross-check, starting its reason with same host —.

Standard HANDOFF format

### 2026-10-05 21:10 — A — H-A-07
- Host: Claude Code
- Action: cross-checked D-B-014; removed exact duplicates and rebuilt group candidates
- Result: duplicates explain most of the gap; see D-A-003 and the review on D-B-014
- Artifacts: experiments/e027_dedup_groups.py; outputs/e027.csv
- Discovery updates: D-B-014 (review), D-A-003
- Review verdict: D-B-014 HOLD
- Resource: CPU
- Other executor: none
- New plan items: H-A-08, H-A-09

When handoff.md becomes hard to scan, archive older completed entries under docs/ and leave a short summary/link in the root handoff.

An event records what happened and points to knowledge; it never restates it. Result is one line that names the discovery, Artifacts lists paths only, and there is no next-step line — each slot's Next action in its active section is the only one.

What goes where

InformationHomeElsewhere
Claim, numbers, interpretationdiscoveries.mdhandoff Result names the discovery ID
Verification state and reasonsdiscoveries.md (Cross-check, Reviews)handoff Review verdict names the ID and verdict only
Who, when, which host, what was donehandoff.md event—
Files produced or changedhandoff.md Artifactsdiscovery Evidence cites what reproduces the claim
Current best and other project-wide valueshandoff.md Shared state, citing a discoverynever in plan.md
What to test nextplan.md Next testhandoff names the plan item ID

Consistency check

Run the bundled checker at session start, after a take-over, and before closing:

python <installed-skill>/scripts/check_project.py .

It reports mismatches between the files: an active agent without its plan or handoff section, a plan item that handoff names as queued but that is missing from plan.md, a finished or retired item still in the queue, champion-style values written into plan.md, a stale or uncited Current best, and cross-check states that do not match their reviews. Each problem is tagged with the agent that owns it; agents fix their own and list others under Open consistency issues.


Adaptive workers and compute routing

Start with two workers (A, B) when parallelism is useful. Add workers only while independent high-value work and actual resource headroom remain.

Each PLAN item declares:

Resource: CPU | GPU | EITHER
Parallel: YES | NO
Other: NONE | <external executor>

Example:

Other: Kaggle

Routing order:

  1. Idle GPU → highest-priority compatible GPU/EITHER item.
  2. Idle CPU → highest-priority compatible CPU/EITHER item.
  3. One local resource busy, the other idle → fill the idle one with worthwhile independent work.
  4. Both local resources saturated → eligible work may overflow to Other when available and authorized.
  5. Never run low-value work merely to keep hardware busy.
  6. Scale down when RAM pressure, I/O contention, duplicated work, or lower throughput appears.

Worker count is not a goal. Useful throughput is the goal.

Claims and waiting for busy compute

  • Claim before you launch. Active compute in handoff Shared state lists who holds which resource: GPU — A (H-A-04, since 2026-10-05 18:20). Before a heavy job, check both the claims and real usage (nvidia-smi, task manager); add the claim and launch in one step, and remove it in the step that records the job's end. When two sessions both see an idle GPU, the claim is what stops them from launching together.
  • Busy resource → wait, but keep working. Do not launch alongside a job unless the item is Parallel: YES and measured free memory and load clearly fit. Meanwhile: run an item that fits an idle resource (or a permitted Other executor) → do work that needs no heavy compute (cross-checks, preparing and smoke-testing the waiting experiment, analysis and follow-up planning, rescoring) → only if nothing is left, record Blocker: waiting for GPU (held by A …) and re-check at an interval that matches the running job.
  • Stale claims. A claim held by a released slot, or for an item no longer in plan.md, is stale; the checker reports it. A session never releases its slot while its own job is still running.

Repository layout

research-orchestrator-skill/
├── README.md
├── README.ko.md
├── README.zh-CN.md
├── README.ja.md
├── LICENSE
├── .gitignore
├── plugin.json
├── .agents/plugins/marketplace.json
├── .claude-plugin/
│   ├── plugin.json
│   └── marketplace.json
├── .codex-plugin/plugin.json
├── assets/readme/
│   ├── hero.svg
│   ├── cross-host-check.svg
│   └── take-over.svg
├── examples/cv-leakage-study/
│   ├── agents.md
│   ├── plan.md
│   ├── discoveries.md
│   └── handoff.md
├── scripts/validate_release.py
└── skills/
    └── research-orchestrator-skill/
        ├── SKILL.md
        ├── agents/openai.yaml
        ├── scripts/
        │   ├── init_research_orchestrator.py
        │   └── check_project.py
        ├── templates/
        │   ├── AGENTS.md.template
        │   ├── PLAN.md.template
        │   ├── DISCOVERIES.md.template
        │   └── HANDOFF.md.template
        └── assets/icon.svg

Design principles

  • Minimal shared state — four coordination documents, no per-agent folder hierarchy.
  • Independent exploration — unfinished agent plans remain separated.
  • Shared evidence — completed findings flow through discoveries.
  • Cross-host verification — each discovery is checked once by a different host, not by every session.
  • Live queue only — completed work does not accumulate in PLAN.
  • Evidence-backed retries — returning discoveries state evidence and improvements.
  • Adaptive concurrency — worker count follows useful work and compute headroom.
  • Host-agnostic agents — A, B, C, ... are slots any host can resume; hosts are recorded in handoff.
  • Safe shared edits — re-read before patching shared files.

Validation

Run the built-in consistency check before publishing changes:

python scripts/validate_release.py

A release should pass all of these checks:

  • Plugin and marketplace manifests parse as valid JSON and share one version.
  • No legacy skill name remains.
  • Skill frontmatter contains only name and description.
  • PLAN, DISCOVERIES, and HANDOFF examples in every README, SKILL.md, the templates, and the worked example use the exact field order.
  • Discoveries record their Host and a valid Cross-check state; reviews use <Host> (<Agent>) with CLOSED, HOLD, or CHALLENGED, and never come from the source host (unless marked same host —).
  • Resource values are CPU/GPU/EITHER in PLAN and CPU/GPU/Other/none in HANDOFF.
  • Real handoff events list their discovery updates and new plan items or say none — <reason>; concrete Priority values match the formula.
  • The worked example and a freshly initialized project pass check_project.py.
  • Agent names are A-Z or AA-ZZ; IDs follow H-<Agent>-NN and D-<Agent>-NNN.
  • Every README has the language switcher, its relative links resolve, and translations keep the same images and code-block structure.
  • The worked example's agents.md matches what the initializer generates today.
  • Initializer writes LF files, rejects duplicate or non-standard agent names, and never overwrites existing project files.

License

MIT.

Host documentation