copilot-autoresearch
A Copilot CLI extension that recreates the pi-autoresearch workflow: an autonomous experiment loop where the agent tries an idea, benchmarks it, keeps what works, reverts what doesn't, and repeats — without stopping until you tell it to.
Try an idea, measure it, keep what works, discard what doesn't, repeat forever.
Prerequisites
- A current GitHub Copilot CLI release with plugin extension support.
- Git, Node.js, and Bash available in the repository where experiments run.
Autoresearch executes benchmark commands and intentionally creates Git commits or reverts experiment changes in the active repository. Use it only in a repository whose changes you are prepared to let the agent modify.
Install
copilot plugin marketplace add scaryrawr/scarypilot
copilot plugin install copilot-autoresearch@scarypilot
Restart Copilot or run /clear after installing or updating the plugin.
Approve the extension access prompt the first time it loads.
Migrating from the standalone extension
Remove any legacy user extension at
~/.copilot/extensions/copilot-autoresearch before enabling this plugin, then
restart Copilot or run /clear. Loading both copies registers the same three
tool names; the second copy will be rejected by Copilot CLI.
What it adds
Tools
| Tool | Description |
|---|---|
init_experiment | One-time session config — name, primary metric, unit, direction. Writes a config header to .auto/log.jsonl. |
run_experiment | Runs any shell command, times wall-clock duration, captures output, parses METRIC name=value lines, retains overflow output in a bounded temporary file until the next run, and runs .auto/checks.sh after a passing benchmark when present. |
log_experiment | Records keep / discard / crash / checks_failed. On keep, auto-runs git add -A && git commit. On other statuses, auto-reverts code changes while preserving .auto/**. Computes a session confidence score after 3+ runs. |
Skills
| Skill | Description |
|---|---|
autoresearch-create | Creates .auto/prompt.md and .auto/measure.sh, establishes the baseline, and starts the loop. |
autoresearch-finalize | Groups kept experiments into clean, independently reviewable branches and verifies their combined tree. |
Slash command
/autoresearch <text> enter autoresearch mode and start (or resume) the loop
/autoresearch off leave autoresearch mode
/autoresearch clear delete current and legacy session logs and turn the mode off
/autoresearch export open a local live dashboard in your browser
/autoresearch status print a rehydration summary of current session state
Auto-resume
When in autoresearch mode and the agent goes idle after logging at least one
new experiment, the extension waits 800 ms then sends a follow-up
Run the next iteration now… prompt with a deterministic rehydration summary
attached. It matches upstream's 200-turn safety ceiling and stops after more
than 20 consecutive discards or crashes. On CLI resume, persisted mode is
restored before the extension joins and existing runs become the scheduler
baseline instead of being mistaken for newly logged work.
Live run status
While a benchmark is running, the extension shows a transient status message with elapsed time and the latest non-empty output line. Completed experiment summaries remain in the session timeline.
Live browser dashboard
/autoresearch export starts a loopback-only HTTP server, opens the dashboard
in the default browser, and refreshes it through server-sent events as results
are logged. The dashboard includes summary cards, a metric trend, confidence,
and the complete run table.
The server stops when autoresearch is turned off, cleared, or the session ends.
Project files
New sessions use the same .auto/ layout as pi-autoresearch. Legacy-only flat
autoresearch.* sessions remain readable; once a current .auto/ session
artifact is introduced, the current layout consistently takes precedence.
| File | Purpose |
|---|---|
.auto/log.jsonl | Append-only run log. The source of truth across restarts. |
.auto/prompt.md | Living session document — objective, metrics, files in scope, what's been tried. |
.auto/measure.sh | Optional benchmark script. When present, run_experiment rejects commands that don't invoke it. |
.auto/checks.sh | Optional correctness gate (tests, types, lint). Runs after every passing benchmark. |
.auto/ideas.md | Optional ideas backlog for promising deferred optimizations. |
.auto/config.json | Optional workspace-contained workingDir and maxIterations configuration. |
.auto/runtime/<session-id>.json | Internal per-session sidecar preserving the run/checks boundary and explicit mode state without leaking activation across sessions. |
Confidence scoring
After 3+ experiments, the extension computes
|best_improvement_over_baseline| / MAD(metric_values). ≥2.0× is likely real,
1.0–2.0× is above noise but marginal, <1.0× is within noise — re-run to
confirm before keeping. Advisory only; never auto-discards.
Revisiting discards
After each logged experiment, log_experiment reminds the agent to consider
whether the latest result invalidates a previous discard's rollback reason.
When retrying a discarded idea because an assumption changed, the agent sets
asi.revisits_run to the earlier run number; the tool result and the live
dashboard then show a ↻ Revisiting #N badge. Verification reruns to resolve
measurement noise are a separate case and don't set this field.
Differences vs pi-autoresearch
Parity is tracked against pi-autoresearch 1.8.1. The .auto/ contract,
experiment lifecycle, benchmark and checks gates, iteration limits,
failure guard, finalization workflow, and live browser reporting are preserved
where Copilot exposes an equivalent extension API.
The Copilot CLI extension surface is narrower than pi's, so a few things are adapted:
- No repository-provided iteration hooks. Executable
.auto/hooks/*.shfiles are not loaded or run because a checked-out branch must not gain code execution merely by activating autoresearch. - No persistent status widget / dashboard expand-collapse / fullscreen
overlay. A transient status message updates while an experiment runs;
/autoresearch statusprints the full rehydration summary on demand. - No fullscreen terminal overlay.
/autoresearch exportinstead opens the same experiment data as a live local browser dashboard. - No compaction summary injection. Copilot CLI does not expose a
session_before_compacthook to extensions, so the deterministic rehydration summary is included in every auto-resume prompt instead. - No configurable keyboard shortcuts. Copilot CLI extensions can't bind keys. Upstream 1.7.0 also makes Pi shortcuts opt-in; use the Copilot slash command subcommands instead.
- Tools cannot be removed dynamically from Copilot's model schema. They
remain discoverable while mode is off, but reject execution until
/autoresearch <goal>activates the loop.
Credits
This extension is a port of pi-autoresearch by @davebcn87 — all credit for the original autoresearch concept, workflow design, and file format goes to that project. This is simply a Copilot CLI adaptation of the same idea.
Develop and test
cd extensions/copilot-autoresearch
npm install
npm run build
npm run typecheck
npm run test
bash ../../skills/autoresearch-finalize/tests/finalize-smoke.sh
Copilot CLI injects its bundled @github/copilot-sdk when it loads the
extension. The package dependency pins the SDK version used for local
development and tests. The committed dist/ bundle includes other runtime
dependencies; plugin users do not need to run npm install or build. Build in
a source checkout after changes and commit the bundle and manifest.
Resources
- Copilot CLI plugin reference
- pi-autoresearch website
- Source migration record
- pi-autoresearch MIT license
This plugin was migrated from
scaryrawr/copilot-autoresearch
at the revision recorded in NOTICE.md.