Skip to content

scaryrawr/copilot-autoresearch

v0.2.6

Run measurable, resumable experiment loops with Copilot CLI tools, commands, and skills.

copilot-autoresearch

A Copilot CLI extension that recreates the pi-autoresearch workflow: an autonomous experiment loop where the agent tries an idea, benchmarks it, keeps what works, reverts what doesn't, and repeats — without stopping until you tell it to.

Try an idea, measure it, keep what works, discard what doesn't, repeat forever.

Prerequisites

  • A current GitHub Copilot CLI release with plugin extension support.
  • Git, Node.js, and Bash available in the repository where experiments run.

Autoresearch executes benchmark commands and intentionally creates Git commits or reverts experiment changes in the active repository. Use it only in a repository whose changes you are prepared to let the agent modify.

Install

copilot plugin marketplace add scaryrawr/scarypilot
copilot plugin install copilot-autoresearch@scarypilot

Restart Copilot or run /clear after installing or updating the plugin. Approve the extension access prompt the first time it loads.

Migrating from the standalone extension

Remove any legacy user extension at ~/.copilot/extensions/copilot-autoresearch before enabling this plugin, then restart Copilot or run /clear. Loading both copies registers the same three tool names; the second copy will be rejected by Copilot CLI.

What it adds

Tools

ToolDescription
init_experimentOne-time session config — name, primary metric, unit, direction. Writes a config header to .auto/log.jsonl.
run_experimentRuns any shell command, times wall-clock duration, captures output, parses METRIC name=value lines, retains overflow output in a bounded temporary file until the next run, and runs .auto/checks.sh after a passing benchmark when present.
log_experimentRecords keep / discard / crash / checks_failed. On keep, auto-runs git add -A && git commit. On other statuses, auto-reverts code changes while preserving .auto/**. Computes a session confidence score after 3+ runs.

Skills

SkillDescription
autoresearch-createCreates .auto/prompt.md and .auto/measure.sh, establishes the baseline, and starts the loop.
autoresearch-finalizeGroups kept experiments into clean, independently reviewable branches and verifies their combined tree.

Slash command

/autoresearch <text>     enter autoresearch mode and start (or resume) the loop
/autoresearch off        leave autoresearch mode
/autoresearch clear      delete current and legacy session logs and turn the mode off
/autoresearch export     open a local live dashboard in your browser
/autoresearch status     print a rehydration summary of current session state

Auto-resume

When in autoresearch mode and the agent goes idle after logging at least one new experiment, the extension waits 800 ms then sends a follow-up Run the next iteration now… prompt with a deterministic rehydration summary attached. It matches upstream's 200-turn safety ceiling and stops after more than 20 consecutive discards or crashes. On CLI resume, persisted mode is restored before the extension joins and existing runs become the scheduler baseline instead of being mistaken for newly logged work.

Live run status

While a benchmark is running, the extension shows a transient status message with elapsed time and the latest non-empty output line. Completed experiment summaries remain in the session timeline.

Live browser dashboard

/autoresearch export starts a loopback-only HTTP server, opens the dashboard in the default browser, and refreshes it through server-sent events as results are logged. The dashboard includes summary cards, a metric trend, confidence, and the complete run table. The server stops when autoresearch is turned off, cleared, or the session ends.

Project files

New sessions use the same .auto/ layout as pi-autoresearch. Legacy-only flat autoresearch.* sessions remain readable; once a current .auto/ session artifact is introduced, the current layout consistently takes precedence.

FilePurpose
.auto/log.jsonlAppend-only run log. The source of truth across restarts.
.auto/prompt.mdLiving session document — objective, metrics, files in scope, what's been tried.
.auto/measure.shOptional benchmark script. When present, run_experiment rejects commands that don't invoke it.
.auto/checks.shOptional correctness gate (tests, types, lint). Runs after every passing benchmark.
.auto/ideas.mdOptional ideas backlog for promising deferred optimizations.
.auto/config.jsonOptional workspace-contained workingDir and maxIterations configuration.
.auto/runtime/<session-id>.jsonInternal per-session sidecar preserving the run/checks boundary and explicit mode state without leaking activation across sessions.

Confidence scoring

After 3+ experiments, the extension computes |best_improvement_over_baseline| / MAD(metric_values). ≥2.0× is likely real, 1.0–2.0× is above noise but marginal, <1.0× is within noise — re-run to confirm before keeping. Advisory only; never auto-discards.

Revisiting discards

After each logged experiment, log_experiment reminds the agent to consider whether the latest result invalidates a previous discard's rollback reason. When retrying a discarded idea because an assumption changed, the agent sets asi.revisits_run to the earlier run number; the tool result and the live dashboard then show a ↻ Revisiting #N badge. Verification reruns to resolve measurement noise are a separate case and don't set this field.

Differences vs pi-autoresearch

Parity is tracked against pi-autoresearch 1.8.1. The .auto/ contract, experiment lifecycle, benchmark and checks gates, iteration limits, failure guard, finalization workflow, and live browser reporting are preserved where Copilot exposes an equivalent extension API.

The Copilot CLI extension surface is narrower than pi's, so a few things are adapted:

  • No repository-provided iteration hooks. Executable .auto/hooks/*.sh files are not loaded or run because a checked-out branch must not gain code execution merely by activating autoresearch.
  • No persistent status widget / dashboard expand-collapse / fullscreen overlay. A transient status message updates while an experiment runs; /autoresearch status prints the full rehydration summary on demand.
  • No fullscreen terminal overlay. /autoresearch export instead opens the same experiment data as a live local browser dashboard.
  • No compaction summary injection. Copilot CLI does not expose a session_before_compact hook to extensions, so the deterministic rehydration summary is included in every auto-resume prompt instead.
  • No configurable keyboard shortcuts. Copilot CLI extensions can't bind keys. Upstream 1.7.0 also makes Pi shortcuts opt-in; use the Copilot slash command subcommands instead.
  • Tools cannot be removed dynamically from Copilot's model schema. They remain discoverable while mode is off, but reject execution until /autoresearch <goal> activates the loop.

Credits

This extension is a port of pi-autoresearch by @davebcn87 — all credit for the original autoresearch concept, workflow design, and file format goes to that project. This is simply a Copilot CLI adaptation of the same idea.

Develop and test

cd extensions/copilot-autoresearch
npm install
npm run build
npm run typecheck
npm run test
bash ../../skills/autoresearch-finalize/tests/finalize-smoke.sh

Copilot CLI injects its bundled @github/copilot-sdk when it loads the extension. The package dependency pins the SDK version used for local development and tests. The committed dist/ bundle includes other runtime dependencies; plugin users do not need to run npm install or build. Build in a source checkout after changes and commit the bundle and manifest.

Resources

This plugin was migrated from scaryrawr/copilot-autoresearch at the revision recorded in NOTICE.md.