Skip to content

zning1994/brainharness-autoresearch

v0.3.1MIT

Autonomous prompt optimization for AI agent skills. Runs controlled experiments to find better prompt variants using the Karpathy autoresearch pattern.

brainharness-autoresearch

Previously named brainforge-autoresearch; earlier names include openclaw-autoresearch and autoresearch. Use the current installation commands below.

Autonomous skill prompt optimizer based on Karpathy's autoresearch methodology.

Official website: brainharness.si.

License: MIT ClawHub brainharness

What it does

Define what "good" means as binary evals, then let an agent loop to optimize your skill prompt automatically. The optimizer reads your current prompt, runs it against eval cases, analyzes failures, mutates the prompt, tests the mutation, and keeps or discards it -- repeating until evals pass or the experiment budget is exhausted.

Based on Andrej Karpathy's autoresearch and Ole Lehmann's Claude Code adaptation. Works with any AI coding agent: OpenClaw, Claude Code, Cursor, Cline.

Quick Start

# Claude Code (via brainharness marketplace)
/plugin marketplace add zning1994/brainharness
/plugin install brainharness-autoresearch@brainharness

# Universal (npx skills)
npx skills add zning1994/brainharness-autoresearch

# OpenClaw ClawHub
npx clawhub@0.23.3 install brainharness-autoresearch

# Standalone
git clone https://github.com/zning1994/brainharness-autoresearch
cd brainharness-autoresearch
python autoresearch.py --target ./my-skill/SKILL.md --evals eval.json

How it works

                        +------------------+
                        |  Read SKILL.md   |
                        +--------+---------+
                                 |
                        +--------v---------+
                   +--->|  Run baseline    |
                   |    +--------+---------+
                   |             |
                   |    +--------v---------+
                   |    | Analyze failures |
                   |    +--------+---------+
                   |             |
                   |    +--------v---------+
                   |    |  Mutate prompt   |
                   |    +--------+---------+
                   |             |
                   |    +--------v---------+
                   |    |   Test mutation  |
                   |    +--------+---------+
                   |             |
                   |    +--------v---------+
                   +----+ Keep or discard  |
                        +------------------+

Creating eval.json

Minimal example with one rule eval and one LLM eval:

[
  {
    "name": "includes_greeting",
    "input": "Say hello to the user",
    "type": "rule",
    "rule": "contains",
    "expected": "hello"
  },
  {
    "name": "tone_is_professional",
    "input": "Draft a project update email",
    "type": "llm",
    "criteria": "The output uses a professional tone with no slang or emojis"
  }
]

See eval-guide.md for details.

Eval types

TypeRuleDescription
ruleregexMatch output against a regex pattern
rulecontainsOutput must contain the expected string
rulenot_containsOutput must not contain the string
rulebanned_phrasesOutput must not contain any listed phrase
ruleword_countOutput word count within min/max range
llm--LLM judges output against freeform criteria

CLI Reference

FlagDefaultDescription
--target./SKILL.mdPath to the skill prompt file to optimize
--evals./eval.jsonPath to the eval definitions file
--providerauto-detectLLM provider: minimax, openai, anthropic
--runs3Number of runs per eval case per experiment
--max-experiments10Maximum optimization iterations
--dashboardoffGenerate an HTML dashboard of results
--output-dir./resultsDirectory for output artifacts
--verboseoffPrint detailed logs to stderr

LLM Providers

Auto-detection order:

  1. MINIMAX_API_KEY
  2. OPENAI_API_KEY
  3. ANTHROPIC_API_KEY

Custom endpoint support via OPENAI_BASE_URL (works with any OpenAI-compatible API).

Output

FileDescription
results.tsvTab-separated eval scores per experiment
changelog.mdHuman-readable log of each mutation and its effect
results.jsonFull structured results for programmatic use
dashboard.htmlVisual dashboard (when --dashboard is set)
SKILL.md.baselineBackup of the original prompt before optimization

Credits

License

MIT

Codex and OpenAI plugins

The root plugin.json provides a portable package using the same skill source as Claude. Once the repository marketplace is available, install with codex plugin add brainharness-autoresearch@brainharness. This plugin is not yet listed in the public ChatGPT plugin directory.

When creating an upload archive, materialize symlinks inside skills/brainharness-autoresearch/ and include all referenced files.

Autoresearch requires a local Python environment and your model API credentials. Installing the plugin does not provision either, and model calls can incur charges. Ordinary ChatGPT execution is not verified.