zning1994/brainharness-autoresearch
v0.3.1MIT
Autonomous prompt optimization for AI agent skills. Runs controlled experiments to find better prompt variants using the Karpathy autoresearch pattern.
Changelog
All notable changes to this project will be documented in this file.
Format based on Keep a Changelog.
[0.3.1] - 2026-10-04
Changed
- Add a portable OpenAI plugin manifest alongside the Claude manifest.
- Keep skill instructions shared and make bundled resources reachable from the skill directory.
- Document installation scope and public ChatGPT listing status.
[0.3.0] - 2026-10-04
Changed
- Rename
brainforge-autoresearchtobrainharness-autoresearchand align repository links and installation names with BrainHarness. - Update distribution metadata and documentation; preserve skill behavior and historical release notes.
[0.2.5] - 2026-04-23
Changed
- Renamed repo and skill:
openclaw-autoresearch→brainforge-autoresearch - SKILL.md
namefield:autoresearch→brainforge-autoresearch - Updated README, plugin.json, and SKILL.md metadata to point at the new repo URL
- Added to the brainforge Claude Code plugin marketplace
- Published to ClawHub as
brainforge-autoresearch; old slugopenclaw-autoresearchmerged in as redirect
Compatibility
- GitHub: old URL
https://github.com/zning1994/openclaw-autoresearchauto-redirects - ClawHub:
openclaw skills install openclaw-autoresearchcontinues to work via ClawHub slug merge/redirect - No behavior changes to the optimizer itself
0.2.4 - 2026-03-29
Fixed
- SKILL.md:
pass/fail→pass_description/fail_descriptionin LLM eval docs and examples (field names didn't match actual API) - SKILL.md:
contains/not_containsparameter documented asvalue(string) →values(list) to match actual usage - SKILL.md: LLM eval table now includes required
typeandnamefields
Changed
- SKILL.md self-optimized via autoresearch (2 rounds: MiniMax 48.6%→68.1%, Claude 80.6%→91.7%)
- Procedure rewritten: added 3-step summary upfront, renamed steps for clarity, added actionable result-review guidance
- Example section restructured: removed "confirm target" preamble, simplified to "create eval → run → review" flow
- Added
self-eval.jsonfor self-optimization (dogfooding eval definition)
Tested
- Self-optimization: Round 1 (MiniMax M2.7, 15 exp, 3 keep) and Round 2 (Claude, 15 exp, 1 keep) independently discovered same doc bugs
0.2.3 - 2026-03-25
Added
--timeoutCLI flag to control HTTP timeout per LLM call (default: 180s)- Large prompts (e.g., 14KB style guides) need longer timeouts for mutation calls
- Previously hardcoded, now user-configurable
Changed
- Default HTTP timeout increased from 60s to 180s (was causing mutation failures on Anthropic API with large prompts)
- Timeout is now stored on the provider instance, passed through
detect_provider()
0.2.2 - 2026-03-25
Fixed
.gitignore: removed incorrect entries (.git,.github), addedautoresearch-*/and*.baseline- README: added ClawHub badge and link, fixed install slug to
openclaw-autoresearch
0.2.1 - 2026-03-25
Fixed
- ClawHub security flag: changed
requires.env(all required) torequires.anyEnv(any one suffices) - Declared optional env vars
OPENAI_BASE_URL/OPENAI_API_BASEin metadata
0.2.0 - 2026-03-25
Fixed
- MiniMax extended thinking support — models with thinking blocks (M2.7, Claude) no longer crash the parser
- Fallback to thinking content when text block is missing (thinking exhausts
max_tokens) - LLM judge
max_tokensincreased from 16 to 256 (thinking models need headroom) - Rule type field: accept both
"rule"and"check"keys for backwards compatibility contains/not_contains: accept both"values"(list) and"value"(string)- LLM eval fields: accept both
"pass_description"/"fail_description"and"pass"/"fail" test_inputs: accept both plain strings and{"name": ..., "input": ...}objectscontainsrule now supports"match": "all"mode (require all values present)- Convergence counter double-increment bug fixed
regexrule now usesre.MULTILINEfor correct^/$matching- CJK-aware word counting in
word_countrule (counts Chinese/Japanese/Korean characters individually)
Added
--modelCLI flag to override the default model per provider- JSON fallback:
runs_per_experimentandmax_experimentsfrom eval.json used when CLI defaults unchanged - Shared
_extract_text()helper for Anthropic-style response parsing across all providers
Tested
- Successfully optimized
brain-searchskill: 37.5% → 54.2% pass rate in 5 experiments with MiniMax M2.7
0.1.0 - 2026-03-24
Added
- Core
autoresearch.pyscript — autonomous skill prompt optimizer- Zero external dependencies (Python 3.9+ stdlib only)
- LLM providers: MiniMax, OpenAI, Anthropic (auto-detect from env vars)
- Custom endpoint support via
OPENAI_BASE_URL --modelflag to override default model per provider
- Hybrid eval system
- Rule-based evals:
regex,banned_phrases,word_count,contains,not_contains containssupportsmatch: "all"andmatch: "any"modes- LLM-as-judge evals with binary YES/NO scoring
- CJK-aware word counting for Chinese/Japanese/Korean text
- Rule-based evals:
- Experiment loop
- Baseline measurement before any changes
- One mutation per experiment (targeted, not bulk rewrites)
- Automatic keep/discard based on score comparison
- Convergence exit: 95%+ pass rate for 3 consecutive experiments
- Budget cap via
--max-experiments
- Output artifacts
results.tsv— tab-separated score logchangelog.md— detailed mutation history with reasoningresults.json— structured data for toolingdashboard.html— self-contained Chart.js dashboard (opt-in via--dashboard)SKILL.md.baseline— original skill backup
SKILL.md— agent instructions for OpenClaw / Claude Code / Cursor / Clineeval-guide.md— practical guide for writing binary eval criteria- Example eval configs:
weekly-report.json,search-skill.json - Compatible with
npx skills add(Vercel Labs skills ecosystem)