m4bwav/evergreen
Self-maintaining, self-proving skills and knowledge: adaptive web research on four tracks (the subject, the skills, plugins, MCP servers and knowledge graphs built for it, how others use agents on the same goal, and how they test that the job was done), never-teach-twice learnings capture, plugin-owned codemaps, a portable AI-preferences profile, and an eval suite per skill (trigger, action, outcome, decoys, baseline) that passes on evidence outside the transcript, with a bounded tuning loop that researches the subject's own testing when a skill fails, and a worth check that warns while a skill is made or changed when it reads as mostly cost and decides KEEP, TRIM or CUT from the with-versus-without A/B. Skills say what they need (tools, packages, keys, servers, models, accounts, and access to databases, telemetry, cloud roles and APIs) and what for: on a new machine or agent harness, or when something is missing, the agent explains it, installs what is safe after saying so, hands the user exact steps for the rest, drafts least-privilege access requests without ever granting itself access, and records the recipe that worked and how each attempt went for that environment. An end-of-session wrap-up harvests the session transcript and keeps only what will save time or tokens or raise quality, routed to its narrowest home. Converts existing skills and makes new ones evergreen by default. Lives in one public git repository (https://github.com/m4bwav/evergreen-protocol): every install is a clone that pulls to stay current and, only if its owner said yes at install, sends improvements to the protocol, skills and scripts back as pull requests it opens itself (log entries go straight to master from a maintainer's clone), with email as the fallback route.
Changelog: evergreen
Every change to the plugin (README.md, protocol/, skills/, scripts/, templates/, profile/, evals/), newest first, with the reason. Reasons cite RESEARCH.md findings (R-), LEARNINGS.md lessons (L-) and TESTS.md runs (T-). State in evergreen.json. Protocol: protocol/PROTOCOL.md.
Entry shape: ### C-YYYYMMDD-n · date · one-line summary, then because:, files:, and what changed. Cite section headings, not line numbers.
C-20261008-2 · 2026-10-08 · evergreen-wrapup works in a desktop or web chat through a local connector
- because: the owner's correction (plain desktop chat reaches local files and processes through connectors such as Desktop Commander, so the wrap-up should say how to work there instead of being written off)
- files: skills/evergreen-wrapup/SKILL.md (Step 1)
- A chat's own transcript is not on disk, so the harvest falls back to the conversation, which a chat usually still holds whole; the script, the stores and every edit go through the connector's process and file tools by host path, as PORTABILITY.md already describes for sandboxes.
C-20261008-1 · 2026-10-08 · Protocol 1.17, version 0.16.0: evergreen-wrapup, an end-of-session wrap-up that harvests the transcript, keeps only what pays, routes it to its narrowest home and fixes the costliest scripts and habits
- because: T-20261008-1 (KEEP: +67 points on action-6 at 0.7x the cost); the owner's request (a researched wrap-up command that consolidates a session's learnings into knowledge bases, skills, plugins and other stores, adding only what is helpful, and records and improves the scripts and processes that save time and tokens or raise quality); R-20261008-1 to R-20261008-4
- files: skills/evergreen-wrapup/SKILL.md (new), scripts/evergreen_wrapup.py (new:
wrapup [--session ID|PATH] [--json] [--top N]), scripts/evergreen.py (registers the module), scripts/evergreen_worth.py and evergreen_ab.py (--caseadds up repeated flags and acceptsa|b:CaseGlobs,case_match; L-019 recurred twice this session), LEARNINGS.md (L-019 updated), scripts/test_evergreen.py (classWrapup, four tests; 115 pass), evals/evals.json and evals/cases (trigger-12, trigger-13, decoy-7, outcome-3, action-5, action-6), evals/fixtures/wrapup-session and wrapup-long (new: session transcripts and the skills they used), TESTS.md (T-20261008-1), protocol/PROTOCOL.md (1.17; §5 End of session), README.md, AGENTS.md, CLAUDE.md, plugin.json and .claude-plugin/plugin.json (0.16.0), RESEARCH.md (R-20261008-1 to R-20261008-4), evergreen.json (counts) - Capture in the moment stays the rule; the wrap-up is the sweep for what the moment missed. The harvest reads the transcript rather than the model's recall, because compaction drops early turns and recall favours what the model said over what failed: user corrections after the first turn, refused tool calls (kept apart from errors, since a refusal is a preference), errors grouped by the line that names them with whether a retry followed, repeated command shapes (syntax, plumbing, heredoc bodies and inline code left out), results over about 2,000 tokens, calls of a minute or more, skills used, files changed per repository with the stores found there, web research and compactions. It writes nothing. The skill gates every candidate on four questions (material, non-derivable, recurring, verified), searches before writing, and routes through the skill that owns each store (evergreen-learn and -tune, everlast-capture, evergreen-setup, a plugin's curate skill), preferring a check or a script to a rule and a file read on demand to an always-loaded one. Script and process fixes are made now when small and testable, measured before and after; larger ones become a plan with the estimated saving. Unattended runs only propose edits to always-loaded files.
- Knowledge bases no skill or plugin owns (the owner asked that the wrap-up also update them, by judgement rather than codified paths: write where the session or the user makes the place clear, never search the disk for one, note in the report what had no clear home, ask when it matters, and codify a destination only once a usage pattern shows it): the harvest's K section lists markdown outside any repository that the session edited or read, grouped by vault (
.obsidian,graphify-out) or folder; files in a git repository nested inside a vault stay with the repository, and the harness's own config and memory are left out. A first draft with a registry and disk discovery was dropped the same day at the owner's word. Testtest_notes_touched_outside_repositories_are_reported_by_vault_or_folder.
C-20261006-1 · 2026-10-06 · Version 0.15.1: ready for GitHub Copilot CLI and awesome-copilot (root plugin.json; skills link the unit files in the public repository)
- because: user request (list evergreen in awesome-copilot;
vally lintrejects SKILL.md links that leave the skill folder, and the intake reads only a root or.github/plugin/manifest; the maintainer chose links to the public repository over copies) - files: plugin.json (new, Agent Plugins 1.0, 10 of the 18 keywords), .claude-plugin/plugin.json (0.15.1), skills/*/SKILL.md (Maintenance: the shared-unit line)
- The fifteen skills' links to
../../RESEARCH.md,CHANGELOG.md,LEARNINGS.mdandTESTS.mdnow point to those files on github.com/m4bwav/evergreen-protocol (master). The sentence still says they sit at the plugin root, so an agent with the plugin installed reads the local copy;evergreen.py lintstill finds each file name.
C-20261003-1 · 2026-10-03 · Ready for the Claude plugin directory: README discloses hooks, git, files and data flows; manifest links; pinned launcher text; installers downloaded before they run
- because: user request (submit evergreen to the Claude plugin directory; claude.com/docs/plugins/pre-submission-checklist, read 2026-10-03)
- files: README.md (new sections Hooks, git and files outside the project; Privacy; One public trunk, many clones: the closing paragraph now matches the public repository), .claude-plugin/plugin.json (documentationUrl, supportUrl, privacyPolicyUrl), setup/RECIPES.md (ollama, uv and claude: download the install script, then run it), agents/evergreen-researcher.md, protocol/PROTOCOL.md §4 and skills/evergreen-new/SKILL.md (
npx skills@1.7.0), protocol/TESTING.md (uvx skillsaw==0.21.0), protocol/LEARNINGS-FORMAT.md, protocol/SETUP.md, skills/evergreen-learn, evergreen-merge and evergreen-worth SKILL.md (wording) - The README now names each hook and what it runs, every commit, push, pull request and email the plugin can send and when, the files it writes outside a project, every network route, and the credentials involved (the user's own git and gh logins; an SMTP app password the user sets for the email route); hooks run only in Claude Code and Cowork. The directory's scanner pins launchers and flags piped installers and the word "pass" before a code span, so the pins were tested (
npx skills@1.7.0 --help,uvx skillsaw==0.21.0 --help), the five piped installer recipes became download-then-run pairs that the recipe parser still reads as two commands (each download tested), and six "pass--flag" phrasings now say "add" or "exceed". No script change, so the version stays 0.15.0. Checks: the self-test (111 tests, 1 skipped),claude plugin validate(one warning: the repository's own CLAUDE.md is not plugin context, as intended), 142 files, none over 256 KiB, no binaries or system files, valid names, noexport-ignore,export-substorfilterattributes.
C-20261001-2 · 2026-10-01 · Version 0.13.1: a partial baseline pass ("without the skill 4 of 5 passed") reads as failed, not "already known"; lite names a recorded verdict as the last measurement
- because: L-029 (the first dated baselines written by
worth --ab, on obsidian-notes, put a false mode 2 "already known" in every lite scan, because any "passed" counted as a full pass) - files: scripts/evergreen_worth.py (
BASELINE_COUNT,baseline_status,render_lite), scripts/test_evergreen.py (test_case_lint_finds_what_would_waste_an_ab), .claude-plugin/plugin.json (0.13.1), LEARNINGS.md (L-029) - The first "N of M passed" in a note that speaks of the run without the skill decides: all passed is "passed", fewer is "failed".
C-20261002-2 · 2026-10-02 · Protocol 1.16, version 0.15.0: access needs (a database, Application Insights, a cloud role, an API scope, a repository, a VPN) are checked by a read-only probe, walked through step by step, requested from an admin by a drafted least-privilege request, and logged attempt by attempt
- because: the owner's request (the same dependency checking for access, such as database access or access to Application Insights: record how to get it and how it went afterwards, and step the user into getting access as best it can); R-20261002-6 to R-20261002-9; T-20261002-2
- files: scripts/evergreen_setup.py (kind
accessandcheck_access: the probe through the shell with a 20 s limit and no input, the exit code and one redacted line kept, the cause named as sign-in (401), permission (403), network, or a missing tool or extension; recipe steps (indented numbered lines under a recipe line, which must now start at column 0);--request,--attempt,--result,--route,--took,--skip-access; the Attempts table, the last attempt shown beside a missing need, and a 'GRANTED?' line when a probe passes while the latest attempt is pending;stamp_verified; access is neverself), scripts/test_evergreen.py (classSetupAccess, six tests; 111 pass), protocol/SETUP.md (§1 Attempts, theaccesskind row, §5 Access new, §6 a never-self-grant rule; later sections renumbered), protocol/PROTOCOL.md (1.16; §2 row; §12 bullet), setup/RECIPES.md (## Access: azure-signin, appinsights-read, postgres-read, sqlserver-read, github-repo), templates/SETUP.md.template (access check guidance, step recipes, Attempts section), templates/SKILL.md.template, templates/MAINTENANCE-SECTION.md.template, templates/MAINTENANCE-POINTER.md.template and templates/MAINTENANCE.md.template (a refused permission is a Step 0 signal too; setup by hand covers access), skills/evergreen-setup (description: access triggers, 1,016 characters; Step 3b), evals/evals.json, evals/cases and evals/fixtures (trigger-11, decoy-6, action-4 with fixtureneeds-access), README.md, AGENTS.md, .claude-plugin/plugin.json (0.15.0), RESEARCH.md (R-20261002-6 to R-20261002-9, Current understanding, two open questions), evergreen.json (counts, tests, the worth verdict now KEEP over action-3 and action-4) - The agent can check access but never obtain it alone, getting it often takes another person and days, and how it went is the part the next person most needs, so access gets three things installs did not: a probe that names why it failed (sign in, ask for a role, fix the network, install a tool), recipes written as steps the agent walks the user through one at a time, and an Attempts log (pending, granted, refused, how long it took) whose pending rows are resolved by re-probing, the out-of-band pattern MCP's authorization spec also uses. The request draft carries the least-privilege fields from Microsoft's and OWASP's agent guidance: who, which resource, the narrowest read role, the narrowest scope, the duration, and the probe that proves it. The agent never runs a grant, role assignment, PIM activation or SQL
GRANT; agents denied a permission tend to grind on, so a 403 means stop, report, draft, log (R-20261002-8). People, internal links, servers and tenant ids stay out of public books.
C-20261002-1 · 2026-10-02 · Protocol 1.15, version 0.14.0: skills say what they need and what for, explain a missing piece instead of failing on it, install what is safe after saying so, and record the recipe that worked on each environment they meet
- because: the owner's request (when a skill is missing a dependency, whether setup steps, a command line tool or an AI model, tell the user how to install it and what it is for; do it yourself where possible and say what you are about to do; add the setup data for each new agent harness and operating system met; help with scripts where they help; fold in what other skills that do this have learned); R-20261002-1 to R-20261002-5; T-20261002-1
- files: protocol/SETUP.md (new: the file, kinds, environment keys and recipe order, when to run, the explain-act-prove-record procedure, safety rules, helper scripts, prior art), protocol/PROTOCOL.md (1.15; companion list; §2 anatomy row; §3 Step 0 bullet; §12 new), scripts/evergreen_setup.py (new:
setupwith--harness,--os,--json,--script,--log,--init,--record ... --env --how --tags --verified --to unit|store|plugin), scripts/evergreen.py (registers the module;SATELLITES: SETUP.md links the main file both ways, not every log), scripts/test_evergreen.py (classSetup, eleven tests; 105 pass), setup/RECIPES.md (new: the public shared book for Python, git, gh, Node, ffmpeg, Ollama, uv and the Claude CLI; winget ids confirmed withwinget showon 2026-10-02, the rest the vendors' documented routes, none yet tagged verified), SETUP.md (new: the plugin's own four needs, all optional; the owner's environment logged), templates/SETUP.md.template (new), templates/SKILL.md.template, templates/MAINTENANCE-SECTION.md.template and templates/MAINTENANCE-POINTER.md.template (Step 0 sentence; the pointer form also carries setup by hand), templates/MAINTENANCE.md.template (§Setup), skills/evergreen-setup (new), skills/evergreen-new (Step 3), skills/evergreen-convert (Step 2), skills/evergreen-learn (routing: install recipes go to SETUP.md), evals/evals.json and evals/cases (trigger-9, trigger-10, decoy-5, outcome-2), README.md, AGENTS.md, CLAUDE.md, .claude-plugin/plugin.json (0.14.0), evergreen.json (files.setup,setup.envs), RESEARCH.md (R-20261002-1 to R-20261002-5, Current understanding, three open questions) - A unit with dependencies keeps
SETUP.md: a Needs table (id, kind, check, what it is for, required or optional with the fallback), recipes per environment key (an operating system, a harness, a package manager on PATH, a combination such aslinux/apt, orany; most specific first, verified first) and the environments met.evergreen.py setupchecks command (with alternatives and a minimum version), python, env (never printing the value), file (globs), url and ollama needs without side effects, hands mcp, account and manual needs back to the agent, and classes each missing needself(a command recipe with no admin, manual, large, secret or paid tag: the agent says what it will run and runs it),user(exact steps for the user; every secret) ornone(find the vendor's route, say where it came from). Recipes come from the unit, then the store's private book, then the shared book. Step 0 runs the check only on a newos/harnesskey or a failure that names a missing piece, so an ordinary use still costs one read. - From the research: the check reports every missing need before anything changes, with exit codes 0, 1, 2 (R-20261002-2); package names only from a recipe or the vendor's documentation, never from memory or a workspace file, and never a changed index (R-20261002-3, R-20261002-4); secrets to the harness's secure store; explain rather than hide, the one point where evergreen departs from OpenClaw (R-20261002-5); harness detection in the vendor libraries' order, with
--harnessbecause inherited variables can name the outer tool (R-20261002-1). Recording into the shared book refuses a recipe that names a home or drive path, andverifiedstamps carry<os>/<harness>, never a hostname, so the book stays publishable.
C-20261001-1 · 2026-10-01 · Protocol 1.14, version 0.13.0: worth --lite (a free quick scan with a case lint) and worth --heavy (probe, A/B, top-ups inside the margin, spend cap); the A/B harness fixes from the first real obsidian-notes run
- because: the owner's request (a quick lite scan to detect issues and a heavy mode for thorough testing; the tested skill updated with what was learned and every token saving found); L-030 (a relative skill path turned the run folder into the plugin dir: no skill loaded, writes refused, the checker not found, a passing baseline graded 0 of 3), L-031 (five case-authoring traps no run reports), L-029 updated (
process_onlyreadandas an alternative) - files: scripts/evergreen_ab.py (absolute skill, unit and checker paths;
case_promptfills<fixture>,<vault>,<repo>,<folder>,<dir>from the case'sfilesorfixture,--varfirst;parse_streamkeeps the init event's skills and refused tool calls;environment_faultclasses a runenvironmentwhen the with arm lacks the skill, the without arm has it, writes are refused as sensitive, or a usage limit hits, and such runs leave the counts;grade_casegradesregex on the answer:together with the evidence;run_ab(append=True)adds runs to an existing aggregate-result.json; the probe fills placeholders), scripts/evergreen_worth.py (_route_onlyreadsand/all/sequence as required andor/anyas alternatives;case_lintand itscaselines in every report;render_liteand--lite, which also shows the last recorded verdict, and a recorded A/B now satisfies the 'never baselined' gap;--heavy,--max-usd,--no-probe,--force,--append,--arm with(re-measure an edited skill against the baseline already in--out),HEAVY_RUNS5 andHEAVY_MAX_USD15; no mode 3 or 4 when both arms fail everything), scripts/test_evergreen.py (three tests added or extended, 94 pass), skills/evergreen-worth/SKILL.md (Lite or heavy; value-case rules in Step 3; environment runs and hand re-grading; Step 5 feeds the skill what the run taught, with a token-saving list; description 995 characters), protocol/TESTING.md (§8 the two modes and the environment class), protocol/PROTOCOL.md (1.14), README.md, AGENTS.md, .claude-plugin/plugin.json (0.13.0), LEARNINGS.md (L-030, L-031, L-029) - Lite answers "is anything obviously wrong" in seconds for free, and catches the problems that would waste a paid run; heavy answers "is it worth it" with a measured verdict and stops at a spend cap. Both were proven on obsidian-notes on 2026-10-01: the lite case lint flagged every authoring problem found by hand (and nothing on the fixed suite), and the heavy run fired the skill in every with-arm run once paths were absolute.
C-20260930-2 · 2026-09-30 · Protocol 1.13, version 0.12.0: worth detects useless skills at will: real usage from transcripts, recorded baselines, process graders, five failure modes, catalog triage, a knowledge probe, a headless A/B that runs on native Windows, and a person's FIX or SUPERSEDED ruling
- because: the owner's request (evaluate any skill for uselessness at will, after the 2026-09-30 study of 47 public skills); L-026 (headless runs logged with the IDE's entrypoint inflated usage), L-027 (what
--setting-sources projectisolates), L-028 (peers from disabled plugins gave a false overlap warning), L-029 (most recorded baselines were predictions, and many gains rested on process graders); L-019 (value-case A/Bs skipped on native Windows since 2026-09-13); T-20260930-2 - files: scripts/evergreen_worth.py (transcript usage split interactive / subagent / headless,
enabled_plugins,installed_as, skill age,baseline_status,process_only,evidence_report,failure_modes,evidence_gaps,set_manual,render_triage; several targets and--triage,--days,--probe,--ab,--set;worth_flagsadds FIX, SUPERSEDED and UNUSED), scripts/evergreen_ab.py (new: the probe and the headless A/B with an evidence grader for trace, sequence, file, file_contains, command, file_changed and file_unchanged,or/and/all/any, placeholders, prose tool names and ' exits 0' checks; writesaggregate-result.jsoninclaude plugin eval's shape), scripts/test_evergreen.py (classWorthUseful, seven tests;test_audit_shows_test_flags_and_counts_failing_unitsnow checks its unit today, since a fixed 2026-09-01 made it due, and the suite red, from 2026-10-01), skills/evergreen-worth/SKILL.md (description and Steps 1 to 4), skills/evergreen-test/SKILL.md (Step 3: the Windows harness and dated baselines), skills/evergreen-audit/SKILL.md (worth flags and triage), protocol/TESTING.md (§6 harness row; §8 usage, evidence, probe, rulings, failure modes, triage), protocol/PROTOCOL.md (1.13), README.md, AGENTS.md, .claude-plugin/plugin.json (0.12.0), LEARNINGS.md (L-026 to L-029), TESTS.md (T-20260930-2) - A skill is judged useless in one of five ways, each now recorded with its evidence: 1 never fires (loaded, older than 14 days, no interactive or subagent call in 60 days), 2 already known (recorded baselines passed, or the probe named most key anchors), 3 no gain, 4 worse (from the A/B), 5 superseded (a ruling).
gapsnames what is unmeasured. Test traffic no longer counts as use. The A/B harness isclaude -pin fresh temp folders with--setting-sources project --no-session-persistence; it says when the skill never fired in the with arm or fired in the without arm, and counts evidence it cannot check as ungradable rather than failed. Rulings set with--setsurvive later--recordruns.
C-20260930-1 · 2026-09-30 · Protocol 1.12, version 0.11.0: evergreen-worth, a check that warns while a skill is made or changed when it reads as mostly cost, and decides KEEP, TRIM or CUT from the with-versus-without A/B
- because: the owner's request (a warning when a skill being made or updated is mostly useless, so no time or tokens go into skills the model is only slowed down by); R-20260930-1 to R-20260930-4 (39 of 49 public skills gave no gain; over 60 percent of skill bodies is not actionable and trimming raised quality; cost regressions are the largest class of skill-induced failure; context that repeats the repo hurts); TESTING.md §7 already said "a skill whose baseline keeps winning is retired", but nothing measured it or weighed gain against cost; T-20260930-1
- files: scripts/evergreen_worth.py (new:
worth,worth-hook,--wrap,--record; thresholds as named constants), scripts/evergreen.py (registers the module's commands like evergreen_sync;unit_flagsaddsworth:CUT,worth:TRIM,worth:SUSPECTfrom a recorded verdict; the session-start line never shows them), scripts/test_evergreen.py (classWorth, nine tests), scripts/evergreen-hook.sh and scripts/evergreen-hook.ps1 (worth/-Worth: returns before Python unless the payload names a SKILL.md), hooks/hooks.json (a secondPostToolUseentry, matcherWrite|Edit|MultiEdit, timeout 10), skills/evergreen-worth/SKILL.md (new), skills/evergreen-new (Step 4: static reading on the draft, A/B verdict at hand-over), skills/evergreen-refresh (step 5), skills/evergreen-tune (Step 6), skills/evergreen-convert (Step 4), skills/evergreen-test (Step 5 hands the retirement candidate to evergreen-worth), protocol/TESTING.md (§8 Worth), protocol/PROTOCOL.md (1.12; §8 and §11), templates/MAINTENANCE.md.template and templates/MAINTENANCE-POINTER.md.template (one sentence each), evals/evals.json and evals/cases (trigger-7, trigger-8, decoy-4), README.md, AGENTS.md, CLAUDE.md, .claude-plugin/plugin.json (0.11.0), RESEARCH.md (R-20260930-1 to R-20260930-4, Current understanding, two testing-track queries), TESTS.md (T-20260930-1) - Static reading, seconds and free: listing and body cost in estimated tokens; the share of prose sentences with a specific anchor (code, path, number, URL, flag, quoted phrasing, a name mid-sentence); general-advice sentences; repeats; hard imperatives; overlap with the repo's README, AGENTS.md or CLAUDE.md; bundled files the skill never names; the nearest installed skill by TF-IDF cosine of the descriptions (smoothed IDF, so two lone skills still compare; not warned when the two name each other);
--against <rev>for what an edit added. It says LEAN, CHECK or SUSPECT and never claims a skill helps. The A/B readsclaude plugin evalresults (action and outcome cases only, the latest per case), compares pass rates against a noise margin of twice the standard error (at least one run's worth), and weighs the gain against the cost, turn and time ratios. A static SUSPECT and the A/B verdicts go to the user; retiring a skill needs the user's explicit yes.--wrapbuilds a throwaway plugin around a standalone skill so the same harness can run it. - The hook is a new matcher on the existing
PostToolUseevent, not a new event (AGENTS.md, Ask first). Measured on the owner's Windows machine: 57 ms for an edit to any other file (the shell returns before probing for Python), 0.4 s for a SKILL.md. It warns once per warning per session and stays silent on a LEAN reading. - A first pass over the owner's installed skills showed the checks separating: a padded example skill read 0 percent specific with ten general-advice sentences (SUSPECT), the plugin's own fourteen skills 63 to 88 percent, two installed skills over 5,000 body tokens, and one bundled file never named; two false positives found on that pass (sibling skills that already name each other; files under a folder the skill names only as part of another path) were fixed and have tests.
C-20260929-1 · 2026-09-29 · Headless eval runs share the session's filesystem: commit before a suite, end it with git status, review what the runs wrote
- because: L-025 (from wikiwright:L-017, wikiwright:T-20260928-3 and wikiwright:T-20260929-1)
- files: protocol/TESTING.md (§6, a paragraph after the preference line), templates/MAINTENANCE.md.template (§Tests (skills)), LEARNINGS.md (L-025)
claude plugin evalgives each run a throwaway workspace;claude -pruns and the tester agent do not, so a skill-arm run that follows "capture learnings" edits the unit's source, and a run can leave untracked files there thatgit diffmisses. The rule: commit the source first, create nothing a case could see, finish withgit status --shortof the source, and review each eval-written change before committing, recording it under side effects in theT-entry. Protocol text clarified, no behaviour change in the scripts, so no version or protocol bump;skills/evergreen-test/SKILL.mdis untouched (it defers to TESTING.md §6), so no suite re-run is owed.
C-20260926-1 · 2026-09-26 · evergreen-audit and evergreen-refresh descriptions split by scope: the whole catalog against one named unit
- because: the owner's request (resolve the pair's selection conflict without losing a trigger); context-health's
selectioncheck on the owner's install scored the two descriptions at TF-IDF cosine 0.49, over its 0.45 "possibly confusable" line, with "due" and "unit" carrying almost all of the overlap; skill-selection research (arXiv 2606.30775) found that editing both sides of a confused pair adds under 0.5 percent over editing one, so each side got the smallest edit that separates it; T-20260926-1 - files: skills/evergreen-audit/SKILL.md (description), skills/evergreen-refresh/SKILL.md (description), TESTS.md (T-20260926-1)
- Audit now opens with a status overview of the whole catalog and calls its triggers catalog-wide ('evergreen status', 'what's stale', 'anything due', 'audit my skills', 'check the skills'); it still hands what is due to evergreen-refresh and ends "Refreshing a single unit is evergreen-refresh." Refresh now opens with one named unit and keeps every trigger about that unit ('refresh X', 'update X's research', 'is X stale', 'is this still current', 'check for changes to X'), the past-due report, the contradiction flag, and working through the whole list an audit hands over; it ends "Listing what is due across the catalog is evergreen-audit." No trigger phrase was dropped, nothing moved to
when_to_use(only Claude Code reads it), and the skill bodies are unchanged. The pair scores 0.30 on the same catalog; audit is 686 characters, refresh 489.
C-20260923-17 · 2026-09-23 · Protocol 1.11, version 0.10.0: the interval rule halves on a major change, cuts by a third on a change and grows by a quarter when quiet
- because: R-20260923-12 (a sweep of twenty candidate step sizes on
bench_intervals.py: for 4 to 6 percent more checks over the seven middle classes the chosen steps hold a stale material claim 10 percent less of the time and see a material change 10 percent sooner, on every seed and horizon tried, and the gap to even spacing at the same count falls from 3.7 to 2.4 points); R-20260923-10 (the 1.10 steps reacted too hard under memoryless change); the owner's decision of 2026-09-23 (quality and reliability first, cost also a concern) - files: scripts/evergreen.py (
RULE,compute_next), scripts/bench_intervals.py (--rule; the header names the steps), scripts/test_evergreen.py (the interval-rule cases), protocol/INTERVALS.md (intro, hand rule steps 2, 3 and 5, worked example, migration, how the rule compares), protocol/PROTOCOL.md (version line), templates/MAINTENANCE.md.template and templates/MAINTENANCE-POINTER.md.template (the hand rule), README.md (§How a unit behaves, step 3), .claude-plugin/plugin.json (0.10.0), TESTS.md (T-20260923-4), RESEARCH.md (R-20260923-12; the change-rate anchor and the promotion-on-one-major-change questions resolved), evergreen.json (counts) - The step sizes are module constants (
RULE), so the benchmark scores a candidate on the shipped code path:--rule major_div=4,change_div=2,quiet_mul=1.5reproduces the 1.10 numbers of R-20260923-10 exactly. Immediate promotion on one major change stays, but with halving it fires only when the halved interval falls below the tier's floor (glacial, or a unit already near its floor); otherwise the unit pins at the floor and the second material change promotes it through the streak rule. A windowed change-rate anchor was scored and rejected. Existing units need no migration: the rule reads the stored interval and applies the new steps at the next check.
C-20260923-16 · 2026-09-23 · Version 0.9.1: only the evergreen plugin itself becomes the registry's plugin_root
- because: L-024 (a
checkedon another plugin unit, dandy and then everlast, repointedplugin_rootaway from the evergreen plugin) - files: scripts/evergreen.py (
register), scripts/test_evergreen.py (test_register_sets_plugin_root_only_for_the_evergreen_plugin), .claude-plugin/plugin.json (0.9.1) registersetsplugin_rootonly when the unit's folder holdsscripts/evergreen.pyandprotocol/PROTOCOL.md. Other plugin units still register as units. An install whose registry already points elsewhere is repaired by the nextcheckedon the evergreen plugin, or by editingplugin_rootinregistry.json.
C-20260923-15 · 2026-09-23 · evergreen.py eval-export: claude plugin eval case folders written from evals/evals.json
- because: R-20260923-7 (the harness is documented and preferred, and its case format is separate from
evals.json); T-20260923-3; the open question on keeping both by hand, now resolved - files: scripts/evergreen.py (
eval-exportcommand;eval_export,export_case,case_skill,skill_input_match), scripts/test_evergreen.py (test_eval_export_writes_plugin_eval_case_folders), protocol/TESTING.md (§3, §6), skills/evergreen-test (Step 3), README.md (§What you get, §Using the scripts), RESEARCH.md (the open question resolved), TESTS.md (T-20260923-3) - One folder per case:
prompt.mdwith the case's runs, turn and time limits and its tools (read-only for trigger, decoy and outcome cases, which keeps them runnable on native Windows; Bash, Write and Edit for action cases, which need--allow-toolsand a sandbox), plus graders: a trigger istool_usedonSkillwhose JSON input names the skill (with or without aplugin:prefix), a decoy the same withmax: 0and a name prefix so a whole plugin stays quiet, trace, file and log evidence becometool_used,file_existsandregex, and an outcome'sregex ...:expectations becomeregexgraders with the rest judged by onellmgrader. Command evidence andoralternatives have no counterpart (no code graders; every grader is required), and the command says so. Existing folders are kept unless--force, because they are often tuned by hand (this plugin's own were, before this command existed). Part of 0.9.0.
C-20260923-14 · 2026-09-23 · Protocol 1.10, version 0.9.0: claim-level recheck dates, typed links, search by meaning, upkeep at session start, a benchmark of the schedule, and four corrections
- because: the owner's request ("update and improve evergreen with everything learned", addressing, where possible without reducing performance or usefulness, no meaning-based search, coarse tracking of time and links, upkeep that depends on the agent, and no benchmark score); C-20260923-1 to C-20260923-13; R-20260923-11
- files: .claude-plugin/plugin.json (0.9.0), protocol/PROTOCOL.md (version line), evergreen.json (the four volatile claims restated as objects with
checkedandrecheck_daysafter the use-time check of R-20260923-11; the check recorded withchecked; counts; the suite run withtested), scripts/test_evergreen.py (a merge fixture that assumed the real state was last checked before 2026-09-12 now uses a date that stays in the future) - For a unit, 1.10 means: a volatile claim can carry its own check date, and Step 0 re-checks only the due ones (plain strings keep their old meaning); a
Related:line can say how two documents relate, and the lint checks it;evergreen.py searchfinds a log entry by what it says, in every registered unit, for the write-time gate and for refresh step 3; the session-start hook also reports due claims and learnings pastconsolidate_every; every refresh records what it cost;scripts/bench_intervals.pyscores the schedule. Left alone, because the owner decides them (AGENTS.md, Ask first): the interval rule, magnitude thresholds, tier bounds and defaults, hook events, and the entry ID grammar. What the benchmark suggests for the rule is in RESEARCH.md's open questions.
C-20260923-13 · 2026-09-23 · The scripts write LF on every platform: a state write on Windows no longer flips evergreen.json to CRLF
- because: L-020 (updated: the scripts themselves broke its rule); found when this release's own
checkedturned every line of evergreen.json into a CRLF line, the same whole-file diff a publish commit showed on 2026-09-22 - files: scripts/evergreen.py (
write_lf, used bysave_state,save_registry, the templates and scaffold writes and the install prompt; the use log appends withnewline="\n"), scripts/test_evergreen.py (test_state_and_scaffold_files_are_written_lf), LEARNINGS.md (L-020) Path.write_texttranslates\nto CRLF on Windows and itsnewline=argument is Python 3.10+, sowrite_lfwrites bytes. With* -textin .gitattributes nothing normalised the result, so each check recorded on Windows rewrote the whole state file and eachinitthere scaffolded CRLF companions; merges between a Windows clone and any other saw whole-file conflicts in evergreen.json. evergreen_sync.py's merge already wrote bytes with the file's own line endings; its other writes go to the store and the outbox, not to a unit.
C-20260923-12 · 2026-09-23 · Suite: trigger cases for evergreen-learn and evergreen-audit; the decoy rubric judges behaviour, not vocabulary; first run under the documented claude plugin eval
- because: T-20260923-2, L-023, R-20260923-7; this release edited four skills (refresh, learn, test, audit), and learn and audit had no trigger case
- files: evals/evals.json (trigger-4, trigger-5), evals/cases/trigger-4 and evals/cases/trigger-5 (prompt.md, graders/fires.md, graders/outcome.md), evals/cases/decoy-1, decoy-2 and decoy-3 (graders/outcome.md), TESTS.md (T-20260923-2), LEARNINGS.md (L-023; L-019 updated)
- Triggers 18 of 18 runs with the plugin and 0 of 18 without; decoys quiet 18 of 18. The action and outcome cases were not run this time: the action cases grant Bash, which the harness refuses on native Windows (L-019), and no edit changed the actions they exercise.
C-20260923-11 · 2026-09-23 · Merge and publish hygiene: commit messages keep the first changed path whole, and lint reports a duplicate entry ID
- because: L-022;
.gitattributesand PROTOCOL §10 saidlintreports a duplicate ID after a union merge, and it did not: a merge of this release into a scratch clone of the owner's private fork produced one (the fork relabels the heading of T-20260913-1 with its host name, and an entry inserted just above that heading makes the union driver keep both copies) - files: scripts/evergreen_sync.py (
changed_pathsparses the porcelain status columns withPORCELAIN_RE), scripts/evergreen.py (check_linksreports an entry ID defined twice in one file), scripts/test_evergreen.py (test_changed_paths_keep_the_first_path_whole,test_duplicate_entry_ids_after_a_union_merge_are_reported), LEARNINGS.md (L-022) - The file list in a publish commit comes from
git status --porcelain; the sharedgit()helper strips its output, so a first line " M LEARNINGS.md" became "M LEARNINGS.md" and a fixed slice cut it to "EARNINGS.md". Staging was never affected (git add -A).
C-20260923-10 · 2026-09-23 · AGENTS.md in Claude Code 2.1.277: keep the @AGENTS.md import; a prose pointer loads nothing
- because: R-20260923-9
- files: protocol/PORTABILITY.md (§The lowest common denominator; the Claude Code row), templates/CLAUDE.md.snippet (a comment on why the import stays; evergreen-publish added to the skill list), CLAUDE.md (the thirteen skills and three subagents; what the SessionStart hook now reports), RESEARCH.md (Current understanding)
- The snippet and this repository's own CLAUDE.md already import, so no converted repo has to change; evergreen-convert and evergreen-new write no CLAUDE.md of their own (README points at the snippet), so they needed no edit.
C-20260923-9 · 2026-09-23 · A benchmark for the refresh schedule: scripts/bench_intervals.py
- because: R-20260923-1 (the survey's open problem 12.3; no published benchmark for refresh policies), R-20260923-10, T-20260923-1; the owner's request (no benchmark score)
- files: scripts/bench_intervals.py (new), scripts/test_evergreen.py (
test_bench_intervals_smoke), protocol/INTERVALS.md (§How the rule compares), README.md (§What you get, §Using the scripts, §Testing), AGENTS.md (§Commands and structure line), RESEARCH.md (R-20260923-10; open questions) - It imports
new_state,compute_nextandparse_whenfrom evergreen.py (the smoke test asserts it runs the same function) and compares the rule with a fixed interval at the rule's own check count, the tier's start interval held fixed, a fixed 14 days, a FreshCache-style constant per class and an oracle, on seeded synthetic units: Poisson changes with mean gaps from 2 to 365 days, a shifting class and a bursty class, magnitudes from a stated mix. Stdlib only, about a second for the default run. The rule is unchanged; the open questions carry what the numbers suggest.
C-20260923-8 · 2026-09-23 · Upkeep that does not wait for the agent: the session-start audit reports due claims and learnings due for consolidation, and every refresh records its cost
- because: R-20260923-3, R-20260923-1; the owner's request (upkeep that depends on the agent)
- files: scripts/evergreen.py (
unit_flags,consolidation_due,learning_entries,consolidate_limit;audit --briefshows a verify-at-use unit only when claims are due and adds one closing line for due claims and one for consolidation;statusand the full audit showclaims due Nandconsolidate:N>M), scripts/test_evergreen.py, agents/evergreen-researcher.md (the closing## Cost:line), skills/evergreen-refresh (Step 2; Step 4 puts the cost line in the--note), skills/evergreen-audit (the flags and what to do about them), skills/evergreen-learn (the consolidation flag), protocol/PROTOCOL.md (§4 steps 2 and 7, §6, §7), protocol/LEARNINGS-FORMAT.md (§Budgets and consolidation), templates/MAINTENANCE.md.template (refresh step 7), README.md, CLAUDE.md - The existing SessionStart hook runs the same
audit --brief, so no hook event was added; it stays silent when nothing needs attention (a verify-at-use unit with nothing due now prints nothing, where 1.9 printed it every session), and the extra cost is one small read of each unit's LEARNINGS.md. No research cap was added, because a cap would cut research quality; the cost line makes the spend visible, and a budget is an open question.
C-20260923-7 · 2026-09-23 · evergreen.py search: BM25 over every registered unit's log entries, for the write-time gate and refresh step 3
- because: R-20260923-4; the owner's request (no meaning-based search)
- files: scripts/evergreen.py (
searchcommand;search_units,search_corpus,split_entries,search_tokens,stem,bm25,entry_title), scripts/test_evergreen.py (test_search_ranks_log_entries_across_registered_units), skills/evergreen-learn (Step 3), skills/evergreen-refresh (Step 3), protocol/PROTOCOL.md (§4 step 3, §5), protocol/LEARNINGS-FORMAT.md (§Write-time gate), templates/LEARNINGS.md.template, templates/MAINTENANCE.md.template, README.md - Entries are the
###sections of LEARNINGS, RESEARCH, CHANGELOG and TESTS (archives included; template examples inside comments excluded) of the units inEVERGREEN_HOME/registry.json, else the unit around the current folder, else the plugin. Tokens: lowercase words with light stemming (s, es, ed, ing, no stem under three letters) plus code-like tokens kept whole (evergreen.py,verify_at_use,L-021); a heading term counts three times. One line per hit: score, unit, entry ID, heading,path:line;--jsonfor scripts. No index file and no dependency: it reads the logs each time, milliseconds at this scale.
C-20260923-6 · 2026-09-23 · Typed Related: lines, the grammar Everlast adopts too, and a lint for them
- because: R-20260923-5; the owner's request (coarse tracking of links)
- files: protocol/PROTOCOL.md (§9), scripts/evergreen.py (
parse_related,split_related,related_problems,unit_markdown;lintreports an unknown label, a target that does not exist and a wikilink on a Related line, and skips code fences, inline code and comments), scripts/test_evergreen.py (test_typed_related_lines_are_linted), templates/AGENTS.md.snippet, templates/MAINTENANCE.md.template (§Links and budgets) Related: supersedes [title](path); builds on [title](path), [title](path); see also [title](path), with the labels supersedes, superseded by, contradicts, builds on and see also; an unlabelled link counts as see also, so every older line passes, and a unit with no Related line is fine. The links stay relative markdown links, so graph editors show the same graph.
C-20260923-5 · 2026-09-23 · Claim-level recheck dates: a volatile claim can carry checked and recheck_days, and Step 0 re-checks only the due ones
- because: R-20260923-2, R-20260923-11; the owner's request (coarse tracking of time)
- files: scripts/evergreen.py (
claim_text,claim_recheck_days,claim_due,claims_due,claims_view, theclaimscommand;freshnesscounts due claims;statusandauditprintclaims due N), scripts/test_evergreen.py (test_volatile_claims_accept_both_shapes,test_claims_command_lists_stamps_and_adds,test_brief_audit_reports_due_claims_and_consolidation_only_when_due), protocol/PROTOCOL.md (§2 table, §3), protocol/INTERVALS.md (the verify-at-use wording in §Tier migration), templates/MAINTENANCE.md.template, templates/MAINTENANCE-POINTER.md.template, templates/MAINTENANCE-SECTION.md.template, templates/SKILL.md.template, templates/AGENTS.md.snippet (Step 0), skills/evergreen-refresh (Step 4), skills/evergreen-audit (Step 1, Step 3), README.md - A plain string keeps its meaning (due at every use), so every existing unit behaves as before until its claims are stamped;
claims <unit> --stamp dueturns a re-checked string into an object dated today. The interval rule never reads the claims, andchecked --use-timeis unchanged; evergreen_sync's state merge copies the list whole from the newer state, whatever its shape.
C-20260923-4 · 2026-09-23 · The counters come from ACE; the promote-and-retire rule extends it
- because: R-20260923-8
- files: protocol/PROTOCOL.md (§6), protocol/LEARNINGS-FORMAT.md (the evidence line; §Promotion and retirement), RESEARCH.md (Current understanding; an open question on a minimum count before retiring)
- Wording and attribution only; the rule is unchanged.
C-20260923-3 · 2026-09-23 · claude plugin eval is documented again (Claude Code 2.1.269) and preferred where it runs
- because: R-20260923-7
- files: protocol/TESTING.md (§3: the harness reads its own case folders, generated from
evals.json; §6: the row rewritten from the docs and moved back to first), protocol/PORTABILITY.md (Claude Code row), skills/evergreen-test (Step 3: prefer the plugin-eval harness, the grader mapping,--no-publish, WSL2 for cases that grant a shell; the baseline arm in Step 3.1; Step 5 wording), README.md (§Testing), AGENTS.md (§Commands), RESEARCH.md (Current understanding; the open question resolved)
C-20260923-2 · 2026-09-23 · SkillsBench v4 numbers
- because: R-20260923-6
- files: protocol/PROTOCOL.md (§4, Ranking the tooling track), RESEARCH.md (Current understanding)
- "About 16 points, a sixth of tasks worse, self-generated skills add nothing" becomes 33.9 to 50.5 percent (+16.6 points), 13 of 87 tasks worse, and self-generated skills 8.1 to 11.5 points below no skills; R-20260917-2 keeps the earlier version's numbers as its own record.
C-20260923-1 · 2026-09-23 · L-021 from the private fork, carried with its ID; L-020's stray control characters written as escapes
- because: L-021 (a general lesson from the owner's private fork, 2026-09-22: an entry inserted before the first
### C-lands inside the header'sEntry shape:line); L-020 itself (a learning about LF endings that held two literal carriage returns) - files: LEARNINGS.md
- L-021 is copied byte for byte so the fork merges cleanly, and this release's CHANGELOG entries were anchored on the newest real heading, as it says. In L-020, two carriage returns and a line break that had replaced
\rand\ninside backquotes are now the two-character escapes, so LEARNINGS.md is LF-only again.
C-20260918-3 · 2026-09-18 · Protocol 1.9: indexes first, with a shape, a budget, tiering, an exclusion list, an optional single back-link and a way to prove they pay; version 0.8.3
- because: the owner's request (indexes that link most docs and are a net positive for agents and people alike, without costing agent performance; back-links only where they earn their place); R-20260918-2
- files: protocol/PROTOCOL.md §9 (the linking paragraph became seven rules), templates/AGENTS.md.snippet, RESEARCH.md (R-20260918-2), .claude-plugin/plugin.json (0.8.3)
- The 1.8 text asked for an index per folder and links between documents; 1.9 says what an index line looks like (
[title](path): when to read it), how big an index may be (about 200 lines, read on demand, never pasted into the always-on file), when a folder index is redundant, what to exclude, that a back-link to the index is optional and at most one line, and how to test an index (paired runs, median of five, first correct open within three tool calls).
C-20260918-2 · 2026-09-18 · Protocol 1.8: documents link each other (index per folder, Related lines, relative markdown links, no wikilinks); version 0.8.2
- because: the owner's request (agent-written markdown should read as a linked graph for people in Obsidian and for later agents, if only through an index file per folder); no script change
- files: protocol/PROTOCOL.md §9 (new paragraph; version line), templates/AGENTS.md.snippet (one bullet), .claude-plugin/plugin.json (0.8.2)
- §2 already required companions to link each other; §9 now extends the rule to every markdown folder (an index that links its files) and to cross-document references (relative markdown links, a
Related:line for dependencies), and says why wikilinks stay out of shared repositories.
C-20260918-1 · 2026-09-18 · Provenance tiebreaker: prefer skills written by the latest frontier models at their highest reasoning setting, when readable or inferable
- because: the owner's request; consistent with R-20260917-2 (quality signals beyond popularity) and with the evidence-first rule, which this never outranks
- files: protocol/PROTOCOL.md §4 (Ranking the tooling track), templates/RESEARCH.md.template (search plan line), agents/evergreen-researcher.md
- Where to read it:
modelorreasoningfields in evals.json, eval_metadata.json or frontmatter; README or changelog credits; commit messages and pull requests; TESTS.md harness entries. Read beats inferred; a style-only inference is a guess.
C-20260917-5 · 2026-09-17 · Protocol 1.7: a refresh answers five questions about the area (newest, most used, most discussed, converged thinking, practitioner-built) and ranks tooling by tier, velocity and a quality gate; version 0.8.1
- because: the owner's question ("does the protocol research the latest in a skill's area and the latest skills of that type; it should also look at the most popularly used thinking and skills, the most discussed thinking, and skills created for work in the area"); R-20260917-2
- files: protocol/PROTOCOL.md §4 (the five questions; "Ranking the tooling track": tiered sources with the tier recorded, velocity not totals with meta-skills excluded, the practitioner quality gate, paired-eval preference with one to three adopts per topic, the supersession sweep), templates/RESEARCH.md.template (new queries: most used, practitioner test, supersession, most discussed and converged thinking), agents/evergreen-researcher.md, skills/evergreen-refresh/SKILL.md
- Existing units' search plans gain the new queries at their next refresh (§4 step 1 already adds missing template lines and logs the edit).
C-20260917-4 · 2026-09-17 · Protocol 1.6: the install asks once whether learnings may go upstream; no means nothing ever leaves the machine; version 0.8.0
- because: the owner's request ("ask the user on install whether they want to do pull requests on learnings; if not, updates only flow from public to the local copy, no pull requests or emails or anything"); R-20260917-1 (opt-in is the norm after the Go and Rust decisions, silent phoning home is reported as a defect, agent pull requests need a named consenting human)
- files: scripts/evergreen.py (
contribute yes|no|status;contribute_setting()with precedenceEVERGREEN_CONTRIBUTE>DO_NOT_TRACK>CI> registry, undecided = None; the answer, its date and the asking version stored in the store's registry.json, never in the repository; the brief session-start audit nudges once per session until decided), scripts/evergreen_sync.py (contribute_gateinpublishandnotify:noand undecided send nothing unattended, an explicit call explains; pull requests open as drafts with a consent line naming the reviewer), scripts/test_evergreen.py (gate cases, DO_NOT_TRACK, brief-audit nudge; fixtures answer yes), protocol/PROTOCOL.md §10 (the condition, wording, storage, precedence, no re-ask on version bumps), README §How it works, templates/INSTALL-PROMPT-GIT.txt and INSTALL-PROMPT.txt (the question the installing agent must ask, withnowhen unanswered), skills/evergreen-audit, evergreen-publish, evergreen-notify - Existing installs are undecided after this update, so nothing leaves them until the owner answers; the maintainer's own install records yes. Not done: an eval case that proves an agent with
contribute nonever callsgh pr create(the unit test covers the script; TESTING.md's trace assertion is the follow-up).
C-20260917-3 · 2026-09-17 · This repository is the official version; private forks pull from it and contribute only general lessons
- because: the owner's request ("the public one is the official version; the private one is just in case I need private customizations I don't want to share; changes should mostly go from the public one to the private one, not the other way around, unless we happen on a meaningful lesson we want to pull request into the public one")
- files: protocol/PROTOCOL.md §10 (the fork model: official repository → fork by
pull; fork → official only as a pull request carrying the general change, never the fork's private content; the scrubbed-snapshot route is retired), README.md §One public trunk - The owner's own private fork now tracks this repository as
upstreamwith a shared history, so a routine update isgit pull upstream master; the one-off scrub that produced the first public snapshot is no longer part of the flow.
C-20260917-2 · 2026-09-17 · Cross-platform by default and scripts-when-they-save-tokens (protocol 1.5 §8); portability fixes; CI on three operating systems
- because: the owner's requests ("make sure the protocol scripts are cross platform compatible", "skills and plugins that use it should prefer cross platform compatibility where it isn't onerous", "prefer scripts if they would be a net reduction in token consumption without adding noise or mistakes", "have the pipeline test the scripts on different OS"); L-020 (the sh hook shipped with CRLF endings and would not have run on Mac or Linux); a portability audit of the scripts run 2026-09-17
- files: protocol/PROTOCOL.md §8 (two new defaults: cross-platform wherever not onerous, and a script wherever it is a net token saving without noise or mistakes), protocol/PORTABILITY.md §Cross-platform by default (the checklist table), skills/evergreen-new §Step 2, skills/evergreen-convert, skills/evergreen-test (second-platform run or
untested elsewhere), all skills'EGline and both INSTALL-PROMPT templates (pythonon Windows,python3on macOS and Linux), README §Install (prerequisites), .gitattributes (*.sh text eol=lf; the whole tree normalised to LF), scripts/evergreen-hook.sh and .ps1 (each interpreter candidate is probed with-c "import sys", so the Windows Store stub and the macOS Command Line Tools stub are skipped), scripts/evergreen.py (stdout and stderr reconfigured to UTF-8; the use-log payload read as UTF-8 bytes; git output decoded as UTF-8;.shfiles normalised to LF byshipped_bytesandexport;audit --brieffrom the hook walks the cwd two levels deep instead of four), scripts/evergreen_sync.py (UTF-8 on every git and gh call;core.quotepath=offso non-ASCII paths appear unescaped), scripts/test_evergreen.py (test_sh_hook_is_lf_everywhere), .github/workflows/tests.yml (new: the suite, lint, links and a hook smoke run on ubuntu, macos and windows with Python 3.9 and 3.13, on every push and pull request to the public trunk) - No second operating system was reachable from the development machine (the LAN Mac refuses ssh, no WSL, no Docker), so the Mac and Linux runs are CI's job from this version on; Docker was considered and rejected (no macOS containers, Windows containers are heavy) in favour of GitHub-hosted runners, which are free for the public repository. The first CI run (public trunk, run 35280055167) failed 5 of 6 jobs and found two things the development machine hid: the git-transport fixtures committed a seed repository without a git identity (runners have none; the test copies were empty), and
Path.write_text(newline=)is Python 3.10+, which broke the bundle writer on 3.9; both fixed the same day and the workflow sets a git identity. macOS with 3.13 passed first time. Not done from the audit: zip permission bits for extracted.shfiles (hooks are invoked assh file, so it does not matter yet) and the BOMpack.ps1writes into INSTALL.txt (cosmetic).
C-20260917-1 · 2026-09-17 · Protocol 1.5: the trunk is the public repository; improvements go back as pull requests from every clone; version 0.7.0
- because: the owner's request ("make my evergreen protocol repo public ... make a new version of the evergreen protocol that references the public repo and mentions automatically doing pull requests with any improvements to the protocol to keep updated with the changes to the public repo version"); the privacy audit of 2026-09-17 (the private trunk's history carried the owner's machine profile, so the public repository started from a scrubbed single-commit snapshot)
- files: protocol/PROTOCOL.md §10 (renamed "One public trunk, many clones"; three new rules: pull requests are automatic for protocol, skill, script, template, agent, hook and README changes from every clone including the maintainer's; staying current is a pull; private facts never reach the public trunk and a clone holding real ones publishes through a scrubbed snapshot), README.md §Install and §One public trunk, AGENTS.md, skills/evergreen-publish/SKILL.md §Rules, .claude-plugin/plugin.json (0.7.0, homepage and repository)
- The public trunk is https://github.com/m4bwav/evergreen-protocol. Its tree is the
pack --shareform of this one: no ownerprofile/(a blank starter unit, profile C-20260917-1, keeps routing and the test suite working),owner@example.comand%USERPROFILE%\.evergreenin the config,<plugin root>for absolute paths,owner-pcfor the hostname, "the owner" in quoted requests; the scrub map is kept outside the repository. The script's existing mechanics already do the pull-request half for any clone of the public repository (publishfrom a non-maintainer clone,publish --branchfrom the maintainer's); what 1.5 adds is the rule that protocol-level improvements always take that route, and the scrubbed-snapshot route for a private working copy. Not yet in the script: apublish --publicthat builds the scrubbed snapshot and opens the pull request in one step; until then the owner's private copy publishes through the scrub by hand.
C-20260913-2 · 2026-09-13 · First run of the plugin's own suite: harness cases under evals/cases, baselines filled, two action cases marked redundant
- because: T-20260913-1 (10/10, action-1 and action-2 also pass without the skill); L-019 (Windows harness limits); the owner's request ("give the plugin an eval run")
- files: evals/cases//prompt.md and graders/*.md (new: the suite in
claude plugin evalform, generated from evals/evals.json; harness cases carry no shell tool on Windows), evals/evals.json (baselinefilled for every case), TESTS.md (T-20260913-1), LEARNINGS.md (L-019), evergreen.json (testsblock), .gitignore (evals/results/) - The suite ran for the first time: triggers fire 12 of 12 with the plugin and 0 of 12 without, decoys stay quiet 18 of 18, the outcome case passes its regex 3 of 3, and both action cases leave the evidence they name in every run. Both action cases also pass with the skill absent, because their prompts hand the agent the script; the next edit to the suite should sharpen them (TESTS.md T-20260913-1 says how). No skill was changed, so no tuning loop ran.
C-20260913-1 · 2026-09-13 · Git transport: the trunk is a repository, every install a clone; publish pushes master or opens a PR; protocol 1.4, version 0.6.0
- because: the owner's request ("instead of emailing stuff around, the protocol will be put into a git repo. If changes are made by a maintainer of the repo just push them to master, but if they are made by others then updates will be pull requests"); L-013 and L-018 (attachments through a mail connector are capped per call, so the email route needed a browser and a Drive picker to carry a bundle at all)
- files: scripts/evergreen_sync.py (new git section:
update_transport,git_config,git_role,commit_message,publish,pull,git_status_summary,publishandpullsubcommands;notifydispatches topublishunless the transport is email;whereprints the transport, remote, role, ahead/behind and last publish), scripts/evergreen.py (maybe_notifyhonoursgit.auto, help text), scripts/test_evergreen.py (GitTransport: three tests against a bare repository; three stale fixtures repaired: the renumbering test used dates the real CHANGELOG now occupies, the diverged-trunk test set alast_checkedthe state already had, the email test now pins the email transport), evergreen.config.json (update.transport,gitblock;_note), .gitattributes (merge=unionon the six append-only logs), scripts/evergreen-hook.sh (comment), skills/evergreen-publish (new), skills/evergreen-notify (fallback framing), skills/evergreen-merge (§Pull requests, §Email bundles, §Step 4), skills/evergreen-diff §Step 2 and §Step 4, skills/evergreen-pack intro, templates/INSTALL-PROMPT-GIT.txt (new), protocol/PROTOCOL.md (intro version, §4 step 7, §7, §10 rewritten), protocol/PORTABILITY.md (§EVERGREEN_HOME layout, Claude Code row, §Self-updates), README.md (§What you get, §How a unit behaves item 6, §Install, §Using the scripts, §One trunk, many clones, §Packaging), AGENTS.md, .claude-plugin/plugin.json (0.6.0), evergreen.json (topic), profile/CHANGELOG.md and profile/ENVIRONMENTS.md (three cross-unit references qualified aslocal-delegate:L-005/L-006, which the plugin's own lint test had been failing on) - The email route was built when the plugin had no shared host: a bundle per change, a transport ladder, an agent with a browser when nothing else worked, and a bespoke merge at the trunk. A git repository does all of that with tools every machine already has. The trunk is now
git.upstreamin the config (a private GitHub repository, branchmaster); every install is a clone.publishcommits the plugin's tree with a subject naming the new entries, asks the remote which role the clone has (a dry-run push;git.roleorEVERGREEN_GIT_ROLEpins it), rebases and pushesmasterfor a maintainer, or pushesupdate/<env>-<stamp>and opens a pull request withghfor anyone else; a rebase that conflicts becomes a pull request, never a forced push.pullfast-forwards a clone and says when a reinstall is due. The six append-only logs merge by union through.gitattributes, so concurrent entries need no tooling.checked,bump,testedand the SessionEnd hook go through the samenotifyentry point, which ispublishunlessupdate.transportisemail; the owner's editing clone stays quiet on the hook as before (notify.report_trunk). The email route, bundles,mergeandpackremain for machines with no access to the host and for Cowork's.plugin. Suite:python scripts/test_evergreen.pypasses 62 (one skipped). The plugin's own eval suite gains a trigger case and a decoy forevergreen-publishand has still not been run (T-20260904-1); a run is owed.
C-20260910-2 · 2026-09-10 · Notify: the connector route is capped by per-call output; the Drive-picker route is preferred when a browser is signed in
- because: L-018 (bundle 20260909-2229:
changes.patchas base64 overflowed one Gmail-connector call; Claude in Chrome with the Drive picker sent both attachments in one message) - files: skills/evergreen-notify/SKILL.md §Step 3 (route order and route 1 sizing); LEARNINGS.md (L-018)
- Route 2 (Claude in Chrome, "Insert files using Drive" from the outbox mirror, both files as attachments) is preferred whenever it exists; route 1 is limited to bundles whose every attachment is under ~16 KB raw, with
changes.patchsplit like the archive. A suite re-run for evergreen-notify is owed (the plugin's own suite has not been run yet).
C-20260910-1 · 2026-09-10 · Refresh: claude plugin eval demoted to undocumented, SessionEnd budget wording, Codex and Copilot payloads, testing pointers
- because: R-20260910-1 (docs withdrawn), R-20260910-2 (SessionEnd budget adjustable; #63360 closed), R-20260910-4 (Codex
transcript_path, CopilotagentStop), R-20260910-5 (skills can be net negative), R-20260910-6 (skillsaw), R-20260906-3 (skill-eval-harness and evals-skills pointers carried since 2026-09-06) - files: protocol/TESTING.md §6 (first and last rows) and §7; protocol/PORTABILITY.md §Per-environment notes (Claude Code, Cowork, Copilot and Codex rows) and §Update email: which transport where (mailer and use-log paragraphs); protocol/PROTOCOL.md §10 (Report); README.md §Testing; RESEARCH.md (Current understanding, Open questions, Search plan, findings R-20260910-1 to -7)
- Claim edits only, no behavior change. TESTING.md §6 now reads skill-creator's runner, then the tester agent, with
claude plugin evalkept as a row that applies only where the command runs and its flags stamped "as documented 2026-09-04"; the SessionEnd wording says the 1.5 s budget is a default that a per-hooktimeoutcan raise to 60 s, and that the mailer detaches by choice so it never blocks exit; the Codex and Copilot rows carry their real payload fields and the Codex docs URL; §7 states the negative effect sizes so a losing baseline retires a skill; the cross-harness row points at skill-eval-harness,validate-evaluatorand skillsaw.skills/evergreen-test/SKILL.mdis untouched (it defers to TESTING.md §6), so no suite re-run is owed. Installed copies need a reinstall to pick up README.md.
C-20260904-1 · 2026-09-04 · Testing pillar: every skill proven by evidence, a tuning loop with a research gate, a fourth research track, the use log; protocol 1.3, version 0.5.0
- because: the owner's request ("I build skills/plugin suites and they don't totally work as expected ... one skill was meant to delegate to another machine but it never did ... test out and fine tune new skills as they're created ... also when existing skills fail ... whenever a test fails and it's been a while since research was done, research should be done ... testing and improvement techniques should be studied for the subject"); L-015; R-20260904-1 through R-20260904-5
- files: protocol/PROTOCOL.md (intro, §1 principle 5, §2 anatomy and IDs, §3 failing-tests line, §4 four tracks and the re-test step, §5, §8, new §11), protocol/TESTING.md (new), protocol/INTERVALS.md §Test failures are a signal, protocol/PORTABILITY.md (Claude Code and Cowork rows, use-log hook); skills/evergreen-test and skills/evergreen-tune (new), agents/evergreen-tester.md (new), agents/evergreen-researcher.md (testing track, response for testing findings), skills/evergreen-refresh §Step 2 and §Step 3, skills/evergreen-new §Step 2 and §Step 4, skills/evergreen-convert §Step 1 to §Step 4, skills/evergreen-audit §Step 1 and §Step 3, skills/evergreen-learn §Step 1, every skill's Maintenance footer; templates/TESTS.md.template and templates/evals.json.template (new), templates/RESEARCH.md.template (testing track, TESTS.md link), templates/CHANGELOG.md.template, templates/LEARNINGS.md.template, templates/SKILL.md.template, templates/MAINTENANCE-SECTION.md.template, templates/MAINTENANCE.md.template §Tests, templates/MAINTENANCE-POINTER.md.template, templates/AGENTS.md.snippet, templates/CLAUDE.md.snippet; scripts/evergreen.py (
T-ids,testsblock andfiles.testsinnew_state, scaffold writes TESTS.md and evals/,test-init,tested,failed,flag --clear-failing,bump --tests,use-log,uses, test flags in status and audit, lint for the testing track, the evals file,led to:and the TESTS.md budget), scripts/evergreen_sync.py (TESTS.md in LOG_FILES and the digest,T-in the entry regex and the merge union,testsblock in state merge,usesin where), hooks/hooks.json (PostToolUse on Skill), scripts/evergreen-hook.sh and .ps1 (usemode), scripts/test_evergreen.py (12 new tests, 59 total); README.md, AGENTS.md, .claude-plugin/plugin.json (0.5.0, tester agent), RESEARCH.md (Current understanding, Open questions, Search plan testing track), evals/evals.json and TESTS.md (the plugin's own suite) - A refreshed unit could still be a skill that does not fire on the user's phrasing, or fires and narrates its action without performing it, or does the job by a route it was told not to use. Refresh cannot see any of that; only a run in a fresh context with evidence outside the transcript can. Every skill now carries
evals/evals.json(skill-creator's format pluskind,decoy,evidence,baseline,runs) with trigger, action and outcome cases, decoys, and a baseline without the skill, and aTESTS.mdlog of runs (T-ids, double-linked). A test passes on a tool call in the trace, a file, a marker or a record, never on the reply.tests.failinginevergreen.jsonis a Step 0 signal likecontradiction.evergreen-testwrites and runs the suite (harness per environment in TESTING.md §6; theevergreen-testeragent is the fallback everywhere, one case per fresh context);evergreen-tuneis the bounded loop: reproduce, classify (undertrigger, overtrigger, no-op, fallback, wrong-outcome, environment, harness), write the learning, research gate (evergreen.py failed: research first whenlast_checkedis older than half the interval, 3 to 30 days, and always for no-op and fallback), smallest edit by class, re-run, three iterations. Research gains a fourth track, testing (how work on the subject is verified and how skills for it are tuned), in every plan, template, the researcher agent and lint; a refresh that edits a skill re-runs its suite beforechecked. Claude Code'sPostToolUsehook on theSkilltool writes every invocation with its transcript path toEVERGREEN_HOME/uses.jsonlso a failure in use can be traced to what the skill did;evergreen.py usesreads it. New skills are handed over only with a run suite; converted skills get one scaffolded. Protocol version 1.3.
C-20260903-5 · 2026-09-03 · Research covers three tracks: subject, tooling, practice; share packs for other people; version 0.4.0
- because: the owner's request ("research should be both: for the latest and most popular skills, plugins, scripts, knowledge graphs ... as well as into the subject ... and how others are using AI for the goals"); R-20260903-6 (the discovery surfaces and how to rank them)
- files: protocol/PROTOCOL.md §1 (fourth principle), §2 (RESEARCH.md row), §4 (the three tracks, how tooling findings are judged, steps 1 to 4), §8 (tooling pass before a new skill is written); templates/RESEARCH.md.template §Search plan and §Findings log, templates/MAINTENANCE.md.template §Refresh, templates/MAINTENANCE-POINTER.md.template, templates/AGENTS.md.snippet, templates/INSTALL-PROMPT.txt (
{tag},{marketplace},{owner}placeholders); agents/evergreen-researcher.md (tracks, per-track queries, Track and Suggested response fields, Tracks covered line); skills/evergreen-refresh §Step 2 and §Step 3, skills/evergreen-new §Step 2 and §Step 3, skills/evergreen-convert §Step 1 and §Step 3, skills/evergreen-pack §Step 1 and §Step 2; scripts/evergreen.py (lintchecks the plan has all three tracks,share_config,shipped_bytes,iter_plugin_files(share),pack(share),install_prompt(share),pack --share); README.md §How a unit behaves §Design notes §Packaging §Using the scripts, AGENTS.md, RESEARCH.md (Current understanding, Search plan in three tracks), scripts/test_evergreen.py - Also:
copy_to_outbox_mirrornow mirrors the composed attachments (the mail-safe archive and the install prompt), not only the digest, patch and manifest (L-014), since the mirror is what an agent attaches from when no transport works; andfiles_from_zipskips the archive's ownINSTALL.txtandINSTALL-PROMPT.txt, which sit beside the plugin folder rather than in it, so a baseline taken from an archive no longer reports them deleted on the next diff. - A unit was current in its subject and blind to its ecosystem: it could hold every fact right while a well-used skill, plugin or MCP server quietly replaced its procedure. Every search plan now has a subject track (the goal and the latest thinking), a tooling track (skills, plugins, MCP servers, scripts, knowledge graphs, with the discovery surfaces and ranking from R-20260903-6), and a practice track (how others use agents on the same goal). Every refresh runs at least one query per track, findings carry their track, and a tooling finding is scored on the same rubric: a maintained, well-used tool that does the unit's job is a supersession at m 0.6 and up, with a response of adopt, point, or note. Installing anything still needs the user's yes.
lintreports a plan that is missing a track, so the fleet migrates as units come due rather than in one sweep. Separately,pack --shareproduces the copy that goes to someone else: noprofile/, a config with no address andnotify.autooff, marketplace namedmy-local, and no baseline or tag, since a share pack is a gift and not a report.
C-20260903-4 · 2026-09-03 · One-paste install: INSTALL-PROMPT in every archive and at the top of every update email, full plugin attached, split pieces; version 0.3.0
- because: L-012 (the work-machine copy was never installed: the email carried instructions, not an installable unit); L-013 (a Gmail connector moves attachments as base64 through the model, so a 150 KB zip does not fit one call); the manifest fixes from installing 0.2.0 on Claude Code (
agentsmust list files;hooks/hooks.jsonis auto-loaded) - files: templates/INSTALL-PROMPT.txt (new), scripts/evergreen.py (
install_prompt,split_file,pack(split_kb, after), INSTALL.txt rewritten around the prompt,pack --split), scripts/evergreen_sync.py (composeopens with the prompt and attaches the mail-safe pack built into the bundle;attach_pack,pack_split_kbdefaults), evergreen.config.json (attach_pack,pack_split_kb), README.md §Install §Packaging §One trunk §What you get, skills/evergreen-pack §Step 1 §Step 2 §Step 3, skills/evergreen-notify §Rules §Step 3, scripts/test_evergreen.py, .claude-plugin/plugin.json (0.3.0,agentsas a file list, nohookskey) - The recipient of an update now needs one action: save the attachments to a folder and paste the prompt with that path. The prompt has the agent join
.partNNpieces, unzip,unmail, place the folder under amark-localmarketplace (reusing an existing one), reinstall, verify withclaude plugin list, and run the first audit.packwrites the prompt inside the archive and beside it;composeputs it first in the body, ahead of the digest, and attaches the whole plugin (mail form) so the email is self-sufficient;--splitandpack_split_kbproduce pieces for routes with a small attachment ceiling.after=Falseletscomposepack without moving the baseline.
C-20260903-3 · 2026-09-03 · Review fixes before packing 0.2.0: hostile patches refused, config protected, idempotent edge inserts, per-unit renumbering, quiet trunk
- because: independent review of 0.2.0 (two blockers, sixteen should-fixes); L-011
- files: scripts/evergreen_sync.py (
safe_rel,inside,PROTECTED,CODE_DIRS,--allow-code,read_text_any,advance_baseline_with,is_trunk,--from-hook, per-unit rename pass,apply_hunksalready-applied rule,wherebaselines,baseline_dirkeyed by plugin path, mail-form names infiles_from_zip,git addof written files only,CREATE_NO_WINDOWand BOM for PowerShell, Graph cache check, compose URL length cap, EOL-only changes skipped), scripts/evergreen.py (_usable_homefallback,EVERGREEN_PLUGIN), scripts/evergreen-hook.sh (backslash root,py -3first on Windows), hooks/hooks.json (no matcher, no timeout), .gitattributes, templates/MAINTENANCE-POINTER.md.template, evergreen.config.json (report_trunk), skills/evergreen-merge §Step 1 and §Step 2, skills/evergreen-notify §Step 2, protocol/PORTABILITY.md §EVERGREEN_HOME, README.md, scripts/test_evergreen.py (five new tests, 45 total) - A patch could write outside the trunk (
a/../x,.git/hooks/...) and could rewritenotify.tothroughevergreen.config.json; both are now refused, and code directories need--allow-code. Second merges no longer duplicate edge insertions; renumbering respects each unit's own id space (plugin vsprofile/);--mark-sentadvances the baseline to the bundle's snapshot rather than the live tree; the trunk stays quiet on the session-end hook unlessreport_trunk;notify.autois honoured on every unattended path; a configured store on a missing drive falls back to~/.evergreen.
C-20260903-2 · 2026-09-03 · Pointer mode: units outside the plugin resolve the protocol through the registry instead of carrying a copy
- because: user request (skills should not have to change when the plugin does); L-009
- files: scripts/evergreen.py (
resolve_protocol,registerwritesplugin_root,init --pointer,scaffold), templates/MAINTENANCE-POINTER.md.template, protocol/PROTOCOL.md (intro, §2, §10), protocol/PORTABILITY.md §Four linking modes, skills/evergreen-convert/SKILL.md §Step 1 and §Step 2, skills/evergreen-new/SKILL.md §Step 3, README.md §Making things evergreen evergreen.json.protocolmay be the wordplugin; the script resolves it to the running plugin or toregistry.json.plugin_root, andlinksaccepts it. The pointerMAINTENANCE.mdsays how to find the plugin and keeps only Step 0 and the hand rule as a fallback. Convert and new default to it for units outside the plugin; full standalone stays for units that travel without the plugin. Cross-unit ID references are written qualified (context-health:L-003) and skipped by the link check; the two such references inprofile/were qualified.
C-20260903-1 · 2026-09-03 · Self-update email and trunk merge: baseline, diff, notify, merge, where; SessionEnd hook; version 0.2.0
- because: user request (email the changes to the owner whenever the plugin updates itself; skills to diff, send, and merge into the trunk); R-20260903-1 through R-20260903-5
- files: scripts/evergreen_sync.py (new), scripts/evergreen.py (
packwrites MANIFEST.json, takes the baseline, keeps the zip, commits and tags;checkedandbumpcallmaybe_notify;--no-notify;--no-git; sync subcommands registered), scripts/evergreen-hook.sh and .ps1 (endmode), hooks/hooks.json (SessionEnd), evergreen.config.json (notifyblock), .gitignore, skills/evergreen-diff, skills/evergreen-notify, skills/evergreen-merge (new), skills/evergreen-refresh §Step 4 and §Step 5, skills/evergreen-learn §Step 3, skills/evergreen-audit §Step 3 and §Install check, skills/evergreen-pack §Step 2 and §Step 3, protocol/PROTOCOL.md §4 §7 §10, protocol/PORTABILITY.md §Per-environment notes and §Update email, README.md, scripts/test_evergreen.py, .claude-plugin/plugin.json - The trunk is the
sourcefolder, now a git repo. Installs baseline themselves from the archive'sMANIFEST.json.diffwrites an update bundle (digest of new C-/R-/L- entries and state deltas, git-style patch with blob ids, manifest with the full state file) toEVERGREEN_HOME/outbox/.notifyemails it tonotify.tothrough outlook (classic COM), graph (Graph PowerShell, delegated Mail.Send, cached token), smtp (Gmail app password), and on request compose-url; the outbox always keeps a copy and mirrors it intooutbox_copy_to; the same content hash is never sent twice; agents mark their own sends with--mark-sent.mergeat the trunk unions whole log entries (renumbering colliding IDs and rewriting references), mergesevergreen.jsonas data (history union, newer schedule wins; reconstructs the incoming state from the manifest, from the base blob in git, or from the trunk copy), applies other files exactly, viagit apply --3way, or with fuzz 2, and writes.rej/.incomingon conflict; it is idempotent, advances the baseline, and commits. The Claude Code SessionEnd hook runsnotify --if-changed --detachbecause SessionEnd hooks share a 1.5 s budget. Only the plugin reports; other units do not.
C-20260902-2 · 2026-09-02 · Email-safe packaging: pack --mail and unmail
- because: L-007 (Gmail rejected the zip)
- files: scripts/evergreen.py (
MAIL_BLOCKED,pack(mail=True),unmail), scripts/test_evergreen.py, skills/evergreen-pack/SKILL.md, README.md §Packaging pack --mailwritesevergreen-<version>-mail.zipwith blocked script types stored as.txtand an INSTALL.txt that explains the restore;unmail <folder>renames them back. The.pluginis omitted from the mail form since it would be blocked the same way. Version 0.1.2.
C-20260902-1 · 2026-09-02 · Angle-bracket placeholder removed from the pack skill's description; lint checks for it
- because: L-006 (Cowork "Save plugin" validation failure)
- files: skills/evergreen-pack/SKILL.md (frontmatter), scripts/evergreen.py (
lintfrontmatter check), templates/SKILL.md.template (description placeholder), scripts/test_evergreen.py - Description now says "a versioned evergreen zip".
lintreads the main file's frontmatter and reports any<...>indescriptionorwhen_to_use. Version 0.1.1.
C-20260901-4 · 2026-09-01 · Self-packaging: pack, export, evergreen-pack skill
- because: user request (one archive to email or install elsewhere; 7-Zip friendly)
- files: scripts/evergreen.py (
pack,export,iter_plugin_files), scripts/pack.ps1, skills/evergreen-pack/SKILL.md, README.md §Install and §Packaging packwritesevergreen-<version>.zip(folder inside plus INSTALL.txt) andevergreen.plugin(flat, for Cowork) with Python's stdlib zip;pack.ps1uses 7-Zip or Compress-Archive when Python is absent.export <repo>copies the plugin into<repo>/.agents/keeping its layout so the skills' relative links keep resolving for Copilot, Codex, Cursor, and Gemini. Archives exclude caches and old archives.
C-20260901-3 · 2026-09-01 · Review fixes: freshness in every skill, absolute script paths, escalation out of slow tiers, verify-at-use state, config portability
- because: independent review (six blockers, ten should-fixes); L-003, L-004
- files: scripts/evergreen.py, scripts/test_evergreen.py, all skills/*/SKILL.md, protocol/INTERVALS.md §hand rule and §migration, protocol/PROTOCOL.md §2 §3 §4, templates/MAINTENANCE.md.template, templates/LEARNINGS.md.template, evergreen.config.json, hooks/hooks.json, scripts/evergreen-hook.ps1, README.md
- Every plugin skill now opens with a Step 0 that reads the plugin's own
evergreen.json(the lazy check the protocol promised, which the plugin's own skills lacked). Skills refer to the script by absolute path (EGshorthand) instead of../../from a shell whose cwd is the user's project.audit --jsonreturned no units (rows appended after the short-circuit); fixed.verify_at_useunits showed as permanently STALE; they now clearnext_dueand reportn/a [verify-at-use]. A major change (m ≥ 0.6) that would clamp below a tier's min now promotes one tier immediately, and streak promotion lands at the new tier's min, soslow→moderatetakes one check instead of months. Custombounds_daysare dropped with a note on migration.evergreen.config.jsontakes a per-OS map and a Windows drive-letter path is ignored on Mac and Linux (it used to create a literalD:\...folder in the cwd).discover()depth limit now works for relative roots.check_linksandlintignore HTML comments, and the LEARNINGS template example uses{{DATEID}}, so back-dated conversions pass.--strictworks after the subcommand;--eventaccepts quoted labels andTtimes;--append-maintenancekeys onevergreen.jsonrather than any## Maintenanceheading. Hand rule now says to setlast_checkedto today and drops the whole-day rounding claim; worked example recomputed. Hook fires on every SessionStart matcher;.ps1no longer assigns to$args.evergreen-mapandevergreen-newdescriptions carry negative clauses so they defer to repo-specific skills and skill-creator. PROTOCOL §2 states the three-companion carve-out for codemaps and profiles.
C-20260901-2 · 2026-09-01 · Entry IDs use compact dates so the link checker can find them
- because: L-001
- files: templates/RESEARCH.md.template, templates/CHANGELOG.md.template, scripts/evergreen.py (scaffold subs)
- Templates emitted
R-2026-09-01-1, which theR-\d{8}-\d+regex never matched; added a{{DATEID}}placeholder (YYYYMMDD) and switched the ID placeholders to it.
C-20260901-1 · 2026-09-01 · Plugin created, version 0.1.0
- because: user request; R-20260901-1 through R-20260901-8
- files: everything
- Protocol (PROTOCOL, INTERVALS, LEARNINGS-FORMAT, CODEMAP-FORMAT, PORTABILITY), six skills (refresh, learn, map, convert, new, audit), two subagents (researcher, mapper),
scripts/evergreen.pywith tests, templates for every companion file plus repo snippets, a SessionStart hook, and the portable profile (AI-PREFERENCES, ENVIRONMENTS). Tierfast, 14-day start interval; the plugin is its own first unit.