Skip to content

m4bwav/evergreen

v0.16.0MIT

Self-maintaining, self-proving skills and knowledge: adaptive web research on four tracks (the subject, the skills, plugins, MCP servers and knowledge graphs built for it, how others use agents on the same goal, and how they test that the job was done), never-teach-twice learnings capture, plugin-owned codemaps, a portable AI-preferences profile, and an eval suite per skill (trigger, action, outcome, decoys, baseline) that passes on evidence outside the transcript, with a bounded tuning loop that researches the subject's own testing when a skill fails, and a worth check that warns while a skill is made or changed when it reads as mostly cost and decides KEEP, TRIM or CUT from the with-versus-without A/B. Skills say what they need (tools, packages, keys, servers, models, accounts, and access to databases, telemetry, cloud roles and APIs) and what for: on a new machine or agent harness, or when something is missing, the agent explains it, installs what is safe after saying so, hands the user exact steps for the rest, drafts least-privilege access requests without ever granting itself access, and records the recipe that worked and how each attempt went for that environment. An end-of-session wrap-up harvests the session transcript and keeps only what will save time or tokens or raise quality, routed to its narrowest home. Converts existing skills and makes new ones evergreen by default. Lives in one public git repository (https://github.com/m4bwav/evergreen-protocol): every install is a clone that pulls to stay current and, only if its owner said yes at install, sends improvements to the protocol, skills and scripts back as pull requests it opens itself (log entries go straight to master from a maintainer's clone), with email as the fallback route.

Evergreen

A giant evergreen pine tree whose branches hold glowing scrolls and small tools, a gardener robot gently pruning it, misty forest

Self-maintaining, self-proving skills and knowledge for AI agents. An evergreen unit (a skill, a doc, a codemap, a preferences profile) re-researches its topic from primary sources on an adaptive schedule, writes down what it learns so nobody teaches the agent the same thing twice, logs every change with its reason, keeps a map of any codebase it explores, and, for skills, carries the tests that prove it triggers, acts, and gets the result right, with a tuning loop for when it does not. It works in Claude Code, Cowork, and any agent that reads markdown (Copilot, Codex, Cursor, Gemini CLI), because the whole paradigm is plain files plus one small JSON state file. Scripts are conveniences, not requirements.

This plugin is its own first unit: see RESEARCH.md for the evidence behind the design, CHANGELOG.md for what changed and why, LEARNINGS.md for lessons, and evergreen.json for its schedule (tier fast, checked every couple of weeks).

What you get

PieceWhat it does
templates/INSTALL-PROMPT-GIT.txtThe one-paste install prompt: the agent clones (or pulls) the repository, registers the marketplace, installs. INSTALL-PROMPT.txt is the archive form, shipped in every pack and update email.
protocol/The paradigm. PROTOCOL.md is the spec; INTERVALS.md the refresh math; TESTING.md how a skill is proven and tuned; SETUP.md how a skill gets what it needs on each machine; LEARNINGS-FORMAT.md, CODEMAP-FORMAT.md, PORTABILITY.md the details.
skills/evergreen-refreshRe-research a due unit on four tracks, apply delta edits, re-test if the skill changed, reschedule.
skills/evergreen-learnCapture a correction, repeated error, workaround, environment fact, preference, or failed test into the right file.
skills/evergreen-wrapupEnd-of-session wrap-up: harvest the session transcript (evergreen.py wrapup: corrections, refused calls, errors, repeated commands, token sinks, slow calls, files and stores touched, and notes or vaults outside any repository that the session used), drop what fails a four-question usefulness gate, send each survivor to its narrowest home through the skill that owns it (or straight into a knowledge base no skill owns, when the session makes clear which), and fix the scripts and habits that cost the most time or tokens.
skills/evergreen-mapBuild and update the plugin's own codemap of any repo you explore.
skills/evergreen-convertMake an existing skill or doc evergreen without rewriting it; skills get a test suite scaffolded.
skills/evergreen-newCreate new skills evergreen by default, handed over only with a run suite.
skills/evergreen-testWrite and run a skill's eval suite: trigger prompts and decoys, action cases proven by evidence outside the transcript, outcome cases, baseline without the skill.
skills/evergreen-tuneFix a skill that failed a test or failed in use: reproduce, classify, learn, research if stale, smallest edit, re-run. Three iterations, then report.
skills/evergreen-worthIs a skill worth its tokens? Warns while a skill is made or changed when it reads as mostly cost (general advice the model already follows, bloat, repeats, a near copy of another skill, an edit that added little that is specific), then decides KEEP, TRIM or CUT from the with-versus-without A/B.
skills/evergreen-setupGet a skill running here: check what it needs (tools, packages, keys, servers, models, MCP servers, accounts), tell the user what each missing piece is for, install what is safe after saying so, hand over exact steps for the rest, re-check, and record the recipe that worked for this operating system and agent harness.
setup/RECIPES.md, SETUP.mdThe shared install recipes for tools many skills need (Python, git, gh, Node, ffmpeg, Ollama, uv, the Claude CLI), keyed by OS, harness and package manager, and step-by-step access routes (Azure sign-in, Application Insights read, PostgreSQL and SQL Server read, a private GitHub repository); and the plugin's own needs in the same format.
skills/evergreen-auditTable of every unit: staleness, drift, test status, links, budgets. Run after install.
skills/evergreen-publishPush the plugin's self-changes to the trunk repository: straight to master from a maintainer's clone, an update branch plus a pull request from anyone else's; pull brings a clone up to date.
skills/evergreen-mergeReview and land what other clones published: merge pull requests on the host, resolve the rare conflict with the delta-edit rule, then pull and reinstall. Also folds an email bundle for the fallback route.
skills/evergreen-packOne archive for Cowork (.plugin), a backup, a share copy for someone without repository access, or an export into a repo's .agents/.
skills/evergreen-diffWhat this clone changed and has not yet published: digest plus unified diff against its baseline.
skills/evergreen-notifyThe email fallback for a machine that cannot reach the repository: bundle, transport ladder (Outlook, Graph, Gmail SMTP, a mail connector, Chrome), outbox.
agents/evergreen-researcher (research pass in its own context), evergreen-mapper (repo sweep), evergreen-tester (one eval case in a fresh context, strict trace report).
scripts/evergreen.pyState, scheduling, scaffolding, audit, link and lint checks (typed Related: lines included), volatile claims with their own check dates (claims), search by meaning over every unit's log entries (search, BM25), test records (test-init, tested, failed) and claude plugin eval case folders written from evals.json (eval-export), the use log (use-log, uses), export, pack. Stdlib only. evergreen_sync.py adds publish, pull, where, baseline, diff, and the email fallback (notify, merge); evergreen_worth.py adds worth (is a skill worth its tokens) and the worth-hook body. test_evergreen.py covers the math, the tests layer, the sync and the new commands. pack.ps1 packages without Python.
scripts/bench_intervals.pyA deterministic benchmark of the refresh schedule: the real interval rule, imported, against fixed and per-class schedules and an oracle on synthetic units; prints a markdown table (see Testing).
templates/Companion-file templates (including TESTS.md and evals.json), the pointer and standalone MAINTENANCE.md forms, and snippets for AGENTS.md, CLAUDE.md, copilot-instructions.md.
profile/Portable preferences (AI-PREFERENCES.md) and environment facts (ENVIRONMENTS.md). Learnings-driven, no web research.
hooks/Claude Code hooks: SessionStart prints what needs attention into context (units due for refresh, verify-at-use units with claims due, units whose learnings passed consolidate_every; silent when nothing does); SessionEnd publishes any self-update to the trunk repository in a detached process; PostToolUse on the Skill tool writes the use log; PostToolUse on `Write
evals/, TESTS.mdThe plugin's own suite and run log; every skill unit carries the same pair.

How a unit behaves

  1. Every use starts with one small read of evergreen.json. Below the due date, nothing else happens.
  2. Past the due date (or when a learning has flagged a contradiction), the unit says so in one line, finishes the user's task, then refreshes in the same session: a few scoped searches of primary sources, findings logged with magnitudes, the main file edited in place, the change logged with its reason, the suite re-run if a skill changed. Every refresh covers four tracks: the subject (the goal and the latest thinking on reaching it), the tooling built for it (skills, plugins, MCP servers, scripts, knowledge graphs, ranked by real use rather than by listicle), how other people are using agents on the same goal, and how they test that the job was done. A maintained, well-used tool that does what the skill's own procedure does counts as a superseded claim, the same as a fact going wrong.
  3. The next interval adapts: halve on a major change, cut by a third on a real change, hold on cosmetic churn, grow by a quarter when nothing changed, all inside the tier's bounds. Topics can be promoted or demoted between tiers, capped by known upcoming events, or switched to "verify at use" when too volatile to schedule; then each volatile claim carries its own check date and only the claims that are due get re-checked at use.
  4. Whenever the user corrects the agent, the same error repeats, or a fact about the environment turns up, a learning is written immediately with its trigger and hypothesis, gated against duplicates (found by meaning, across every unit, with search), and later promoted into the main file or retired with a reason. The session-start hook says when a unit's learnings are due for consolidation.
  5. Any repo explored gets a codemap in the plugin's store, stamped with the git sha, every claim marked verified or inferred, and a log of which questions it has answered. The map says, every time, that it may be stale, incomplete, or confused.
  6. At install the plugin asks one question: may this install send its learnings back as pull requests? no means nothing ever leaves the machine and updates only flow in with pull; yes enables the next step and can be revoked with evergreen.py contribute no. When the plugin changes itself on an install that said yes, it commits the change with a subject naming the new entries and pushes it to the trunk repository: straight to master when the clone is a maintainer's, as an update branch and a pull request when it is anyone else's. The owner merges pull requests on the host; every other clone pulls. One repository, fixed in config; nothing is ever force-pushed; email remains the fallback for a machine that cannot reach it.
  7. Every skill carries a suite (evals/evals.json, skill-creator's format plus evergreen fields) with trigger prompts and decoys, action cases that pass only on evidence outside the transcript (a tool call in the trace, a file, a marker, a remote record; never the reply saying "done"), outcome cases, and a baseline of what a fresh context does without the skill. Runs are logged in TESTS.md; a failing case is a Step 0 signal like a contradiction. When a case fails, or the skill fails in use, the tuning loop reproduces it in a fresh context, classifies it (undertrigger, overtrigger, no-op, fallback, wrong-outcome, environment, harness), writes the learning, researches the subject's own testing and tooling when the unit's research is older than half its interval (always for no-op and fallback), makes the smallest edit for the class, and re-runs; three iterations, then it reports what it tried.
  8. Every skill is also weighed (evergreen-worth, TESTING.md §8): a static reading of what it costs (its listing line in every session, its body on every use) against how much of it is specific, run by a hook whenever a SKILL.md is written, warns when a skill being made or updated is mostly advice the model already follows; the with-versus-without A/B on its action and outcome cases then says KEEP, TRIM or CUT, with the cost, turn and time ratios. The same command ranks a whole catalog (--triage), counts real use from the transcripts with test runs left out, reads which recorded baselines already pass, probes what a fresh model knows without the skill (--probe), runs the A/B headless on any OS (--ab), and names the failure mode: never fires, already known, no gain, worse, or superseded. Two modes wrap it: --lite, a scan in seconds with no model calls that also lints the value cases for problems that would waste a paid run, and --heavy, the probe, the A/B and top-up runs inside the noise margin under a spend cap, ending in a pull request that feeds the skill what the run taught. Most published skills add nothing measurable (39 of 49 in one 2026 benchmark), so the warning comes before the investment.
  9. Every skill says what it needs (SETUP.md, PROTOCOL.md §12): each tool, package, key, server, model or account, what it is for, and how to install it on each operating system and agent harness met so far. On a new machine or harness, or when a step fails for a missing piece, evergreen.py setup checks everything at once and the agent explains what is missing and why, installs what is safe after saying so, hands the user exact steps for admin rights, accounts, large downloads and secrets, checks again, and records the recipe that worked, so each environment the skill meets makes the next install easier. Access works the same way: a skill that needs a database, Application Insights, a cloud role, an API scope or a private repository checks it with a read-only probe, says whether a failure is a sign-in, a missing permission, the network or a missing tool, walks the user through the recipe's steps, drafts a least-privilege request for an administrator (setup --request), never grants itself anything, and logs how each try went (setup --attempt: pending, granted, refused, how long it took), so the next person knows who grants it and which role worked. Recipes that help anyone go to the shared book and upstream; package names come only from recipes or the vendor's documentation, never from memory.

Releases

Official releases are on the repository's Releases page: each vX.Y.Z tag builds evergreen-<version>-share.zip (the plugin folder), evergreen-share.plugin (for Cowork) and the one-paste install prompt, after the suite passes on the release runner. The version is .claude-plugin/plugin.json; the reasoned history is CHANGELOG.md.

Install

Prerequisites: git, Python 3.9+ (python on Windows, python3 on macOS and Linux; on macOS that means the Xcode Command Line Tools or a python.org install), and gh for pull requests. Nothing else: the scripts are stdlib only and the hooks are POSIX sh, which Claude Code on Windows runs through Git Bash.

From the repository (the normal way). Open Claude Code and paste templates/INSTALL-PROMPT-GIT.txt (the repository URL is already in it). The agent clones the repository under a local marketplace root (~/claude-plugins unless one already exists), or pulls if the clone is there, registers the marketplace, installs or reinstalls with claude plugin, verifies with claude plugin list, and runs where and the first audit. Updating later is python <clone>/scripts/evergreen.py pull followed by a reinstall when it says skills or scripts changed. The repository is public (https://github.com/m4bwav/evergreen-protocol); cloning needs no account, and gh auth login once lets the clone open pull requests with its improvements.

From an update email or a packed archive (fallback). Save the attachments to one folder and paste the archive form of the prompt (INSTALL-PROMPT.txt, inside every archive and at the top of every update email), replacing its one placeholder with that folder's path. An archive install is not a git clone, so it reports by email until it is replaced by a clone.

By hand (Claude Code). claude plugin install takes a marketplace entry, not a bare path, so give it one: put the plugin folder under some ROOT, write ROOT/.claude-plugin/marketplace.json as {"name": "mark-local", "owner": {"name": "..."}, "plugins": [{"name": "evergreen", "source": "./evergreen"}]}, then:

claude plugin marketplace add ROOT
claude plugin install evergreen@mark-local --scope user
claude plugin list

Claude Code caches a copy; after changing the source (an edit, or a pull that touched skills or scripts), claude plugin marketplace update mark-local, then uninstall and install again (a same-version update does not recopy). The manifest's agents is a list of files and hooks/hooks.json is picked up on its own (naming it in the manifest is a duplicate-hooks error). The SessionStart hook runs scripts/evergreen-hook.sh (Claude Code on Windows ships with Git Bash, so sh is there); if hooks on your machine go through PowerShell, point one at scripts/evergreen-hook.ps1 instead (example in that file). The hook prints nothing when nothing is due.

Then run the evergreen-audit skill once in a new session: it confirms the store, creates the registry, and refreshes the plugin's own research if it sat unused.

Cowork (desktop): install the evergreen.plugin file produced by evergreen-pack (or python scripts/evergreen.py pack), or add the folder through Customize. Cowork keeps a copy; edits made by refreshes go to the source path recorded in evergreen.json, and the audit tells you when to reinstall. Cowork's sandbox cannot see host paths outside the selected folder, so the store and the script run through Desktop Commander there, or the protocol is followed by hand.

Other agents (Copilot, Codex, Cursor, Gemini CLI, Windsurf): python scripts/evergreen.py export /path/to/repo copies the plugin into <repo>/.agents/ with its layout intact (skills at .agents/skills/evergreen-*, protocol and scripts beside them, so every relative link still resolves). Then paste templates/AGENTS.md.snippet into the repo's AGENTS.md, and add templates/CLAUDE.md.snippet as CLAUDE.md so Claude Code sees the same rules. Gemini CLI needs .gemini/settings.json → {"context": {"fileName": "AGENTS.md"}}.

Store location: set EVERGREEN_HOME (environment variable) or evergreen.config.json at the plugin root, which takes a string or a per-OS map ({"home": {"nt": "D:\\Evergreen", "posix": "~/.evergreen"}}). Default is ~/.evergreen. A Windows drive-letter path is ignored on Mac and Linux rather than creating a stray folder. The store holds the registry, codemaps, and units that have no other home.

Using the scripts

python scripts/evergreen.py audit --checks          # everything, with link and lint problems
python scripts/evergreen.py status <unit>           # one line
python scripts/evergreen.py next <unit> --m 0.4     # dry-run the interval rule
python scripts/evergreen.py checked <unit> --m 0.4 --note "what changed; cost: searches 6 · fetches 3 · tool calls 14"
python scripts/evergreen.py claims <unit> [--due] [--stamp 1,3|due|all] [--add "claim"] [--recheck-days N]   # volatile claims and their check dates
python scripts/evergreen.py search "query" [--unit x] [--kinds learnings,research,changes,tests] [-n 5] [--json]   # BM25 over every unit's log entries
python scripts/evergreen.py init <dir> --name x --topic "..." --tier moderate --append-maintenance [--standalone]
python scripts/evergreen.py map-init <repo>         # codemap unit in the store
python scripts/evergreen.py drift <map> [--update-sha]
python scripts/evergreen.py links <unit>  |  lint <unit>  |  flag <unit> --contradiction "why"   # lint also checks typed Related: lines
python scripts/evergreen.py test-init <unit>       # add evals/evals.json, TESTS.md and the tests block to an existing skill
python scripts/evergreen.py eval-export <unit> [--out <dir>] [--force]   # write `claude plugin eval` case folders from evals/evals.json
python scripts/evergreen.py tested <unit> --passed 5 --failed 1 --failing action-1 --harness tester --note "..."
python scripts/evergreen.py failed <unit> --case action-1 --class no-op --note "..."   # failure in use; says whether research comes first
python scripts/evergreen.py uses [--skill x] [--days 7]   # recent skill invocations from the use log (Claude Code hook)
python scripts/evergreen.py worth <skill|plugin|folder> [--against HEAD] [--results <dir>] [--record] [--wrap <dir>]   # is it worth its tokens? static reading + with/without A/B
python scripts/evergreen.py worth <skill> --lite                       # quick scan, no model calls: cost, use, verdict, case lint
python scripts/evergreen.py worth <skill> --heavy --model sonnet --out <dir> [--max-usd 15]   # probe + A/B + top-ups inside the margin
python scripts/evergreen.py worth <repo> <repo> ... --triage                     # a catalog ranked by cost per use, with usage, baselines, failure modes and gaps
python scripts/evergreen.py worth <skill> --probe | --ab [--case 'action-*'] [--runs 3]   # what a fresh model already knows; headless A/B on any OS
python scripts/evergreen.py worth <skill> --set FIX|SUPERSEDED|... --why '...'     # record a person's ruling
python scripts/evergreen.py setup <unit> [--harness codex] [--script <path>] [--log]   # what it needs here: checks, recipes for this OS and harness
python scripts/evergreen.py setup <unit> --record <id> --env linux/apt --how "`sudo apt-get install -y x`" --tags admin --verified   # what worked
python scripts/evergreen.py setup <unit> --request <id>     # least-privilege access request draft for an administrator
python scripts/evergreen.py setup <unit> --attempt <id> --result pending|granted|refused --route any --took "2 days"   # how getting it went
python scripts/evergreen.py export <repo>          # copy into <repo>/.agents/ for other agents
python scripts/evergreen.py pack [--out <dir>]     # evergreen-<version>.zip + evergreen.plugin
python scripts/evergreen.py pack --mail            # Gmail-safe zip (scripts stored as .txt); unmail <dir> restores
python scripts/evergreen.py pack --share [--mail]  # copy for someone else: no profile/, no address in the config
python scripts/evergreen.py where                  # root, store, env name, baseline, git remote/role/ahead/behind, last publish, notify readiness
python scripts/evergreen.py publish [--dry-run] [--branch] [--message "..."]   # commit + push to master (maintainer) or branch + PR; runs by itself after checked/bump/tested and at session end
python scripts/evergreen.py pull [--dry-run]       # fast-forward this clone from the trunk; says when a reinstall is due
python scripts/evergreen.py baseline [--from <pack zip>]   # snapshot this install as the base for diff
python scripts/evergreen.py diff                   # what this clone has not published yet (UPDATE.md, changes.patch, manifest.json)
python scripts/evergreen.py notify [--if-changed] [--dry-run] [--transport smtp]   # git: same as publish; email route when update.transport is email
python scripts/evergreen.py merge <bundle | changes.patch | saved-email.txt> [--dry-run]   # email route: fold a bundle into the trunk
python scripts/test_evergreen.py                    # self-test
python scripts/bench_intervals.py [--seed N] [--days 730] [--json]   # benchmark the refresh schedule on synthetic units

Run the script by absolute path; the shell's working directory is usually your project, not the plugin. Every command fails soft (prints a note, exits 0) unless --strict follows the subcommand, so hooks never break a session.

Making things evergreen

  • An existing skill: say "make X self-maintaining" (evergreen-convert). It adds the companions, migrates any hand-rolled refresh state, appends the Step 0 and Maintenance sections, and registers the unit. Skills outside the plugin point at it (protocol: "plugin", a short MAINTENANCE.md that says where to look), so a protocol update never means editing every skill; only skills that travel to machines without the plugin get the full self-contained copy.
  • A new skill: just ask for a skill; evergreen-new makes it evergreen unless you say "plain skill", and hands it over with its suite run.
  • A skill that misbehaves: say what it did not do ("the delegation skill never delegated"); evergreen-tune reproduces it in a fresh context, classifies the failure, researches the subject's testing and tooling if the unit is stale, fixes, and re-runs. "Test this skill" or "prove X works" runs evergreen-test.
  • A repo: paste the AGENTS.md snippet; docs and codemaps in docs/ can carry their own evergreen.json and MAINTENANCE.md.
  • A knowledge doc: evergreen.py init <dir> --kind doc --main <file>.md ....

Design notes

Five principles settle most edge cases: research beats recall on anything time-sensitive (model knowledge is a stale snapshot on fast-moving topics); lazy and never blocking (staleness is checked on use, the task comes first); delta edits, never wholesale rewrites; subject and ecosystem both, since a skill can be right about the world and wrong about the tools; evidence, not claims, since a skill that narrates an action reads exactly like one that performed it. The evidence for each, and for the interval math, the learnings format, the testing rules, and the codemap rules, is in RESEARCH.md with sources.

Testing

Why a separate pillar: refresh keeps a skill current, and a current skill can still fail in three quiet ways. It does not trigger on the user's phrasing (the most common failure in 2026 evals). It triggers and narrates the action without performing it, which the transcript cannot distinguish from success. Or the tool it depends on moved, and it falls back to another route. protocol/TESTING.md covers the case kinds, the evidence rule, the harness per environment (claude plugin eval, documented since Claude Code 2.1.269 and preferred where it runs, then skill-creator's runner, the evergreen-tester agent, a headless CLI, promptfoo and friends as options), how a suite is judged (three runs per case, with-versus-without), and the bounded tuning loop with its research gate. The suite format is skill-creator's evals/evals.json plus a few fields, so its runner and description improver work on the same file. Budgets keep files short: main files under 200 lines, active learnings under 200, codemaps under 150; overflow is archived, not deleted, so the reasoning lineage survives.

The schedule has a benchmark of its own. scripts/bench_intervals.py imports the interval rule from evergreen.py and runs it on synthetic units whose facts change as a Poisson process with mean gaps from 2 to 365 days (plus a regime-shifting class and a release-like bursty class), next to a fixed interval with the same number of checks, the tier's start interval held fixed, a fixed 14 days, a FreshCache-style constant per class, and an oracle that checks at every change. First results (seed 1, 730 days, TESTS.md T-20260923-1, RESEARCH.md R-20260923-10): the rule spends checks roughly in proportion to how often a unit changes (about 214 a year for the 2-day class, 1 for the 365-day class, 43 on average), and its share of time holding a stale material claim (17.4 percent) matches a fixed 14-day interval (17.6) at 1.7 times the checks, while a fixed interval with the rule's own check count per unit gets 13.6 percent: under memoryless change, reacting to the last check costs about 25 to 70 percent more delay on material changes than spacing the same checks evenly (every class, ten simulated years). The change model is synthetic (real topics cluster changes around releases and are not independent), so read it as a comparison between schedules, not a forecast; the rule is unchanged, and what the numbers suggest is written as open questions for the owner in RESEARCH.md.

Packaging

A clone is the normal install; packing is for Cowork's .plugin, a backup, and share copies. python scripts/evergreen.py pack (or the evergreen-pack skill) writes two files next to the plugin folder: evergreen-<version>.zip, which unzips to an evergreen/ folder plus INSTALL.txt and is the thing to back up or carry to another machine, and evergreen.plugin, the same files flat, which is what Cowork's "Save plugin" button installs (Claude Code installs from the folder and never needs it). Python's built-in zip does the work; nothing to install, and both open in 7-Zip, Explorer, or Finder. No Python? scripts/pack.ps1 does the same with 7-Zip if present or Windows' built-in Compress-Archive.

Emailing it: Gmail rejects .ps1 and other script types even inside a zip, so python scripts/evergreen.py pack --mail writes evergreen-<version>-mail.zip with those files stored as name.ps1.txt. After unzipping, python evergreen/scripts/evergreen.py unmail evergreen renames them back (INSTALL.txt inside says so). Or attach the normal zip via Google Drive instead.

Every archive also carries INSTALL-PROMPT.txt, and pack writes a copy beside it as evergreen-<version>-INSTALL-PROMPT.txt: the one-paste install prompt from §Install, filled in with the version. pack --mail --split 24 additionally cuts the mail zip into evergreen-<version>-mail.zip.part01, .part02, ... for routes with a small per-attachment ceiling (a mail connector that carries attachments as base64 through the model, for one); the prompt tells the recipient to join them in order.

The archive includes profile/ (your preferences and environment facts), so a copy for someone else uses python scripts/evergreen.py pack --share [--mail] instead: evergreen-<version>-share.zip has no profile/, ships a config with no address and notify.auto off so their install mails nobody, and names their marketplace my-local. A share pack never baselines your install and never tags the trunk.

One public trunk, many clones

This repository is the official version. Every install is a clone of it and keeps itself current with pull. If the user agrees to it at install (evergreen.py contribute yes, revocable at any time), the install also sends improvements to the protocol, skills, scripts or templates back as draft pull requests in the user's name, keeping the user informed: every send is announced in the session with its URL, and nothing is sent until the user has said yes (protocol 1.7, PROTOCOL.md §10). A private fork is for customisations you do not want to share (a filled-in profile/, a real address in the config, private units): it pulls from here routinely and sends a pull request back only when it hits a lesson worth sharing, carrying that lesson alone. On an install that agreed, when the plugin changes itself (a refresh, a learning, an environment fact, a test run), evergreen.py publish commits the plugin's tree with a subject that names the new entries and pushes: a clone the host lets write to master (a maintainer's) pushes there after a rebase; any other clone pushes update/<env>-<stamp> and opens a pull request with gh against the trunk. The role comes from the remote (a dry-run push), or is pinned per machine with git.role / EVERGREEN_GIT_ROLE; publish --branch sends any change through a pull request, which is the rule for scripts, hooks and the protocol even from a maintainer's clone. It runs after checked, bump and tested on the plugin and from Claude Code's SessionEnd hook (the owner's editing clone stays quiet there, notify.report_trunk). The append-only logs merge by union through .gitattributes, so two clones adding entries keep both; a rebase that still conflicts turns into a pull request rather than a forced push. evergreen.py pull fast-forwards a clone and says when the tool's cached copy needs a reinstall. Set-up and failure cases are in skills/evergreen-publish; which route suits which machine is in protocol/PORTABILITY.md.

Email is the fallback for a machine that cannot reach the repository (update.transport: "email", or EVERGREEN_UPDATE_TRANSPORT=email): notify builds an update bundle and emails it to notify.to through classic Outlook, Microsoft Graph or Gmail SMTP, or leaves it in EVERGREEN_HOME/outbox/ for an agent; at the trunk, evergreen.py merge folds the bundle, patch or saved email in (entry union with ID renumbering, evergreen.json as data, git 3-way for the rest; never evergreen.config.json, code only with --allow-code). Details in skills/evergreen-notify and skills/evergreen-merge.

In this public repository evergreen.config.json carries a placeholder address and profile/ is an empty template. A private fork is where a real address and a filled-in profile belong; a pack --share archive carries neither.

Hooks, git and files outside the project

The plugin's hooks run only in Claude Code and Cowork; other agent products get the skills alone. Each hook runs scripts/evergreen-hook.sh, which looks for a working Python (py -3, python or python3) and runs scripts/evergreen.py. It never stops a session: on any error it exits with no output.

  • SessionStart runs audit --brief. It reads the evergreen units on this machine and prints the ones due for a refresh, a claim check or a consolidation. It makes no network call and prints nothing when nothing is due.
  • PostToolUse on the Skill tool appends one line (time, skill name, session id, transcript path, working folder, machine name) to uses.jsonl in the evergreen home folder. It prints nothing and sends nothing.
  • PostToolUse on Write, Edit and MultiEdit returns at once unless the file is a SKILL.md. For a SKILL.md it runs the static worth reading, which is local, and may give the model a one-line warning.
  • SessionEnd starts a detached notify --if-changed, which does nothing unless this install answered yes to contribute. With that yes and the default git route, it commits the plugin's own changed files in the plugin's clone and pushes them: to master of the trunk repository when this clone may write there (a maintainer's), otherwise to a new update/<machine>-<time> branch with a draft pull request opened by gh pr create against m4bwav/evergreen-protocol. It finds out which by a dry-run push. With the email route instead (update.transport: "email"), it emails an update bundle to the address in evergreen.config.json, The shipped config has a placeholder address and notify.auto off, so the hook sends no email until you set both.

The same publish runs when you or the agent run publish, notify, checked, bump or tested on the plugin itself. pull fetches the trunk and fast-forwards the clone. evergreen-merge merges pull requests on the trunk with gh and pulls. Nothing is ever force-pushed.

Files written outside the project: the evergreen home folder (~/.evergreen by default; evergreen.config.json or EVERGREEN_HOME moves it), which holds the registry, the contribute answer, codemaps, baselines, the use log, the email outbox and the local book of install recipes; and the evergreen files (RESEARCH, CHANGELOG, LEARNINGS, TESTS, evergreen.json, evals) of each skill or doc you make evergreen, wherever that unit lives. Update bundles for the email route go to the outbox. pack writes zip and .plugin archives only when asked.

Privacy

Everything runs on your machine. The scripts are Python standard library only, with no analytics and no telemetry. They read and write markdown and JSON in the evergreen home folder, in the plugin's clone and in the units you register.

Data leaves your machine by these routes only:

  • GitHub, through your own git and gh: pulls from the trunk repository, and, only after you answer yes to contribute, pushes and pull requests carrying the plugin's file diffs (never transcripts). Branch names, commit messages and pull request text carry the machine's name (its host name, or EVERGREEN_ENV). Setting DO_NOT_TRACK or CI turns contribution off.
  • Email, only if you switch the update route to email and set a recipient: through classic Outlook on Windows (its COM interface), Microsoft Graph (the Graph PowerShell module's own sign-in, with the Mail.Send scope), or SMTP (Gmail's server by default) with the user name and app password you put in EVERGREEN_SMTP_USER and EVERGREEN_SMTP_PASS or in a password file. The compose-url transport, when you choose it, opens a Gmail compose window in your browser for you to send by hand.
  • Web research: when a unit is refreshed, the agent reads public web pages with its own web search and fetch tools. The scripts fetch nothing for this.
  • Set-up checks (evergreen.py setup), only when run: a GET request to each URL a unit's SETUP.md names, a request to your Ollama server (OLLAMA_HOST, localhost:11434 by default), and the read-only access probes the unit's SETUP.md names, run in your shell with your own sign-ins. A probe's output is not kept; one redacted error line is.
  • Paid model runs (worth --probe, --ab, --heavy), only when asked: they run the claude command line with your Claude sign-in, or with ANTHROPIC_API_KEY for --blind, and cost model usage.

What is kept: everything above stays in the evergreen home folder, the plugin's clone and the unit folders. worth --triage reads your Claude Code transcripts on this machine to count real skill use; it sends nothing. Nothing goes to the plugin's author except a contribution you agreed to.

Credentials: evergreen stores none of its own. Git and GitHub use the login git and gh already have. The SMTP route reads an app password that you create and put in an environment variable or a file; an agent never writes it. The Graph route checks that the Graph module's token cache exists and leaves the token to that module. The environment variables it reads are its own settings (EVERGREEN_HOME, EVERGREEN_ENV, EVERGREEN_CONTRIBUTE, EVERGREEN_GIT_ROLE, EVERGREEN_UPDATE_TRANSPORT, EVERGREEN_NOTIFY_TRANSPORTS, EVERGREEN_PLUGIN, the three EVERGREEN_SMTP_ variables), DO_NOT_TRACK, CI, OLLAMA_HOST, CLAUDE_BIN, CLAUDE_CONFIG_DIR, ANTHROPIC_API_KEY (checked for --blind only), and standard system variables (PATH, TEMP, TMP, LOCALAPPDATA, WSL_DISTRO_NAME). setup also checks whether variables a unit needs are set, without printing their values.

License

MIT.