m4bwav/evergreen
Self-maintaining, self-proving skills and knowledge: adaptive web research on four tracks (the subject, the skills, plugins, MCP servers and knowledge graphs built for it, how others use agents on the same goal, and how they test that the job was done), never-teach-twice learnings capture, plugin-owned codemaps, a portable AI-preferences profile, and an eval suite per skill (trigger, action, outcome, decoys, baseline) that passes on evidence outside the transcript, with a bounded tuning loop that researches the subject's own testing when a skill fails, and a worth check that warns while a skill is made or changed when it reads as mostly cost and decides KEEP, TRIM or CUT from the with-versus-without A/B. Skills say what they need (tools, packages, keys, servers, models, accounts, and access to databases, telemetry, cloud roles and APIs) and what for: on a new machine or agent harness, or when something is missing, the agent explains it, installs what is safe after saying so, hands the user exact steps for the rest, drafts least-privilege access requests without ever granting itself access, and records the recipe that worked and how each attempt went for that environment. An end-of-session wrap-up harvests the session transcript and keeps only what will save time or tokens or raise quality, routed to its narrowest home. Converts existing skills and makes new ones evergreen by default. Lives in one public git repository (https://github.com/m4bwav/evergreen-protocol): every install is a clone that pulls to stay current and, only if its owner said yes at install, sends improvements to the protocol, skills and scripts back as pull requests it opens itself (log entries go straight to master from a maintainer's clone), with email as the fallback route.
Status overview of the whole evergreen catalog (skills, docs, codemaps, profiles) in one table: tier, last check, due date, staleness, drift, test status (untested, overdue, failing), link integrity and budget problems, then hands what is due to evergreen-refresh. Use on catalog-wide questions such as 'evergreen status', 'what's stale', 'anything due', 'audit my skills', 'check the skills', right after installing or updating the evergreen plugin (it may have sat on a shelf), at the start of work in a new environment, or whenever a session-start hook reports due units. Also the place to register units the plugin does not know about. Refreshing a single unit is evergreen-refresh.
Convert an existing skill, plugin, knowledge document, or repo doc into a self-maintaining evergreen unit: adds RESEARCH.md, CHANGELOG.md, LEARNINGS.md, evergreen.json, for skills a TESTS.md and evals suite, a Maintenance section in the main file, and either a link to the plugin's protocol or a standalone MAINTENANCE.md companion. Use on 'make X self-maintaining', 'convert X to evergreen', 'add research refresh to X', 'make this skill keep itself current', 'evergreen-ify', or when auditing skills that still refresh ad hoc (a hand-rolled state file, a 'refresh every N days' note) and should join the paradigm.
Show what the evergreen plugin changed about itself since it was installed or last reported: an update bundle with a human digest (new changelog, research and learning entries, state changes, file list) and a unified diff against the install baseline. Use on 'what changed in evergreen', 'generate the diff', 'diff the plugin', 'what did the plugin update', 'build the update bundle', 'anything to send home', or before evergreen-notify and evergreen-merge. Also to set or repair the baseline ('rebaseline', 'the diff shows the whole upgrade').
Capture a lesson so no one has to teach the agent the same thing twice: user corrections ('no, actually...', 'I told you before', 'remember that', 'stop doing X'), the same error or failed approach happening a second time, a discovered workaround, an environment fact (path, port, tool quirk, what works where), or a stated preference about how work should be done. Use proactively the moment any of these happens, not at the end of the session, and for 'save that as a learning', 'note that for next time', 'consolidate the learnings', or an end-of-task reflection. Routes the entry to the right file (skill, codemap, environment profile, or preferences profile) and mirrors a pointer into native memory.
Build or update the plugin's own codemap of a repository or system while exploring it, so the exploration is never repeated. Use when the user asks how an unfamiliar or recurring codebase, service, or system is put together ('walk me through this repo', 'what talks to what', 'where is Y configured', 'how does X handle auth' in a repo you will plausibly see again), when you are about to grep around an unfamiliar repo, or on 'map this repo', 'update the codemap', 'is the map stale'. Before answering from an existing map, check drift first: maps can be out of date, incomplete, or confused. Not for one-off questions in a repo with no future, and not a replacement for a repo-specific skill that already covers that repo (use both: the specific skill for the task, this one to keep the map current).
Land what other evergreen clones published into the trunk repository: review and merge the plugin's update pull requests on the host (gh), resolve the rare conflict with the delta-edit rule (the logs merge by union; a duplicate entry ID is renumbered), then pull every clone and reinstall where skills or scripts changed. Also folds an email update bundle (a folder, a changes.patch, or the update email saved as text) into the trunk for the fallback route. Use on 'merge the update', 'merge the evergreen PR', 'review the pull requests from the plugin', 'the rebase conflicted', 'apply the patch from work', 'pull in the changes from the email', 'sync the trunk', or when a pull request or update email from another machine has arrived.
Add the self-maintaining layer to a skill being created (research refresh on an adaptive schedule, learnings capture, changelog, an eval suite that proves it triggers and acts, double-linked companion files) and pick its tier. Use whenever the user asks to create, write, draft, or scaffold a skill, command, or plugin skill ('make me a skill for X', 'new skill', 'turn this into a skill'), unless they explicitly ask for a plain skill with no maintenance. The authoring itself is skill-creator's job (or create-cowork-plugin's for a whole plugin); this skill runs alongside, after the SKILL.md draft exists, and adds evergreen.
The email fallback for publishing the evergreen plugin's self-update when a machine cannot reach the trunk git repository (update.transport email): mail the digest and diff to the owner (notify.to in evergreen.config.json) through the script's ladder (classic Outlook COM, Microsoft Graph, Gmail SMTP with an app password), then agent-driven routes (a Gmail connector, Claude in Chrome with the Drive picker), then the outbox. Use on 'email me the changes', 'send the evergreen update by email', 'did the update email go out', 'the notify failed', 'set up the update email', 'this machine cannot push', or when a session-end hook reports an unsent bundle. Pushing to the repository is evergreen-publish.
Package the evergreen plugin into a single archive to email, back up, or install elsewhere: a versioned evergreen zip (folder inside, opens in 7-Zip or Explorer), an email-safe variant that Gmail accepts, and evergreen.plugin (flat, for Cowork's Save plugin button). Use on 'zip up evergreen', 'package the plugin', 'give me one file I can email myself', 'gmail blocked the zip', 'make a .plugin file', 'export evergreen for another machine', or 'bundle it for Copilot/Codex/Cursor' (which uses export instead of pack). No zipper to install: Python's built-in zip does it, with a PowerShell fallback.
Publish the evergreen plugin's self-changes (refreshes, learnings, environment facts, test runs) to the trunk git repository: a maintainer's clone commits and pushes straight to master, anyone else's clone pushes an update branch and opens a pull request; then bring a clone up to date with pull. Use on 'push the evergreen changes', 'publish the update', 'sync evergreen', 'open a PR for the plugin changes', 'pull the latest evergreen', 'is my evergreen clone behind', 'did the push go out', 'the publish failed', or when evergreen-refresh, evergreen-learn, evergreen-test, or a session-end hook reports unpublished changes. Email is the fallback route (evergreen-notify) for a machine that cannot reach the repository.
Refresh one named evergreen unit's research from primary web sources and reschedule its next check adaptively; also works through the whole list an audit hands over. Use when the user asks about that specific skill or doc: 'refresh X', 'update X's research', 'is X stale', 'is this still current', 'check for changes to X', when a skill or doc reports it is stale or past due, or when a learning has flagged a contradiction in it. Listing what is due across the catalog is evergreen-audit.
Get a skill running on this machine: check what it needs outside itself (a tool, a package, an API key, a local server, an AI model, an MCP server, or access to a database, Application Insights, a cloud role or an API), tell the user what each missing piece is for, install what is safe after saying so, walk the user through sign-ins and access requests, re-check, and record what worked and how it went for this operating system and agent harness. Use whenever a skill fails for a missing piece ('command not found', 'No module named', connection refused, model not found, a missing MCP tool) or a refused permission (401, 403, 'AuthorizationFailed', 'permission denied for table'), on the first use of a skill on a new machine or harness, and on 'what does X need', 'set up X', 'I don't have access to', 'get me access to the database', 'X does not work on my laptop'. Also when writing a skill's SETUP.md. Not for installing a dependency into the user's own project, or granting access in the user's own systems.
Write and run the eval suite that proves a skill or plugin works: trigger prompts and decoys, action cases proven by evidence outside the transcript (a tool call in the trace, a file, a marker, a remote record), outcome cases, and a baseline without the skill. Use when a skill has just been written (with evergreen-new or skill-creator), after a refresh or a tune edited a skill, on 'test this skill', 'does X actually trigger', 'run the evals', 'prove X works', 'regression test the skills', 'write tests for X', or when the audit lists untested or overdue units. When a case fails, hand over to evergreen-tune. Not for testing application code; that is the repo's own test suite.
Fix a skill or plugin that failed a test or failed in use: it did not trigger, it described an action without performing it (said it delegated, never delegated), it did the job by a forbidden route, or it produced the wrong result. Reproduces the failure in a fresh context, classifies it, writes the learning, researches the subject's testing and tooling when the unit's research is stale, makes the smallest edit, and re-runs the suite. Use on 'X didn't do Y', 'the skill never delegated', 'the skill didn't trigger', 'fix this skill', 'tune X', 'why did X fail', 'X keeps doing it wrong', or when evergreen-test reports a failure or a unit shows failing tests. Not for application bugs; that is the repo's job.
Judge whether a skill is worth its tokens, one skill or a catalog, as a lite scan in seconds or a heavy measured run: its cost per session and per use, how often it really fires (tests excluded), whether baselines pass without it, what a fresh model already knows, then the with-versus-without A/B that decides KEEP, TRIM or CUT, plus a person's FIX or SUPERSEDED ruling. Use while making or updating a skill ('is this skill worth it', 'am I overdoing this skill', 'does this skill actually help', 'is this skill useless', 'should I cut this skill', 'is it bloated', 'trim this skill', 'which of my skills are dead weight', 'audit my skills for uselessness', 'is X superseded', 'quick check of my skills', 'test this skill thoroughly'), after evergreen-new, evergreen-refresh or evergreen-tune edits a SKILL.md, and when the edit hook prints [evergreen-worth]. Not for proving a skill triggers and acts (evergreen-test), fixing a failure (evergreen-tune), or rewording descriptions (skill-tidy).
End-of-session wrap-up that turns what a work session learned into lasting improvements, keeping only what will save time or tokens or raise quality: harvest the transcript (corrections, refused calls, errors, repeated commands, token sinks, slow calls, files and stores touched), drop what fails a usefulness gate, then send each survivor to its narrowest home (a test, hook or script; a skill's LEARNINGS or a fix to the skill; project docs and handoff; a knowledge base, including vaults and notes folders no skill owns; the user profile; an install recipe) and improve the scripts and processes that cost the most. Use when the user says 'wrap up', 'wrap-up', 'wrapup', 'end of session', 'consolidate what we learned', 'retro', 'compound this', 'what should we keep from this session', 'before I close this', or at the end of a long session that changed skills, scripts or several repos. Not for capturing one lesson mid-task (evergreen-learn), a handoff or project note alone (everlast-capture), or merging a unit's LEARNINGS entries (evergreen-learn consolidation).