Skip to content

m4bwav/evergreen

v0.16.0MIT

Self-maintaining, self-proving skills and knowledge: adaptive web research on four tracks (the subject, the skills, plugins, MCP servers and knowledge graphs built for it, how others use agents on the same goal, and how they test that the job was done), never-teach-twice learnings capture, plugin-owned codemaps, a portable AI-preferences profile, and an eval suite per skill (trigger, action, outcome, decoys, baseline) that passes on evidence outside the transcript, with a bounded tuning loop that researches the subject's own testing when a skill fails, and a worth check that warns while a skill is made or changed when it reads as mostly cost and decides KEEP, TRIM or CUT from the with-versus-without A/B. Skills say what they need (tools, packages, keys, servers, models, accounts, and access to databases, telemetry, cloud roles and APIs) and what for: on a new machine or agent harness, or when something is missing, the agent explains it, installs what is safe after saying so, hands the user exact steps for the rest, drafts least-privilege access requests without ever granting itself access, and records the recipe that worked and how each attempt went for that environment. An end-of-session wrap-up harvests the session transcript and keeps only what will save time or tokens or raise quality, routed to its narrowest home. Converts existing skills and makes new ones evergreen by default. Lives in one public git repository (https://github.com/m4bwav/evergreen-protocol): every install is a clone that pulls to stay current and, only if its owner said yes at install, sends improvements to the protocol, skills and scripts back as pull requests it opens itself (log entries go straight to master from a maintainer's clone), with email as the fallback route.

evergreen-test

Write and run the eval suite that proves a skill or plugin works: trigger prompts and decoys, action cases proven by evidence outside the transcript (a tool call in the trace, a file, a marker, a remote record), outcome cases, and a baseline without the skill. Use when a skill has just been written (with evergreen-new or skill-creator), after a refresh or a tune edited a skill, on 'test this skill', 'does X actually trigger', 'run the evals', 'prove X works', 'regression test the skills', 'write tests for X', or when the audit lists untested or overdue units. When a case fails, hand over to evergreen-tune. Not for testing application code; that is the repo's own test suite.

Read SKILL.md at the source

Pinned to revision 946faeac0810, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.