Skip to content

lynxtwo/scaffold-kit

v0.4.0FSL-1.1-MIT

Plan broadly, build narrowly, test reality, expand deliberately. A planning interview that produces an Architecture Document, an Engineering Document, a Decision Log and one active Slice Brief that bounds what an AI agent may build.

The Scaffold Kit

Version 0.4 (Draft). September 26, 2026.

A document scaffolding system for AI-assisted software development. It works for any kind of build: app, service, CLI tool, library, database system, game, automation, or site. It works for any experience level, from a first project to a senior engineer converting judgment into specifications.

The model behind it:

Plan broadly. Build narrowly. Test reality. Expand deliberately.

You do not build the whole product first. You design the whole puzzle, then build one valuable, production-quality section of it, with clean places for future pieces to connect. Then you repeat, one slice at a time.

Start here: plan an AI coding project

Scaffold Kit turns a project description into four planning documents: architecture, engineering rules, decisions, and one active slice brief. It writes documents first. Product code starts only after the owner approves the slice brief.

Read the overview and starter example, or use the files directly:

  1. Review this source and the FSL-1.1-MIT license. Version 0.4 is a draft, not a claim of universal host compatibility.
  2. Give your assistant the Conductor and load each of the four templates as its phase starts. Rough project descriptions are welcome.
  3. Ask: “Use Scaffold Kit 0.4 to plan a small reading-list app. Start with triage and questions. Write planning documents only; do not build product code.”
  4. Review the decisions, assumptions, and first slice. Build only after the active brief records your approval.

This reading-list prompt is an illustrative starter, not an evaluation result. For existing code, Anti-Dark-Code provides codebase mapping and verification evidence; Scaffold Kit turns that evidence into planning decisions and a build boundary.

Installation depends on your host's current discovery rules. The Codex verification record includes failed file-resolution and discovery checks; packaging is not acceptance on every host. The manual file workflow avoids assuming a plugin has loaded.

Revision note: v0.4 (Draft)

Version 0.4 applies the findings from a second field run, a real two-founder product built slice by slice under v0.3 for six weeks, and packages the kit as an agent skill. Two lessons from the run: the kit bounded scope but never bounded verification effort, so a proof campaign that served no requirement ran for days inside a slice; and the Decision Log grew past nine hundred lines because receipts and run output were pasted into entries. Changes: a verification bound and a two-vocabulary rule in the Engineering Document; a Gate column and an unlisted-inputs row in the slice brief's acceptance table; a SLICE STATE block; a rule that evidence lives in ledgers and entries link it; a spoken build gate; three new failure modes to refuse; a brownfield procedure that names the map artifacts anti-dark-code produces; a stdlib audit script that mechanizes the Phase 6 checklist; and the layout below, which every skill host can load. Templates lost their numeric prefixes; the mapping is in the table.

Revision note: v0.3 (Field-Tested)

Version 0.3 applies the fifteen findings from a full stress-test run: a real T2 project with payments and personal-data risk flags, full interview mode, thirty-three sections, twenty-six logged decisions. New mechanics: a fast-run mode with mandatory checkpoints at irreversible decisions and risk-flag sections; a named spike pattern for closing Open decisions by evidence; a slice growth tally the Conductor must present at slice selection; an optional build-order table in the slice brief; and a product-context companion convention for pricing hypotheses, validation thresholds, and staged plans. Codified from the run: sections may share a turn within the question cap; readbacks on tap surfaces use Confirmed or adjust-in-reply; multi-select forms suit precedence questions; documents materialize at phase boundaries; mid-interview amendments supersede through the log; and the Decision Block gains an optional AI-recommendation-differed line, used seven times in the test. Adjustments: the audit rule now permits Open decisions inside the slice when their closing spike is scheduled before dependent work; risk-flag depth may be scheduled to a named trigger with full depth as the default; and the triage timeline field accepts staged notes. The run also validated the core loop: seven of twenty-six decisions overrode the AI recommendation, and every override produced a more specific design.

Revision note: v0.2 (Audited)

Version 0.2 incorporates a full audit of the 0.1 draft. Fixes: a wrong cross-reference in the Engineering template, a status vocabulary mismatch between the Conductor and the Decision Log (Superseded was missing), and duplicate entries in the hygiene word lists. Structural changes: the First Slice Brief became the Slice Brief, one template for every slice with filled briefs numbered SLICE-001 onward, plus an expansion loop in the Conductor; the Decision Log is now written continuously with a completeness check phase; triage asks whether code already exists and routes existing codebases through a mapping pass first; and a shape adaptations table remaps sections for games, libraries, CLIs, data systems, automations, and sites.

What is in the kit

Everything lives under skills/scaffold-kit/, which is the skill directory every host loads.

FileWas (v0.3)What it is
SKILL.mdnewThe entry point: what the kit produces, which template each phase loads, the gates that do not move.
references/conductor.md01-CONDUCTOR.mdThe protocol. Hand this to any AI along with your project description. It runs the interview and fills out everything else.
assets/templates/ARCHITECTURE.md02-TEMPLATE-ARCHITECTURE.mdThe Architecture Document (ADD). The whole puzzle: modules, interfaces, data flow, technology, extension points.
assets/templates/ENGINEERING.md03-TEMPLATE-ENGINEERING.mdThe Engineering Document (EDD). The rules for placing pieces: requirements, data model, security, standards, verification, definition of done.
assets/templates/DECISION-LOG.md04-TEMPLATE-DECISION-LOG.mdThe Decision Log. Every significant choice, its reasoning, and its revisit trigger, kept inside the project instead of scattered across old chats.
assets/templates/SLICE-BRIEF.md05-TEMPLATE-SLICE-BRIEF.mdThe Slice Brief. The narrow section being built right now, specified so an agent can build it without touching the rest.
references/claude-addendum.md06-CLAUDE-ADDENDUM.mdOptional. Claude Code specifics: CLAUDE.md wiring, permissions, verification hooks.
references/working-with-anti-dark-code.mdnewHow the kit and the anti-dark-code skill divide the work on an existing codebase and at verification.
scripts/kit_audit.pynewMechanical audit of a filled document set. Standard library only.

One convention rides alongside the files: prior research and business planning (pricing hypotheses, validation thresholds, staged rollout) lives in a PRODUCT-CONTEXT.md companion. The Conductor mines it during the interview instead of letting it rot in chat history.

Quick start

As a skill: install skills/scaffold-kit/ where your host reads skills (.agents/skills/scaffold-kit/ for Codex, Cursor, Copilot and Gemini; .claude/skills/scaffold-kit/ for Claude Code), or add this repository as a plugin marketplace in Claude Code (claude plugin marketplace add LynxTWO/scaffold-kit, then claude plugin install scaffold-kit@lynxtwo). Then describe what you want to build in a paragraph. Rough is fine. The interview exists to sharpen it.

By hand: give your AI skills/scaffold-kit/references/conductor.md and the four templates under skills/scaffold-kit/assets/templates/, then describe the project. Answer the questions, or choose fast-run at triage and review filled defaults in batches instead. The AI fills out the Architecture Document, the Engineering Document, the Decision Log, and the first Slice Brief. After the owner approves the active Slice Brief, implementation can build only that slice.

Time expectations: a T1 project interview usually fits in one sitting. T2 takes a few sessions. T3 takes longer, and should. The Conductor's state block makes pausing and resuming safe.

The flow

Your description
      |
      v
CONDUCTOR: intake and triage  ->  Triage Card (type, tier, team, experience, risk, starting point)
      |
      v
Architecture interview        ->  ARCHITECTURE.md    (the puzzle)
      |
      v
Engineering interview         ->  ENGINEERING.md     (the rules)
      |
      v
Decision completeness check   ->  DECISION-LOG.md    (the reasoning)
      |
      v
Slice selection               ->  SLICE-001-[name].md (the build boundary)
      |
      v
Audit pass, then build the slice. Verify with evidence.
Then loop: next slice brief, expand deliberately.

Why documents at all

AI agents did not make engineering discipline obsolete. They made undisciplined engineering expensive at machine speed. A "minimum" viable product quietly becomes the permanent foundation, and every new feature forces the agent to reopen the finished puzzle: trace dependencies, reinterpret old decisions, migrate the database, patch around shortcuts that were never supposed to survive.

These documents are how judgment becomes legible to an agent. The agent gets explicit context, documented boundaries, clear acceptance criteria, and a verification standard where evidence counts and "the agent said it worked" does not.

When not to use the kit

A true throwaway spike does not need scaffolding: something you build in an afternoon to answer one question and then delete. Delete it on schedule and the kit was never needed. The dangerous middle is the spike everyone knows is temporary, right up until real users depend on it and nobody is allowed to replace it. If a spike survives its question, it enters the kit through triage like any other existing codebase.

Scaling honesty

The kit scales to the project. Triage assigns a tier, and sections collapse or expand to match. A weekend prototype is never dragged through enterprise ceremony, and a system that touches money or personal data is never allowed to skip the sections that matter. Risk flags override tier.

Versioning the documents

Documents produced by this kit start at v0.1 Draft. When a document survives a full audit pass, bump it and mark it Audited, for example v0.2 (Audited). The Conductor explains the audit pass. Filled documents are living files: they change through the Decision Log, not through silent edits.

Tests and audit

python3 -m unittest discover -s skills/scaffold-kit/tests exercises the audit script. python3 skills/scaffold-kit/scripts/kit_audit.py --docs <your docs dir> audits a filled document set. evals/ holds the behavioral trial set and its evidence.

License

FSL-1.1-MIT. Copyright 2026 Daniel Boyd (LynxTWO). LICENSE.md has the complete terms.