Skip to content

pooyahayati/vibe-coding-skill

v1.4.0MIT

Risk-adaptive software engineering operating skill for reliable AI-assisted vibe coding, architecture, staged implementation, project intelligence, verification, security, and handoff.

Changelog

Unreleased

1.4.0 — Reliable routing, native release security and local runtime

  • Accept native Trivy's Windows path separators by binding filesystem reports to the canonical absolute host path. Keep other targets and altered immutable image identities rejected; add the actual platform reproduction to the existing release-gate tests.

  • Add a Head-owned, conditional product installation/runtime policy: prefer Compose for suitable self-hosted products, offer user-selected local Docker execution, retain native/WordPress delivery routes, and verify the documented clean-start/core-workflow and required data persistence. Route the policy through planning, build, verification and delivery; document safe local defaults, prerequisites and honest target-specific evidence. No new runtime module, specialist or fixed test suite.

  • Keep security native to the project during development and require installed native Trivy before publication. Add planning preflight and local release scans to the existing Trivy adapter, inspecting fresh target-bound JSON findings instead of treating exit zero or compatibility smoke as approval. Bind required scan evidence to final delivery; gate the Skill's own extracted ZIP before automated publication. Preserve ordinary offline installation, proportional checks and existing specialists; add no scanner ecosystem, model evaluation or host auto-installation.

  • Fix real-project runtime routing from Markdown prose/examples: share source-type marker gates across project and affected-file scans, ignore document path/directory signals, match complete extensions, and keep metadata descriptions out of executable platform/browser evidence. Read browser markers from embedded script bodies rather than page prose. Ignore clear excluded-domain mentions when selecting task packs and required specialists while retaining positive signals, actual source surfaces, explicit/structured selection, conservative risk floors and retained commitments. Add focused reproductions and an extracted-package router/planner route; preserve the existing portable layout and dependencies.

1.3.1 — Targeted reliability hardening

  • R4: Reconcile benchmark/scenario-output guidance with the implemented release gate: no model evaluation or provider key is required for ordinary use or publication, and former credentialed workflow/trigger claims are obsolete. Keep optional collection/aggregation tooling explicitly full-source and separately requested; remove unavailable runner commands from portable guidance, retain its included offline scorer, and distinguish declared policy scores from task contracts/receipts and product outcomes. Update related bootstrap/delivery references and portable documentation. No executable/schema/workflow changes, new tests, model runs, version bump or installation.

  • R3: Treat . and equivalent project-relative forms consistently as project-root scope in execution-plan ownership, retained-contract containment and drift checks. Share scope normalization/coverage with the existing behavior-contract helper; block root/nested ownership across different writers while allowing disjoint scopes. Preserve literal bracket/hidden names, exact retained-contract fingerprints, legacy planner directory shorthand and material drift triggers; reject absolute/parent-escaping coverage. Add the confirmed bound-plan and drift regressions and an extracted-package CLI route. No version bump, new runtime module, dependency, model evaluation or broader path refactor.

  • R2: Preserve platform markers in oversized sources using bounded prefix reads. Bound area-header discovery and invariant extraction as well as project/affected scans; prioritize affected headers and enforce actual byte budgets before reading. Expose truncated/unreadable/skipped inspection with bounded path samples, separate routing uncertainties, Head assessment guidance and plain-CLI warnings. Count real inspected input rather than silent empty placeholders; retain platform evidence gating and local routing. Add large-plugin, generic-PHP, unreadable and read-budget regressions plus an extracted-package routing check. No version bump, model evaluation or automatic risk escalation.

  • R1: Default fresh structured task state and --new-task to schema-3 receipt-backed completion; reserve schema 2 for retained in-flight legacy tasks, including older snapshots without a workflow field. Keep explicit legacy migration and prevent downgrade across capture/handoff/resume. Preserve required specialist assignment IDs against the retained contract before state replacement and plan validation; show commitments in handoffs and require recorded Head reconciliation for removal. Add focused regressions for both confirmed failures, positive compatibility/manual routes and the extracted portable runtime. No version bump, model evaluation or change to authorization.

1.3.0 — Local skill feedback and portable workflow reliability

  • Add an evidence-triggered Head-owned skill-feedback route for confirmed/suspected Vibe instruction, resource, tooling and composition defects. Prepare minimal English reports outside product source, deduplicate locally, preserve attribution and safe task progress, and invite optional user-reviewed email to hayatipooya@gmail.com. No automatic sending, telemetry, product-UI notice, model evaluation or bypass of required controls; include the reference in portable/install validation.

  • Complete D1 guidance for usable behavior-led handoffs, current schema-3 versus declared legacy completion, extracted-package evidence limits, risk-appropriate maintainer checks and stable/source installation choices. Preserve workflow/project-size tables, important links and full author attribution; correct stale implementation descriptions and the contributor newline typo. No runtime behavior, version bump, model evaluation or additional project gate.

  • Add I1 integration checks that build/extract the portable ZIP and exercise light, receipt-backed and cross-boundary specialist routes through its runtime entrypoints. Run them in the existing OS matrix; preserve legacy/manual evidence, scoped freshness and separate Head acceptance without adding runtime instructions, models or host dependencies.

  • Accept literal bracket-containing project paths in specialist Git baselines, assignments/returns, evidence inputs/exclusions, relative artifacts and receipt references; retain wildcard, traversal and link protections. Add regressions for dynamic/catch-all routes, unrelated assignment scopes, real execution receipts, missing/deleted files and stale acceptance.

  • Check bracket-named skill resources referenced in prose or inline code instead of silently skipping them. Preserve literal paths when relocating prose/commands and URL-encode only Markdown link destinations. Advance the package-adapter policy to revalidate cached packages under the corrected dependency extraction.

1.2.0 — Bound outcomes, receipts and specialist acceptance

  • Fix fresh Head preflight rejecting the valid 1.1.0 package because prose mentions the specialist-owned vibe-head-contract.md. Require that contract only for head-delegated packages. Share dependency extraction between validation and relocation: ignore illustrative/literal Markdown, retain executable command dependencies, check nested owner-relative resources and inline links with angle destinations/titles, and preserve fragments and URL-encoded paths during relocation. Advance the adapter fingerprint so prior cached validation is rechecked. Regressions cover the real portable Head package, missing delegated resources, cold/cached adoption, resource syntax and relocation containment without modifying the product or global installation.

  • Implement S2 scoped assignments and separate Head acceptance against actual project changes, local receipts, required-finding history and producer/consumer contract decisions. Bind stage/source/task identities and delivery inputs, defer instruction replacement during active assignments, and prevent schema-3 Done from bypassing retained specialist acceptance. Analysis-stage acceptance stays historical; current-delivery stage checks remain fresh. Preserve tiny/legacy routes and existing authorization; no new model evaluation or blanket test gate.

  • Implement S1 instruction-compatibility assessments separately from specialist currency/package checks. Provide bounded resource diffs and an explicit Head decision for authority, platform, dependencies, verification and scope/authorization. Bind cached acceptance to observed source, installed resources and Head/domain controls; reject mismatches and unresolved conflicts. Only affected active delegation blocks; unchanged acceptance is reused without another assessment or model evaluation. Task-return acceptance is implemented by S2.

  • Implement E2 schema-3 completion against independently retained task contracts and local receipts. Resolve bounded task-local references/digests, validate identity/origin/obligations, recompute scoped input/artifact/context bindings and retain risk/evidence-family rules. Keep schema-2 reports explicitly declared evidence; retain receipt workflow activation through state/handoff/resume and reject downgrade. Support actual manual observations without invented execution; imported remote claims require a separate authorized resolver and remain unverified here.

  • Implement E1 local execution receipts with explicit argument-list commands, bounded timeouts, before/after scoped input and exact-artifact hashes, honest non-pass/unavailable results, external local storage and optional WordPress ZIP version binding. Do not retain process output or environment dumps; reject escaping/link paths and incomplete collection. Package the collector and exercise its actual behavior across supported operating systems. Schema-2 completion remains unchanged; receipt resolution is E2.

  • Implement B2 optional version-1 task-contract validation and propagation through execution-plan drafting/validation, local state, handoff and resume. Preserve stable criteria/behavior/evidence obligations and risk floors; compare edited plans/state against retained outcomes. Label schema-2 results as reported, invalidate changed source/contract context, and retain explicit Head reconciliation without treating it as permission or receipt proof. Legacy/light routes remain supported; package the new shared helper.

  • Implement B1's observable-behavior guidance: reuse acceptance criteria, express actor/action/result and only material failure/protected-state expectations, keep tiny tasks inline and persist durable/high-risk decisions in existing appropriate locations. Choose checks from those outcomes; schema-2 reports use existing descriptions and evidence links until B2/E2 implement structured propagation/receipts.

  • Specify P0's shared behavior, receipt, specialist assignment/return and Head-acceptance contracts, trust limits and explicit completion-schema migration. B2/E1/E2/S1/S2 implement the retained-contract, receipt and acceptance routes; schema-2 compatibility remains available for explicitly reported legacy evidence.

  • Link the design from relevant operating references and maintain the roadmap's implementation/merge status separately from release and installation.

1.1.0 — Lifecycle and scoped specialists

  • Consolidate the Head workflow into Discover, Define, Plan, Design, Build, Verify, Review and Ship with explicit stage outputs and conditional specialist ownership.
  • Register application security, interface-contract and systematic-debugging specialists alongside the sole UI/UX specialist; add multi-domain, stage-aware routing and domain-only authority contracts.
  • Package only selected upstream skills and referenced approved shared resources, rewrite/validate resolved relative links, and inject a small Head-delegation contract without copying specialist methods.
  • Expand the README with workflow ownership/activation, specialist responsibilities/permissions, and separate Graphify/Trivy purpose and use tables.

Earlier specialist composition

  • Keep Vibe as the engineering Head and require the sole UI/UX specialist, pooyahayati/UI-UX-Skill, for affected product interfaces across app, website, dashboard, and WordPress settings work.
  • Add offline specialist selection to context plans and a scoped, latest-source resolver/installer with daily inventory, provenance, protected local edits, atomic replacement/rollback, and Head reload handling.
  • Keep specialist methods in the specialist; preserve Head scope, architecture, risk, acceptance, and final delivery controls. No fixed specialist revisions or model evaluations.

1.0.0 — Reviewed release and software-based publication checks

  • Remove model-evaluation requirements and automatic credentialed workflows from publication. Stable releases retain the complete same-commit software validation baseline.
  • Add a reproducible, installation-checked portable ZIP and SHA-256 release assets.
  • Preserve unresolved risk facts, normalize bounded aliases, recognize the reported Persian destructive request, and carry structured risk into execution planning.
  • Keep custom sibling projects and WordPress plugins distinct; accept ordinary runtime facts without requiring a specialist pack.
  • Require a required completion outcome and optionally bind completion to independently retained approved criteria.
  • Execute WordPress benchmark behavior from the exact ZIP with a trusted PHP harness; reject comment-only and unsafe-output implementations. Report real install preconditions without claiming database freshness.
  • Retain offline maintainer regressions for delivery-envelope validation, recomputed success, scenario/check identity, and simulated-evidence classification. Maintainer research utilities do not qualify releases and are excluded from the portable runtime.
  • Preserve earlier maintainer CLI adapter fixes for option placement, sandboxed shell checks, and parser/version/isolation preflight; no model campaign runs automatically.
  • Read selected-release dependency metadata, isolate offline CLI stubs on Windows/POSIX, retain immutable-baseline change evidence after agent commits/renames, normalize fixture line endings, and let compatibility CI reach the supported fallback resolver.

0.10.7 — Real-delivery benchmark framework and stable dual-evidence gate

  • Added the Real Delivery Benchmark as a second evidence class alongside the existing behavior-selection benchmark.
  • Added control/treatment benchmark arms so Skill impact is measured within the same Agent rather than by ranking Agents against one another.
  • Added provenance-bound delivery result envelopes with fixture, grader, catalog, schema, Skill-tree, baseline-commit, final-tree, diff, changed-path, dependency-change, and raw-output evidence.
  • Added hidden graders outside the Agent workspace and made grader execution operate on immutable snapshots so grader-side artifact generation cannot contaminate captured Agent changes.
  • Added five representative deterministic delivery scenarios: tiny local copy fix, brownfield duplicate-save regression, contained CSV export feature, mixed-monorepo API-only normalization with sibling protection, and a WordPress installable-artifact change.
  • Added deterministic tests proving each representative broken baseline fails and each known-good implementation passes its hidden grader.
  • Added real Codex and Claude Code delivery adapters with bounded workspace editing, provider credential redaction, credential-leak detection, timeout/error evidence, and strict missing-run handling.
  • Added a credentialed Real Delivery Benchmark GitHub Actions workflow and exact-SHA benchmark trigger path.
  • Hardened the stable release gate so stable releases require both complete real-agent behavior evidence and complete real-delivery evidence for the current Skill/catalog/schema identity, with no missing/duplicate/invalid delivery runs and no treatment regressions.
  • Aligned GitHub Release metadata with internal release governance: pre-1.0/RC releases are now published as prereleases.
  • Attempted the real behavior and real-delivery benchmarks on the exact Phase 9D main commit. Both provider credential preflights reported no configured non-interactive credentials, so no real-agent success is claimed and the stable gate remains correctly blocked until credentials are configured.

0.10.6 — Language-independent routing and semantic completion evidence

  • Added config-driven structured context facts for runtime, platform, capability, and concern routing so known semantics no longer depend on natural-language keywords.
  • Added explicit Context Router CLI inputs for structured runtime/platform/capability/concern facts and rejects unsupported fact values instead of silently ignoring them.
  • Propagated structured context facts into execution planning so routing and planning operate from the same semantic inputs.
  • Added regression coverage proving equivalent English/Persian greenfield WordPress/WooCommerce payment tasks select the same capability packs and workflow floor when supplied the same structured facts.
  • Replaced free-form completion evidence diversity with a controlled semantic evidence taxonomy and evidence families.
  • Made Tier 2/3 evidence diversity count semantic families rather than raw labels, preventing multiple verification-test labels from manufacturing independent evidence.
  • Added regressions that reject unknown evidence kinds and block higher-risk completion when all passing evidence belongs to only one semantic family.
  • Kept canonical and portable routing, completion, configuration, references, and Skill contracts synchronized.

0.10.5 — Path-aware context routing for large and mixed repositories

  • Made capability routing project-area aware so platform evidence stays local to the affected application/package in common monorepo layouts such as apps/, packages/, services/, plugins/, themes/, and scoped WordPress plugin/theme directories.
  • Switched known-path routing to bounded path-scoped source scanning, added directory-level affected-path support, and kept the wider project scan only as the fallback when impact paths are still unknown.
  • Refined change-scope detection so sibling applications under the same top-level container can correctly become cross-boundary and trigger execution planning without making repository size itself a planning requirement.
  • Made repository-root and subtree-applicable AGENTS.md files explicit persistent instructions while excluding unrelated sibling agent instructions and preserving the standard README.md fallback context.
  • Expanded context observability with scan strategy/areas plus separate Skill core, selected reference, selected pack, persistent-context upper-bound, and combined context upper-bound byte metrics.
  • Added regressions for same-container mixed monorepos, directory-only paths, local platform evidence after more than 700 unrelated files, scoped AGENTS.md, and sibling-application execution planning.

0.10.4 — Release orchestration reliability

  • Fixed release orchestration for GitHub Actions eventual consistency by trusting the successful triggering workflow_run event when the runs API still reports that same run with a temporary null conclusion; release-gate workflow changes now retrigger the full path-filtered baseline.
  • Hardened live dependency contracts so rate-limited/unavailable optional GitHub source-health evidence passes only when the dependency decision explicitly degrades to REVIEW REQUIRED with a repository.health_unavailable signal; false ACCEPT remains a contract failure.

0.10.3 — Independent review remediation and delivery hardening

  • Added structured risk facts for operation, environment, data sensitivity, and change boundary, with supplemental English/Persian text signals and safer destructive-operation matching.
  • Centralized tier policy generation so router escalation rebuilds the final tier, approval requirement, reasons, and required controls from one source of truth.
  • Upgraded completion reports to schema v2 with explicit criterion IDs/descriptions/required status/evidence links and required evidence that blocks completion when it fails.
  • Converted the toolchain resolver tests to discoverable unittest cases and added regressions for package imports, timeout fallback, and portable Graphify version lookup.
  • Made Graphify compatibility version detection work from the portable Skill without a root VERSION file.
  • Made toolchain compatibility timeouts structured failures so the resolver can attempt the verified last-known-good fallback, and switched child execution to the active Python interpreter.
  • Gated WordPress REST routing on actual WordPress platform evidence so generic REST/API terminology does not inject WordPress/PHP context into unrelated backends.
  • Added explicit task change-scope reporting and changed execution-plan requirements to depend on Tier 2/3 risk or demonstrated cross-boundary scope rather than repository size alone.
  • Kept large-repository project intelligence active while allowing local low-risk work to stay on the light route.
  • Reduced the canonical/portable SKILL.md entry point from roughly 480 lines to a 215-line routed core and added a 260-line validation guard against maintainer-detail creep.
  • Added explicit execution-scoped toolchain sessions so compound operations resolve each managed tool once and reuse the exact selected version without a process-global cache.
  • Made Trivy filesystem execution target-aware: container fallback now bind-mounts the host target read-only and generic Trivy command construction rejects ambiguous unmounted fs scans.
  • Allowed Graphify and Trivy compatibility contracts to use an exact-version native executable when available instead of requiring uvx or Docker unconditionally.
  • Added regressions for same-session version reuse, new-session re-resolution, Windows/space-containing bind mounts, native compatibility runtimes, and explicit filesystem command routing.
  • Added a bounded technology-selection contract: existing capability/stack first, at most 2–3 realistic options when a material choice remains, constraint-based selection, official support/version verification when decision-critical, and explicit rejection of the nearest alternative.
  • Added a stage interaction contract with concrete inputs, outputs, exit criteria, readiness status, limitations, evidence, and next-step reporting while keeping routine implementation details autonomous.
  • Replaced test-count/coverage-quantity thinking with behavior/failure/boundary-based selection, deduplication against existing coverage, stage-specific execution, and an explicit verification stopping rule.
  • Added behavior-contract eval scenarios for existing-stack technology decisions and proportionate bug verification, including regressions against rewrite-by-default, universal language rankings, fixed coverage quotas, duplicate tests, and unrelated E2E repetition.
  • Added a WordPress delivery workflow that distinguishes discovery/bootstrap/implementation/security/lifecycle/UI/WooCommerce/delivery stages and requires verification against the exact installable ZIP for release/distribution work.
  • Added wordpress_artifact.py to build deterministic single-root plugin ZIPs, verify embedded slug/version and SHA-256, perform safe install-shape smoke checks, and drive WP-CLI install/upgrade/deactivate/reactivate checks against exact artifacts.
  • Added guarded uninstall execution for disposable environments and explicit project-specific data-retention verification; deactivation is never treated as uninstall.
  • Added a real WordPress CI contract that builds two fixture releases, installs the first exact ZIP, upgrades to the second exact ZIP, and verifies activation/version behavior in a disposable WordPress runtime.
  • Added observable code-structure invariants for clear responsibility, data/boundary ownership, entry validation, diagnosable failures, useful separation of decision logic from side effects, evidence-before-abstraction, and minimum safe refactoring in existing codebases.
  • Made advanced structure concern-triggered: external integrations require timeout/failure/retry effects, multi-step data changes require transaction/consistency boundaries, multitenancy requires ownership/isolation, and performance structure requires measurement before optimization.
  • Explicitly rejected interface-per-class, repository-per-table, service-per-function, DTO/layer ceremony, message/event buses, microservices, and Clean/DDD/layered architecture when no current requirement justifies them.
  • Added mandatory behavior-contract eval coverage for small local structure, external integration structure, and multidomain module boundaries, including anti-overengineering regressions.
  • Added proportional Plugin Check guidance, conditional HPOS/Cart-Checkout Blocks checks, and an explicit Composer/Packagist fallback because the bundled dependency guard does not claim Composer coverage.
  • Hardened release readiness so RC/pre-1.0 releases require Validate Skill, Cross Platform Smoke, Real World Repository Validation, Agent Skills Spec Compatibility, Tool Contract Tests, and WordPress Artifact Contract on the exact target commit.
  • Consolidated PR/release tool validation into Tool Contract Tests while keeping dedicated Graphify/Trivy workflows for scheduled/manual compatibility monitoring; removed the duplicate Live Integration Contracts workflow.
  • Added release-trigger guarantees for path-filtered validation workflows, orphan-reference detection, and broader merged-branch cleanup for refactor/docs/release work branches.
  • Corrected stale validation documentation so real-world routing coverage and bounded conservative eval escalation match current implementation.

0.10.2 — Latest-compatible-stable toolchain runtime binding

  • Replaced normal operating-version pins for Graphify and Trivy with latest-compatible-stable resolution.
  • Added a shared resolver that tests the latest published stable candidate and falls back to a verified last-known-good version when compatibility fails.
  • Bound runtime execution to the exact resolved version instead of silently invoking whichever local executable happens to be installed.
  • Added exact Graphify execution through a matching local binary or an isolated uvx --from graphifyy==<version> runtime, and exact Trivy execution through a matching local binary or versioned official container image.
  • Added execution-scoped session pins so a resolved tool version remains stable for the current run without re-resolving mid-execution.
  • Removed misleading integration-guard claims that merely detecting an installed executable meant the runtime version was pinned.
  • Packaged the resolver, runtime adapter, and Graphify/Trivy compatibility contracts in the portable skill and extended install/sync validation so portable installations cannot omit these runtime dependencies.
  • Removed fixed current tool versions from normal user-facing documentation while retaining machine-managed last-known-good fallback state and historical release evidence.
  • Expanded CI coverage for latest-compatible-stable resolution, fallback behavior, exact runtime binding, portable installation, cross-platform smoke tests, and live Graphify/Trivy compatibility contracts.
  • Kept the credentialed Codex/Claude Code benchmark deferred in Issue #12; deterministic, cross-platform, tool-contract, and pinned real-repository validation remain the active pre-1.0 release baseline.

0.10.1 — Real-world routing validation hardening

  • Added pinned real-world routing validation against a large Django repository and the official WooCommerce Stripe gateway.
  • Extended real-world validation to execute the Context Router and Execution Plan and record project complexity, selected capability packs, pack-reduction ratio, selected context size, integration points, project-intelligence preservation, and execution-plan requirements.
  • Added selectivity regressions showing that small WooCommerce admin UI work avoids payment/security/external-HTTP concern packs while payment webhook/refund work composes the sensitive packs and escalates to Tier 3.
  • Expanded CI triggers so routing configuration, capability packs, execution planning, and project-intelligence changes automatically run the real-world validation suite.
  • Kept product repositories clean while validating pinned external repositories at exact commits.
  • Deferred the credentialed blind Codex and Claude Code benchmark to future work in Issue #12; it is not a blocker for the current pre-1.0 release path.

0.10.0 — Composable context routing and execution coordination

  • Added an evidence-based Context Router that keeps project complexity separate from task risk and selects minimum sufficient context instead of loading broad references by default.
  • Added composable Capability Packs for PHP, WordPress, WooCommerce, browser JavaScript, WordPress REST, external HTTP, payments, web security, and web performance.
  • Kept Vibe Core authoritative for scope, risk, project intelligence, evidence, recovery, and delivery; domain packs only add relevant constraints, risks, integration points, and checks.
  • Added interaction rules so overlapping concerns can raise the workflow floor without making small UI/copy tasks inherit unnecessary payment/security ceremony.
  • Added Context Plan metrics for selected packs/references, approximate context bytes, integration points, and preserved invariant/risk/project-intelligence coverage.
  • Added lightweight Execution Planning for medium/large projects, Tier 2/3 work, or multi-boundary tasks with explicit workstreams, dependencies, integration ownership, completion levels, and plan-drift approval triggers.
  • Kept one implementation agent as the default and required explicit ownership/stable contracts before parallelization.
  • Changed the existing-project startup protocol to route context before broad architecture/domain loading while preserving project-wide invariants and large-project impact analysis.
  • Added WordPress/WooCommerce constraints for public platform APIs, HPOS/modern compatibility surfaces, plugin/theme interoperability, authorization/input/output boundaries, request/query cost, remote-call failure modes, and payment-state integrity.
  • Added deterministic regressions for tiny WooCommerce UI tasks, overlapping WooCommerce payment flows, WordPress REST plus external APIs, browser-JS overlap, large generic repositories, execution-plan dependency validation, and plan drift.
  • Added portable packaging and installation/structure validation for routing config, packs, router, planner, and routing reference documentation.

0.9.3 — Installation completeness hardening

  • Expanded the offline installation check to require the complete portable runtime script set, including release_readiness.py.
  • Required every reference document directly used by the Skill to be present before installation can pass.
  • Added checks for runtime configuration, eval catalog/schema, Agent metadata, and project templates.
  • Added JSON parse validation for runtime config/eval files.
  • Added regression coverage so incomplete portable installations cannot silently pass installation validation.

0.9.2 — Stable benchmark orchestration hotfix

  • Re-triggered validated release evaluation when the manual Real Agent Benchmark completes successfully, so stable readiness can advance after benchmark evidence arrives without requiring a redundant base-CI rerun.
  • Bound stable readiness to the exact current eval scenario ID list rather than accepting an arbitrary positive scenario count.
  • Required each stable Agent row to report expected, completed, and passed scenario counts equal to the current catalog size.
  • Added regression coverage for benchmark-triggered release orchestration, scenario-set mismatch, and per-Agent scenario-count mismatch.

0.9.1 — Evidence correctness and release provenance hotfix

  • Rejected malformed acceptance-criteria entries instead of silently ignoring non-object or non-boolean completion claims.
  • Required Tier 2/3 completion reports to declare a target commit and excluded passing evidence bound to a different revision.
  • Tightened Tier 3 evidence timestamps to require timezone-aware ISO-8601 values.
  • Hardened benchmark aggregation so complete evidence requires valid runner envelopes, matching Agent/scenario provenance, and exactly one consistent Skill/version/hash identity set.
  • Prevented unexpected per-scenario runner exceptions from aborting the remaining benchmark scenarios; internal failures are now recorded as failing raw evidence and execution continues.
  • Changed release readiness to use the latest check result for each required workflow, so an older success cannot mask a newer failure on the same commit.
  • Changed stable benchmark resolution to consider the latest Real Agent Benchmark run instead of falling back to an older successful run when a newer run failed.
  • Added regression coverage for malformed completion claims, revision mismatch, ambiguous timestamps, benchmark envelope identity, runner exception isolation, and latest-run release semantics.

0.9.0 — Adaptive evidence and release readiness

  • Hardened the real-agent benchmark workflow so Codex and Claude Code failure paths are isolated, aggregation still runs, raw evidence is preserved, and aggregate failures cannot be masked by tee.
  • Added bounded benchmark tier policies that distinguish preferred behavior, legitimate conservative escalation, under-classification, and over-engineering.
  • Added per-scenario tier ceilings with explicit rationale where conservative escalation is permitted.
  • Added risk-aware completion evidence provenance: higher-risk Done claims require stronger, revision-bound, and timestamped evidence.
  • Added deterministic beta, RC, and stable release-readiness gates.
  • Required stable releases to have complete, fully conformant real Codex and Claude Code benchmark evidence.
  • Bound stable benchmark evidence to the current Skill version, portable Skill tree, eval catalog, and output-schema hashes.
  • Wired the validated release workflow to the release-readiness gate while keeping repository branch protection outside the release criteria by project decision.
  • Extended regression, failure-injection, portable-package, Agent Skills compatibility, and cross-platform coverage for the new policies.
  • Kept missing real-agent runs as missing evidence rather than treating them as success; the remaining real-agent evidence requirement stays tracked separately for 1.0 readiness.

0.8.0 — Real-agent benchmark execution and evidence hardening

  • Added strict machine JSON Schema for live-agent behavior contracts.
  • Added real Codex CLI and Claude Code benchmark adapters with blind prompt generation.
  • Added provenance envelopes containing agent version, model override, Skill version, timing, hashes, and workspace-integrity evidence.
  • Added raw stdout/stderr preservation outside the repository.
  • Added strict benchmark completeness checks so missing scenarios or agents cannot produce a conformance rate.
  • Added credential-safe preflight that records only credential variable names, never values.
  • Added isolated temporary project-scoped Skill installation for benchmark runs.
  • Added manual credential-gated GitHub Actions workflow for full Codex + Claude Code evaluation.
  • Fixed non-JSON scorer output and added schema/integrity validation.
  • Added regression coverage for prompt blindness, provenance integrity, missing evidence, and secret redaction.
  • Kept actual Agent performance claims blocked until genuine raw runs exist.

0.7.0 — Dependency intelligence hardening

  • Upgraded Dependency Guard from baseline registry/OSV checks to risk-adaptive multi-signal evidence.
  • Added explicit dependency necessity and purpose as separate judgment fields.
  • Added deps.dev package/version, license, advisory, deprecation, related-project, and verified-attestation evidence.
  • Added GitHub source-repository health checks with authenticated mode when a token is available.
  • Added maintenance/release-age and deprecated-package review signals.
  • Added package-name similarity / typo-squatting signals using explicit trusted names and discovered project dependencies.
  • Added project license allow/deny policy evidence without pretending to provide legal advice.
  • Added Tier 2/3 evidence requirements and Tier 3 provenance-attestation review.
  • Added deterministic unit/failure-injection coverage and live provider contract coverage.

0.6.1 — Audit hotfix and governance hardening

  • Fixed literal escaped-newline corruption in SKILL.md and added semantic validation to prevent recurrence.
  • Clarified the project-intelligence graph policy while keeping machine-generated graph data local-only.
  • Removed the recommended Claude project-local checkout and hardened repository purity against local Vibe Skill checkouts.
  • Pinned core GitHub Actions by commit SHA and made Graphify/Trivy contract tests read approved versions dynamically.
  • Added automatic latest-stable Trivy compatibility proposals.
  • Added official Agent Skills reference-validator compatibility workflow pinned to a known upstream commit.
  • Hardened last-known-good recording with offline installation validation and isolated target validation before upgrades.
  • Added validated-release automation and safe cleanup of unchanged merged work branches.
  • Extended GitHub traceability snapshots and dry-run mutations to Milestones and GitHub Projects.
  • Improved private security-reporting guidance.
  • Removed stale hard-coded internal version identity from dependency/tool requests.

0.6.0 — Recovery, resume, cross-platform, and installation hardening

  • Added repository-first resume context generation that does not depend on chat history.
  • Added safe local-state inspection, quarantine, and recovery.
  • Added offline installation validation for canonical and portable skill installs.
  • Added last-known-good recording, local upgrade planning, and explicit rollback support.
  • Added Linux, macOS, and Windows portability smoke tests across supported Python versions.
  • Added portable-install validation in CI.
  • Documented Python 3.10+ baseline, offline behavior, recovery semantics, and rollback limits.

0.5.0 — Graph provider contract, GitHub traceability, and project state automation

  • Added a provider-neutral graph contract with Graphify as the default adapter.
  • Added local shadow-source graph refresh so Graphify output never needs to live in the user project working tree.
  • Added working-tree fingerprint freshness, including uncommitted and untracked non-ignored source changes.
  • Added local graph query/path/explain routing with stale-graph blocking by default.
  • Expanded Graphify compatibility contracts to cover explicit graph queries and incremental update.
  • Added GitHub traceability snapshots, requirement mapping, verification, and dry-run-by-default Issue creation.
  • Added local project-state capture, drift detection, and handoff snapshots without repository churn.
  • Refactored doctor and integration guard to consume shared graph/GitHub adapters.
  • Added end-to-end tests for graph shadowing, GitHub traceability, and local state automation.

0.4.1 — Local-only tooling and clean project repositories

  • Moved Vibe Coding operational state from project-local .vibe/ to ~/.vibe-coding/projects/<project-id>/.
  • Added a Local Workspace Manager for graph, security, test artifacts, benchmarks, caches, worktrees, and state.
  • Added clone-local exclusions through .git/info/exclude without modifying project .gitignore.
  • Added Repository Purity Gate to block tracked/staged Graphify, Trivy, benchmark, coverage, and Vibe state artifacts.
  • Updated bootstrap, doctor, integration guard, stale-graph tests, and failure injection to use external local state.
  • Clarified that real product tests remain in Git while generated test/coverage reports remain local.
  • Added worktree-safe Git exclude path resolution.

0.4.0 — Real-project validation, failure injection, and agent benchmark

  • Added representative project validation fixtures and deterministic runner.
  • Added failure-injection coverage for prompt injection, hallucinated dependencies, vulnerable packages, stale graphs, and unsupported Done claims.
  • Added a structured completion evidence gate.
  • Added vendor-neutral aggregation for real Codex/Claude Code behavior contracts.
  • Added validation and benchmarking guidance without fabricated agent results.
  • Added CI coverage for project validation and failure injection.
  • Cleaned duplicate changelog heading.

0.3.1 — Live agent behavior evaluation

  • Added a vendor-neutral JSON behavior contract for real Codex/Claude Code evaluation.
  • Added deterministic scoring of agent tier, approval requirements, required controls, and forbidden actions.
  • Added machine-checkable control contracts to every eval scenario.
  • Added regression tests for passing and failing live-agent outputs.

0.3.0 — Risk classifier, full dependency adapters, and hardened tool contracts

  • Added deterministic, explainable risk classification with Tier 0–3 workflow floors.
  • Added Maven Central, NuGet, and Go module support to Dependency Guard.
  • Added read-only GitHub/Graphify/Trivy integration health gate.
  • Added executable scenario evals against the risk classifier.
  • Added network-backed registry/OSV smoke tests.
  • Added Trivy contract testing and approved-version toolchain tracking.
  • Added unit tests for risk classification and stale-graph integration behavior.

0.2.0 — Bootstrap, dependency guard, evals, and portable packaging

  • Added non-destructive adaptive project bootstrap.
  • Added executable dependency guard with PyPI/npm/crates.io registry checks and OSV verification.
  • Added deterministic behavior-contract eval catalog.
  • Added unit tests for dependency decisions and bootstrap safety.
  • Added portable Agent Plugin packaging for Codex-compatible discovery.
  • Added package drift validation to CI.

0.1.0 — Initial core

  • Added the risk-adaptive Vibe Coding Skill operating model.
  • Added Tier 0–3 workflow classification.
  • Added Graphify project-intelligence policy and automated compatibility proposal workflow.
  • Added Trivy/OSV/package-registry security and dependency guidance.
  • Added persistent project-state and traceability rules.
  • Added doctor, change-budget, Graphify compatibility, and skill-validation utilities.
  • Added GitHub Actions validation.