Skip to content
v3.11.0MIT

A Codex engineering toolkit with orchestration, task ledgers, solve loops, policy-gated effects, async GitHub Stack Merge, catalogs, and GitHub-native stacked delivery.

api-design

Use when designing or reviewing an API — REST, GraphQL, gRPC, or a library interface. Covers consistency, resource modeling, versioning and evolution, error contracts, pagination, and idempotency so the interface stays usable and changeable after it ships.

caching-strategies

Use when designing or debugging a cache — choosing a caching pattern, setting TTLs and invalidation, picking a layer (client/CDN/app/DB), or fixing staleness, a low hit rate, or a stampede. Complements performance-profiling (which finds the bottleneck); this decides whether and how to cache it.

code-review-rubric

Use when reviewing code or a pull request to apply a consistent quality bar. Provides a severity-ranked rubric covering correctness, security, tests, readability, and maintainability, plus how to write feedback that gets fixed. See CHECKLIST.md for the full pre-merge checklist.

concurrency-and-parallelism

Use when writing or reviewing concurrent code — threads, async/await, shared state, locks, and parallel work — or diagnosing a race condition, deadlock, or heisenbug that only appears under load. Covers the patterns that make concurrency correct and the traps that make it intermittently wrong.

conventional-commits

Use when writing git commit messages and structuring commits. Covers the Conventional Commits format, choosing the right type/scope, breaking-change notation, and how to split work into small, atomic, reviewable commits.

doctor

Use before orchestration, stack delivery, or a release when the host and repository need a read-only capability and merge-readiness preflight. Run the Forge Doctor CLI, interpret pass/warn/fail/unknown evidence, and never mutate policy or install tools.

error-handling

Use when designing how code handles failure — where to catch, what to propagate, fail-fast vs recover, retries and idempotency, and error types. Covers building robust failure paths without swallowing bugs or over-engineering for impossible cases.

feature-flags

Use when adding, rolling out, or cleaning up feature flags — gating a change, doing a progressive/canary release, building a kill switch, or removing a stale flag. Covers flag types, safe rollout, and the discipline that keeps flags from becoming permanent tech debt.

forge-catalog

Use when choosing which Forge agent, skill, command, bundle, or workflow should handle a request; when comparing Forge capabilities; when building a focused install surface; or when a user asks "what should I use?" before starting work. Routes to the smallest useful Forge capability instead of loading everything.

git-workflow

Use when working with git beyond writing a commit message — branching, rebasing vs merging, resolving conflicts, recovering lost work, and bisecting to find a bad commit. Covers what is safe on shared branches and what is not.

iterate-to-done

Use when running a solve-loop over a task ledger — repeatedly pick the next ready task, dispatch it, verify against acceptance criteria, update the ledger, and repeat until done or blocked. Covers the loop, its stop conditions, and how to drive it recurring or in the background with the harness (/loop, the scheduler, background agents).

observability

Use when adding logging, metrics, or tracing to code, or designing how a service is monitored. Covers the three pillars, structured logging, what to measure (RED/ USE), useful alerts, and the SLO mindset — so failures are debuggable after the fact.

open-source-contribution

Use when preparing or reviewing an upstream open-source contribution: reproduce an issue, follow maintainer rules, run exact-head checks, prepare a focused PR and reconcile review feedback. Covers fork identity, DCO/CLA gates and local evidence.

orchestration

Use when driving a large, multi-part task end to end with multiple models — planning at a high tier, decomposing into a task ledger, and delegating each piece to the right specialist at the right model (plan with Opus/Fable, implement with Sonnet, mechanical work with Haiku). Covers the conductor loop and how to delegate. See MODEL-ROUTING.md for the tier policy.

performance-profiling

Use when investigating a performance problem — slow endpoints, high latency, memory/CPU pressure, or N+1 queries. A measure-first method to find the real bottleneck and verify the gain, instead of guessing at optimizations.

policy

Use when an orchestrated run may create an external effect or mutate a protected resource. Defines versioned action envelopes, declarative profiles, scoped one-use approvals, staged previews, pre-effect re-evaluation, and privacy-safe decision receipts for Forge mutation adapters.

prompt-engineering

Use when authoring or improving Claude Code agents, skills, slash commands, or CLAUDE.md instructions. Covers how to write a triggering description, structure a system prompt, scope tools, and apply progressive disclosure so the model actually uses what you build. See PATTERNS.md for operating-discipline techniques distilled from production agent prompts.

pull-request-authoring

Use when opening a pull request — writing the description, sizing the change, and making it easy and fast to review. Covers PR structure, what reviewers need, and how to keep PRs small enough to actually get good review.

refactoring-catalog

Use when improving the structure of existing code without changing behavior — identifying code smells and applying the right named refactoring safely. Covers the discipline of behavior-preserving change; see CATALOG.md for the smell→fix reference.

root-cause-debugging

Use when diagnosing a bug, failing test, crash, or wrong output. A disciplined, hypothesis-driven method for finding the true root cause from evidence instead of guessing or masking symptoms. Use before writing any fix.

safe-database-migrations

Use when writing or reviewing a database schema migration, especially on a live system. Covers the expand/contract pattern for zero-downtime changes, avoiding long locks, safe backfills, and always having a rollback path.

stacked-changes

Use when a feature is too large for one reviewable pull request; when creating, restacking, reviewing, repairing, or landing dependent GitHub pull requests; or when choosing between vanilla git/gh, Graphite, and Sapling for a stacked-change workflow.

task-ledger

Use when turning a plan into trackable tasks and working them to done — a lightweight issue tracker for an orchestrated run. Covers the task format (acceptance criteria, status, dependencies, assigned agent + model) and three backends: local markdown files (default), GitHub issues via gh, or Jira/Linear via MCP. See TEMPLATE.md for the issue format.

technical-writing

Use when writing developer-facing documentation — READMEs, API references, architecture docs, ADRs, runbooks, and docstrings. Covers structure by document type, writing for the reader's task, and keeping docs accurate against the code.

test-driven-development

Use when implementing a feature or fixing a bug test-first: write a failing test, make it pass with the minimum code, then refactor. Covers the red-green-refactor loop, what is worth testing, and the common pitfalls that make TDD backfire.

threat-modeling

Use when designing or reviewing a feature for security before it ships — to systematically identify what can go wrong and where the defenses must be. Covers the STRIDE method, trust boundaries, and turning threats into concrete controls. Defensive use only.