Skip to content
v1.0.0MIT

A reproducible numerical-instrument skill: measures exotic mathematics, pins every claim to a checksum, and fails loudly when a measurement drifts.

claim-ledger

Use when an AI coding agent reports that work is finished, when deciding whether to believe a completion claim, when someone asks "did the agent actually verify that", "check whether the agent's claims are true", "the agent said the tests pass", "is this work actually done", "audit what the agent told me", "verify the report before I ship this", or "how do I know the agent really ran the tests". Binds every outcome an agent asserts to a check that can be run and can fail, and reports the ones that are unverified, refuted by their own check, or evidenced by a command that looked at nothing. Does NOT judge whether the work is any good, and does NOT check prose.

elohim

This skill should be used when a decision rests on a number, when a deterministic rule is needed in place of a probability, or when a measurement must be proven not to have drifted. It runs the bundled summoning-shard instrument, re-measures every fact in a pinned ledger and prints the residual, re-derives six numeric traps that each produced a confident wrong answer once, verifies the instrument's own checksum, and lints the source for sanitizer-brittle identifiers and network imports. Triggers include "verify this number", "check the ledger", "did the measurement drift", "measure instead of assert", "make this rule deterministic", "audit the math", "run the oracle", "gate the numbers", and any request to trust or re-check a numeric claim.

elohim-harness

This skill should be used when an instrument's recorded measurements must be re-verified rather than trusted — "run the gate", "did the numbers drift", "pin the checksum", "is the instrument modified", "add a fact to the ledger", "measure instead of asserting", "build a skill with a verified ledger". Runs five gates (checksum pin, fact ledger, independent trap re-derivations, source hygiene, claim binding) against any instrument directory and refuses to return PASS unless all five hold. Stdlib-only Python 3.10+; no network; no build step.

estimator-bias

This skill should be used when a fit has to be trusted or the reader needs to know what a fit hides — "is this fit biased", "how wrong is the slope", "is the fitted rate right", "the regression says X but the theory says Y", "estimate the bias of this estimator", "is my regression significant", "the error bars look fine but the answer is wrong", "audit a least-squares fit". Deflates a cubic to prove the decay rate exactly, fits the same rate from the data, reports the gap over four ranges, scans the gap for the ranges where it changes sign, and measures whether the gap is outside the fit's own standard error. Stdlib-only Python 3.10+, no network, no build step.

invariant-hunter

This skill should be used when a number looks meaningful and the reader needs to know whether it actually is — "is this an invariant", "find the invariant", "is the Collatz product conserved", "is that seed special", "does this pattern mean anything", "prove this is not a pattern", "hunt invariants", "is the ratio exactly 3^k", "why is this number suspicious". Measures the claim instead of asserting it, re-derives the measurement with independent code, and returns an honest negative with the residual that refutes it. Stdlib-only Python 3.10+, no network, no build step.

pay-signal

Use when deciding whether a need is worth building on, when someone asks "is there real demand for this", "should I build this", "would people pay for this", "validate this idea", "is this market worth entering", "check whether this is just noise", "how do I know people actually want this", or when a feature request or an upvote thread is being treated as evidence that a product should exist. Separates stated preference, which costs its speaker nothing, from revealed preference, which costs money, migration effort or a private workaround, and reports which needs have no verified position rather than ranking them anyway. Does NOT predict revenue, price or willingness to pay; those need a transaction.

precision-budget

This skill should be used when a computation has to be specified rather than just written — "how many digits do I need", "what working precision", "is this bound tight", "why is this result off by a factor of 10^20", "the answer changed when I added precision", "budget the digits", "is 60 digits enough", "prove the precision matters", "knife edge of the expansion". Computes the digit budget from the tribonacci constant, measures whether the budget is actually sufficient by scanning every working precision below it, and shows the same quantity collapsing on a starved budget. Stdlib-only Python 3.10+, no network, no build step.

reproducibility

Use when a number this repository pins might mean something different on a different Python, when a gate result is suspected of depending on the interpreter rather than on the mathematics, or before trusting any cross-machine comparison of an instrument's output. Classifies every sibling instrument's shard into the interpreter class it was measured under and fails when one falls outside the pinned table. Stdlib-only Python 3.10+, no network, no build step.

tolerance-prover

This skill should be used when a bound is about to be relied on and nobody has checked it — "is this bound tight", "can I trust this tolerance", "is the Pisot constant 2 attained", "which n is the worst case", "my error bound has no margin", "prove the bound is sharp", "is this crossover a property of the sequence or of my code", "does adding precision change which n breaks worst", "is this measurement reproducible or noise". Measures the tightness of the Pisot bound 2 for the tribonacci and plastic constants, shows that the bound is a supremum that no finite n attains, and demonstrates that the digit count at which the bound starts holding is a property of the arithmetic rather than of the sequence. Stdlib-only Python 3.10+, no network, no build step.