experiment-design-and-readout
Design and read randomised experiments and A/B tests. Always use this skill when someone asks how many users or how long a test needs to run, the smallest effect a given sample can detect, whether a lift or difference between variants is real or just noise (even when given only the counts for each variant), whether to ship the winning variant, whether an uneven split such as 52 percent to 48 percent matters, or why a result keeps changing while the test runs, even when the question sounds like a quick yes or no. Also use it for choosing the randomisation unit, power, guardrail metrics, peeking and stopping early, several variants or metrics at once, and novelty effects. Not for comparing groups that were not randomised, not for deciding whether X caused Y in observational data, not for offline evaluation of a recommender, and not for explaining what an A/B test is when no test is in play.
Pinned to revision a18d88341e79, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/experiment-design-and-readout/SKILL.md
- skills/experiment-design-and-readout/scripts/ab_test_plan.py
Every link opens the file at its source, pinned to the revision this page describes.