Skip to content

moheetsubudhi-isb/statistics-toolkit

v1.0.0MIT

Statistics skills for the moments that decide things at work: designing and reading A/B tests, comparing groups, estimating with a margin of error, setting control limits, checking causal claims, trusting a regression, giving honest prediction ranges, and modelling counts and rates.

causal-claim-check

Pressure-test a claim that one thing caused another before anyone acts on it. Always use this skill when someone asks whether X causes or drives Y, whether a correlation is causal or just confounding, whether a change, campaign, feature or policy really worked, what to control for when estimating an effect, whether an observational result can be trusted as a causal effect, whether an estimate was biased because the effect shrank or grew after adding a control, when an experiment is needed instead of regression with controls, or whether the groups being compared were alike to begin with, even when the question sounds like a simple yes or no. Also use it for selection bias, confounders and omitted variable bias. Not for designing or reading a randomised A/B test, not for reading the coefficients of a prediction model, and not for checking regression assumptions.

control-limits-and-error-costs

Set the threshold for a monitored measurement or quality check, balancing false alarms against missed problems. Always use this skill when someone asks where to put an alert limit, control limit, pass or fail cut-off or significance level for a process or metric; how costly a false alarm is compared with a miss; the chance a batch, pack, shipment or day breaches a specification or legal limit, even as a quick calculation from an average and spread; what target average keeps output inside a limit; or how quickly a shift in the process would be detected. Also use it for control charts, acceptance sampling, type I and type II errors, power to detect a shift, and choosing alpha from business costs. Not for setting a classifier's probability cut-off, not for deciding alert volumes from a review team's capacity, and not for sizing an A/B test.

count-and-rate-models

Model outcomes that are counts or rates, such as orders per day, defects per batch, claims per policy, visits per store or units sold per product and month. Always use this skill when someone asks how to model or forecast a count, why ordinary regression gives negative or fractional counts, how to handle different exposure (days on shelf, population, customer-months) with an offset, what a rate ratio or the exponential of a coefficient means, or why a count model's standard errors look too small (overdispersion), even when the question sounds like a quick regression. Also use it for Poisson, quasi-Poisson and negative binomial models, excess zeros with hurdle or zero-inflated models, repeated units, and comparing count models with deviance and AIC. Not for yes or no outcomes, not for judging forecast error size, and not for choosing between linear and logistic models.

estimate-with-margin-of-error

Estimate a true average or rate from a sample and say how precise it is, or work out how big a sample is needed. Always use this skill when someone asks for a confidence interval or margin of error, what a survey, audit, inspection or quality sample says about the whole population, how many records, customers or units to sample for a target precision, whether a sample is big enough to trust, whether to use a t or z interval for a small sample, or how to explain a confidence interval to a manager, even when the question sounds like a quick check. Also use it for intervals for proportions with small counts or rates near zero or one, finite populations, and sampling pitfalls such as convenience samples and non-response. Not for comparing two groups, not for sizing an A/B test, and not for explaining what a confidence interval means when no data is in play.

experiment-design-and-readout

Design and read randomised experiments and A/B tests. Always use this skill when someone asks how many users or how long a test needs to run, the smallest effect a given sample can detect, whether a lift or difference between variants is real or just noise (even when given only the counts for each variant), whether to ship the winning variant, whether an uneven split such as 52 percent to 48 percent matters, or why a result keeps changing while the test runs, even when the question sounds like a quick yes or no. Also use it for choosing the randomisation unit, power, guardrail metrics, peeking and stopping early, several variants or metrics at once, and novelty effects. Not for comparing groups that were not randomised, not for deciding whether X caused Y in observational data, not for offline evaluation of a recommender, and not for explaining what an A/B test is when no test is in play.

group-difference-test

Check whether two groups, or two points in time, differ by more than chance in data you already have. Always use this skill when someone asks whether region A really differs from region B, whether one customer segment (such as app users and web users) spends more than another, whether an average or rate changed before versus after a change, whether a new process, supplier, branch or cohort performs differently, whether the spread differs between groups, whether a statistically significant difference is big enough to matter, or which test to use, even when data is attached and the question sounds like a quick calculation. Also use it for paired versus independent samples, confidence intervals for a difference, small or skewed samples, and comparing three or more groups. Not for designing or reading a randomised A/B test, not for deciding whether the difference was caused by the group, and not for estimating a single average or rate.

prediction-interval-reporting

Give an honest range around a single prediction or forecast. Always use this skill when someone asks for an interval, band or range around one predicted value, such as the likely price of this house, this customer's spend or next month's demand for one product; how much error to expect on an individual forecast; the difference between a confidence interval and a prediction interval; or whether a stated 95 percent range really covers 95 percent of outcomes, even when a point estimate seems enough. Also use it for intervals from regression, residual-quantile and conformal methods, coverage checks on held-out data, and why ranges widen away from typical inputs. Not for judging average error metrics such as MAE or RMSE, not for simulating a plan with several uncertain inputs, and not for estimating an average or rate from a sample.

regression-diagnostics

Read a regression output table and check whether it can be trusted for inference. Always use this skill when someone shares regression output and asks how to read the estimates, standard errors, t values, p-values, R-squared, adjusted R-squared or F test; whether a coefficient is significant and can be trusted; whether a model with a low R-squared but a significant F test is any good; whether to use robust or clustered standard errors, for example because the same customers or stores appear in many rows; what a residual or Q-Q plot shows; or whether outliers or influential points (leverage, Cook's distance) drive the fit, even when the question sounds routine. Also use it for fan-shaped residuals, unequal variance and the Breusch-Pagan test, heavy tails and curved residual patterns. Not for judging how large prediction errors are, not for deciding whether X causes Y, and not for multicollinearity, VIF or regularisation.