moheetsubudhi-isb/ml-toolkit
Data-science skills for framing ML problems, auditing data, engineering features, reducing dimensions, clustering, model selection and validation, classification and regression metrics, tree ensembles, linear and logistic models, anomaly detection, and text embeddings.
Find unusual transactions, sensor readings, log events, users or records, and choose the detection method: z-score or IQR rules, robust statistics, Mahalanobis distance, isolation forest, local outlier factor, or seasonal baselines for time series. Use whenever someone wants to detect fraud, suspicious payments, equipment faults, sudden spikes or drops, bot or attack traffic, or unusual behaviour with few or no labelled examples; set the contamination rate or anomaly score cut-off; turn a team's daily review capacity into an alert threshold; explain why a record was flagged; or test a detector against a handful of confirmed cases. Not for cleaning outliers from a training dataset before modelling, not for segmenting customers, and not for setting the threshold of a supervised classifier.
Judge a classifier by what its errors cost, and set the decision threshold. Use when someone reports accuracy, precision, recall, F1, ROC-AUC or PR-AUC and asks what counts as good; when a model looks accurate but misses the cases that matter; when classes are imbalanced, as in fraud, churn, loan default, defects or medical screening; when choosing a probability cutoff; when comparing two classifiers; when reading a confusion matrix; when a review team can handle only so many alerts a day; or when checking whether predicted probabilities are calibrated. Not for choosing the algorithm or designing cross-validation, and not for ranking metrics for recommender systems.
Group customers, products, stores, sessions or any records into segments, and pick the clustering method that fits the data and the business use. Use whenever someone wants to segment, cluster or group similar items, build personas or cohorts from behaviour, run k-means, hierarchical clustering, DBSCAN or Gaussian mixtures, choose the number of clusters, read an elbow plot, silhouette score or Calinski-Harabasz index, profile or name clusters, build an RFM segmentation, cluster mixed categorical and numeric data, or check whether segments someone else built are real. Not for reducing the number of columns before modelling, and not for predicting a label that already exists.
Reduce many features to fewer, and choose between PCA, Fisher linear discriminant analysis, feature selection and 2-D visualisation methods. Use when a dataset has dozens or hundreds of correlated columns; when someone asks whether to use PCA, how many principal components to keep, or what the loadings mean; how to visualise high-dimensional data or embeddings; whether to use t-SNE or UMAP; how to remove noise or multicollinearity; when to use LDA to separate classes; or which features to keep or drop. Not for grouping records into segments, and not for creating new features from raw data.
Design, transform and debug model features so the model learns the signal that matters. Use when someone asks what features to build; how to encode categorical variables or high-cardinality IDs; how to handle dates, times, text, locations or event histories; how to create ratios, lags, rolling windows or per-customer aggregates without leaking the future; how to log-transform or scale skewed columns, impute missing values or pick defaults; whether target encoding is safe; or when a model underperforms, looks strangely complex, or misbehaves on certain inputs and a feature may be the cause. Not for auditing whether a dataset is usable at all, and not for reducing many existing columns to fewer.
Build, regularise and interpret linear regression, logistic regression and multinomial (softmax) models used for prediction. Always use this skill when someone asks what a regression coefficient or an odds ratio means, or why a coefficient changed size or flipped sign when another feature was added, even when the question sounds like a quick one. Also use it for standardising features before comparing coefficients; dummy variables and the dropped category; multicollinearity or VIF; Ridge vs Lasso vs Elastic Net and choosing alpha; polynomial or interaction terms; logistic regression vs a perceptron vs linear discriminant analysis; gradient descent that will not converge; or whether a linear model is enough. Not for deciding whether X causes Y or whether an experiment's result is significant, not for choosing a regression error metric, and not for explaining what these terms mean when no model or data is in play.
Check whether a dataset is fit for machine learning, and catch the problems that silently ruin models. Always use this skill when someone asks which columns leak the target, whether a feature will actually be available at prediction time, why accuracy looks suspiciously high, or whether training and serving data differ, even when the question sounds like a quick check and even before the data is shared. Also use it when someone shares or describes a dataset for modelling and asks whether it is usable, what is wrong with it or what to clean; or mentions missing values, duplicates, sentinel codes like 999 or -1, outliers, skewed or mixed-scale columns, class imbalance, label quality, how the target was defined, label delay, or post-outcome columns. Not for designing new features, and not for choosing or tuning a model.
Decide whether, and how, machine learning should solve a business problem before any data work starts. Use when someone asks whether ML or AI can help with a problem; wants to predict, classify, forecast, rank, recommend, detect anomalies or group things but hasn't pinned down how; needs to choose between rules and a model; must define the target or label, or find a proxy label when true outcomes are missing; wants to pick the success metric and tie it to business value; or needs to judge whether an ML project is feasible and worth funding. Not for a dataset that is already chosen and ready for cleaning, feature building, model tuning or threshold setting.
Choose a model family, design validation that proves the model will generalise, and diagnose overfitting or underfitting. Use when someone asks which algorithm to use; how to split training and test data; about cross-validation, or stratified, grouped or time-based splits; hyperparameter tuning or grid search; regularisation, tree depth or pruning; k for k-nearest neighbours; bias and variance; learning curves; why test performance is much worse than training; whether more data would help; or why a model that validated well failed after launch. Not for choosing a decision threshold or business metric for a classifier, not for cleaning or auditing the dataset, and not for explaining what a term such as cross-validation or overfitting means when no dataset or model is in play.
Judge whether a regression or forecast model's errors are good enough, and choose the error metric the business should see: MAE, RMSE, MAPE, WAPE, R-squared, bias, or a cost-weighted error. Use whenever someone asks whether an RMSE, MAE or MAPE value is good; why R-squared is high but predictions are still useless, or negative on test data; which metric to report for price, demand, delivery-time, sales or salary predictions; why MAPE explodes when actual values are small; how to compare a model with a naive, last-period or average baseline; how errors differ across segments or ranges; whether over- and under-prediction cost the same; or how to read a residual plot for a predictive model. Not for classification metrics or thresholds, not for ranking or recommender metrics, and not for testing whether a coefficient is statistically significant.
Turn text into numbers for search, matching, deduplication, classification or clustering, and choose between TF-IDF, embeddings and simpler methods. Use whenever someone asks how to find similar tickets, products, documents, CVs or reviews by their text; detect near-duplicate records from names, addresses or descriptions; build TF-IDF vectors, n-grams or stop-word lists; choose word, sentence or document embeddings (word2vec, GloVe, fastText, sentence transformers or a hosted embedding API); decide whether embeddings are worth the cost over TF-IDF; handle tokenisation, subwords, spelling variants or several languages; measure cosine similarity; or store and search vectors for semantic search. Not for choosing how to cluster the resulting vectors into segments, and not for encoding ordinary categorical columns.
Build, tune and explain decision trees and tree ensembles: random forest, bagging, AdaBoost, gradient boosting, XGBoost, LightGBM and CatBoost. Use whenever someone asks whether to use a single tree, a random forest or boosting; how to set max_depth, min_samples_leaf, n_estimators, learning rate, max_features or subsample for one of these models; gini or entropy; how to prune a tree; why a tree overfits or a boosted model is unstable; what the out-of-bag score means; how far to trust feature importance, or when to use permutation importance or SHAP instead; how to turn a tree into business rules; or whether a regression tree beats linear regression. Not for designing the validation split or reading bias-variance in general, and not for choosing a classification threshold.