Skip to content

moheetsubudhi-isb/ml-toolkit

v1.0.0MIT

Data-science skills for framing ML problems, auditing data, engineering features, reducing dimensions, clustering, model selection and validation, classification and regression metrics, tree ensembles, linear and logistic models, anomaly detection, and text embeddings.

model-selection-and-validation

Choose a model family, design validation that proves the model will generalise, and diagnose overfitting or underfitting. Use when someone asks which algorithm to use; how to split training and test data; about cross-validation, or stratified, grouped or time-based splits; hyperparameter tuning or grid search; regularisation, tree depth or pruning; k for k-nearest neighbours; bias and variance; learning curves; why test performance is much worse than training; whether more data would help; or why a model that validated well failed after launch. Not for choosing a decision threshold or business metric for a classifier, not for cleaning or auditing the dataset, and not for explaining what a term such as cross-validation or overfitting means when no dataset or model is in play.

Read SKILL.md at the source

Pinned to revision a18d88341e79, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.