Skip to content

moheetsubudhi-isb/ml-toolkit

v1.0.0MIT

Data-science skills for framing ML problems, auditing data, engineering features, reducing dimensions, clustering, model selection and validation, classification and regression metrics, tree ensembles, linear and logistic models, anomaly detection, and text embeddings.

tree-and-ensemble-models

Build, tune and explain decision trees and tree ensembles: random forest, bagging, AdaBoost, gradient boosting, XGBoost, LightGBM and CatBoost. Use whenever someone asks whether to use a single tree, a random forest or boosting; how to set max_depth, min_samples_leaf, n_estimators, learning rate, max_features or subsample for one of these models; gini or entropy; how to prune a tree; why a tree overfits or a boosted model is unstable; what the out-of-bag score means; how far to trust feature importance, or when to use permutation importance or SHAP instead; how to turn a tree into business rules; or whether a regression tree beats linear regression. Not for designing the validation split or reading bias-variance in general, and not for choosing a classification threshold.

Read SKILL.md at the source

Pinned to revision a18d88341e79, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.