Skip to content

moheetsubudhi-isb/ml-toolkit

v1.0.0MIT

Data-science skills for framing ML problems, auditing data, engineering features, reducing dimensions, clustering, model selection and validation, classification and regression metrics, tree ensembles, linear and logistic models, anomaly detection, and text embeddings.

dimensionality-reduction

Reduce many features to fewer, and choose between PCA, Fisher linear discriminant analysis, feature selection and 2-D visualisation methods. Use when a dataset has dozens or hundreds of correlated columns; when someone asks whether to use PCA, how many principal components to keep, or what the loadings mean; how to visualise high-dimensional data or embeddings; whether to use t-SNE or UMAP; how to remove noise or multicollinearity; when to use LDA to separate classes; or which features to keep or drop. Not for grouping records into segments, and not for creating new features from raw data.

Read SKILL.md at the source

Pinned to revision a18d88341e79, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.