Skip to content

moheetsubudhi-isb/ml-toolkit

v1.0.0MIT

Data-science skills for framing ML problems, auditing data, engineering features, reducing dimensions, clustering, model selection and validation, classification and regression metrics, tree ensembles, linear and logistic models, anomaly detection, and text embeddings.

clustering-and-segmentation

Group customers, products, stores, sessions or any records into segments, and pick the clustering method that fits the data and the business use. Use whenever someone wants to segment, cluster or group similar items, build personas or cohorts from behaviour, run k-means, hierarchical clustering, DBSCAN or Gaussian mixtures, choose the number of clusters, read an elbow plot, silhouette score or Calinski-Harabasz index, profile or name clusters, build an RFM segmentation, cluster mixed categorical and numeric data, or check whether segments someone else built are real. Not for reducing the number of columns before modelling, and not for predicting a label that already exists.

Read SKILL.md at the source

Pinned to revision a18d88341e79, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.