distribution-key-and-partitioning
Choose how a large table is spread across nodes and split into partitions in an MPP warehouse or distributed store, and fix the skew and data movement that make queries slow. Use whenever someone asks which distribution key, DISTKEY, DISTRIBUTED BY, hash distribution, sort key, partition key, shard key or clustering key to use; whether to hash, round-robin or replicate a table; why one node or partition is much bigger or slower than the rest; why joins shuffle or broadcast huge amounts of data; or how to partition a big fact table by date, region or tenant, in Redshift, Synapse, Greenplum, BigQuery, Snowflake, Databricks, Spark, Cassandra or similar. Also use when reviewing warehouse table DDL before it goes live. Not for choosing which kind of database to use, and not for pipeline or data quality design.
Pinned to revision a18d88341e79, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/distribution-key-and-partitioning/SKILL.md
- skills/distribution-key-and-partitioning/scripts/skew_check.py
Every link opens the file at its source, pinned to the revision this page describes.