pipeline-and-quality-design
Design how data moves from source systems to its consumers, and how its quality is proven along the way. Use whenever someone asks whether to use ETL, ELT or ETLT; batch, micro-batch or streaming; copying data or using federation or virtualisation; how to lay out bronze, silver and gold (medallion) layers; what data quality checks, reconciliation, SLAs or freshness alerts a pipeline needs; how to write a data contract between producer and consumer teams; how to capture lineage or set up a data catalog; how to classify data as public, internal, confidential or restricted and handle personal data under GDPR or similar rules; or why a dashboard number disagrees with the source system. Also use for data mesh and data product ownership. Not for choosing a database, not for distribution or partition keys, and not for judging whether a dataset is fit to train a machine learning model.
Pinned to revision a18d88341e79, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/pipeline-and-quality-design/SKILL.md
- skills/pipeline-and-quality-design/scripts/dq_contract_check.py
Every link opens the file at its source, pinned to the revision this page describes.