[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"glossary-data-lineage::en":3,"gloss-cluster-data-lineage::en":23,"gloss-next-data-lineage::en":9},{"slug":4,"category":5,"name":6,"definition":7,"meta_desc":8,"faq":9,"schema_markup":9,"related":10},"data-lineage","data-infra","Data Lineage","Data lineage is the map of where data comes from and where it goes: which source produced a value, which transformations touched it, and which downstream tables, dashboards, or models depend on it. It answers two crucial questions — \"where did this number come from?\" (upstream) and \"what breaks if I change this column?\" (downstream).\n\nFor SaaS builders, lineage is what turns debugging a wrong dashboard from a day of Slack archaeology into a five-minute trace. When a KPI looks off, you follow the lineage back through each pipeline step to find where it diverged. Before a schema migration, you check downstream lineage to see who you'll break.\n\nModern tools capture lineage automatically. dbt builds a dependency graph from your SQL models; column-level lineage tools (in DataHub, OpenMetadata, or warehouse-native features) parse queries to trace individual fields. Practical note: column-level lineage is far more useful than table-level for impact analysis, but harder to maintain — invest in it for your most business-critical metrics first, and let automated parsing do the heavy lifting rather than manual diagrams.","Data lineage maps where a value came from and what depends on it — answering \"where did this number come from?\" and \"what breaks if I change this column?\"",null,[11,14,17,20],{"slug":12,"name":13},"data-catalog","Data Catalog",{"slug":15,"name":16},"data-contract","Data Contract",{"slug":18,"name":19},"data-pipeline","Data Pipeline",{"slug":21,"name":22},"reverse-etl","Reverse ETL",[24,28,31,34,37,41,44,47,50,53,57,60],{"slug":25,"category":5,"name":26,"updated_at":27},"acid","ACID","2026-08-24T02:46:37+00:00",{"slug":29,"category":5,"name":30,"updated_at":27},"ann-search","ANN Search",{"slug":32,"category":5,"name":33,"updated_at":27},"backpressure","Backpressure",{"slug":35,"category":5,"name":36,"updated_at":27},"batch-processing","Batch Processing",{"slug":38,"category":5,"name":39,"updated_at":40},"bm25","BM25","2026-08-24T02:46:38+00:00",{"slug":42,"category":5,"name":43,"updated_at":27},"cache","Cache",{"slug":45,"category":5,"name":46,"updated_at":27},"cap-theorem","CAP Theorem",{"slug":48,"category":5,"name":49,"updated_at":27},"change-data-capture","Change Data Capture (CDC)",{"slug":51,"category":5,"name":52,"updated_at":27},"chroma","Chroma",{"slug":54,"category":5,"name":55,"updated_at":56},"chunk-overlap","Chunk Overlap","2026-08-24T03:30:02+00:00",{"slug":58,"category":5,"name":59,"updated_at":27},"columnar-storage","Columnar Storage",{"slug":61,"category":5,"name":62,"updated_at":27},"connection-pooling","Connection Pooling"]