[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"glossary-data-catalog::en":3,"gloss-cluster-data-catalog::en":23,"gloss-next-data-catalog::en":9},{"slug":4,"category":5,"name":6,"definition":7,"meta_desc":8,"faq":9,"schema_markup":9,"related":10},"data-catalog","data-infra","Data Catalog","A data catalog is a searchable inventory of the data assets across your company — tables, columns, dashboards, pipelines — enriched with metadata like descriptions, owners, freshness, and how each field is defined. Think of it as a library catalog for your data: instead of asking a colleague \"which table has the real revenue number?\", people search the catalog and get an authoritative answer.\n\nFor SaaS teams, a catalog becomes valuable once you have more than a handful of tables and more than one person querying them. It combats \"tribal knowledge\", reduces duplicated metrics, and helps new hires self-serve. Popular options include open-source DataHub, OpenMetadata, and Amundsen, plus commercial tools baked into warehouse platforms.\n\nPractical note: a catalog is only as good as the metadata in it. The winning approach is automated harvesting — pulling schemas, lineage, and usage stats directly from your warehouse and BI tools — supplemented by lightweight human descriptions and ownership tags. Don't try to document everything by hand; prioritize the tables people actually query.","A data catalog is a searchable inventory of your tables, columns, dashboards, and pipelines, enriched with owners, definitions, and freshness metadata.",null,[11,14,17,20],{"slug":12,"name":13},"data-contract","Data Contract",{"slug":15,"name":16},"data-lineage","Data Lineage",{"slug":18,"name":19},"data-warehouse","Data Warehouse",{"slug":21,"name":22},"schema-registry","Schema Registry",[24,28,31,34,37,41,44,47,50,53,57,60],{"slug":25,"category":5,"name":26,"updated_at":27},"acid","ACID","2026-08-24T02:46:37+00:00",{"slug":29,"category":5,"name":30,"updated_at":27},"ann-search","ANN Search",{"slug":32,"category":5,"name":33,"updated_at":27},"backpressure","Backpressure",{"slug":35,"category":5,"name":36,"updated_at":27},"batch-processing","Batch Processing",{"slug":38,"category":5,"name":39,"updated_at":40},"bm25","BM25","2026-08-24T02:46:38+00:00",{"slug":42,"category":5,"name":43,"updated_at":27},"cache","Cache",{"slug":45,"category":5,"name":46,"updated_at":27},"cap-theorem","CAP Theorem",{"slug":48,"category":5,"name":49,"updated_at":27},"change-data-capture","Change Data Capture (CDC)",{"slug":51,"category":5,"name":52,"updated_at":27},"chroma","Chroma",{"slug":54,"category":5,"name":55,"updated_at":56},"chunk-overlap","Chunk Overlap","2026-08-24T03:30:02+00:00",{"slug":58,"category":5,"name":59,"updated_at":27},"columnar-storage","Columnar Storage",{"slug":61,"category":5,"name":62,"updated_at":27},"connection-pooling","Connection Pooling"]