A data catalogue is a centralised, searchable inventory of an organisation’s data assets enriched with metadata, descriptions, lineage and usage information. It helps users discover, understand, trust and govern data by linking technical metadata with business context and policies. Modern catalogues often automate metadata harvesting and use collaboration features to capture institutional knowledge.

Overview

  • A data catalogue acts as the connective tissue of a data platform, providing a single place to find and understand data assets across databases, warehouses, lakes and pipelines. It harvests technical metadata automatically, augments it with business glossaries, ownership and quality indicators, and exposes search, lineage and access-request workflows. By making data discoverable and trustworthy it accelerates analytics while supporting governance and compliance.

Key aspects

  • Automated harvesting of technical metadata from sources
  • Business glossary linking data to organisational terminology
  • Lineage tracking from source through transformation to consumption
  • Search, tagging and collaboration to capture tribal knowledge
  • Integration with governance policies and access controls

Applications

  • Self-service data discovery for analysts and scientists
  • Data governance and stewardship programmes
  • Regulatory compliance and audit traceability
  • Impact analysis through end-to-end lineage
  • Onboarding and knowledge retention across teams

Provenance