Active Metadata and the Rise of the Modern Data Catalog
The first generation of data catalogs answered “what data do we have?” Almost nobody asked. The current generation answers “can I use this, and what happens if I do?” — and that question gets asked constantly.
Contents
Why the First Generation Underperformed
Between roughly 2016 and 2021, a great many enterprises bought a data catalog, crawled their estate, populated tens of thousands of assets and watched adoption flatline within two quarters.
The post-mortems usually blame change management. That is partly right and mostly a deflection. The structural problem was that a passive catalog asked users to leave the tool they were working in, go somewhere else, search for something, read a description written by someone who no longer worked there, and then return. The value had to exceed that friction. Usually it did not.
Three specific failure modes recur:
- Documentation decay. Descriptions written during onboarding were accurate for six months. Nothing forced their update, so the catalog silently drifted from reality and users learned not to trust it.
- Coverage as the goal. Programmes optimised for assets catalogued, which is a metric a crawler can satisfy without a human ever benefiting.
- No consequence. A classification in the catalog had no effect on what the warehouse permitted. Governance described policy; it did not apply it.
What Active Metadata Actually Means
“Active metadata” is a term that has been used loosely enough to lose meaning. The useful definition is narrow: metadata is active when other systems consult it at runtime and behave differently as a result.
That single property produces most of the difference between generations.
| Passive catalog | Active metadata platform | |
|---|---|---|
| Population | Crawl on a schedule; humans write descriptions | Continuous ingestion, including query logs and usage telemetry |
| Accuracy over time | Degrades; no forcing function | Self-correcting; usage reveals what is real |
| Where users meet it | A separate web application | Embedded in BI, notebooks, IDEs, Slack, ticketing |
| Classification | A label recorded | A label that propagates into masking and access policy |
| Quality | A score on a dashboard | A gate that can pause a downstream pipeline |
| Access | Documented owner; request by email | Request routed and provisioned through workflow |
The test
If switching your catalog off for a week would inconvenience people but break nothing, it is passive. If it would block provisioning, halt pipelines and stop policy enforcement, it is active.
Usage Signals Beat Curation Effort
The most consequential shift is that modern platforms treat usage as a primary metadata source rather than an afterthought.
Query log ingestion — the approach Alation built its reputation on and which most platforms now offer in some form — reveals things curation never will:
- Which tables are actually queried, versus which are catalogued and untouched
- Who the real subject matter expert is, based on who queries an asset most and whose queries others copy
- Which joins are conventional, which reveals implicit relationships nobody documented
- Which assets are deprecated in practice long before anyone marked them so
This inverts the curation problem. Instead of asking a team to document 4,000 assets, you identify the 120 that carry 90% of the query volume and curate those properly. The rest can wait, and much of it should never be curated at all because nothing depends on it.
In practice this is the fastest route we know to a catalog people trust. Trust comes from the top results being right, not from coverage being complete.
Where the Platforms Actually Differ
Vendor comparisons tend to become feature-matrix exercises that obscure the real differences. From delivery experience, the meaningful distinctions are these.
Collibra
The strongest fit where governance process is the primary requirement — regulated industries with formal stewardship, defined approval chains and audit obligations. Its BPMN workflow engine is genuinely powerful, which is both the reason to choose it and the reason implementations overrun: it is possible to model a process far more elaborate than the organisation will actually follow. Design the workflow for the behaviour you can sustain, not the one your policy describes.
Microsoft Purview
Compelling where the estate is predominantly Azure and where the governance requirement is entangled with information protection and compliance. Sensitivity labels propagating from Purview through Microsoft 365 and into the data estate is a capability that is hard to replicate elsewhere. Less strong as a business glossary and stewardship platform; often deployed alongside rather than instead of a dedicated catalog.
Informatica CDGC
Natural where Informatica is already the integration backbone, because scanner coverage and lineage fidelity into existing pipelines are strong out of the box. Cloud-native posture has improved substantially. The curation model rewards deliberate design early.
Atlan
Built around the embedded-experience thesis — governance surfaced in Slack, Jira and the BI tool rather than in a destination application. Its personas-and-purposes access model is a genuinely different way of thinking about catalog permissions. Strongest where the organisation is already collaborative and tooling-modern; less differentiated where the requirement is formal regulatory stewardship.
Alation
Deepest heritage in behavioural analysis and query log intelligence, with trust flags and endorsements that make the human signal explicit. Strong where analyst self-service is the primary use case. Policy centre capability has matured considerably.
The unhelpful truth
All five are capable of supporting a successful programme, and all five have been bought by organisations whose programmes then failed. Platform choice is not the variable that determines the outcome; operating model and adoption are. Choose for fit with your estate and your stewardship reality, then stop deliberating.
Why AI Made This Urgent
Active metadata was a good idea for a decade. It became urgent in the last two years for a specific reason: AI systems consume data at machine speed and without human judgement in the loop.
A human analyst who stumbles on a data set with a misleading name will often notice something is wrong. A feature pipeline will not. A retrieval system embedding a document repository will not notice that a third of the documents are superseded drafts. The judgement that used to sit between the data and its use has been removed, and metadata is what has to replace it.
This has three practical consequences:
- Classification must be enforced, not recorded. A data set marked as restricted from AI training needs the feature store to refuse it, because no human will check.
- Lineage must extend past the model boundary. Data to feature to model version to decision. Catalogs that stop at the warehouse leave the most consequential hop undocumented.
- Agents need catalog entries too. As autonomous agents proliferate, the catalog becomes the natural home for the agent registry — owner, purpose, permitted tools and data, authority tier, approval gates, audit trail. Several platforms are moving in this direction; the organisations ahead of the vendors are modelling agents as first-class catalog assets today.
Getting Value Without a Two-Year Programme
A sequence that produces something usable inside a quarter:
- Ingest usage before curating anything. Query logs first. Let the data tell you what matters rather than asking people to guess.
- Curate the top decile properly. Full definitions, owners, quality metrics, lineage. Ignore the long tail entirely for now.
- Embed one surface. Push catalog context into whichever tool your analysts actually live in. One well-integrated surface beats five partial ones.
- Activate one control. Pick a single enforcement path — classification driving masking, or quality gating a pipeline — and make it real. This is the moment the catalog stops being a library.
- Then expand. Coverage becomes worth pursuing once each new asset inherits a working control plane.
A catalog that describes your data is a reference work. A catalog that governs your data is infrastructure.
Published by KRISID · 26 March 2026. This paper reflects our delivery experience and publicly available sources at the time of writing. It is general guidance, not legal advice — regulatory obligations vary by jurisdiction and by how a system is used.
Choosing or Rescuing a Catalog?
We implement and recover catalog programmes across Collibra, Purview, Informatica CDGC, Atlan and Alation — and we will tell you honestly if the platform is not your problem.
Email contact@krisid.com