Amundsen vs Collibra.
Amundsen and Collibra both anchor in catalog & discovery — 7 dimensions differ, 1 hold. Below: posture, coverage diff, and capability matrix.
The Lyft-born OSS catalog that invented search-first discovery — historically important, but development has largely stalled since 2024.
Enterprise data-and-AI governance incumbent: catalog, glossary, workflow stewardship, lineage, and a separate ML data-quality module.
Teams that already run Amundsen and need to understand what they have, or teams with a genuinely minimal requirement — usage-ranked table search and ownership tracking, nothing more — who are comfortable owning a codebase that is no longer moving.
Large, regulated enterprises — banks, insurers, pharma, public sector — that need a governance-first control plane: a real CDO function, formal stewardship, a business glossary, policy enforcement, and auditable lineage for regulations like BCBS 239, GDPR, SOX, HIPAA, and the EU AI Act.
What each is betting on.
Honest read: maintenance mode. Created at Lyft, open-sourced October 2019, joined LF AI & Data as an incubation project in August 2020. Development has slowed to near-zero — the last release was databuilder 7.5.1 (August 2024) and the last commit to the monorepo (April 2025) moved a maintainer to emeritus status. Stemma, the managed-Amundsen company founded by an Amundsen co-creator, was acquired by Teradata in 2023 and discontinued as a standalone product. The repo is not archived, but treat this as software that is no longer evolving.
Independent and active as of mid-2026. Founded 2008 in Brussels by VUB researchers; one of the original category-defining governance incumbents. Itself an acquirer, not a target — Raito (access management), Husprey (SQL notebook), and Deasy Labs (unstructured/AI metadata) in 2025, on top of OwlDQ (2021, now the Data Quality & Observability module). Last disclosed private valuation USD 5.25B (2021).
Each tool's current strategic narrative, verbatim from its profile.
Spec sheet diff.
| Amundsen | Collibra | |
|---|---|---|
| Vendor | LF AI & Data Foundation | Collibra |
| Deployment | Self-hosted only | SaaS only |
| License | Open source | Proprietary |
| Pricing | OSS · paid tiers | Contact sales |
| Free tier | Yes | No |
| OSS self-host | Yes | No |
| Founded | 2019 | 2008 |
| HQ | — | Brussels, Belgium |
Full Amundsen pricing → Full Collibra pricing →
Both share Primary cluster: Catalog & discovery · dbt integration: Plugin · OpenLineage: Consumer · Status: ● active
Each tool's center of gravity.
| Cluster | Amundsen | Collibra |
|---|---|---|
| Quality & testing | 0/3 | 2/3 |
| Lineage & metadata | 1/3 | 3/3 |
| Catalog & discovery | 3/3primary | 3/3primary |
Scored 0–3 per cluster on the same rubric across all tools. A 0 means the cluster isn't the tool's focus, not that the feature is absent. See the methodology.
Where they cover different ground.
The declared feature set.
8 of 8 declared features differ — listed first.
These are each tool's self-declared key_features; a blank dot means
undeclared, not impossible.
| Feature | Amundsen | Collibra |
|---|---|---|
| Data Contracts Quality & testing | ||
| ML Anomaly Detection Quality & testing | ||
| Business Glossary Catalog & discovery | ||
| PII Auto-Classification Catalog & discovery | ||
| Column-Level Lineage Lineage & metadata | ||
| Reverse Impact Analysis Lineage & metadata | ||
| Table-Level Lineage Lineage & metadata | ||
| Transformation Lineage Lineage & metadata |
Where they disagree.
Catalog & discovery
8 of 9 differ| Amundsen | Collibra | |
|---|---|---|
| Business glossary | ||
| NL search | ||
| Data contracts | ||
| Governance flows | ||
| Access requests | ||
| PII auto-classify | ||
| Tag propagation | ||
| Free self-host |
Lineage & metadata
4 of 7 differ| Amundsen | Collibra | |
|---|---|---|
| Column-level | ||
| Cross-system | ||
| Reverse impact | ||
| Historical |
When to pick each.
Teams that already run Amundsen and need to understand what they have, or teams with a genuinely minimal requirement — usage-ranked table search and ownership tracking, nothing more — who are comfortable owning a codebase that is no longer moving. The core idea still holds up: index tables, dashboards, and people into Elasticsearch, rank results by query-log usage so the tables analysts actually trust float to the top, and keep the UX focused on the single question "which table should I use?" Databuilder's pull model is plain Python, so extending it from an existing Airflow deployment is straightforward for a data platform team.
Large, regulated enterprises — banks, insurers, pharma, public sector — that need a governance-first control plane: a real CDO function, formal stewardship, a business glossary, policy enforcement, and auditable lineage for regulations like BCBS 239, GDPR, SOX, HIPAA, and the EU AI Act. Collibra is strongest where governance process and accountability matter more than developer ergonomics, and where a single vendor for catalog plus governance plus lineage plus data quality plus AI governance is preferred over best-of-breed point tools.
What each does best.
Amundsen stands out for
- Pioneered usage-ranked, search-first discovery — PageRank-style ranking from query logs remains a genuinely good idea that successors copied
- Deliberately small surface area — analysts get table search, owners, column stats, and previews without a governance platform's learning curve
- Apache-2.0 under neutral LF AI & Data governance, with nothing held back for a paid tier
- Databuilder is plain-Python ETL — custom extractors are easy to write and schedule from an existing Airflow deployment
Collibra stands out for
- The deepest governance and stewardship tooling in the cluster — a configurable workflow engine, business glossary, policies, ownership, and audit trails purpose-built for regulated enterprises
- Broad single-vendor footprint — catalog, lineage (table and column, OpenLineage-aware), an ML data-quality module (from the OwlDQ acquisition), privacy, and AI governance under one platform
- Strong automated lineage with root-cause and downstream impact analysis at table, column, and report level, with in-line transformation context
- A mature, analyst-recognised leader with 100+ catalog integrations and a large regulated-enterprise customer base
Tools both also compete with.
All Amundsen alternatives, scored →All Collibra alternatives, scored →
A note on this comparison.
Every capability value above traces to Amundsen or Collibra's own structured spec, which links back to its source — nothing here is averaged or smoothed across the two.
Notice something inaccurate? Send a correction.