DataHub.
Founded 2021 · Palo Alto, CA
Status · ● active
Verified · ● 2mo ago
Apache-2.0 metadata platform with a serious managed counterpart — strongest event-driven architecture and column-level SQL lineage in OSS.
Where it fits — and where it doesn't.
Engineering-led data platforms that want an open, extensible metadata layer they can shape to their stack — with a credible managed escape hatch (DataHub Cloud) when self-hosting Kafka, Elasticsearch, and the graph store stops being fun.
Particularly strong for organisations that already think in events: DataHub's Kafka-based Metadata Change Log makes it a natural fit for shops that want metadata to flow the same way data does. The SQL parser is genuinely best-in-class in the OSS catalog space, with SQLGlot-based column-level lineage benchmarked at 97–99% accuracy on standard corpora — materially better than competing parsers. A good fit also for teams wiring DataHub into AI agents via the native MCP server.
You don't have platform engineers — self-hosting DataHub is real work (multiple stateful components, ingestion containers, careful upgrade paths) and the managed product is the way out, not optional.
Avoid also if you want a polished governance and glossary surface out of the box; OSS DataHub is improving but lags Atlan and OpenMetadata on stewardship UX. Finally, avoid if your buyer is a non-technical CDO — the product story is more developer-shaped than executive-shaped.
The honest scorecard.
- Best-in-class column-level SQL lineage parser (SQLGlot-based, benchmarked at 97–99% accuracy on standard corpora)
- Event-driven Kafka MCL architecture — metadata changes are a stream, not a snapshot, which composes well with downstream consumers
- Native OpenLineage consumer endpoint plus dedicated Spark and Airflow plugins
- Open-core model with a credible managed product (DataHub Cloud) means buyers can start free and graduate without a re-platforming
- Strong AI-agent story — native MCP server, Claude / Cursor / LangChain integrations announced through 2024–2025
- OSS deployment is heavy — Kafka, MySQL or Postgres, Elasticsearch, and a graph store (JanusGraph or Neo4j) all need running and tuning
- Steeper learning curve than OpenMetadata for ingestion configuration and the entity model
- No published Cloud pricing — sales-led motion despite the OSS on-ramp
- Some advanced governance features (auto-classification, data quality agents, blast-radius analysis) are Cloud-only
- Glossary and stewardship UX still trails Atlan; DataHub is more developer-shaped than steward-shaped
What DataHub actually is.
What DataHub actually is
DataHub is two things sold as one. The OSS is an Apache-2.0 metadata platform with a Kafka-based event log (Metadata Change Log, or MCL) at its core, an Elasticsearch-backed search layer, a graph store for lineage, and 80+ ingestion connectors. The managed counterpart, DataHub Cloud, takes the same code and adds proprietary features: auto-classification, data quality agents that run anomaly detection on warehouse tables, blast-radius analysis for change impact, and a polished AI/MCP layer.
The defining technical fact is the SQL parser: DataHub uses SQLGlot to extract column-level lineage from query logs, and the parser benchmarks materially better than competitors (97–99% vs. low-90s). For organisations whose lineage is primarily computable from SQL, that gap is real.
Where it fits against the alternatives
Against openmetadata, the trade is architecture and audience. DataHub’s stack (Kafka + graph DB) is heavier to operate but more event-native. OpenMetadata’s stack (Postgres + Elasticsearch) is simpler to run but pull-only. DataHub’s lineage parser is technically stronger; OpenMetadata ships features faster (Multi-Domain, Data Contracts GA, Data Quality as Code all landed quickly through 2024–2025). Engineering-led shops tend to pick DataHub; steward-led shops tend to pick OpenMetadata.
Against atlan, DataHub is the OSS counterpoint. Atlan has the more polished governance UX and the bigger enterprise GTM; DataHub has the open license, the stronger lineage parser, and a credible OSS-to-managed graduation path. Buyers with a hard OSS requirement skip Atlan; buyers with a hard “no platform team” requirement skip OSS DataHub.
On the AI / MCP narrative
DataHub Cloud was one of the earlier catalogs to ship a native MCP server and document deep integrations with Claude, Cursor, Cortex, CrewAI, and LangChain. The bet is the same as Atlan’s — that catalogs are the natural metadata surface for agentic AI. The execution is more developer-shaped on DataHub: the MCP server is real, but the product packaging around it is less front-and-centre than Atlan’s “Context Layer” pitch. For shops that want to drop a catalog into an agent stack and get to work, DataHub Cloud is a credible choice.
How to evaluate it
The honest test is to stand up the OSS for a few weeks against a representative subset of your stack — one warehouse, one orchestrator, one BI tool — and form an opinion on (a) operating cost, (b) lineage accuracy, and (c) ingestion-config maintainability. If the operating cost is unacceptable but the lineage and metadata model fit, DataHub Cloud is the natural escape. If the operating cost is acceptable and you can live with somewhat less governance polish than Atlan or OpenMetadata, OSS DataHub is one of the strongest open-core data infrastructure projects in production today.
All capabilities by cluster.
Quality & testing
Secondary · strength 2/3Catalog & discovery
Primary · strength 3/3Lineage & metadata
Secondary · strength 3/3Where it plugs in.
Native warehouse support
Orchestrators & pipeline tools
The honest pricing breakdown.
Free tier DataHub Core (OSS) is free under Apache-2.0; self-host with your own infrastructure. DataHub Cloud (managed) is contact-sales.
Sales-only tier DataHub Cloud (managed)
Full DataHub pricing breakdown — model, cost factors, alternatives by price →
What it doesn't do.
Compares the output of a model change against production before the pull request is merged — showing row-level and aggregate differences. Shifts data quality left into the development workflow. Datafold is the category-defining tool here; dbt's own cloud offering has added similar capabilities. Requires production-scale compute on a development branch, which has cost implications.
Warehouse-Native Monitoring →Monitors tables directly in the warehouse via query log parsing or scheduled metric queries — independent of the pipeline that produced the data. Catches issues regardless of which tool wrote the data, including ingestion-layer problems dbt can't see. Trade-off against dbt-native testing: reactive rather than preventive, and adds warehouse cost.
Drill into one capability.
Other key features
If not DataHub, then what?
Common alternatives
Quick answers.
- Is DataHub open source?
- Yes. DataHub is open source under the Apache-2.0 license, and can be self-hosted at no license cost.
- How much does DataHub cost?
- DataHub does not publish list pricing — it is sales-led, so you request a quote. A free tier is available: DataHub Core (OSS) is free under Apache-2.0; self-host with your own infrastructure. DataHub Cloud (managed) is contact-sales.
- How is DataHub deployed?
- DataHub can run as managed SaaS or be self-hosted.
- Does DataHub work with dbt and my warehouse?
- It has a native dbt integration. DataHub supports snowflake, bigquery, redshift, databricks, postgres, plus 7 more.
More catalog & discovery tools
Provenance.
Last verified 2026·05·08 against vendor documentation and, where possible, hands-on trial. Spot something off? Send a correction →