Data Stack Index / v 02.06
Verified 2026·07·03
Send a correction
§ Capability · Data observability

PII auto-classification.

Automatic detection and tagging of PII, PHI, and PCI data in warehouse columns — without manual schema annotation.

Buyer importance · Important
Tools with this 10 of 24
Open-source options 2
Across clusters 3
Last updated 2026·07·03
01
What this is

What counts as PII auto-classification?

PII auto-classification scans column values (or schemas, or both) and flags fields that look like personal data — emails, names, government IDs, payment data — without requiring a steward to tag every column by hand. Useful at scale, where manual classification is the bottleneck. Trade-off: false positives are common on synthetic-looking strings, and false negatives are common on internal IDs that happen to be sensitive. Most tools layer auto-classification with a human review step.

02
Tools with this capability

10tools, grouped by primary cluster.

Cluster · Catalog & discovery 7 tools

Alation

Alation

SaaS / Self-host

The incumbent that defined the data catalog — behavioral search, deep governance, and strong column-level lineage.

Pricing
Contact sales
First strength
Category-defining catalog with behavioral, usage-ranked search and pioneering natural-language search

Atlan

Atlan

Hybrid

Enterprise catalog and governance plane positioned as the AI context layer — connectors, lineage, contracts, and an MCP server for agents.

Pricing
Contact sales
First strength
Polished UX and onboarding — consistently scores top in analyst rankings on time-to-value relative to peers

Collibra

Collibra

SaaS

Enterprise data-and-AI governance incumbent: catalog, glossary, workflow stewardship, lineage, and a separate ML data-quality module.

Pricing
Contact sales
First strength
The deepest governance and stewardship tooling in the cluster — a configurable workflow engine, business glossary, policies, ownership, and audit trails purpose-built for regulated enterprises

DataHub

Acryl Data

OSS SaaS / Self-host

Apache-2.0 metadata platform with a serious managed counterpart — strongest event-driven architecture and column-level SQL lineage in OSS.

Pricing
OSS · free
First strength
Best-in-class column-level SQL lineage parser (SQLGlot-based, benchmarked at 97–99% accuracy on standard corpora)

OpenMetadata

Collate

OSS SaaS / Self-host

Apache-2.0 unified metadata platform with a deliberately simple stack — discovery, lineage, quality, and contracts in one project.

Pricing
OSS · free
First strength
Highest connector count in the OSS catalog space (120+) — particularly strong on dashboards, ML, and pipeline systems

Secoda

Secoda (Atlassian)

SaaS / Self-host

AI-native data catalog, lineage, and observability from Toronto — acquired by Atlassian in December 2025 to power Rovo AI.

Pricing
Contact sales
First strength
AI-native search and assistant as the primary interface — natural-language data questions across the catalog, plus purpose-built agents for search, documentation, observability, and governance

Select Star

Select Star (Snowflake)

SaaS

Automated data catalog with column-level lineage parsed from query logs — being acquired by Snowflake to power Horizon Catalog.

Pricing
Contact sales
First strength
Automated column-level lineage parsed from warehouse query logs is the differentiator — proven at very large scale (Block's case study covers lineage automation at exabyte scale)
03
Compare tools with this capability

Head-to-head, side by side.

04
Other capability hubs

Drill into a different capability.

How this list is built.

Inclusion here is one boolean on each tool's structured profile — if a tool you'd expect is missing, the field is recorded false or not yet verified, never an editorial call. See the methodology for how each field is sourced.