Secrets of Snowflake Migration Success Ebook
Learn how moving from legacy solutions to a modern cloud data platform provides greater simplicity, performance and cost efficiency.
Unity Catalog locks customer data into Databricks by design, which leaves customers with duplicate data and fragmented governance just to write to their Apache Iceberg™ tables.
Customers want choice: interoperable data governed by an interoperable catalog. Iceberg brings choice at the data layer — the promise of any engine, one copy of data, minimal switching cost. However, Unity Catalog restricts interoperability for the tables it doesn't manage, creating lock-in.
Unity Catalog claims to have an open foundation that lets any engine read and write data to tables managed by Unity Catalog, but it restricts writes to tables not managed in Unity Catalog. Databricks engines can read that data but can't write to it, forcing teams to create and maintain multiple copies of data, reconcile access across systems, and navigate complexity just to write to those tables.
The open source software (OSS) version of Unity Catalog also sits at the earliest governance tier of LF AI & Data, the Linux Foundation's umbrella for open source AI and data projects. More than two years after Databricks contributed it in June 2024, it is still at Sandbox, a stage that doesn't yet require independent governance, and Databricks remains the only established supplier of that catalog.4,3
Unity Catalog leads to lock-in and restricts interoperability for tables it doesn't manage. The table below summarizes the documented limitations.
| Unity Catalog constraint | What it means |
|---|---|
| Forced dependence | Customers have to copy their data into Databricks because Unity Catalog does not write to foreign tables and can only read those tables, leading to data duplication and complexity.1,2 Customers also have less flexibility because they cannot replace Unity Catalog with an independent Iceberg catalog such as Polaris or AWS Glue as the governing catalog, or use a different managed version of Unity Catalog OSS. |
| Databricks control | Databricks owns the commercial version of Unity Catalog and controls the roadmap and the contribution model for the OSS version.3 Unity Catalog OSS sits at Sandbox, the earliest of LF AI & Data's four hosting stages.4 |
| One-way interoperability | Interoperability only exists when Unity Catalog is managing the tables. Reads go through catalog federation — Databricks' feature for connecting directly to external catalogs. Databricks engines can't write to foreign tables.1 |
| Security and trust risks | Credential vending isn't supported on Databricks workspaces using default storage.5 Customers must grant Databricks standing access to the authorized storage location.8 |
Yes. Unity Catalog locks in customer data with Databricks because once customers adopt Unity Catalog, switching catalogs could require reengineering their data pipelines and access controls.6 This happens because it restricts writes to tables customers don't manage in Unity Catalog. And once customers adopt Unity Catalog, leaving isn't easy, because customers can't easily swap in any Iceberg-compatible catalog, and even Unity Catalog OSS is not truly community-driven.3
Yes. Customers have to copy their data into Databricks to modify it because Unity Catalog can only read data managed by other catalogs, not write to it.1 This can lead to multiple copies of data and fragmented governance.
No other vendor offers a managed version of Unity Catalog OSS.3 Customers cannot replace Unity Catalog with an independent Iceberg catalog such as Polaris or AWS Glue as the governing catalog, or use a different managed version of Unity Catalog OSS.
In license, yes; in project governance, no: Databricks owns the commercial catalog, and more than 99% of attributable commits to the OSS version come from Databricks.3
Unity Catalog OSS sits at Sandbox, the earliest of LF AI & Data's four hosting stages (Sandbox, Incubation, Graduation, Emeritus). Sandbox projects aren't yet required to have an independent Technical Steering Committee with multiple contributing organizations; that requirement only applies starting at Incubation.4 Databricks controls the roadmap and the contribution model, so, at this early stage, "open source" doesn't imply shared stewardship.7
Open format but closed catalog: To use Apache Iceberg managed or foreign tables, Databricks requires Unity Catalog, which is in turn steered by Databricks.13
No. Interoperability only exists when Unity Catalog is managing the tables.6 Two issues drive this: one-way interoperability, and security and trust risks.
No. Databricks engines can't write to foreign tables.1 Reads go through catalog federation instead, Databricks' feature for connecting directly to an external catalog.
Catalog federation requires a long-lived storage credential set up in advance, not one issued per query, giving Databricks standing access to the customer's Iceberg storage rather than a credential scoped to a single request.8 This runs counter to cloud security best practice, which recommends temporary credentials over long-term ones.9
Turning on Unity Catalog's federation path makes foreign tables read-only. Spark itself reads from and writes to the customer's externally managed Iceberg tables via open libraries, until Unity Catalog enters the picture.10
Catalog federation gives Databricks read access to externally managed Iceberg tables, not write access. Writes through that path aren't supported.8,11
Only for tables Unity Catalog manages, and not on Databricks workspaces using default storage.5 Databricks uses temporary credentials to allow other engines to read from and write to Unity Catalog-managed Iceberg tables.12
The reverse isn't true. To read data governed by another Iceberg catalog, customers must grant Databricks standing access to the authorized storage location (not a short-lived, request-scoped credential), which is a security and compliance risk worth evaluating because long-standing credentials may increase the risk exposure.8
Not a good one: Customers must copy their data into Databricks whenever they need to modify it from Databricks.1
Catalog federation isn't a workaround either. It's Databricks' read-only path to tables governed by another catalog, so writes through that path aren't supported.8,11
It also adds risk. Instead of getting a short-lived credential from the owning catalog for each query, catalog federation relies on a storage credential that an admin configures in advance. Databricks then reads the files directly from storage, so there's no record of the read at the owning catalog.8,13
Not for Iceberg. To use Apache Iceberg managed or foreign tables, Databricks requires Unity Catalog, which is in turn controlled by Databricks.13
No, not as your governing catalog. Databricks requires Unity Catalog to govern and write Iceberg tables; you can't replace it with Polaris or AWS Glue (Databricks can only connect to an external catalog such as AWS Glue for read-only access), and there's no managed alternative to Unity Catalog OSS.13,2,3
Only when Unity Catalog manages them. Databricks uses temporary credentials to allow other engines to read from and write to Unity Catalog-managed Iceberg tables, but Databricks engines can't write to foreign tables.2
No. Catalog federation gives Databricks read access to externally managed Iceberg tables, not write access. Writes through that path aren't supported.8,11
Yes. Catalog federation requires a long-lived storage credential set up in advance, not one issued per query, giving Databricks standing access to the customer's Iceberg storage rather than a credential scoped to a single request.14,3