SHORT ANSWER
Unity Catalog is the governance layer of Databricks. One metastore per region holds catalogs, schemas and tables, and every permission, lineage record and audit log lives there. A solid setup uses one catalog per environment, schemas per medallion layer, account-level groups synced from your identity provider, and governed tags with attribute-based policies. A typical Hive metastore migration takes 4 to 12 weeks.
You need to know who can see which data, where a number came from and who touched a table last. On Databricks, Unity Catalog answers all three. It controls access to tables, files, models and functions, records lineage automatically and logs every access for audits. A good setup comes down to four decisions: metastore and storage layout, catalog structure, permission model and how you classify sensitive data. This guide walks through each, then covers migration from the legacy Hive metastore.
How does the Unity Catalog hierarchy work?
Unity Catalog uses a three-level namespace, catalog.schema.object, under a metastore:
Metastore: the top-level container, one per region, attached to all workspaces in that region.
Catalog: the main unit of isolation, typically per environment, business unit or domain. Catalogs can be bound to specific workspaces.
Schema: a group of related objects inside a catalog, for example a medallion layer or a subject area.
Objects: tables, views, materialised views, volumes (for files), functions and registered ML models.
Around this hierarchy sit storage credentials and external locations, which grant governed access to cloud storage, and connections for query federation to external databases. Give each catalog its own managed storage location so production and development data land in separate storage accounts or buckets.
What catalog layout works for most enterprises?
A layout that works for most enterprises combines environments at the catalog level with medallion layers at the schema level:
Catalog | Schemas | Bound to workspaces | Who writes | Who reads |
|---|---|---|---|---|
dev | bronze, silver, gold, sandbox | Development | Data engineers | Data engineers |
test | bronze, silver, gold | Test | CI/CD service principal | Engineers, testers |
prod | bronze, silver, gold | Production (and read-only in analytics) | Production service principal only | Analysts and BI on gold, engineers on silver |
prod_sandbox | one schema per team | Analytics | Analysts in their own schema | The owning team |
shared | reference, external | All | Data platform team | Everyone |
In a larger organisation with domain ownership, add the domain to the catalog name (prod_finance, prod_supply_chain) and give each domain team ownership of its catalogs. Don’t create a catalog per project. Catalogs are hard to rename, and projects come and go.
How should you set up groups and service principals?
Sync groups from your identity provider. Use SCIM provisioning (or automatic identity management on Azure) from Microsoft Entra ID or Okta to create account-level groups. Never grant permissions to individual users.
Grant on catalogs and schemas, not tables. Privileges inherit downwards, so grant USE CATALOG, USE SCHEMA and SELECT at the schema level for readers, and MODIFY or CREATE TABLE for writers.
Run production jobs as service principals. Only service principals should write to production catalogs, and only through CI/CD pipelines.
Set clear ownership. Make a group, not a person, the owner of each catalog and schema, so access keeps working when people leave.
Use workspace-catalog bindings. Bind the prod catalog read-write only to the production workspace and read-only to analytics workspaces, so development work cannot touch production data.
What we see in delivery: govern before the first login
On a HIPAA-aligned data platform we built on Azure and Databricks for a nonprofit human-services organization, Unity Catalog carried the governance load from the start: central access control, lineage and role-based rights for teams working with protected health information. We followed Databricks’ Security Reference Architecture, kept the whole environment on private networking with no public exposure, and trained the client’s teams on secure data handling.
The lesson: set up groups, ownership and access rules before the first data scientist logs in. Retrofitting permissions onto tables people already use is slower and riskier than starting governed.
How do lineage and auditing work?
Unity Catalog captures table and column-level lineage automatically for queries, notebooks, jobs and dashboards running on Unity Catalog compute. You can view it in Catalog Explorer or query the system.access.table_lineage and system.access.column_lineage system tables. The system.access.audit table records who accessed what and when, which is the evidence auditors ask for under frameworks such as ISO 27001, HIPAA or DORA.
How do tags, classification and ABAC fit together?
Tagging is where governance scales. Unity Catalog now offers three features that work together, all generally available in 2026:
Governed tags: account-level tag definitions with allowed values and control over who can apply them, for example a sensitivity tag with the values public, internal, confidential and pii.
Data classification: automatic scanning that detects sensitive data such as emails or national identifiers and suggests or applies tags.
ABAC policies: row filter and column mask policies defined once at catalog or schema level and applied to every table or column with a matching tag.
For example, one policy can mask all columns tagged pii for everyone except members of a privacy-approved group, across hundreds of tables. You can still use table-level row filters and column masks for special cases. For regulated data, combine this with the controls in our HIPAA-aligned data platform checklist.
How do you migrate from the Hive metastore?
From 30 September 2026, new Databricks workspaces are Unity Catalog-only, without the Hive metastore, DBFS root and mounts, or no-isolation shared clusters. Existing workspaces aren’t affected, but staying on the Hive metastore means missing out on lineage, ABAC, system tables and most new features. There are two main paths:
Approach | How it works | Best for |
|---|---|---|
Upgrade tables (UCX, SYNC, CTAS, DEEP CLONE) | Register or copy tables into Unity Catalog. UCX automates assessment, group migration and table upgrades | A clean, permanent move |
Hive metastore federation | Mounts the Hive metastore as a foreign catalog governed by Unity Catalog | Large estates that need a gradual, phased migration |
A typical migration runs in five steps: assess with UCX, migrate groups to account level, create the target catalogs and external locations, upgrade tables and views, then repoint jobs, notebooks and BI connections. For a mid-sized workspace (a few hundred tables and jobs) expect 4 to 12 weeks, with most of the effort in updating and testing code rather than moving data. Our data platform migration checklist covers the testing and cutover steps.
Dates and feature status reflect Databricks announcements at the time of writing. This guide is general information, not legal advice.
How RUBICON helps with Unity Catalog
We design catalog layouts and permission models, set up tagging and ABAC, and run Hive metastore migrations together with your team. Our work includes a HIPAA-aligned data platform on Azure and Databricks. We’re a Databricks Partner and a Microsoft Solutions Partner for Cloud & AI Platforms, about 55 people with 40+ engineers, working in CET under ISO 27001:2022 certified processes.
See our data engineering services. If you’re planning a Unity Catalog setup or migration, our architects can map the layout with you.
Frequently asked questions
How many metastores do I need in Unity Catalog?
One per region in which you run Databricks workspaces. All workspaces in that region attach to the same metastore, and you separate environments and business units with catalogs and workspace-catalog bindings rather than with separate metastores. If you need data from another region, share it through Delta Sharing instead of copying it.
Should I use one catalog per environment or per domain?
Most enterprises start with one catalog per environment (dev, test, prod) and schemas per layer or domain, because that maps cleanly to workspace bindings and CI/CD. Large organisations with a data mesh approach often add the domain to the catalog name, for example prod_sales, so each domain gets its own ownership and permissions.
Is the Hive metastore being retired on Databricks?
Databricks is phasing it out for new workspaces. From 30 September 2026, new workspaces are provisioned as Unity Catalog-only, without the Hive metastore, DBFS root and mounts, or no-isolation shared clusters. Existing workspaces keep working, but new features increasingly require Unity Catalog, so it makes sense to plan your migration now.
What is ABAC in Unity Catalog?
Attribute-based access control lets you write row filter and column mask policies based on governed tags, for example masking every column tagged as personal data for everyone outside a given group. Policies apply automatically to all tables that carry the tag, so you don't define security table by table. ABAC and governed tags became generally available in 2026.
Related case study

HIPAA Aligned Data Platform on Azure & Databricks
How RUBICON Delivered a Secure, Scalable Foundation for Healthcare Data, Analytics and Machine Learning
More resources
