Databricks HIPAA compliance: a checklist for HIPAA-aligned data platforms

Databricks HIPAA compliance checklist for Azure data platforms: BAAs, encryption, private networking, access control, audit logs and de-identification of PHI.

Databricks HIPAA compliance: a checklist for HIPAA-aligned data platforms

Databricks HIPAA compliance checklist for Azure data platforms: BAAs, encryption, private networking, access control, audit logs and de-identification of PHI.

Databricks HIPAA compliance: a checklist for HIPAA-aligned data platforms

Databricks HIPAA compliance checklist for Azure data platforms: BAAs, encryption, private networking, access control, audit logs and de-identification of PHI.

IN THIS GUIDE

No headings found on page

SHORT ANSWER

A HIPAA-aligned data platform on Azure and Databricks needs six things: a signed business associate agreement for every service that touches PHI, encryption in transit and at rest, private networking, role-based access limited to the minimum necessary data, audit logs for every PHI access, and a documented risk assessment. HIPAA has no official certification, so you prove alignment through safeguards, documentation and regular review.

You want the same analytics and machine learning as every other team, but your data includes protected health information (PHI). That changes how you build. Safeguards have to sit in the architecture from day one, because adding private networking or access rules to a live platform later means rework and downtime. This checklist walks through what a HIPAA-aligned platform on Azure and Databricks needs, area by area.

This guide is general information, not legal advice. Bring your compliance and legal teams in early and keep them involved as the platform grows.

What does HIPAA require from a data platform?

HIPAA doesn’t name technologies. The Security Rule asks for administrative, physical and technical safeguards that protect the confidentiality, integrity and availability of PHI. For a cloud data platform, that comes down to six areas: contracts, security controls, access, audit, de-identification and written procedures. The checklist follows that order.

Contracts and scope

  • Sign a BAA with every provider. Every cloud and software service that stores or processes PHI needs a business associate agreement.

  • Use only covered services. A BAA lists specific services. Keep PHI workloads on those and nothing else.

  • Cover your partners. Any development or support partner that may see PHI signs a BAA too.

  • Map the data flow. Write down where PHI enters, where it’s stored, where it’s transformed and where it leaves.

Security safeguards

  • Encrypt data in transit with TLS and at rest, with customer-managed keys where your policy requires them.

  • Use private networking and private endpoints so no data service is reachable from the public internet.

  • Enforce single sign-on and multi-factor authentication through your identity provider.

  • Separate development, test and production. PHI lives in production only.

  • Patch and harden compute, and limit who can create or change clusters.

What this looks like in a real build

We built a HIPAA-aligned data intelligence platform on Azure and Databricks for a U.S. nonprofit in human services. When we started, data scientists worked in local environments with no shared, governed place to collaborate. That’s slow, and it’s risky when the data is PHI.

We followed Databricks’ Security Reference Architecture and built a fully private Azure environment: hub-and-spoke networking, private endpoints, private DNS and VPN-only access with no public exposure. Unity Catalog took over governance, lineage and permissions, and Pulumi defined the infrastructure as code so every environment could be rebuilt the same way. The result was a shared, governed workspace for data scientists and engineers, and machine learning models running in production in the cloud.

The lesson we’d pass on: settle the network and identity model before the first pipeline. Once nothing is publicly exposed and every permission runs through one catalog, each new use case starts on safe ground instead of reopening the security discussion.

Access control and minimum necessary

  • Grant access by role, through groups, never to individual users.

  • Apply row-level and column-level security so people see only the data their job needs.

  • Mask or tokenise direct identifiers for analysts who don’t need them.

  • Review access on a fixed schedule and remove it as soon as someone changes role or leaves.

Most analytics questions don’t need names or record numbers. Start from that assumption and the group of people with access to identifiable data stays small, and so does the scope of every audit.

Audit and monitoring

  • Log every access to PHI, including queries, exports and admin changes.

  • Send logs to a central, tamper-resistant store with a defined retention period.

  • Alert on unusual activity, such as large exports or access outside working hours.

  • Track lineage so you know which reports and models use PHI.

Logs nobody reads don’t protect anything. Name a person who reviews alerts every week, and keep a record that they did.

De-identification for analytics and AI

  • Build de-identified datasets for broad analytics, research and model training.

  • Use the HIPAA Safe Harbor method or an expert determination, and record which one you used.

  • Store re-identification keys separately, with tight control over who can use them.

  • Test machine learning and AI outputs to make sure they don’t reveal PHI.

Policies and procedures

  • Run and document a security risk assessment, and repeat it regularly and after major changes.

  • Keep an incident response and breach notification procedure, and test it.

  • Train everyone who works on the platform in handling PHI.

  • Document backups, disaster recovery and data retention.

Who is responsible for what?

Azure and Databricks secure the platform. You and your partner secure how it’s configured and used. The provider gives you the tools, and switching them on and checking them is your job.

Area

Cloud provider

Your organisation and partner

Physical data centre security

Yes

No

Platform infrastructure security

Yes

Configuration

Network and identity configuration

Tools provided

Yes

Access control and data permissions

Tools provided

Yes

Audit logging

Tools provided

Enable, store and review

Policies, training and risk assessment

No

Yes

Moving healthcare data off a legacy system? Pair this list with our data platform migration checklist. If you’re still picking a partner, read how to choose a Databricks consulting partner.

How RUBICON helps with HIPAA-aligned data platforms

We design and build data platforms that handle PHI on Azure and Databricks, from private networking and Unity Catalog governance to MLOps pipelines and training for your team. RUBICON has about 55 people, 40+ of them engineers, and is ISO 27001:2022 certified, a Databricks Partner and a Microsoft Solutions Partner for Cloud & AI Platforms. See our healthcare work and data engineering services.

If you’re planning a platform with health data, our architects can walk through your setup with you.

Frequently asked questions

Is there a HIPAA certification for software or platforms?

No. The U.S. Department of Health and Human Services doesn't certify products or vendors. You show HIPAA alignment through your safeguards, policies, risk assessments and business associate agreements. Some organisations add third-party audits or frameworks such as HITRUST to give customers and partners extra assurance, but none of these is an official HIPAA certificate.

Are Azure and Databricks HIPAA compliant?

Microsoft Azure and Databricks both sign business associate agreements and support HIPAA workloads, but only for covered services that you configure correctly. On Databricks that usually means turning on the Compliance Security Profile. Responsibility is shared: the provider secures the platform, and you configure access, encryption, networking, logging and data handling.

Can we use real patient data for development and testing?

Avoid it. Use de-identified or synthetic data in development and test environments, and keep identifiable data in production, where only the people who need it can reach it. If a test really needs production-like data, create a de-identified copy with a documented method and treat the re-identification key as PHI.

What is the difference between de-identified and pseudonymised data?

Under HIPAA, de-identified data has had identifiers removed through the Safe Harbor method or an expert determination, and it's no longer PHI. Pseudonymised data swaps identifiers for codes but can still be linked back to a person. You must keep treating it as PHI, with the same access controls, encryption and audit logging.

Related case study

Case study image showcase

HIPAA Aligned Data Platform on Azure & Databricks

How RUBICON Delivered a Secure, Scalable Foundation for Healthcare Data, Analytics and Machine Learning

More resources

If you're putting PHI on Azure and Databricks, our engineers can check your architecture against this list with you.
If you're putting PHI on Azure and Databricks, our engineers can check your architecture against this list with you.
If you're putting PHI on Azure and Databricks, our engineers can check your architecture against this list with you.