Databricks cost optimisation: 12 ways to cut your bill

Databricks cost optimisation: 12 practical ways to cut your bill, from jobs compute and spot workers to system tables, compute policies and budget alerts.

Databricks cost optimisation: 12 ways to cut your bill

Databricks cost optimisation: 12 practical ways to cut your bill, from jobs compute and spot workers to system tables, compute policies and budget alerts.

Databricks cost optimisation: 12 ways to cut your bill

Databricks cost optimisation: 12 practical ways to cut your bill, from jobs compute and spot workers to system tables, compute policies and budget alerts.

IN THIS GUIDE

No headings found on page

SHORT ANSWER

To cut a Databricks bill, fix the patterns that waste the most first: scheduled work on all-purpose clusters, idle clusters without short auto-termination, full reloads instead of incremental processing, oversized SQL warehouses and no spot workers. Track spend in the system.billing.usage table against a 90-day baseline. Then hold the savings with compute policies, serverless usage policies and budget alerts.

Your Databricks bill grew faster than your workloads, and nobody can say exactly why. Databricks costs come from two sources. Databricks Units (DBUs) are billed per second of compute by workload type. The cloud resources underneath (virtual machines, storage and networking) are included in serverless prices, but your cloud provider bills them separately for classic compute. This guide lists 12 techniques, the typical impact of each and how to keep costs under control afterwards.

What are the 12 ways to reduce your Databricks bill?

  1. Run scheduled work on jobs compute, not all-purpose clusters. All-purpose (interactive) compute has a much higher DBU rate than jobs compute. Any notebook that runs on a schedule should run as a Lakeflow Job on a job cluster or serverless jobs compute.

  2. Set short auto-termination on interactive clusters. A default of 10 to 30 minutes stops idle clusters from running overnight and at weekends. Enforce it with a compute policy so nobody can set it to hours.

  3. Use autoscaling with sensible limits. Autoscaling lets clusters shrink when load drops. Set a low minimum and a realistic maximum so one heavy query cannot scale a cluster to dozens of nodes.

  4. Use spot instances for workers. Spot or preemptible VMs are much cheaper than on-demand capacity. Keep the driver on-demand and allow fallback to on-demand so jobs still finish when spot capacity disappears.

  5. Right-size instance types. Match memory-optimised, compute-optimised or storage-optimised instances to the workload, and check the Spark UI for idle cores or spill to disk. Newer instance generations usually give more performance per euro.

  6. Use serverless where it fits, and pick the right performance mode. Serverless SQL warehouses and serverless jobs start in seconds and stop billing when idle. For batch jobs that are not time-critical, the standard performance mode is cheaper than performance-optimised mode.

  7. Tune SQL warehouses. Start with a smaller size, set auto-stop to a few minutes, and scale out with clusters for concurrency rather than scaling up the size. Separate warehouses for BI dashboards and ad-hoc analysis make usage easier to track.

  8. Test Photon per workload. Photon costs more per DBU but can finish SQL and DataFrame workloads much faster. Keep it where the cost per run drops, and switch it off where it does not.

  9. Process data incrementally. Replace full reloads with Auto Loader, Lakeflow Declarative Pipelines, MERGE or change data feed so each run touches only new or changed data. This often has the largest effect on engineering pipelines.

  10. Let Databricks manage table layout. The predictive optimisation feature runs OPTIMIZE, VACUUM and ANALYZE automatically on Unity Catalog managed tables, and liquid clustering (including automatic key selection) reduces the data each query scans.

  11. Clean up storage. Vacuum old file versions, set retention on logs and staging tables, drop unused tables and move rarely read data to cheaper storage tiers. Storage is cheap per gigabyte but grows quietly.

  12. Commit to usage once patterns are stable. Pre-purchase or committed-use agreements with Databricks, and reservations for classic VMs with your cloud provider, give discounts on predictable baseline usage. Do this after the other steps, so you commit to the optimised level.

What impact can you expect from each technique?

The ranges below are typical, not guarantees. Your results depend on how the workspace is used today, so measure a baseline first and show real savings against it.

Technique

Typical saving on affected workload

Effort

Jobs compute instead of all-purpose

30 to 60%

Low

Auto-termination on interactive clusters

10 to 30% of interactive spend

Low

Autoscaling with limits

10 to 25%

Low

Spot workers

20 to 50% of VM cost

Low

Right-sizing instances

10 to 30%

Medium

Serverless, standard mode for batch

10 to 40% for bursty or idle-heavy jobs

Low to medium

SQL warehouse tuning

15 to 40% of SQL spend

Low

Photon tested per workload

0 to 30% (can be negative)

Low

Incremental processing

30 to 80% of pipeline compute

Medium to high

Managed table layout and liquid clustering

10 to 40% of query compute

Low to medium

Storage clean-up

10 to 50% of storage cost

Low

Committed-use discounts

Depends on contract

Low

What we see in delivery: visibility before savings

When we built a HIPAA-aligned data platform on Azure and Databricks for a nonprofit human-services organization, data scientists were moving from local machines into a shared workspace for the first time. We set up Unity Catalog from day one for central governance, lineage and access control, inside a fully private network that followed Databricks’ Security Reference Architecture.

That order matters for cost too. System tables and consistent tagging both depend on Unity Catalog, and without them you can’t tell which team or job is spending what. Teams that try to cut costs before they can see them end up guessing.

How do you monitor Databricks costs with system tables?

Unity Catalog system tables give you a queryable record of usage across the account. The ones that matter most for cost are:

  • system.billing.usage: every billable usage record with SKU, workspace, DBUs, custom tags and identity metadata, including serverless usage.

  • system.billing.list_prices: list prices per SKU over time, so you can join them to usage for an estimated cost (before any negotiated discount).

  • system.compute.clusters and system.compute.warehouses: configuration history, useful for finding clusters without auto-termination or with oversized limits.

  • system.lakeflow tables: job and job run history, to find the most expensive and fastest-growing jobs.

  • system.query.history: SQL warehouse and serverless query history, to find the heavy queries behind SQL spend.

Build a simple cost dashboard on these tables with spend per workspace, SKU, team tag and the top 20 jobs, and review it every week. Databricks also provides a ready-made usage dashboard in the account console that you can import and extend.

Which governance controls keep the savings?

Savings don’t last without guardrails. Three controls do most of the work:

  • Compute policies (previously cluster policies). Limit instance types, maximum workers and auto-termination, require tags such as team and cost centre, and give most users policy-based cluster creation instead of unrestricted access.

  • Serverless usage policies. Serverless compute has no cluster to configure, so these policies (previously called budget policies) attach tags automatically to serverless usage by users or groups, which makes chargeback possible.

  • Budgets and alerts. Account-level budgets track spend by workspace or tag and email owners when thresholds are crossed. Budgets alert but do not stop compute, so pair them with policies and an agreed response process.

Unity Catalog underpins much of this, from system tables to tagging. If you’re still on the Hive metastore, see our Unity Catalog setup guide. Databricks renames and changes these features often, so check the current documentation before you change policies.

Where should you start?

  1. Enable system tables and build a baseline of the last 90 days of spend.

  2. Find the top ten cost drivers by SKU, workspace, job and warehouse.

  3. Apply the low-effort fixes (1, 2, 3, 4 and 7) and measure again after two to four weeks.

  4. Plan the medium-effort work, such as incremental pipelines and table layout, as part of normal delivery.

  5. Introduce compute policies, tagging and budgets so the savings hold.

How RUBICON helps with Databricks cost optimisation

Our data engineers review workspaces, system tables and pipelines, deliver the quick wins, then rebuild the expensive pipelines for incremental processing as part of normal delivery. We’re about 55 people, 40+ engineers, a Databricks Partner and a Microsoft Solutions Partner for Cloud & AI Platforms, and we work under ISO 27001:2022 and ISO 9001:2015 certified processes.

Read how to choose a Databricks consulting partner or see our data engineering services. If your bill keeps climbing, we can go through your system tables with you.

Frequently asked questions

Why is my Databricks bill so high?

The usual causes are interactive all-purpose clusters left running, scheduled jobs on all-purpose compute, oversized or always-on SQL warehouses, full reloads instead of incremental loads, and poorly organised tables that force large scans. Query the system.billing.usage table grouped by SKU, workspace and tag to see which of these applies to you before you change anything.

Is Databricks serverless cheaper than classic compute?

It depends on the workload. Serverless has a higher DBU rate but includes the cloud VM cost, starts in seconds and stops billing when idle, so it's often cheaper for bursty jobs and SQL. For steady, long-running batch jobs, well-tuned classic job clusters with spot workers can still cost less. Serverless jobs also offer a standard mode that trades some latency for a lower price.

How do I see Databricks costs per team or project?

Apply tags consistently and query them from system.billing.usage, joined to system.billing.list_prices for an estimated cost. Compute policies can force tags on classic clusters, and serverless usage policies (previously called budget policies) attach tags automatically to serverless usage. Account-level budgets then send alerts per tag, workspace or team.

Does Photon save money on Databricks?

Often, but not always. Photon consumes DBUs at a higher rate, so it only saves money when queries finish enough faster to offset that. It usually pays off for SQL-heavy work and large aggregations or joins, and less so for code heavy on Python UDFs. Test it on a representative job and compare the total cost per run, not just the runtime.

Related case study

Case study image showcase

HIPAA Aligned Data Platform on Azure & Databricks

How RUBICON Delivered a Secure, Scalable Foundation for Healthcare Data, Analytics and Machine Learning

More resources

If you can't tell where your Databricks spend goes, our data engineers can review your workspace and system tables and hand you a ranked savings plan.
If you can't tell where your Databricks spend goes, our data engineers can review your workspace and system tables and hand you a ranked savings plan.
If you can't tell where your Databricks spend goes, our data engineers can review your workspace and system tables and hand you a ranked savings plan.