SHORT ANSWER
A data platform migration to Azure and Databricks works best in five phases: inventory what exists, agree on target architecture and governance, migrate pipelines in waves by business domain, validate results against the old system, then cut over and decommission the legacy platform. A mid-sized warehouse typically takes four to nine months. Most overruns come from a skipped inventory, manual validation and long parallel running.
Your legacy warehouse is slow, expensive to maintain and blocking the AI work everyone keeps asking about. Moving to a lakehouse on Azure and Databricks fixes much of that, but it touches every report and pipeline in the company. Use this checklist to plan each phase and avoid the mistakes that make migrations run late.
Phase 1: Assessment and inventory
List all data sources, with owners, refresh frequency and data volume.
List all pipelines, stored procedures and ETL jobs, and mark which ones are still used.
List all reports and dashboards, with their users and business criticality.
Write down known data quality issues and undocumented business logic.
Agree on business goals and success metrics for the migration.
Estimate current running costs so you can compare them with the new platform.
Don’t rush this phase. The inventory usually turns up pipelines nobody owns and reports nobody opens. Every one you retire now is one less thing to migrate and validate.
Phase 2: Target architecture and governance
Design the lakehouse layers (bronze, silver, gold) and naming conventions.
Set up workspaces and environments for development, test and production.
Configure Unity Catalog for access control, lineage and data discovery.
Design networking, identity and security, including private endpoints where required.
Set up CI/CD for code and jobs, and infrastructure as code for the platform.
Define cost controls: cluster policies, budgets and monitoring.
Phase 3: Migration in waves
Pick one business domain with high value and manageable complexity for the first wave.
Ingest raw data into the bronze layer, with full history where you need it.
Rebuild the business logic in the silver and gold layers, fixing known issues on purpose rather than copying them.
Add data quality tests to every pipeline.
Repoint that domain’s reports to the new gold tables.
Phase 4: Validation and parallel running
Reconcile row counts, sums and key metrics between old and new automatically.
Have business owners sign off on critical reports.
Run old and new systems in parallel for a fixed, short period.
Test performance with realistic user loads.
Phase 5: Cutover and decommissioning
Agree on a cutover date and tell every user group.
Switch schedules and access to the new platform.
Archive legacy data you must retain, then switch off the legacy systems.
Train users and data teams, and hand over documentation.
Review running costs after the first months and tune clusters and jobs.
Why user testing belongs in a migration plan
For a global chemical and consumer goods company, we built a cloud-native portal on Azure that lets employees and partners browse, request, upload and share data from their data lake. The scale target was at least 1,000 peak users in parallel and more than 10 million files, with 30,000 new files a day. We started with a four-day Lean Inception workshop and worked in two-week sprints. Within seven months of development, users were reading and writing data lake files through a secure portal with single sign-on.
Right after the first release, we ran remote user tests one person at a time, with people from different departments. 80% of them asked for features that were already in development, which told us the roadmap was right. The rest of the feedback shaped the second release. The same habit pays off in a migration: put real users on the new platform before cutover, so surprises show up while the old system still runs.
How long does each phase take?
These are typical durations for a mid-sized warehouse with dozens of sources and a few hundred reports. Phases overlap, so the total is shorter than the sum.
Phase | Typical duration | Main risk |
|---|---|---|
Assessment and inventory | 2 to 4 weeks | Undocumented pipelines and business logic |
Target architecture and governance | 3 to 6 weeks | Security reviews and networking approvals |
Migration in waves | 3 to 6 months | Scope creep from redesigning too much |
Validation and parallel running | 4 to 8 weeks per wave | Numbers that don’t match and nobody can explain |
Cutover and decommissioning | 2 to 4 weeks | Legacy systems that never get switched off |
Plan your budget around waves, not the whole estate. Once the first domain runs on the new platform, your estimate for the remaining waves gets far more accurate.
Common migration mistakes and how to avoid them
Mistake | Consequence | How to avoid it |
|---|---|---|
Skipping the inventory | Surprise dependencies late in the project | Spend two to four weeks on assessment |
Redesigning everything at once | Delays and scope creep | Migrate in waves by business domain |
Manual validation | Distrust in the new numbers | Automate reconciliation |
Long parallel running | Double costs and double work | Set a fixed end date |
Governance added later | Access and security rework | Set up Unity Catalog in phase 2 |
No real users before cutover | Adoption problems after go-live | Run user tests on the first migrated domain |
Before you start, read how to choose a Databricks consulting partner. If you’re still deciding on a platform, compare Microsoft Fabric and Databricks.
How RUBICON helps with data platform migrations
We’re a Databricks Partner and a Microsoft Solutions Partner for Cloud & AI Platforms, certified to ISO 27001:2022. We’ve built data lake platforms on Azure for global chemical and consumer goods companies, and we start every migration with a short assessment so the plan rests on your real estate, not assumptions.
Our data engineering services cover the full path from inventory to cutover. If you’re scoping a move, we can look at your current setup with you.
Frequently asked questions
How long does a data platform migration take?
A mid-sized warehouse with dozens of sources and a few hundred reports typically takes four to nine months, including parallel running and validation. Large enterprise estates can take a year or more, delivered in waves by business domain. The inventory in phase one is what turns that range into a date you can commit to.
Should we lift and shift or redesign?
Usually a mix. Move stable, well-understood pipelines with minimal changes, and redesign the parts that cause problems today, such as slow ETL jobs or unclear business rules. Redesigning everything at once is one of the most common reasons migrations run late, because every report changes at the same time.
How do we prove the new platform gives the same numbers?
Automate reconciliation. Compare row counts, totals and key metrics between the old and new system for every migrated table and report, and run the checks on every load. Then have business owners sign off on the most important reports. People trust the new platform when the numbers match without anyone checking by hand.
What happens to our existing reports?
Reports are repointed to the new data model, domain by domain. Plan this with report owners and keep a list of reports nobody opens anymore. A migration is a good moment to retire them, which cuts the work and the number of things you have to validate.
Related case study

Data Lake Portal for Chemical & Consumer Goods
Digital Transformation: RUBICON Develops a Secure Cloud Native Platform for a Global Chemical & Consumer Goods Company
More resources
