SHORT ANSWER
Change data capture (CDC) detects inserts, updates and deletes in a source database and sends only those changes to other systems, usually within seconds to minutes. Log-based CDC reads the database transaction log, so it puts little load on the source and replaces slow full nightly extracts. Use it when your data must be fresher than a daily batch can deliver.
Change data capture (CDC) is a set of techniques that spot changes in a source system, such as new orders, updated prices or deleted records, and pass only those changes to a target. The target can be a data lakehouse, a search index or another application. Instead of copying a whole table every night, you get a steady stream of small updates.
How does log-based CDC work?
A CDC tool connects to the source database and reads its transaction log (the SQL Server transaction log, the PostgreSQL write-ahead log or the Oracle redo log).
Each committed insert, update or delete becomes a change event with before and after values, a timestamp and an operation type.
The tool publishes events to a stream such as Apache Kafka or Azure Event Hubs, or writes them to a landing zone.
The target applies the changes, usually with a MERGE into Delta Lake tables in the bronze layer of a medallion architecture.
Why does CDC matter?
Full extracts slow down as tables grow, load production databases hard and leave data a day old. CDC cuts the delay from hours to seconds or minutes. That opens up live inventory, fraud checks, supply chain monitoring and operational dashboards. It also catches deletes, which timestamp-based incremental loads often miss, and keeps a change history you can use for audits.
Batch extracts vs. CDC
| Full nightly batch | Log-based CDC |
|---|---|---|
Data latency | Up to 24 hours | Seconds to minutes |
Load on source | High during extract | Low |
Captures deletes | Only by comparison | Yes |
History of changes | No | Yes, every change |
Complexity | Low | Medium |
What does CDC take to run?
CDC brings its own work. Source schema changes must not break the pipeline. Events must stay in order. After an outage, the pipeline has to recover without losing or duplicating changes. You also need to watch the lag between source and target. Managed connectors in Databricks, Fabric or Azure Data Factory take much of this off your team.
Check how fresh the data really needs to be before you add CDC. For a global chemical and consumer goods company, we replaced an off-the-shelf supply chain tool with a custom platform on Azure. Users filter a dataset of more than 40 million rows in near real time, and performance improved by 10x, more in some cases, over the tool it replaced. The speed came from the serving layer (column store indexes and memory-optimized tables in Azure SQL), fed by Azure Data Factory pipelines. Fast queries and fresh data are two separate problems.
How RUBICON helps with change data capture
We build data pipelines on Azure and Databricks for supply chain and operational analytics, and we start by agreeing how fresh each feed has to be. See our data engineering services, or our architects can review your source systems with you.
Related terms
Frequently asked questions
What are the types of change data capture?
There are three main types. Log-based CDC reads the database transaction log. Trigger-based CDC uses database triggers to write changes to a shadow table. Query-based CDC polls tables using a timestamp or version column. Log-based CDC is the usual choice for production systems because it captures deletes and puts the least load on the source.
What tools are used for CDC?
Popular options include Debezium (open source, usually with Apache Kafka), Qlik Replicate, Oracle GoldenGate, Fivetran, AWS Database Migration Service and Azure Data Factory. SQL Server and PostgreSQL have built-in CDC or logical replication. Lakehouse platforms then apply the changes with MERGE operations or declarative CDC pipelines.
Is CDC the same as real-time streaming?
Not quite. CDC captures changes from a database. Streaming moves and processes events continuously. CDC often feeds a stream, for example Debezium publishing to Kafka, but you can also apply changes in micro-batches every few minutes. Many business cases need data that is minutes old, not milliseconds, so check the real requirement first.
Can CDC be used with SAP?
Yes, but SAP needs its own approach. Options include SAP's replication tools, the Operational Data Provisioning framework and certified third-party connectors. Licensing matters, because SAP restricts some forms of direct database access. Check the options and your contract terms before you design the pipeline, not after.
Related case study

Real Time Analytics for Supply Chain Optimization
RUBICON Develops a Custom Real-Time Operation Analytics Supply Chain Management (SCM) Platform for a Global Client
More resources
