Integration glossary

Change data capture (CDC)

Change data capture is a way of noticing every insert, update and delete in a database as it happens and passing only those changes on to the systems that need them.

Why it matters

Without change data capture, keeping two systems aligned means copying whole tables on a schedule. A nightly load moves far more data than actually changed, takes longer as the tables grow, and leaves every other system working from yesterday's figures until the next run.

Change data capture moves only the rows that changed, shortly after they changed. Reports, portals and connected applications stay close to the source, and the load on the source database stays small because nothing is read twice.

How change data capture works

Databases already keep a record of their own changes for recovery and replication, and change data capture reads that record instead of querying the tables over and over. SQL Server has built-in change tracking, MySQL writes every change to its binary log, and PostgreSQL can publish changes through logical replication.

A capture process turns that record into a stream of events: this row was inserted, this one was updated, that one was removed. Each event carries the key of the row, so the receiving side knows exactly which record to touch.

What a capture process has to get right

Capturing changes reliably is harder than it first looks. A good capture process handles a few situations that a simple copy job never meets.

  • It remembers where it stopped, so a restart carries on from the right place without missing or repeating a change.

  • It copes with the source table changing shape, such as a new column added by a vendor upgrade.

  • It keeps changes to the same record in the order they happened.

  • It slows down when the receiving system is busy, so a burst of changes never overwhelms it.

Change data capture in STRAX

STRAX uses change capture on SQL Server, MySQL and PostgreSQL, so updates flow into its hub as they happen and from there to every connected system that needs them. The capture is generated from the mappings you approve, and synchronisation is rate-limited per system.

The hub is additive. It keeps its full record even when a source system removes one, so history is never lost when a row disappears at the source.

See it on your own systems.

A demo takes about an hour and shows these ideas working on systems like yours.