How Replication Lag Tracking Actually Works


A user updates their profile.

The primary commits the change.

But if the application reads from a replica, how does it know whether that replica has caught up?

That is the problem replication lag tracking solves.

The Simple Mental Model

Replication can be thought of as a stream of database changes:

Primary
   |
   | Change 1001
   | Change 1002
   | Change 1003
   | Change 1004
   v
Replica

The primary may be at:

Change 1004

while the replica has applied only:

Change 1002

The replica is behind.

The database maintains replication state so the system can determine how far the replica has progressed.

Received ≠ Applied

A replica typically goes through stages:

Receive → Store → Apply/Replay

This distinction matters.

A replica may have received a change but not yet applied it.

For read routing, the important question is usually:

Has the required change been applied and made visible according to the database's semantics?

How Is Progress Measured?

Different databases expose replication progress differently.

PostgreSQL

Uses WAL positions and replay progress.

MySQL

Uses GTIDs and transaction execution state.

Other databases may expose:

  • Log positions
  • Sequence numbers
  • Replay positions
  • Apply progress
  • Time-based lag

The mechanism differs, but the basic idea is the same:

Measure how far the replica has progressed through the primary's change stream.

Time-Based Lag

You might see:

Replica A: 20 ms
Replica B: 35 ms
Replica C: 8 seconds

Replica C is clearly worth investigating.

But there is an important limitation.

"200 ms lag" does not necessarily answer whether a specific transaction is visible on the replica.

That's why time-based lag alone may not be enough for consistency-sensitive reads.

Position-Based Tracking

Consider a simplified example:

Primary:
Position = 5000

Replica:
Applied position = 4970

The replica is behind.

Later:

Primary:
Position = 5000

Replica:
Applied position = 5000

The replica has caught up to that point.

The actual meaning of these positions depends on the database.

Where Does Lag Tracking Happen?

A typical architecture looks like:

Primary
   |
   | Replication
   v
Replica
   |
   | Replication state
   v
Lag Tracker
   |
   | Freshness state
   v
Read Router

The database maintains replication state.

The lag tracker observes it.

The read router uses it.

For example:

Required position: 5000

Replica A: 4970 ❌
Replica B: 5000 ✅

The router can now choose the appropriate replica.

Monitoring vs Routing

These are two different questions.

Monitoring asks:

Is this replica falling behind?

Routing asks:

Can this replica safely serve this read?

A system can have excellent monitoring and still have incorrect read routing if the freshness information is interpreted incorrectly.

Healthy ≠ Current

This is one of the most common mistakes.

A replica can be:

  • Reachable
  • Accepting queries
  • Passing health checks

…and still be behind the primary.

For example:

SELECT 1;

proves the database is responding.

It does not prove the latest transaction has been applied.

So production systems often track two separate signals:

Availability

and

Freshness

Why Lag Can Become a Scaling Problem

Suppose replicas normally handle most reads.

Then replication falls behind.

The router stops using stale replicas for consistency-sensitive requests.

More reads move back to the primary:

Replication slows
       ↓
Replicas become stale
       ↓
Reads move to primary
       ↓
Primary gets more load
       ↓
Read-scaling capacity decreases

So replication lag is not just a database metric.

It can affect application scalability and correctness.

The Complete Mental Model

1. Primary commits a change
        ↓
2. Replication state advances
        ↓
3. Replica receives the change
        ↓
4. Replica applies/replays it
        ↓
5. Replica updates its progress
        ↓
6. Lag tracker observes the progress
        ↓
7. Freshness information is updated
        ↓
8. Read router uses that information

Key Takeaway

Replication lag tracking answers one fundamental question:

How far has this replica progressed through the primary's changes?

The database provides the replication state.

Monitoring observes it.

Routing uses it.

And that is the machinery behind reliable primary → replica read routing.

Previous Post Next Post