The primary commits the change.
But if the application reads from a replica, how does it know whether that replica has caught up?
That is the problem replication lag tracking solves.
The Simple Mental Model
Replication can be thought of as a stream of database changes:
Primary
|
| Change 1001
| Change 1002
| Change 1003
| Change 1004
v
ReplicaThe primary may be at:
Change 1004while the replica has applied only:
Change 1002The replica is behind.
The database maintains replication state so the system can determine how far the replica has progressed.
Received ≠ Applied
A replica typically goes through stages:
Receive → Store → Apply/Replay
This distinction matters.
A replica may have received a change but not yet applied it.
For read routing, the important question is usually:
Has the required change been applied and made visible according to the database's semantics?
How Is Progress Measured?
Different databases expose replication progress differently.
PostgreSQL
Uses WAL positions and replay progress.
MySQL
Uses GTIDs and transaction execution state.
Other databases may expose:
- Log positions
- Sequence numbers
- Replay positions
- Apply progress
- Time-based lag
The mechanism differs, but the basic idea is the same:
Measure how far the replica has progressed through the primary's change stream.
Time-Based Lag
You might see:
Replica A: 20 ms
Replica B: 35 ms
Replica C: 8 secondsReplica C is clearly worth investigating.
But there is an important limitation.
"200 ms lag" does not necessarily answer whether a specific transaction is visible on the replica.
That's why time-based lag alone may not be enough for consistency-sensitive reads.
Position-Based Tracking
Consider a simplified example:
Primary:
Position = 5000
Replica:
Applied position = 4970The replica is behind.
Later:
Primary:
Position = 5000
Replica:
Applied position = 5000The replica has caught up to that point.
The actual meaning of these positions depends on the database.
Where Does Lag Tracking Happen?
A typical architecture looks like:
Primary
|
| Replication
v
Replica
|
| Replication state
v
Lag Tracker
|
| Freshness state
v
Read RouterThe database maintains replication state.
The lag tracker observes it.
The read router uses it.
For example:
Required position: 5000
Replica A: 4970 ❌
Replica B: 5000 ✅The router can now choose the appropriate replica.
Monitoring vs Routing
These are two different questions.
Monitoring asks:
Is this replica falling behind?
Routing asks:
Can this replica safely serve this read?
A system can have excellent monitoring and still have incorrect read routing if the freshness information is interpreted incorrectly.
Healthy ≠ Current
This is one of the most common mistakes.
A replica can be:
- Reachable
- Accepting queries
- Passing health checks
…and still be behind the primary.
For example:
SELECT 1;proves the database is responding.
It does not prove the latest transaction has been applied.
So production systems often track two separate signals:
Availability
and
Freshness
Why Lag Can Become a Scaling Problem
Suppose replicas normally handle most reads.
Then replication falls behind.
The router stops using stale replicas for consistency-sensitive requests.
More reads move back to the primary:
Replication slows
↓
Replicas become stale
↓
Reads move to primary
↓
Primary gets more load
↓
Read-scaling capacity decreasesSo replication lag is not just a database metric.
It can affect application scalability and correctness.
The Complete Mental Model
1. Primary commits a change
↓
2. Replication state advances
↓
3. Replica receives the change
↓
4. Replica applies/replays it
↓
5. Replica updates its progress
↓
6. Lag tracker observes the progress
↓
7. Freshness information is updated
↓
8. Read router uses that informationKey Takeaway
Replication lag tracking answers one fundamental question:
How far has this replica progressed through the primary's changes?
The database provides the replication state.
Monitoring observes it.
Routing uses it.
And that is the machinery behind reliable primary → replica read routing.
