Imagine an application handling thousands of writes per second.
A user writes at position 5000.
The next read needs a replica at least at 5000.
But the replicas are behind:
Replica A → 4970
Replica B → 4980
Replica C → 4960
Meanwhile, new writes keep arriving:
5000 → 5100 → 5200 → 5300
The important question isn't:
“Has the replica caught up?”
It's:
“Is the replica caught up enough for this read?”
A replica doesn't need to be at the latest primary position for every request. It only needs to satisfy that request's consistency requirement.
When Lag Becomes a Real Problem
If:
Write rate > Replica apply rate
the lag keeps growing.
At that point, routing alone cannot solve the problem.
Sending more reads to the primary may preserve freshness, but increases primary load.
Waiting for replicas preserves replica usage, but increases latency.
What Can You Do?
- Increase replica capacity where possible.
- Allow bounded staleness for reads that don't need the latest state.
- Send consistency-sensitive reads to the primary.
- Use bounded waiting with an explicit fallback.
- Monitor persistent lag as a capacity problem.
The Real Trade-Off
Consistency
↕
Latency
↕
Primary Load
↕
Replica Capacity
The goal isn't to make every replica perfectly current.
The goal is to give each read the consistency it actually requires.
I’ve brought these ideas together in my latest Substack deep dive:
𝗗𝗲𝗲𝗽 𝗗𝗶𝘃𝗲 𝟯: 𝗪𝗵𝗲𝗻 𝗮 𝗦𝘂𝗰𝗰𝗲𝘀𝘀𝗳𝘂𝗹 𝗪𝗿𝗶𝘁𝗲 𝗦𝘁𝗶𝗹𝗹 𝗥𝗲𝘁𝘂𝗿𝗻𝘀 𝗢𝗹𝗱 𝗗𝗮𝘁𝗮
Key Takeaway
Persistent replica lag is ultimately a capacity and consistency problem—not a routing problem.
Good production design makes the trade-off explicit.
Common Questions / FAQ
Can routing fix replica lag?
No. It can avoid stale replicas, but cannot increase replication capacity.
Should every read go to the primary?
No. Only reads requiring stronger freshness may need that path.
Is stale data always bad?
No. Some workloads can safely tolerate bounded staleness.
When is lag a capacity problem?
When replicas consistently apply changes slower than new changes arrive.
If you've dealt with persistent replica lag, share your technical perspective in the comments. Follow and subscribe for more practical distributed-systems content.
