Primary Database Failure and Replica Lag


A primary database goes down.

Writes start failing.

One replica is 500ms behind.
Another is 30 seconds behind.
A third has stopped replicating.

All three may still respond to queries.

But treating them all as simply “database unhealthy” can make the outage worse. 🚨

𝗣𝗿𝗶𝗺𝗮𝗿𝘆 𝗙𝗮𝗶𝗹𝘂𝗿𝗲

When the primary is unavailable, writes need a defined policy:

• Retry
• Fail over
• Temporarily reject writes

A healthy replica is not automatically a new primary.

The failover mechanism must establish which node is authoritative before applications send writes there.

𝗥𝗲𝗽𝗹𝗶𝗰𝗮 𝗟𝗮𝗴

A replica can be available but stale.

For example:

• Replica A → 200ms behind
• Replica B → 8s behind
• Replica C → 40s behind

They may all return successful responses.

But a read that must immediately reflect a recent write should not necessarily go to the 40-second-old replica.

𝗪𝗵𝗲𝗻 𝗟𝗮𝗴 𝗞𝗲𝗲𝗽𝘀 𝗚𝗿𝗼𝘄𝗶𝗻𝗴

100ms → 500ms → 2s → 10s → 30s

That is different from a temporary lag spike.

The replica is falling further behind because it is applying changes slower than the primary produces them.

Redirecting freshness-sensitive reads to the primary can protect consistency—but may also increase primary load. ⚖️

𝗧𝗵𝗲 𝗣𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻 𝗖𝗵𝗲𝗰𝗸

Routing decisions should consider:

• Primary availability
• Replica health
• Replication progress
• Current lag
• Read consistency requirements

The goal is not to find a node that responds.

It is to find a node safe for that particular operation.

𝗖𝗼𝗺𝗺𝗼𝗻 𝗠𝗶𝘀𝘁𝗮𝗸𝗲

Using a healthy replica as a primary fallback without verifying its replication state and authority.

𝗧𝗮𝗸𝗲𝗮𝘄𝗮𝘆

A database node being available does not mean it is current or safe for every read or write.

What’s your perspective on handling primary failure and replica lag? Share it in the comments. 👇

#Databases #DatabaseReliability #DistributedSystems #SystemDesign #BackendEngineering #SoftwareArchitecture #ReliabilityEngineering #Scalability #SoftwareEngineering #DataEngineering

Previous Post Next Post