RDS offers two ways to use more than one instance: Multi-AZ, for high availability, and read replicas, for scaling reads. They solve different problems, and a common interview mistake is treating read replicas as a failover mechanism or Multi-AZ as a way to serve read traffic.
Before you start
You should understand replication and database failover. This article compares the two features.
Step-by-step walkthrough
Step 1: Multi-AZ is synchronous failover
In Multi-AZ, the primary replicates synchronously to a standby in another zone, and the standby serves no traffic. On failure, RDS promotes the standby and updates the endpoint’s DNS, so the application reconnects to the new primary. Availability improves; read capacity does not.
Step 2: Read replicas scale reads asynchronously
A read replica replicates asynchronously, so it can lag the primary. It serves read traffic, offloading reporting and read-heavy queries. Because it is asynchronous, it is not a lossless failover target, and a query on it may see slightly old data.
Step 3: Combine them for both goals
Use Multi-AZ for failover and read replicas for read scale, together if needed. The application must route reads to replicas and writes to the primary, and it must tolerate replica lag. Failover uses the Multi-AZ standby, not a replica.
Worked scenario
The two features serve different patterns.
Multi-AZ: primary -> synchronous standby; failover on failure (no read traffic)
read replica: primary -> asynchronous replica; serves reads (may lag)
both: primary (writes) -> standby (failover) + replicas (reads)Walk through the example
The standby exists to take over, so it does no work until a failure. The replica exists to read, so it takes traffic immediately but may be a little behind. Combining them gives both failover and read scaling, with routing and lag handling in the application.
Common mistake
Treating a read replica as a failover target, which risks losing recently committed writes because replication is asynchronous. Another is expecting Multi-AZ to increase read throughput, which it does not, since the standby is idle.
Verify the behavior
Fail over the Multi-AZ primary and confirm the endpoint reconnects to the new primary. Write on the primary and immediately read a replica to observe lag. Check that reads are routed to replicas and writes to the primary.
Interview exercise
Why is a read replica not a safe failover target?
Answer and reasoning
Because replication is asynchronous, so the replica may be behind the primary and may not have received the most recent commits. Promoting it would lose those writes. Multi-AZ’s standby replicates synchronously, so it holds all committed data and is the correct failover target.
Continue learning
Compare replication in Replication and connections in RDS connections. Read the AWS RDS Multi-AZ documentation and try the AWS interview questions.