A highly available AWS application must keep serving useful requests when a component fails. Running two servers is only the beginning: routing, application state, database behavior, capacity and recovery procedures must all survive the same failure.

This guide develops a reference design for a small transactional web application. It is an engineering walkthrough, not a report of production results. The examples assume one AWS Region, two Availability Zones and a relational database. The aim is to make every redundancy decision testable before adding more infrastructure.

Start with recovery requirements

Write down the user journey that must remain available. For an appointment service, browsing availability and confirming a booking have different failure consequences. A stale listing may be tolerable for a short interval; accepting two bookings for the same slot is a correctness failure. Define success from the user's perspective rather than from the health of an individual virtual machine.

Set a recovery time objective (RTO), the maximum acceptable restoration time, and a recovery point objective (RPO), the acceptable amount of lost data expressed in time. These are targets to validate, not capabilities created by writing them in a design document. If the team requires recovery within minutes but has never restored a backup, the requirement is currently an assumption.

Also record the expected failure scope. A stopped application process, an unavailable zone, accidental deletion and a regional outage need different mechanisms. Treating all four as one problem usually produces a costly architecture with unclear behavior. For this example, the availability design addresses instance and zone failures; backup recovery addresses data mistakes separately.

A practical two-zone architecture

A reasonable starting point is an internet-facing Application Load Balancer with application instances in private subnets across two zones. An Auto Scaling group maintains the desired application capacity. A relational database uses a suitable Multi-AZ configuration, and uploaded objects live outside the application machines. This follows the failure-isolation principle described in the AWS Reliability Pillar.

Reference architecture and its failure responsibilities
LayerResponsibilityQuestion to test
Load balancerRoute to healthy application targetsCan a failed target receive new requests?
Application tierReplace instances and serve from either zoneCan the surviving zone carry peak traffic?
DatabasePreserve committed state and recover serviceDo clients reconnect after failover?
Object storageKeep uploads independent of instance lifetimeCan a replacement instance retrieve existing files?
Recovery processRestore from corruption or deletionHas a restore been verified end to end?

Keep the network diagram honest about outbound dependencies. A private application may still need an identity provider, package registry or third-party API. Draw those paths and identify their failure domains. A design with duplicated application servers can still depend on one egress component or one unavailable external service.

Remove hidden state from application instances

Imagine that a user signs in through instance A and the next request reaches instance B. If authentication state exists only in A's memory, the second request may appear unauthenticated. Sticky sessions can conceal this problem during normal operation, but they do not recover state when A disappears. Choose an explicit session design and test requests across different instances.

Uploaded files, generated exports and scheduled jobs create similar traps. A file written to a local directory vanishes from the user's perspective when traffic reaches another machine. A scheduled task started by every replica may run twice. Move durable files to appropriate shared storage and give background work a clear ownership or deduplication mechanism.

Idempotency matters during recovery. If a client retries a booking request after a timeout, the application should identify whether the original operation committed. A stable operation identifier and a database uniqueness constraint can protect the business action. Merely retrying every failed HTTP request can turn a temporary outage into duplicate side effects.

Understand what database failover actually provides

For a traditional Amazon RDS Multi-AZ DB instance deployment, the standby supports failover and does not serve application read traffic. That differs from other RDS deployment types and read-replica arrangements. Check the exact engine and deployment option rather than assuming all configurations called “Multi-AZ” behave identically. The RDS documentation explains this distinction.

At the application boundary, assume existing connections can fail. Test connection-pool recovery, connection timeouts and bounded retries. The application should reconnect through the database endpoint rather than pinning a resolved address indefinitely. Retrying a read is different from retrying a transaction whose commit result is unknown; the latter needs a business-level reconciliation strategy.

Replication also does not replace backup. An unwanted update can be replicated just as faithfully as a correct one. A recovery exercise should restore to an isolated environment, check the recovered records and prove that the application can use them. Recording only that a backup job completed leaves the most important part untested.

Reserve capacity for the degraded state

Consider an illustrative load test where each application instance can sustain 200 requests per second within the latency target. Two instances per zone provide a nominal total of 800 requests per second. Losing one zone leaves 400. If normal peak traffic is 600, the architecture is distributed but cannot meet its target after that loss. These figures are hypothetical; replace them with measurements from the actual workload.

Auto Scaling can help restore capacity, but replacement instances take time to start and may face capacity constraints. Budget enough headroom to cover the interval between failure and recovery. Account for database connections too: adding application instances can overload the database if every instance opens a large pool.

Health checks deserve their own design. A process being alive does not prove it can serve a booking. Conversely, checking every remote dependency too aggressively can eject all application instances during a brief dependency disturbance. Define what makes a target ready, choose thresholds deliberately and verify how the load balancer behaves when targets become unhealthy.

Run a failure drill with acceptance criteria

Use a controlled environment and realistic traffic. First terminate one application instance and measure user-visible errors, replacement time and queue growth. Then exercise loss of the application capacity in one zone. Test database failover separately so that the observed behavior can be attributed to a specific event rather than several simultaneous changes.

  1. Record the baseline: throughput, error ratio, latency percentiles and connection counts.
  2. Declare the failure being simulated and the exact stop condition for the exercise.
  3. Apply the failure, observe recovery and keep a timestamped event log.
  4. Check business records for missing or duplicated operations.
  5. Restore normal capacity and confirm the service has returned to baseline.
  6. Turn each unexpected result into a tracked correction with an owner.

A drill passes when the application meets its stated objectives, not when every dashboard eventually becomes green. Preserve the evidence with the infrastructure configuration and repeat after changes that affect routing, persistence or capacity. Pair this work with request-level observability so recovery can be measured from the customer's side.

Choose the smallest design that meets the requirement

Multi-region operation adds another set of decisions: data replication, write ownership, routing, deployment coordination and operating cost. Introduce it when a defined requirement justifies those obligations. An untested regional failover plan is weaker than a well-understood single-region recovery process whose limitations are documented.

The useful output of an availability review is a set of proven statements: which failures the service tolerates, how long recovery takes, what data may be lost and how operators know recovery succeeded. Those statements make the architecture reviewable. They also make cost discussions more precise, because resilience capacity has an explicit purpose.