5 min read

Meet a 98% uptime SLA on AWS, and beat it

What 98% uptime really means, and the AWS building blocks that get you there. Managed multi-AZ services, health checks, alarms, and a recovery plan you have tested.

AWS reliability and incident responseAWSReliabilityServerlessDevOps

An SLA is a promise with a number attached. Before you commit to one, you should know what the number costs you in real downtime, and which parts of your stack put that promise at risk. I have run a multi-tenant platform on AWS to an uptime target, and the good news is this. If you lean on managed services and you can tell when you are down, 98% is a floor you clear easily. Here is how you get there, and why the same work pushes you well past it.

What 98% actually allows

Do the math before you promise anything. Uptime percentages translate to real time.

  1. 98% allows about 14.6 hours of downtime a month, or roughly 7.3 days a year.
  2. 99% allows about 7.2 hours a month.
  3. 99.9%, the number most managed AWS services target, allows about 43 minutes a month.

So 98% is a generous target. That is the point. You do not need heroics to hit it. You need an architecture that fails rarely, recovers quickly, and tells you the moment something breaks. The AWS Well-Architected reliability pillar is the reference worth reading alongside this.

Start with managed, multi-AZ services

The fastest way to raise uptime is to stop running the parts that fail. Managed services carry their own high availability, spread across Availability Zones, so a single hardware or zone failure does not take you down.

  1. Lambda and API Gateway run across zones by default. You do not manage servers, so there are no servers to lose.
  2. S3 and DynamoDB are designed for high durability and availability out of the box.
  3. Aurora runs with automatic failover to a standby in another zone. AWS documents how Aurora high availability works.

A serverless stack inherits a lot of reliability for free. That is most of the battle.

Know the moment you are down

You cannot meet an SLA you cannot measure. Two layers tell you the truth.

Health checks from outside

Run Route 53 health checks against your public endpoints. Check the frontend at its root and the backend at a real health route, over HTTPS, every 30 seconds, and only treat it as down after a few consecutive failures so a single blip does not page you. AWS covers creating health checks.

Alarms from inside

Health checks tell you the front door is shut. Alarms tell you why. Watch the signals that move before an outage.

  1. API Gateway 5xx errors and latency.
  2. Lambda errors and throttles.
  3. Database CPU and connection count.
  4. WAF blocked requests, which can mean an attack in progress.

Send every alarm to a place a human will see, such as an SNS topic wired to email or chat. Treat missing data carefully so maintenance windows do not flood you with false alarms. AWS shows how to build an alarm that notifies you.

Design so a small failure stays small

Not every failure should take down the page. Isolate the non-critical parts so they fail quietly. On this site, the view counter and comments call a separate API, and if that API is unavailable the page still renders. The reader never notices. You get this for free when you separate static content from dynamic features, which is one reason I host the site on S3 and CloudFront.

Have a recovery plan you have tested

Uptime is not only about avoiding failure. It is about how fast you recover. A plan you have never run is a guess.

  1. Replicate critical data to a second region so a regional problem is survivable.
  2. Write runbooks for the failures you expect, with the exact commands.
  3. Practice a restore. Backups you have never restored are not backups yet.

Ship changes without breaking the promise

Most outages are self-inflicted, caused by a deploy. Reduce that risk. Require an approval gate for production, run health checks right after a release, and keep rollback one step away. I use short-lived credentials and a gated pipeline, which I cover in deploying to AWS with OIDC.

The takeaway

98% is a floor, not a ceiling. Managed multi-AZ services, health checks and alarms that reach a human, graceful degradation, and a tested recovery plan will clear it comfortably. The same building blocks are what carry you to 99.9%. Pair this with safe deploy practices so your own releases do not become the outage.

Frequently asked questions

How much downtime does a 98% SLA allow?

About 14.6 hours a month, or roughly 7.3 days a year. It is a generous target, which is why a well-built AWS stack clears it without much strain.

Is serverless more reliable than servers I manage?

For most workloads, yes. Managed services like Lambda, API Gateway, S3, and Aurora run across Availability Zones and handle failover for you, so there is less that you can lose or misconfigure.

What single change raises uptime the most?

Knowing when you are down. Health checks from outside plus alarms on 5xx, latency, and errors mean you react in minutes instead of hearing about an outage from a user.

Do I need multi-region to hit 98%?

No. Multi-AZ managed services cover the common failures. Multi-region matters when you target 99.9% and above, or when you must survive a full regional outage.

How do I stop my own deploys from causing outages?

Add an approval gate for production, run health checks after each release, and keep rollback easy. Most outages come from changes, so make changes safe.

Comments

    Leave a comment