CloudWatch defaults exist but nobody reads the alarms
AWS already emits metrics and logs. The gap is routing the right signals to a human. Metrics, alarms, logs, and canaries, in the order you should build them.
Most teams think they need more monitoring. What they actually have is data nobody acts on. AWS already emits metrics for Lambda, API Gateway, and Aurora. Logs land in CloudWatch by default. Alarms exist in the console. And on-call still hears about outages from users first. I have seen this on a 100+ Lambda production platform and on smaller stacks that looked healthy until someone checked the wrong graph. The fix is not enabling every CloudWatch feature. It is building observability in layers, in the right order, so the right signal reaches a human who can act.
Start with metrics that answer one question
A metric is a time-ordered set of data points. Each one lives in a namespace like AWS/Lambda, has a name like Errors, and is identified by dimensions such as FunctionName. That is all you need to start.
Pick metrics that answer a real question, not every metric AWS publishes.
- API Gateway: 5xx errors and latency.
- Lambda: errors, throttles, and duration.
- Aurora or RDS: CPU, connections, and replica lag.
Use p99 for latency, not Average. Averages hide the slow requests that users complain about. AWS documents CloudWatch metrics concepts if you want the full reference.
Basic EC2 monitoring gives you CPU and network. It does not give you memory, disk space, or application logs. If you need those, install the CloudWatch unified agent on the instance. The agent can also listen for StatsD or collectd if your app already emits metrics that way.
For custom business metrics, you have two paths. Call PutMetricData from application code when you want direct publishing. Use Embedded Metric Format in JSON logs when you already write structured logs from Lambda or containers. CloudWatch extracts EMF metrics at ingest, so you skip extra API calls.
Build alarms that reach a human
Metrics on a graph are not observability. An alarm that pages someone is.
Every alarm sits in one of three states: OK, ALARM, or INSUFFICIENT_DATA. That third state is where silent failures hide. If your alarm never has enough data points to evaluate, it never enters ALARM, and nobody gets paged.
Configure evaluation as M out of N. For example, 3 of the last 5 periods must breach before the alarm fires. That cuts noise from a single spike without hiding a real problem.
Treat missing data deliberately. A heartbeat metric that goes quiet should probably be breaching. An error rate with no errors in the window should probably be notBreaching. The default missing behavior is not always what you want.
Send the alarm to a place a human will see. SNS wired to email or chat is the minimum. AWS walks through creating an alarm that sends email.
Alarm actions fire on state transitions, not on every evaluation period. If you need multi-step remediation, route CloudWatch Alarm State Change events through EventBridge to Lambda or Step Functions. Use composite alarms when several related alarms all page for the same incident. Use anomaly detection when a fixed threshold would be wrong every week but the metric still needs a guardrail.
Turn logs into signals, not a storage bill
Lambda writes to /aws/lambda/<function> automatically. API Gateway, ECS, and WAF can too. That does not mean your logs are useful yet.
Logs live in groups and streams. A log group holds retention and permissions. A log stream is the sequence of events from one source. Set a retention period on every log group. Indefinite retention is a cost trap that buys you nothing until an auditor asks for six months of history.
Use Logs Insights for ad hoc investigation during an incident. Run a query, find the stack trace, move on. AWS covers Logs Insights query syntax.
When you need an alert on a log pattern, chain three pieces together.
- Create a metric filter that matches the pattern, such as
ERROR. - Publish a custom metric from the filter.
- Put an alarm on that metric and send it to SNS.
That is the standard path for "tell me when errors spike in logs."
For moving log data elsewhere, pick the right tool. A subscription filter streams matching events in near real time to Lambda, Kinesis Data Streams, or Data Firehose. An export task is a batch job over a time range that lands in S3. Use subscriptions for live pipelines. Use exports for one-off compliance archives.
When the question is "which endpoint or IP is causing most of the 5xx errors," use Contributor Insights on access logs. That is a different question from "did we breach a threshold," and it deserves a different tool.
Test the front door before users do
Metrics and logs tell you what broke after traffic arrives. They do not tell you that checkout stopped working at 3am when nobody was looking.
CloudWatch Synthetics runs canaries on a schedule. A canary is a script that hits your endpoints, validates responses, and publishes metrics to the CloudWatchSynthetics namespace. Pick the blueprint by depth.
- Heartbeat monitoring for a URL that should load.
- API canary for REST calls with status code and body checks.
- Multi Checks for several simple HTTPS endpoints in one canary without a headless browser.
Route 53 health checks answer a narrower question: is this endpoint reachable from DNS? Canaries answer a broader one: does the workflow still work? CloudWatch RUM measures real browser sessions. Use RUM when client-side performance matters. Use canaries when you need proactive synthetic checks. AWS documents CloudWatch Synthetics canaries.
Alarm on SuccessPercent or Duration from the canary metrics. Schedule the scale-up or deploy check so the canary runs after every release, which pairs well with gated deploy practices.
Put it on one dashboard, and stop there
Dashboards visualize data. They do not trigger remediation. That is what alarms are for.
Build one dashboard around services and user outcomes, not one widget per instance. A layout that works on most stacks looks like this.
- Service health: availability, request count, p99 latency, errors.
- Infrastructure: CPU, memory where the agent is installed, throttles.
- Dependencies: database latency, queue depth, downstream errors.
- Active alarms and a Logs Insights widget for recent failures.
Codify the dashboard in Terraform or CloudFormation with AWS::CloudWatch::Dashboard. Manual console dashboards drift the moment someone copies one for staging.
Know what CloudWatch is not
CloudWatch observes operational behavior. It does not audit who changed your infrastructure.
- Who changed this IAM policy? CloudTrail.
- Route an alarm state change to Step Functions? EventBridge.
- Continuously export all metrics to S3 or a third-party tool? Metric Streams.
- Record every AWS API call for compliance? CloudTrail, optionally forwarded to CloudWatch Logs, but CloudTrail is the source of truth.
Confusing these shows up in exam questions and in production incidents. CloudWatch tells you the database CPU spiked. CloudTrail tells you who changed the parameter group an hour earlier.
The takeaway
Defaults are not observability. Pick metrics that answer real questions, wire alarms that reach a human, set log retention before the bill surprises you, and put canaries on the paths users actually take. One dashboard gives incident context. Everything else is noise until it pages someone who can act. For the uptime patterns that sit on top of this stack, read meeting a 98% SLA on AWS. For Lambda-specific signals like Init Duration, read taming Lambda cold starts.
Frequently asked questions
Why do my alarms stay in INSUFFICIENT_DATA?
The alarm does not have enough data points in the evaluation window to decide. Check that the metric is actually publishing, that the period matches the metric resolution, and that dimensions match exactly. An alarm on a metric that never arrives will sit in INSUFFICIENT_DATA forever and never page anyone.
Should I use CloudWatch or CloudTrail to audit IAM changes?
CloudTrail. CloudWatch metrics tell you how a resource is behaving. CloudTrail records who called which AWS API, from where, and when. Use CloudTrail for audit and compliance questions.
How do I alert when ERROR appears in logs?
Create a metric filter on the log group that matches the pattern, publish a custom metric from the filter, and put a CloudWatch alarm on that metric with an SNS action. That turns a log line into a page.
When do I need the CloudWatch agent on EC2?
When basic EC2 metrics are not enough. The default set covers CPU and network, not memory, disk utilization, or custom application logs. Install the unified agent when those gaps matter for your alarms or dashboards.
Canaries vs Route 53 health checks: which do I pick?
Route 53 health checks when you only need to know if an endpoint is reachable and you want to tie that to DNS failover or routing. CloudWatch Synthetics canaries when you need to validate an API response, run a multi-step workflow, or capture screenshots and HAR files on a schedule.