You inherit an alarm set for a Spring Boot service behind an ALB and RDS: `TargetResponseTime` Average above 1 s, `HTTPCode_Target_5XX_Count` above 10, CPU above 80%. It pages at night for single errors, missed a 20-minute outage in which the tasks crashed and traffic stopped, and never caught a slow p99. As the owner of on-call quality, redesign it and explain the CloudWatch mechanics each change depends on.
Each failure maps to a CloudWatch mechanism. A metric is not a live reading; it is an aggregate per period (one minute by default for the ALB's metrics), read through a statistic you choose. Average hides a slow tail inside a fast minute, so latency alarms use a percentile (p99) against the SLO. A raw count of 5xx pages for one error on a quiet night and misses a real error rate at peak, so errors are a ratio (metric math over RequestCount) sustained for several periods. The outage was missed because a period with no requests publishes no data point: the alarm went to INSUFFICIENT_DATA, which is not ALARM, and nobody had set treat missing data. So: alarm on symptoms (error ratio, p99 latency, HealthyHostCount, the load balancer's own 5xx), choose missing-data handling per alarm, drop CPU as a pager, compose related alarms so an incident pages once, and put the runbook in each alarm's description.