An ECS service behind an ALB has its target group health check pointed at `/actuator/health`, which includes the database. During a 10-minute RDS incident the dashboard shows 0 healthy targets, yet some requests still reach the tasks, and ECS keeps replacing tasks that never finish starting. Explain the health check's semantics on an ALB, the fail-open behaviour, and what you would change.
An ALB health check does not restart anything: a target that fails it is removed from rotation, and returns when it passes again. So the check should be the readiness endpoint, and including the database there is correct in principle: no traffic to a task that cannot serve. Two AWS behaviours explain the rest. When every target in a group is unhealthy, the ALB fails open and routes to all of them anyway, which is why requests still arrived at '0 healthy'. And ECS adds a restart on top: a task whose target stays unhealthy is stopped and replaced, so a shared-dependency outage turns into churn, and new tasks that need longer than the health check grace period to start are killed before they are ready. Point the check at /actuator/health/readiness, set a grace period that covers a Spring Boot start, and read '0 healthy targets' during a database incident as the database, not the tasks.