Your team runs RDS PostgreSQL Multi-AZ and treats it as 'automatic HA, nothing for the application to do'. During a failover drill, one service recovers in about a minute while another keeps failing queries for much longer, until its tasks are restarted. Explain what Multi-AZ replication and failover actually are, why the two services behaved differently, and what the application must configure.
Multi-AZ keeps a standby in another availability zone and replicates every write to it synchronously: the primary does not acknowledge a commit until the standby has it, so a failover cannot lose an acknowledged write (and writes are a little slower for it). On failure, RDS promotes the standby and changes the DNS CNAME the endpoint name resolves to. Nothing moves at the connection level, so failover is a reconnect: every client must notice its connections are dead, resolve the name again and connect to the new address. The service that recovered did that. The other one kept using connections to the old primary, or kept a cached DNS answer for the old address, until a restart cleared both. The application's job: a JVM DNS cache that expires quickly, timeouts that fail dead connections fast, a pool that validates and replaces connections, and retries, all proven in a drill.