Three requirements land at once for a public API that sits behind an ALB: a partner bank must allow-list a fixed set of IP addresses for your endpoint, the service's per-client rate limiter keys on the connecting IP and is currently limiting 'everyone at once', and a new gRPC API needs to be exposed. Explain the architectural difference between an ALB and an NLB that decides each requirement, and what you would build.
An ALB is a layer-7 proxy, run as a fleet of nodes per availability zone that AWS scales invisibly. It terminates the client's TCP connection and TLS, parses the HTTP request, and opens a second, separate connection to the target. Two consequences: it has no fixed IP (only a DNS name whose addresses change as the fleet changes), and the target sees the ALB node's address as the source, with the client's in X-Forwarded-For. That explains the rate limiter: it keys on the ALB. An NLB works at layer 4, forwards packets, has one static address per zone (an Elastic IP can be attached) and can preserve the client's source IP. So: fix the limiter to use the client address from the forwarded headers, correctly trusted; give the partner static addresses by putting an NLB (or another fixed-address front door) in front of the ALB rather than giving up HTTP routing; and serve gRPC through the ALB's gRPC target groups, or through an NLB when the application must terminate the connection itself.