Gatewayhard5-8 years

Three years into a platform, the shared API gateway has sixty-one custom filters, one of which runs a blocking JDBC query against the orders table to check whether an order is cancellable before proxying the request. What's wrong here beyond "it's slow", and how should this actually be structured?

There are two separate problems in one filter. First, a blocking JDBC call inside a filter on the gateway's non-blocking runtime stalls one of a small, fixed number of event-loop threads shared across every route — under load, the gateway's latency degrades for requests that never touch orders at all, not just the ones hitting this filter. Second, reading the orders table directly couples the gateway's own deploy to the orders schema: a column rename the orders team makes for their own reason now breaks checkout for everyone, and they can't ship it without coordinating with the platform team. The fix moves the cancellable check into the orders service itself, which already owns that data and can express the result as its own 409 response; the gateway goes back to routing, auth, limits and edge resilience. If genuine per-client response composition is needed, that belongs in a backend-for-frontend owned by the client's own team, not the shared gateway.

The lesson behind it →
More on Gateway