Deployment strategieshard5-8 years

A team wants to send exactly 5% of production traffic to a canary release and watch its error rate before widening. Someone proposes adding a second DNS record for the service's hostname pointing at the canary, expecting roughly 5% of lookups to resolve there. Why doesn't this hold anywhere near 5%, and what would actually work?

DNS answers are cached — by the resolver and by the client — for the record's TTL, anywhere from seconds to hours, and some clients and corporate resolvers ignore a short TTL entirely. Changing which fraction of DNS records point at the canary doesn't redirect existing traffic at all (a client that already resolved and cached the old record keeps using it until that cache expires), and it doesn't even redirect new connections from every client at the same moment, because different resolvers around the world have their own independent, uncoordinated cache states. Worse, once a given client's resolver does cache the canary's IP, that client is pinned to the canary for the rest of its own TTL — there's no way to route just 5% of that one client's subsequent requests elsewhere. A weighted load balancer target group does this correctly because the routing decision happens on the balancer itself, per new connection, against live configured weights — a change from 5% to 25% takes effect for new connections immediately and uniformly, with no cache anywhere in the path working against it.

The lesson behind it →
More on Deployment strategies