reliabilityincident.io ↗

We turned off Pub/Sub and nobody noticed

incident.io moved ~240M messages a day onto a dual-broker load balancer topic by topic, then deleted its NATS cluster in production to prove the failover — and finally switched Pub/Sub off.

incident.io · · 14m read

Title card reading “We turned off Pub/Sub and nobody noticed”, with rows of alert icons on a branching line

Why we picked it · the editor's summary

The mechanism is a load-balancing layer above the brokers: every publish goes to one of two, split evenly, with a circuit breaker that removes a failing broker and a delay-based consumer schedule that drains whichever side has more waiting. That layer is what let incident.io move topic by topic instead of all at once, with 99.99% availability as the constraint it could not spend. The proof ran in production: switch NATS off and watch nothing happen, then do the same to Pub/Sub. What the post does not give is the cost of running two brokers, the latency the layer adds, or how long the whole thing took. Its two recommendations are plain: write your broker requirements down before you pick one, and exercise failover on a schedule so it is a practice rather than a hope.

Read on incident.io ↗

Comments

Loading comments…