← Back to Field Notes
Lane 02 / 5 deployments Routing

Circuit breakers for email: stop the bleeding first

Ilya Kade · July 19, 2026

You would never let a failing database take down your app. Yet most companies let a failing vendor take down their mail.

A breaker watches hard-bounce rate per vendor and trips past threshold — re-routing traffic before the queue backs up.

In five deployments, the median breaker saved eleven minutes of degraded delivery per incident. Diagnosis comes after; stopping the bleeding comes first.

The pattern is the same one you already run on your databases: watch a health signal, open the circuit before the failure cascades, half-close it to test recovery, close it when the signal heals. The only difference is the signal. For a vendor connection, hard-bounce rate per domain is the vitals monitor — it moves before queues back up, and it does not lie about why.

Tuning matters more than the metaphor. A single global threshold trips on Tuesday’s newsletter and sleeps through a real storm; per-domain thresholds cut the false positives by half while keeping the breaker sharp where it counts. An eleven-minute median saving per incident does not sound dramatic until you multiply it by every incident you no longer notice — at our volume, that is the difference between a 97% week and a 99.2% one.

Keep reading