← Back to Field Notes
Lane 03 / 3 evaluations Incident Response

Letting an LLM read the bounce queue

Maren Voss · June 11, 2026

LLMs do not get tired, but they do drift. Prompts change, vendor behaviors change, receiver policies change around them.

The supervision playbook transfers surprisingly well: score every diagnosis, alert on outliers, review the receipts before approval.

What does not transfer is cadence. Human triage can be weekly; bounce storms are measured in minutes, not weeks.

Three evaluations in, the pattern we trust looks like this: the model reads the bounce queue continuously, proposes a diagnosis with the receipts attached — headers, cluster, trend — and a human approves or rejects with one tap, straight from the incident channel. The full audit trail is preserved either way. The 03:14 phone tap from our failover postmortem is this loop in production, not in a slide deck.

The honest limit is scope. The LLM is very good at reading what a rejection says and reasonably good at guessing why; it is not accountable, and it should never be. Accountability stays with the on-call engineer holding the phone. That division — machine reads, human decides — is what makes the bounce queue safe to hand over.

Keep reading