Operations · Monitoring
Alert rules
Get paged before deliverability burns — thresholds on bounce rates, queues, circuits and vendor health, routed to email, Slack or PagerDuty.
Active rules
across 3 channels
Fired (7 d)
both auto-resolved
Silenced
maintenance window
MTTR
min
▼ 6 min vs last month
Rules
| Rule | Condition | Severity | Notify | Status | |
|---|---|---|---|---|---|
Bounce rate spike any vendor, 15 min window |
bounce_rate > 5% for 15m | Critical | PagerDuty · #mailops-ops | ||
Complaint rate high spam complaints, daily |
complaint_rate > 0.1% / day | Warning | Slack #mailops-ops | ||
Webhook queue lagging event processing backlog |
queue_depth > 1000 for 10m | Critical | PagerDuty · email on-call | ||
Circuit open too long vendor failover running |
circuit_open > 30m | Warning | Slack #mailops-ops | ||
Vendor API errors 5xx from any vendor |
vendor_5xx > 2% for 5m | Warning | Slack #mailops-ops | ||
Diagnosis backlog LLM queue awaiting review |
pending_diagnoses > 50 | Info | Email digest |
Notification channels
S
Connected
Slack · #mailops-ops
Warnings + critical
PD
Connected
PagerDuty · MailOps on-call
Critical only
@
Default
Email · on-call rota
Info + digest
Recent alerts
Webhook queue lagging Critical
Depth hit 1,240 after a SendGrid event burst — autoscaling drained it.
Fired Aug 11 · 14:02 · auto-resolved in 22 min · paged on-call
Vendor API errors Warning
Scaleway TEM 5xx at 3.1% for 7 min — circuit opened, failover to SES.
Fired Jul 28 · 09:44 · auto-resolved in 41 min · posted to Slack
No open incidents
Everything below threshold for 16 days straight.
Open system health →