Operations

System health

Components, queues and workers — the pipes behind the pretty dashboards.

All systems nominal-ish
SMTP accept rate
/min
Queue depth
▲ webhook worker lagging
Median routing latency
ms
LLM spend today
$

Components

SMTP ingress
3 nodes · HAProxy fronted
Healthy
Vendor router
rule engine + circuit breakers
Healthy
Webhook ingest
SendGrid / SES / TEM endpoints
Degraded
Diagnosis pipeline
bounce clustering + Anthropic LLM
Healthy
Postgres
primary + 1 replica
Healthy

Queues

outbound.send128 / 5,000
webhook.process lagging1,240 / 5,000
diagnosis.cluster36 / 1,000
llm.inference4 / 100

Webhook worker saturation — 6 h