Alerts
Pages that mean it. Silence that earns trust.
Alert fatigue is a design failure, not a fact of life. TraceLoop alerts on SLO burn rate and scored anomalies, attaches the full investigative context to every page, and tells you weekly which rules are still wasting your engineers' sleep.
Alert Rules
From error budgets to escalation, in one rule editor
Fictional workspace, real interface: the checkout-availability SLO, its burn-rate windows, and exactly who gets paged when the budget burns too fast.
slo · checkout-availability
objective 99.95% · 30-day window · multi-burn-rate
- PAGE 1h burn ≥ 14.4× · fast budget drain → PagerDuty SEV1
- TICKET 6h burn ≥ 6× · slow drain → Jira, business hours
- MUTE maintenance window 02:00–04:00 Sun · auto-suppressed
detector board
anomaly & threshold detectors · scoring every 60s
- FIRING
checkout-svc · p99 latency
score 0.91 - WATCHING
postgres-01 · deadlock rate
score 0.82 - WATCHING
redis-cache · evictions/sec
score 0.64 - NORMAL
kafka · consumer lag order.created
score 0.21 - NORMAL
api-gateway · rps
score 0.18
Incident Investigation
The timeline writes itself
When INC-4417 fired, TraceLoop had already stitched the deploy, the error logs, the similar traces and the alert sequence into one narrative. The humans just made decisions.
incident timeline
INC-4417 · elevated 5xx on checkout-svc · SEV1
-
14:28:04 · alert
burn-rate alert fired · checkout-availability 1h window · 14.2× burn
-
14:28:04 · auto
INC-4417 opened · SEV1 · paged on-call (A. Sharma) via PagerDuty
-
14:28:06 · auto
timeline auto-populated · 187 error logs · 4 similar traces · deploy v2.41.3 (14:17) flagged as suspect
-
14:29:31 · engineer
A. Sharma acknowledged · started war room in #inc-4417
-
14:33:12 · engineer
root cause hypothesis: deadlock loop introduced by order-write refactor
-
14:36:40 · engineer
rollback v2.41.3 → v2.41.2 initiated · canary 10% → 100%
-
14:41:55 · alert
error rate below threshold for 5m · incident status → Monitoring
-
15:11:02 · auto
monitoring window closed · draft retro generated from timeline
incident board
active & recent · auto-linked to traces & deploys
- SEV1 Investigating
Elevated 5xx on checkout-svc (eu-west-2)
INC-4417 · checkout-svc · opened 14:28 UTC · IC: A. Sharma
- SEV2 Identified
Kafka consumer lag > 50k on order.created
INC-4412 · kafka-events · opened 12:51 UTC · IC: P. Whelan
- SEV2 Monitoring
Redis cache eviction rate above baseline
INC-4409 · redis-cache · opened 09:17 UTC · IC: D. Okonjo
- SEV3 Resolved
TLS cert expiring in 7 days — edge-lb
INC-4401 · edge-lb · opened yesterday · IC: E. Vance
noise analytics
last 30 days · pages vs actionable
pages sent41
actionable pages38 (93%)
rules flagged noisy2 · fixes suggested
median acknowledge2m 41s
Capabilities
Alerting your on-call team will thank you for
SLO burn-rate alerts
Multi-window, multi-burn-rate alerting per Google SRE canon: fast burns page immediately, slow burns open tickets. Your error budget drives the urgency.
Anomaly & threshold alerts
Seasonal detectors for metrics that change shape, static thresholds for the ones that must not. Every alert ships with correlated traces, logs and recent deploys.
Routing & escalation
Route by service, team, severity or label to PagerDuty, Opsgenie, Slack or Teams. Escalation chains, snooze windows and maintenance suppression built in.
On-call that learns
Alert noise analytics show which rules woke people up for nothing. Weekly digest recommends threshold and window fixes with one-click apply.
Fewer pages. Better pages.
Burn-rate alerting, incident timelines that assemble themselves, and noise analytics that keep improving both.