TraceLoop

Alerts

Pages that mean it. Silence that earns trust.

Alert fatigue is a design failure, not a fact of life. TraceLoop alerts on SLO burn rate and scored anomalies, attaches the full investigative context to every page, and tells you weekly which rules are still wasting your engineers' sleep.

Alert Rules

From error budgets to escalation, in one rule editor

Fictional workspace, real interface: the checkout-availability SLO, its burn-rate windows, and exactly who gets paged when the budget burns too fast.

slo · checkout-availability

objective 99.95% · 30-day window · multi-burn-rate

HEALTHY
error budget remaining71%
  • PAGE 1h burn ≥ 14.4× · fast budget drain → PagerDuty SEV1
  • TICKET 6h burn ≥ 6× · slow drain → Jira, business hours
  • MUTE maintenance window 02:00–04:00 Sun · auto-suppressed

detector board

anomaly & threshold detectors · scoring every 60s

3 ACTIVE
  • FIRING

    checkout-svc · p99 latency

    score 0.91
  • WATCHING

    postgres-01 · deadlock rate

    score 0.82
  • WATCHING

    redis-cache · evictions/sec

    score 0.64
  • NORMAL

    kafka · consumer lag order.created

    score 0.21
  • NORMAL

    api-gateway · rps

    score 0.18

Incident Investigation

The timeline writes itself

When INC-4417 fired, TraceLoop had already stitched the deploy, the error logs, the similar traces and the alert sequence into one narrative. The humans just made decisions.

incident timeline

INC-4417 · elevated 5xx on checkout-svc · SEV1

RESOLVED IN 13m
  1. 14:28:04 · alert

    burn-rate alert fired · checkout-availability 1h window · 14.2× burn

  2. 14:28:04 · auto

    INC-4417 opened · SEV1 · paged on-call (A. Sharma) via PagerDuty

  3. 14:28:06 · auto

    timeline auto-populated · 187 error logs · 4 similar traces · deploy v2.41.3 (14:17) flagged as suspect

  4. 14:29:31 · engineer

    A. Sharma acknowledged · started war room in #inc-4417

  5. 14:33:12 · engineer

    root cause hypothesis: deadlock loop introduced by order-write refactor

  6. 14:36:40 · engineer

    rollback v2.41.3 → v2.41.2 initiated · canary 10% → 100%

  7. 14:41:55 · alert

    error rate below threshold for 5m · incident status → Monitoring

  8. 15:11:02 · auto

    monitoring window closed · draft retro generated from timeline

incident board

active & recent · auto-linked to traces & deploys

4 OPEN
  • SEV1

    Elevated 5xx on checkout-svc (eu-west-2)

    INC-4417 · checkout-svc · opened 14:28 UTC · IC: A. Sharma

    Investigating
  • SEV2

    Kafka consumer lag > 50k on order.created

    INC-4412 · kafka-events · opened 12:51 UTC · IC: P. Whelan

    Identified
  • SEV2

    Redis cache eviction rate above baseline

    INC-4409 · redis-cache · opened 09:17 UTC · IC: D. Okonjo

    Monitoring
  • SEV3

    TLS cert expiring in 7 days — edge-lb

    INC-4401 · edge-lb · opened yesterday · IC: E. Vance

    Resolved

noise analytics

last 30 days · pages vs actionable

−62% NOISE

pages sent41

actionable pages38 (93%)

rules flagged noisy2 · fixes suggested

median acknowledge2m 41s

Capabilities

Alerting your on-call team will thank you for

SLO burn-rate alerts

Multi-window, multi-burn-rate alerting per Google SRE canon: fast burns page immediately, slow burns open tickets. Your error budget drives the urgency.

Anomaly & threshold alerts

Seasonal detectors for metrics that change shape, static thresholds for the ones that must not. Every alert ships with correlated traces, logs and recent deploys.

Routing & escalation

Route by service, team, severity or label to PagerDuty, Opsgenie, Slack or Teams. Escalation chains, snooze windows and maintenance suppression built in.

On-call that learns

Alert noise analytics show which rules woke people up for nothing. Weekly digest recommends threshold and window fixes with one-click apply.

Fewer pages. Better pages.

Burn-rate alerting, incident timelines that assemble themselves, and noise analytics that keep improving both.