How we ingest 1.2M metric points per second on commodity hardware
Inside the TraceLoop storage engine: columnar compression, pre-aggregation and the write path we rebuilt three times.
read post →Resources
Quickstarts measured in minutes, a query reference with 120+ runnable examples, an alerting handbook written by people who have been paged at 3am — and the engineering blog where we explain how TraceLoop itself works.
Documentation
Ship your first trace in under five minutes with the TraceLoop agent or OpenTelemetry SDK.
Guide · 5 min read →
Configuration, sampling strategies, tail-based sampling and resource attributes for every supported runtime.
Reference · updated weekly →
The TraceLoop Query Language: filter logs, aggregate metrics and stitch traces with one composable syntax.
Reference · 120+ examples →
SLO-based alert design, burn-rate windows, escalation policies and on-call rotation patterns.
Handbook · 42 pages →
Guides & Best Practice
A pragmatic path off per-host licensing: dual-write with OTel, validate parity, then cut over.
Migration guide →
How to pick SLIs, set error budgets and wire them to alerts that never cry wolf.
Best practice →
Keep every interesting trace, drop the noise — a worked example at 900k spans/second.
Deep dive →
Blameless post-mortem templates and the TraceLoop timeline export that writes half the report for you.
Templates →
Engineering Blog
Inside the TraceLoop storage engine: columnar compression, pre-aggregation and the write path we rebuilt three times.
read post →Why correlation IDs are only half the story, and how monotonic timeline alignment changes incident review.
read post →Our seasonal-baseline detector explained in plain English, with the false-positive budget we hold it to.
read post →We are all-in on OTel wire formats. Here is the honest list of where our platform does more than raw collectors.
read post →Platform Status
We run our production estate on TraceLoop and publish the same status view our customers see. 99.99% availability, measured and posted honestly.
platform status
all components · updated 12s ago
Telemetry ingest (eu-west-2)
OperationalTelemetry ingest (eu-west-1)
OperationalQuery & search
OperationalAlerting & routing
OperationalWorkspace UI & API
OperationalIntegration delivery
Degraded — filelog agents (legacy hosts)uptime · 90 days
monthly availability · SLO 99.95%
Subscribe to status alerts from any incident feed.
Read the quickstart, wire up a collector, and ask our engineers anything that the docs don't answer.