Your on-call rotation is not a noise filter
ObsrvHQ learns what normal looks like for every service and pages a human only when it truly matters, turning an alert storm back into signal.
No credit card required
Most alerts are not incidents. Your on-call engineers know it.
Alerts per on-call shift, on average. 11 of them self-resolve before an engineer touches them.
Of on-call hours spent triaging noise, not diagnosing real incidents. Time that does not show up in MTTD or MTTR metrics, but your team feels it.
Of SRE departures cite alert fatigue as a factor, in teams we have spoken with directly. It is a retention problem masquerading as a tooling problem.
Three steps. 48 hours to first suppression.
One YAML block. Your alertmanager stays.
Point ObsrvHQ at your Prometheus endpoint or Datadog API key. It reads your existing alert rules without modifying them. No new data plane, no agent installs.
obsrvhq: source: prometheus endpoint: http://prometheus:9090 rules_path: /etc/prometheus/rules
Per-series, per-hour baseline envelopes
ObsrvHQ watches every time-series for 48 hours and builds a separate envelope for each one. A metric that peaks at 9am Monday is not anomalous at 9:05am Monday. Nothing else in your stack knows that.
Suppress noise. Page signal.
Every incoming alert gets a suppress-or-page decision before it reaches your alertmanager. Metrics within their learned envelope are suppressed. Genuine deviations go through. The on-call engineer sees fewer pages, all of them worth acting on.
Two panels. One job: show you what needs attention.
The baseline view shows each metric against its learned envelope. The decision feed shows every suppress vs. page call in order. No new dashboard to learn, no new query language.
Works with the alerting stack you already run
ObsrvHQ sits upstream of your alertmanager, not instead of it. Your existing routing rules, escalation policies, and on-call schedules stay exactly as they are.
From the engineers who live with on-call
The envelope widens automatically after we push a deploy, then tightens back. I stopped babysitting alert rules after every release. That alone is worth the setup time.
Lead SRE, payments platformWe were at 800 alerts a day across 60 microservices. After 48 hours of learning, that dropped to around 40 pages, and every one of those 40 was a real issue. The correlation grouping means you get one page for a cascading failure, not twelve separate ones.
Staff engineer, distributed systems team, e-commerce platformI pointed it at our Prometheus endpoint and added four lines of YAML. Two days later the dashboard showed 1,200 suppressed decisions and 3 paged. The 3 were all real incidents. We had been living with that noise for a year.
Head of Platform Engineering, early-stage logistics company