Alert noise reduction

Your on-call rotation is not a noise filter

ObsrvHQ learns what normal looks like for every service and pages a human only when it truly matters, turning an alert storm back into signal.

No credit card required

~94% of alerts suppressed on avg, first 24h
< 5 min to integrate with your existing alertmanager
2024 Founded by SREs who ran on-call at scale

Most alerts are not incidents. Your on-call engineers know it.

12

Alerts per on-call shift, on average. 11 of them self-resolve before an engineer touches them.

30-40%

Of on-call hours spent triaging noise, not diagnosing real incidents. Time that does not show up in MTTD or MTTR metrics, but your team feels it.

60%

Of SRE departures cite alert fatigue as a factor, in teams we have spoken with directly. It is a retention problem masquerading as a tooling problem.

Three steps. 48 hours to first suppression.

01 / Connect

One YAML block. Your alertmanager stays.

Point ObsrvHQ at your Prometheus endpoint or Datadog API key. It reads your existing alert rules without modifying them. No new data plane, no agent installs.

obsrvhq:
  source: prometheus
  endpoint: http://prometheus:9090
  rules_path: /etc/prometheus/rules
02 / Learn

Per-series, per-hour baseline envelopes

ObsrvHQ watches every time-series for 48 hours and builds a separate envelope for each one. A metric that peaks at 9am Monday is not anomalous at 9:05am Monday. Nothing else in your stack knows that.

03 / Filter

Suppress noise. Page signal.

Every incoming alert gets a suppress-or-page decision before it reaches your alertmanager. Metrics within their learned envelope are suppressed. Genuine deviations go through. The on-call engineer sees fewer pages, all of them worth acting on.

Two panels. One job: show you what needs attention.

The baseline view shows each metric against its learned envelope. The decision feed shows every suppress vs. page call in order. No new dashboard to learn, no new query language.

Baseline view: api-gateway.request_latency_p99
Anomaly detected Learned normal range (rolling 7d) 00:00 06:00 12:00 18:00 24:00
Decision feed: last 10 events
03:14:02 api-gateway cpu_usage_high SUPPRESSED
03:14:18 cache-svc mem_rss_warn SUPPRESSED
03:15:01 order-worker latency_p99_crit SUPPRESSED
03:15:22 auth-svc error_rate_warn SUPPRESSED
03:16:00 cdn-edge miss_rate_warn SUPPRESSED
03:16:41 db-replica-2 replication_lag SUPPRESSED
03:17:03 payment-svc queue_depth_crit SUPPRESSED
03:18:14 checkout-api error_rate_crit PAGED
03:18:16 db-primary conn_pool_saturation PAGED
03:18:19 ingestion-svc disk_io_critical PAGED

Works with the alerting stack you already run

ObsrvHQ sits upstream of your alertmanager, not instead of it. Your existing routing rules, escalation policies, and on-call schedules stay exactly as they are.

P Prometheus
G Grafana Alerting
D Datadog Monitors
PD PagerDuty
OG Opsgenie

From the engineers who live with on-call

The envelope widens automatically after we push a deploy, then tightens back. I stopped babysitting alert rules after every release. That alone is worth the setup time.

Lead SRE, payments platform

We were at 800 alerts a day across 60 microservices. After 48 hours of learning, that dropped to around 40 pages, and every one of those 40 was a real issue. The correlation grouping means you get one page for a cascading failure, not twelve separate ones.

Staff engineer, distributed systems team, e-commerce platform

I pointed it at our Prometheus endpoint and added four lines of YAML. Two days later the dashboard showed 1,200 suppressed decisions and 3 paged. The 3 were all real incidents. We had been living with that noise for a year.

Head of Platform Engineering, early-stage logistics company

Priced by metric series, not by seat or incident volume.

Free tier is permanent: 500 metric series, no credit card. Pro at $79/mo covers most engineering teams running up to 50 microservices.

Free Pro $79/mo Scale $299/mo