Operator-native network assurance

Prove your network is behaving correctly.

Trigpoint puts a lightweight agent in every POP and cross-measures latency, loss, jitter, DNS and HTTP across your mesh. It learns what normal looks like on every path, and when something drifts, it tells the NOC what broke, where, who's affected, and why — before the customer calls.

  • <1 MB static agent binary
  • 10 s measurement windows
  • 1 node carries a 50-POP mesh
major alert a7f3-02e1 02:15:34 UTC

Latency level shift — atl ↔ chi

8 paths affected · onset 02:13:41 · +18 ms vs nearest mode (≈60σ)

10 20 30 40 ms 01:50 02:00 02:10 02:20 02:30 baseline 18–22 ms · median ± MAD · hour-of-week opened +1 m 53 s rtt p50 · atl→chi · icmp
bgp
0 events ±120 s — clean
is-is
no adjacency or SPF events — clean
ifaces
inferred link set: no drops, optics nominal
trace
path_hash unchanged on 8/8 flows
cause
transport-layer change (carrier / DWDM) · confidence 0.87
impact
14 customers ride atl-core1 ↔ chi-core2

The gap

Device telemetry is a solved problem. Delivery isn't.

Operators have excellent visibility into devices and links, and almost none into what the network actually delivers between locations. Interface counters are nominal, optics are in spec, every BGP session is established — and a customer in Atlanta is quietly getting 36 ms to Chicago on a path that has run at 19 ms for two years. Router health is green; the customer experience is a question mark.

Trigpoint closes that gap by measuring the thing you actually sell — path performance — continuously, from inside your own POPs. The measurement mesh is table stakes. The product is the intelligence layer that explains why a measurement changed.

How it works

Measure. Learn. Detect. Explain.

01

Measure everything, from everywhere

A single static Rust binary — under a megabyte, no runtime dependencies — registers with the control plane over mTLS and executes a signed test schedule. ICMP, UDP with DSCP marking, TCP connect, DNS against the resolvers your customers actually use, HTTP with full TLS phase timing, STAMP with hardware timestamping where the reflector offers it, and Paris traceroute — UDP, keyed on source port — that enumerates the ECMP paths a LAG fabric really hashes on instead of smearing them. Router-destined ICMP is kept honest as control-plane reachability and never folded into dataplane latency, because a CoPP-rate-limited echo measures the control plane, not transit. Agents ship 10-second aggregates and keep raw 1-second samples in a ring buffer, pulled on demand when the correlator wants forensic detail. When the control plane is unreachable, results spool locally — the moments you most need data are the moments the network is broken.

34 ms 50 ms 20 ms 18 ms 20 ms 16 ms lax dal atl chi nyc
Fixed points, cross-measured — the way a country gets surveyed.
02

Learn what normal means, per path, per hour

The magic is never "latency is 42 ms". It's "latency is normally 18–22 ms on this path at this hour". The baseliner builds per-path, per-metric, hour-of-week baselines using robust statistics — median and MAD, not mean and standard deviation, because network latency distributions are heavy-tailed and often multimodal. A path that ECMPs across two fiber routes has two normal latencies; Trigpoint stores both modes and measures deviation from the nearest one. Anomalous windows are excluded from baseline updates, so an incident never teaches the system that broken is normal.

03

Detect deviations without crying wolf

Deviation scoring with hysteresis and multi-window confirmation — no flapping page at 03:14 because one probe crossed a threshold. The detector recognizes distinct signatures: a level shift (the routing-change shape), variance increase with p99 inflating before p50 (congestion), loss onset with flat RTT (hardware or optics, not congestion), and correlated multi-path anomalies — when many paths through one POP degrade together, the POP is the problem, not the paths.

04

Explain it with your own routing state

Given an anomaly window and the affected path set, the correlator queries the event record: did BGP paths change? Did the IGP reconverge? Did an interface on the inferred path show drops, errors, or optical degradation? Did the forwarding path itself change — a one-comparison check, because every traceroute carries a stable path fingerprint? In an MPLS core, where no ttl-propagate collapses the backbone into one invisible hop, live IGP topology is the primary path source rather than traceroute: SPF over the link graph proves which paths crossed a failed link instead of merely coinciding with it in time. The output is a ranked hypothesis with evidence and a confidence score. Correlation is rule-based and explainable first; every root-cause claim ships with the checks behind it, and "unexplained" is an honest, first-class category — never a guess dressed up as an answer.

Where the agent sits

Four planes, one mesh

"Deploy an agent in every POP" hides a design question: attached to what, measuring through what? Trigpoint models four distinct measurement planes, each with its own targets and cadence — and the agent attaches to the production forwarding plane, never the management network, because measuring the OOB network measures the wrong network entirely.

Device

Agent → each local core, aggregation, and PE router — STAMP to line-card reflectors with hardware timestamps.

1 s · localizes a sick device in seconds

Infrastructure

Agent ↔ agent across the backbone, global table, one measured path per attachment pair.

1–10 s primary · 30–60 s cross-attachment

Service

Agent in a customer VRF → remote PE loopbacks, riding the same LSPs and QoS classes your customers do.

30–60 s · the measurement that backs an SLA

External

Agent → transit and peering next-hops, eyeball targets, resolvers, and CDNs, per exit.

30–60 s · per upstream, scorecard-ready

The payoff is localization without inference. Each agent is dual-homed on two routed uplinks — one per aggregation router, each with its own source address — not a LAG that deliberately hides which router forwarded the packet. So when every path sourced via agr2 degrades while the agr1 twins to the same destinations stay clean, the aggregation uplink is the fault and the alert says exactly that: a group-by, not a guess. Path identity carries the attachment, the VRF, and the target's role, so "which layer is sick" becomes a query — and the cheap intra-POP device plane means the extra dimensions barely move the probe budget.

Features

Built the way operators run networks

A full probe suite

ICMP, UDP with per-class DSCP marking, TCP connect, path MTU discovery, DNS, HTTP(S) with DNS/connect/TLS/TTFB phase timing, source-port Paris traceroute, and STAMP with hardware timestamping. STAMP and TWAMP-Light interoperate with the reflectors already in your routers — Nokia, Juniper, and Arista line cards become schedulable vantage points, so aggregation and PE routers get measured without an agent sitting on them.

Baselines that respect reality

Median + MAD per path, metric, and hour of week. Multimodal baselines for ECMP twin-route paths. Operator-declared physics floors (fiber-distance latency) so gross violations alert on day one, while the system is honest about reduced confidence during its learning period.

Alerts that carry their evidence

Every alert states what deviated and by how much, when it started to the sample, the blast radius, each evidence source checked with its result, and a ranked cause with confidence. If Trigpoint can't say most of that, it holds the alert and keeps gathering — a wrong root cause destroys trust faster than a slightly late alert.

Route changes as a first-class signal

Each traceroute's hop sequence hashes to a stable path fingerprint, so detecting a forwarding change is a single comparison — the correlator's cheapest, highest-value input. "Latency shifted and the path changed 40 seconds earlier" is a root cause, not a mystery.

Your control plane as context

BMP feeds from your route reflectors, IS-IS topology via BGP-LS, gNMI interface and optics telemetry — down to per-lane light levels — RPKI validation state over RTR, and customer inventory from your warehouse's curated views. Router-reported link delays are cross-checked against the mesh's own numbers, because a network's self-report and an independent measurement disagreeing is itself a finding. Trigpoint delivers value with the mesh alone and gets dramatically smarter with each feed you attach.

Impact, not just anomaly

Anomalous paths joined against topology and customer-port inventory answer the question that makes a NOC take an alert seriously: which services and which customers are riding the thing that just broke.

An agent you'd let near production

One static Rust binary, under 1 MB, no GC pauses polluting microsecond timing, kernel timestamping where available. Agents are deliberately dumb: they execute signed schedules and never decide what to test, so fleet behavior stays predictable and auditable. Clock-sync quality ships with every sample; one-way numbers are only trusted when the clocks deserve it. Every agent heartbeats its version, clock quality, and spool depth to a fleet view, and upgrades roll out POP-by-POP as a watched canary.

A mesh that doesn't melt

Full mesh is quadratic: 50 POPs is 1,225 paths, 500 is 124,750. Trigpoint tiers the mesh — core paths at 1–10 s, regional at 30–60 s, a background sweep feeding baselines — then escalates frequency and depth automatically where something smells wrong. With an IGP feed it prunes to the distinct link-path cover, testing your links rather than your permutations.

Your data, your metal

Self-hosted first — operators don't ship their topology to a SaaS. ClickHouse stores measurements and routing events side by side so correlation is a SQL join, not a cross-system export. One modest node (8 vCPU, 32 GB, NVMe) carries a 50-POP mesh with years of rollups. NOCs live in Grafana, so Trigpoint meets them there.

Incidents, not alert storms

When forty paths through one link degrade together, that's one event. Trigpoint rolls anomalies sharing a device, link, or attachment into a single incident with a timeline — thirty path alarms become one — and exports it as an evidence pack, the RFO artifact engineers otherwise assemble by hand at 4 a.m.

Maintenance-aware

Declare a window and anomalies inside it are suppressed and annotated, never dropped — detection never stops, paging does. Before you declare, a MOP impact preview names what goes dark downstream — devices isolated, devices left on their last uplink — computed against the live IGP, so a redundant path already lost to another outage is already accounted for. When the window closes, an automatic post-change report answers the question every change ticket begs: did everything return to baseline, per path and device touched?

Replay any incident

Scrub back through a past window and watch it unfold as if live — every anomaly, alert, routing event and maintenance window on one timeline, at 1× to 300×. Onboarding walkthroughs and tabletop drills run on real history instead of a broken lab, and "what did the correlator know, and when?" gets answered by dragging a slider, not by writing SQL.

On-demand bursts

Mid-incident, fire a traceroute storm, a STAMP burst, or an MTU sweep at a suspect region straight from the alert, the API, or a trig CLI — the same adaptive escalation the scheduler runs on its own, now under an engineer's hand. The first tool you reach for at 03:14.

Per-class truth

The mesh runs per QoS class — every EF and AF path has a best-effort twin, so "EF is slow while best-effort is clean" reads as a policer fault, not a transport mystery. And STAMP reflectors echo the DSCP they received: when a hop silently strips markings, latency stays perfect, the SLA quietly dies, and Trigpoint is the only witness — "sent EF, arrived best-effort" is an alert, with the hop range to check. Then a one-shot localization trace reads the DSCP each hop's ICMP quotes back and names the hop pair where the mark changed — the where, not just the whether.

A shared work queue

Acknowledge, assign, escalate, annotate — every alert carries lifecycle state and timestamped notes, attributed to the signed-in engineer and event-sourced beside the verdicts. Ack stops the re-page and shows who owns it; escalate steps severity and re-notifies. The alert list stops being a read-only feed and becomes the queue a whole shift works from — and it travels: the console is mobile-responsive, so on-call acknowledges an alert with one tap from a phone.

Search that reaches everything

One keystroke — Ctrl-K, or /, from anywhere — opens a command palette over everything the platform knows: alerts, incidents, paths, customers, devices, POPs, verdicts, evidence. Type an IP and it longest-prefix-matches to the customers riding it; type a hypothesis and it pulls that alert history and its hit rate. Mid-incident, "what else mentions r3?" is a keystroke, not a scroll through the last fifty rows.

The map, the route, and the what-if

An interactive topology map draws your live IGP graph with open incidents colored onto the exact device or link, interface counters per span, and IGP changes as badges. A path explorer puts the correlator's inferred SPF route beside the observed traceroute and says plainly when they disagree — honest uncertainty as a UI state, not a footnote. And because that SPF engine is right there, "if this link fails, what happens?" is a read-only click — which paths shift, which lose their route, which customers are in the blast radius — before the maintenance window, not during it. Compose a scenario — this router and that span — and the same engine answers for the multi-failure case, the one resilience reviews actually argue about.

A language layer over facts, never around them

Trigpoint can render an alert or a shift handoff into plain prose — but the model renders the evidence pack; it never finds a fault, assigns a cause, or invents a number. Detection, root cause, and impact stay deterministic and checkable. And nothing sensitive leaves un-pseudonymized: a redaction gateway is the single egress, mapping every customer, device, prefix and operator name to a stable token — "8/8 paths via DEV-04, sibling DEV-05 clean" reasons exactly as well as the real sentence, while your topology never leaves the box. Every call is audited after redaction, so "show me exactly what left the box" is one query, not a promise — and the same gateway drafts an RFO from a closed incident and turns a plain-English question into a typed, validated query that runs end-to-end: the IR executes deterministically in the UI as editable filter chips, so a wrong translation is a visibly wrong chip you can fix, never wrong data. Always rendering or translating what's already true, never inventing it. Local-first by default; the cloud, if you allow one, sees only tokens and numbers.

Notifications that respect a bad day

Named destinations and routing rules replace one firehose webhook: PagerDuty Events v2, Slack and Teams cards carrying the narrative line and a deep link, per-incident ServiceNow and Jira tickets (open creates, close resolves), any JSON webhook, and a versioned stream for a SIEM or event bus — routed by hypothesis, severity, scope, customer or POP. In front of all of it sits storm control: per-destination dedup and rate limits, exponential re-notify backoff, meta-incident grouping, and a global circuit breaker that flips to summary mode before the pager melts. PagerDuty even gets the resolve event when an alert auto-closes, and the bridge runs both ways — ack, assign, verdict, or declare a maintenance window straight from Slack, back through the same audited API. The monitoring must never become the incident.

Reports that stand up as evidence

Executive, operational, compliance, transit, routing, and per-customer reports assemble from the one joined store — availability and latency percentiles, acknowledge and resolution times, incident-by-cause, the correlator's own hit rate. Each freezes to an immutable archive: the computed model is stored, not the query, so re-issuing last March's SLA report next year returns byte-identical numbers long after the raw samples aged out. A report that isn't frozen isn't evidence — and a closed incident drafts its own RFO from the timeline the correlator already assembled.

Change validation, before and after

Declare a window with the device and link list from the MOP and suppression stops being a blunt scope: an unrelated failure in the same POP still pages, because the question becomes "does this path actually traverse something you touched?" On declare, the what-if engine pre-computes the blast radius — which paths shift, which lose their route, which customers ride them. When the window closes, a path that didn't recover and crosses a touched element is flagged a rollback candidate; a degradation that crosses nothing you declared is called out as correctly-alerted, not maintenance. MOP errors fail toward alerting, never toward silence.

Built for the review, not just the demo

Viewer, operator, and admin roles map to your LDAP groups and two-factor, enforced at a single API choke point. Every state-changing action — verdicts, workflow, maintenance, rebaselines, logins — lands in one append-only audit stream. Production credentials resolve from env, file, or your CyberArk vault behind one reference syntax, so no secret rides on a command line. And the irreplaceable data — baselines, verdicts, incidents, inventory — has a nightly dump and a restore drill that has actually been run.

The contract with the NOC

Anatomy of an alert

An alert you can't act on is noise with a timestamp. Every Trigpoint alert answers six questions, or it doesn't fire:

  1. What deviated. Metric, path set, and magnitude against the learned baseline — "+18 ms vs the nearest mode, ≈60σ", not "threshold exceeded".
  2. When it started. To the sample, including the gap between onset and confirmation.
  3. Blast radius. POPs, services, and customers affected — from inventory, not guesswork.
  4. Evidence checked. BGP: no change. IS-IS: no events. Interface et-0/0/1: output drops rising. Every source consulted, with its result — including the exculpatory ones.
  5. Ranked likely cause, with confidence. Congestion, route change, hardware, service failure — or, honestly, unexplained.
  6. The drill-down. Links straight to path forensics: raw samples around the window, traceroute history, the routing events that coincided.

And the loop closes: operators confirm or reject every hypothesis, that verdict is recorded next to the evidence, and the correlator's hit rate is itself a tracked metric. The system earns trust the same way an engineer does — by being right, and by being checkable when it isn't.

One platform

Three questions every operator has to answer

Service Trust

"Is the network delivering the experience it should — POP to POP, POP to Internet?"

The continuous measurement mesh, self-learning baselines, anomaly detection, and path forensics. Deploy a container in every POP; get a self-baselining anomaly map of your network.

Routing Trust

"Is the control plane doing what it should?"

Observed BGP checked against RPKI validation state over RTR — an ROV-invalid origin on your prefixes or a customer's becomes an alert. Route-churn and flap scoring name the chronically unstable prefix or adjacency a snapshot hides. External vantage over public RIS feeds answers "does the Internet see us correctly?" while the internal mesh answers "are we seeing ourselves correctly?" — and anycast catchment validation tracks which instance each POP actually reaches.

Configuration Trust

"Are the devices configured the way we intend?"

Golden-config drift detection, compliance and security-posture checks, and digital-twin queries — "if this link fails, what happens to these paths?" — grounded in the same live topology the correlator already keeps. Both halves run today. What-if: read-only failure simulation — single element, a composed multi-failure scenario, or a planned link that doesn't exist yet — over the live IGP graph, plus a resilience audit naming the single points of failure. Drift: running configs snapshotted and hashed against an operator-blessed golden per device, changes surfacing as a red badge on the topology map with the diff attached — on its first day in the lab it caught a test fixture leaving residue behind that the fixture's own revert claimed to clean up. Compliance rules and intent modeling remain the deliberately-later part.

Why another monitoring tool

Inside-out, not outside-in

Enterprise synthetics platforms watch the Internet from the outside: useful if you're a company consuming networks, incomplete if you're the company running one. They don't speak your IGP, they can't see your route reflectors, and they have no idea which customers ride which links.

Trigpoint is built for the other side of the demarc. It understands your AS, your IS-IS topology, your POPs, and your customer inventory — so when a path degrades, the answer isn't a red dot on a world map. It's "the ATL–CHI link changed transport path at 02:13, no IGP or BGP involvement, 14 customers affected, and here is the evidence for each claim."

No product today — commercial or open source — combines a synthetic measurement mesh with the operator's own routing state for root-cause correlation. That's the gap Trigpoint exists to close.

Architecture

Small parts, sharp edges

Storage that matches the questions

Measurement tables cluster by path and time — the shape of "this path, this window". Event tables cluster by time — the shape of "what happened between 02:11 and 02:14". Raw data keeps 14 days, minute rollups 90, hourly rollups two years, because "was it always like this?" is a real question and deserves a real answer.

Deployment without ceremony

Docker Compose for small footprints, Helm for large ones; agents as containers or systemd units. Everything is Rust — the same properties that make the agent trustworthy make the control plane cheap to run. Ingest for a 50-POP tiered mesh is about 38 rows a second; sizing is not the risk.

Where it stands

Being proven the honest way

Trigpoint is in active development and validated continuously against a lab network with thirteen scripted, labeled fault classes: latency steps, bufferbloat congestion, clean packet loss, IS-IS reroutes, dead DNS at 03:14 with the WAN untouched, an aggregation-uplink degrade only one attachment sees, a customer prefix withdrawn upstream, a killed LSP, a starved QoS class, a degrading transit provider, a hop that silently strips DSCP markings, and a resolver that quietly stops validating DNSSEC. The detector catches every one with the correct signature and the correct discrimination — and has held zero false positives on a quiet mesh throughout.

The concept alert fires end-to-end, evidence and blast radius included. An IS-IS reroute is caught as an adjacency-down event within a second, and SPF over the live topology proves which paths crossed the failed link — including a cross-pair set no time-window could have attributed. A customer outage traced to a BMP withdraw event says "routing pull, not fault," naming the prefix and the peers that saw it vanish. Mesh loss joins against interface counters to name the dropping port. A dead LSP surfaces as the alert only a service-plane mesh can produce: the VPN is down and the underlay beneath it is provably clean. A starved EF class is called a class fault because its best-effort twins stay healthy; stripped DSCP markings are caught by the reflector echo with latency perfect; a degrading upstream is named because its sibling transit is clean. A paired probe — one validly-signed name, one deliberately bogus — tells a resolver that stopped validating apart from a zone whose signatures expired, faults that are invisible to every reachability check because resolution keeps "working". Each alert ships as a plain-language narrative — what happened, why we believe it, what to check — with the affected customers attached and the raw evidence one click below.

The loop closes in the tooling that exists today: a NOC web UI with a live mesh matrix, search over everything the platform knows (one keystroke, from alerts to customers to path histories), one-click verdicts, and team workflow — acknowledge, assign, escalate, annotate — so the alert list is a shared work queue, not a read-only feed. Engineers get an interactive topology map with read-only failure simulation, and a path explorer that draws the inferred SPF route beside the observed traceroute and says plainly when they disagree — honest uncertainty as a UI state, not a footnote. Alerts reach humans through PagerDuty, Slack and Teams, open per-incident ServiceNow and Jira tickets, and feed any JSON webhook or a versioned SIEM/event-bus stream — routed by hypothesis, severity, scope, customer or POP, and all of it behind a storm-control layer — dedup, backoff, a circuit breaker — because the monitoring must never become the incident. A trig CLI covers bursts, triage, shift handoffs, and monthly SLA evidence reports; a maintenance calendar with recurring windows keeps planned work from paging anyone; and a tracked per-hypothesis hit rate means the correlator's accuracy is itself a metric, judged by the operators it serves.

The operator plane is production-shaped too: directory (LDAP) sign-in with PingID two-factor and viewer/operator/admin roles, every state-changing action — verdicts, workflow, maintenance, rebaselines, logins — in one audit stream, retention tiering with a rehearsed backup-and-restore drill for the data that matters, and an intelligence layer that survives reboots honestly — a restarted correlator reconciles every claim it left open, so history never shows an alert that nobody closed. The AI layer arrives trust-first: a redaction gateway pseudonymizes every name, prefix and topology fact to a stable token before a byte reaches any model — and it now runs end-to-end on a local open-weight model, on-box, chosen by a measured nine-model bake-off rather than fashion: alert narratives, shift handoffs, plain-English queries compiled to a validated typed IR, watermarked RFO drafts, and "have we seen this before?" recall that pairs deterministic history joins with local-embedding similarity — with the entity-leak check reading zero across every call class, embeddings included, and every outbound prompt byte-identical across restarts.

The newest wave turns that same discipline on the control plane itself. Routing trust now runs in the lab: announcements checked against RPKI over RTR, so an ROV-invalid origin on your prefixes or a customer's becomes an alert; route-churn and flap scoring that names the chronically unstable prefix or adjacency a snapshot hides; an external vantage over public RIS feeds asking whether the Internet still sees your prefixes; and anycast catchment validation that catches an instance flip while the address itself stays up. Alongside it, a reporting spine: evidence-grade SLA reports — measured and outage-based availability, merged-percentile latency, an incident-and-verdict ledger, a method appendix — plus executive, operational, compliance, transit and per-customer templates, every one frozen to an immutable archive. The computed model is stored, not the query, so a re-issued report is byte-identical years later, after the raw samples have aged out — because a report that isn't frozen isn't evidence. Closed incidents draft their own RFO. And change windows now carry the MOP's device list: suppression follows the topology, so an unrelated fault in the same POP still pages; the blast radius is pre-computed when the window is declared; a path that fails to recover across something you touched is flagged a rollback candidate; and a resilience audit names the single points of failure and single-homed customers before the change, not after.

Where these surfaces reach outside the box, the site says so plainly. The paging and ticketing integrations — PagerDuty, Slack, Teams, ServiceNow, Jira — and the warehouse sync in both directions are built against each vendor's published contract and proven against mocks, not yet pointed at a live tenant. The real-iron measurements — STAMP against line-card reflectors, BGP-LS from a live speaker, gNMI optics, router-reported link delay, SR-steered probes — are built behind an already-normalized model, waiting only on hardware to swap a lab stand-in for the real feed. And a documented, versioned read API with a pip-installable client lets a Nornir or Ansible playbook gate a change on "is this path healthy?" without opening the UI. Each surface is named for exactly what it is, because the whole argument of the product is that it tells you the truth.

That bar — a zero-false-positive week followed by caught, explained, customer-attributed degradations — is the standard the production release has to clear. If you run POPs and want a self-baselining anomaly map of your own network, we'd like to talk while the roadmap is still wet.