Trigpoint puts a lightweight agent in every POP and cross-measures latency,
loss, jitter, DNS and HTTP across your mesh. It learns what normal looks like
on every path, and when something drifts, it tells the NOC what broke, where,
who's affected, and why — before the customer calls.
Device telemetry is a solved problem. Delivery isn't.
Operators have excellent visibility into devices and links, and almost none into
what the network actually delivers between locations. Interface counters are
nominal, optics are in spec, every BGP session is established — and a customer
in Atlanta is quietly getting 36 ms to Chicago on a path that has run at 19 ms
for two years. Router health is green; the customer experience is a question mark.
Trigpoint closes that gap by measuring the thing you actually sell — path
performance — continuously, from inside your own POPs. The measurement mesh is
table stakes. The product is the intelligence layer that explains why
a measurement changed.
How it works
Measure. Learn. Detect. Explain.
01
Measure everything, from everywhere
A single static Rust binary — under a megabyte, no runtime dependencies —
registers with the control plane over mTLS and executes a signed test
schedule. ICMP, UDP with DSCP marking, TCP connect, DNS against the
resolvers your customers actually use, HTTP with full TLS phase timing,
STAMP with hardware timestamping where the reflector offers it, and Paris
traceroute — UDP, keyed on source port — that enumerates the ECMP paths a
LAG fabric really hashes on instead of smearing them. Router-destined ICMP
is kept honest as control-plane reachability and never folded into
dataplane latency, because a CoPP-rate-limited echo measures the control
plane, not transit. Agents ship 10-second aggregates and keep raw 1-second
samples in a ring buffer, pulled on demand when the correlator wants
forensic detail. When the control plane is unreachable, results spool
locally — the moments you most need data are the moments the network is broken.
Fixed points, cross-measured — the way a country gets surveyed.
02
Learn what normal means, per path, per hour
The magic is never "latency is 42 ms". It's "latency is normally 18–22 ms
on this path at this hour". The baseliner builds per-path, per-metric,
hour-of-week baselines using robust statistics — median and MAD, not mean
and standard deviation, because network latency distributions are
heavy-tailed and often multimodal. A path that ECMPs across two fiber
routes has two normal latencies; Trigpoint stores both modes and
measures deviation from the nearest one. Anomalous windows are excluded
from baseline updates, so an incident never teaches the system that broken
is normal.
03
Detect deviations without crying wolf
Deviation scoring with hysteresis and multi-window confirmation — no
flapping page at 03:14 because one probe crossed a threshold. The detector
recognizes distinct signatures: a level shift (the routing-change shape),
variance increase with p99 inflating before p50 (congestion), loss onset
with flat RTT (hardware or optics, not congestion), and correlated
multi-path anomalies — when many paths through one POP degrade together,
the POP is the problem, not the paths.
04
Explain it with your own routing state
Given an anomaly window and the affected path set, the correlator queries
the event record: did BGP paths change? Did the IGP reconverge? Did an
interface on the inferred path show drops, errors, or optical degradation?
Did the forwarding path itself change — a one-comparison check, because
every traceroute carries a stable path fingerprint? In an MPLS core, where
no ttl-propagate collapses the backbone into one invisible hop,
live IGP topology is the primary path source rather than traceroute: SPF
over the link graph proves which paths crossed a failed link
instead of merely coinciding with it in time. The output is a ranked
hypothesis with evidence and a confidence score. Correlation is rule-based
and explainable first; every root-cause claim ships with the checks behind
it, and "unexplained" is an honest, first-class category — never a guess
dressed up as an answer.
Where the agent sits
Four planes, one mesh
"Deploy an agent in every POP" hides a design question: attached to what,
measuring through what? Trigpoint models four distinct measurement planes,
each with its own targets and cadence — and the agent attaches to the
production forwarding plane, never the management network, because measuring
the OOB network measures the wrong network entirely.
Device
Agent → each local core, aggregation, and PE router — STAMP to line-card reflectors with hardware timestamps.
1 s · localizes a sick device in seconds
Infrastructure
Agent ↔ agent across the backbone, global table, one measured path per attachment pair.
1–10 s primary · 30–60 s cross-attachment
Service
Agent in a customer VRF → remote PE loopbacks, riding the same LSPs and QoS classes your customers do.
30–60 s · the measurement that backs an SLA
External
Agent → transit and peering next-hops, eyeball targets, resolvers, and CDNs, per exit.
30–60 s · per upstream, scorecard-ready
The payoff is localization without inference. Each agent is dual-homed on two
routed uplinks — one per aggregation router, each with its own source address —
not a LAG that deliberately hides which router forwarded the packet. So when
every path sourced via agr2 degrades while the agr1 twins to the same
destinations stay clean, the aggregation uplink is the fault and the alert
says exactly that: a group-by, not a guess. Path identity carries the
attachment, the VRF, and the target's role, so "which layer is sick" becomes a
query — and the cheap intra-POP device plane means the extra dimensions barely
move the probe budget.
Features
Built the way operators run networks
A full probe suite
ICMP, UDP with per-class DSCP marking, TCP connect, path MTU discovery, DNS,
HTTP(S) with DNS/connect/TLS/TTFB phase timing, source-port Paris traceroute,
and STAMP with hardware timestamping. STAMP and TWAMP-Light interoperate with
the reflectors already in your routers — Nokia, Juniper, and Arista line cards
become schedulable vantage points, so aggregation and PE routers get measured
without an agent sitting on them.
Baselines that respect reality
Median + MAD per path, metric, and hour of week. Multimodal baselines for
ECMP twin-route paths. Operator-declared physics floors (fiber-distance
latency) so gross violations alert on day one, while the system is honest
about reduced confidence during its learning period.
Alerts that carry their evidence
Every alert states what deviated and by how much, when it started to the
sample, the blast radius, each evidence source checked with its result, and a
ranked cause with confidence. If Trigpoint can't say most of that, it holds
the alert and keeps gathering — a wrong root cause destroys trust faster than
a slightly late alert.
Route changes as a first-class signal
Each traceroute's hop sequence hashes to a stable path fingerprint, so
detecting a forwarding change is a single comparison — the correlator's
cheapest, highest-value input. "Latency shifted and the path changed
40 seconds earlier" is a root cause, not a mystery.
Your control plane as context
BMP feeds from your route reflectors, IS-IS topology via BGP-LS, gNMI
interface and optics telemetry — down to per-lane light levels — RPKI
validation state over RTR, and customer inventory from your warehouse's
curated views. Router-reported link delays are cross-checked against the
mesh's own numbers, because a network's self-report and an independent
measurement disagreeing is itself a finding. Trigpoint delivers value
with the mesh alone and gets dramatically smarter with each feed you
attach.
Impact, not just anomaly
Anomalous paths joined against topology and customer-port inventory answer
the question that makes a NOC take an alert seriously: which services and
which customers are riding the thing that just broke.
An agent you'd let near production
One static Rust binary, under 1 MB, no GC pauses polluting microsecond
timing, kernel timestamping where available. Agents are deliberately dumb:
they execute signed schedules and never decide what to test, so fleet
behavior stays predictable and auditable. Clock-sync quality ships with
every sample; one-way numbers are only trusted when the clocks deserve it.
Every agent heartbeats its version, clock quality, and spool depth to a
fleet view, and upgrades roll out POP-by-POP as a watched canary.
A mesh that doesn't melt
Full mesh is quadratic: 50 POPs is 1,225 paths, 500 is 124,750. Trigpoint
tiers the mesh — core paths at 1–10 s, regional at 30–60 s, a background
sweep feeding baselines — then escalates frequency and depth automatically
where something smells wrong. With an IGP feed it prunes to the distinct
link-path cover, testing your links rather than your permutations.
Your data, your metal
Self-hosted first — operators don't ship their topology to a SaaS.
ClickHouse stores measurements and routing events side by side so correlation
is a SQL join, not a cross-system export. One modest node (8 vCPU, 32 GB,
NVMe) carries a 50-POP mesh with years of rollups. NOCs live in Grafana, so
Trigpoint meets them there.
Incidents, not alert storms
When forty paths through one link degrade together, that's one event.
Trigpoint rolls anomalies sharing a device, link, or attachment into a single
incident with a timeline — thirty path alarms become one — and exports it as
an evidence pack, the RFO artifact engineers otherwise assemble by hand at
4 a.m.
Maintenance-aware
Declare a window and anomalies inside it are suppressed and annotated,
never dropped — detection never stops, paging does. Before you declare,
a MOP impact preview names what goes dark downstream — devices isolated,
devices left on their last uplink — computed against the live IGP, so a
redundant path already lost to another outage is already accounted for.
When the window closes, an automatic post-change report answers the
question every change ticket begs: did everything return to baseline,
per path and device touched?
Replay any incident
Scrub back through a past window and watch it unfold as if live — every
anomaly, alert, routing event and maintenance window on one timeline, at
1× to 300×. Onboarding walkthroughs and tabletop drills run on real
history instead of a broken lab, and "what did the correlator know, and
when?" gets answered by dragging a slider, not by writing SQL.
On-demand bursts
Mid-incident, fire a traceroute storm, a STAMP burst, or an MTU sweep at a
suspect region straight from the alert, the API, or a trig CLI — the
same adaptive escalation the scheduler runs on its own, now under an
engineer's hand. The first tool you reach for at 03:14.
Per-class truth
The mesh runs per QoS class — every EF and AF path has a best-effort twin,
so "EF is slow while best-effort is clean" reads as a policer fault, not a
transport mystery. And STAMP reflectors echo the DSCP they received:
when a hop silently strips markings, latency stays perfect, the SLA quietly
dies, and Trigpoint is the only witness — "sent EF, arrived best-effort" is
an alert, with the hop range to check. Then a one-shot localization trace
reads the DSCP each hop's ICMP quotes back and names the hop pair where the
mark changed — the where, not just the whether.
A shared work queue
Acknowledge, assign, escalate, annotate — every alert carries lifecycle
state and timestamped notes, attributed to the signed-in engineer and
event-sourced beside the verdicts. Ack stops the re-page and shows who owns
it; escalate steps severity and re-notifies. The alert list stops being a
read-only feed and becomes the queue a whole shift works from — and it
travels: the console is mobile-responsive, so on-call acknowledges an
alert with one tap from a phone.
Search that reaches everything
One keystroke — Ctrl-K, or /, from anywhere — opens a
command palette over everything the platform knows: alerts, incidents,
paths, customers, devices, POPs, verdicts, evidence. Type an IP and it
longest-prefix-matches to the customers riding it; type a hypothesis and it
pulls that alert history and its hit rate. Mid-incident, "what else mentions
r3?" is a keystroke, not a scroll through the last fifty rows.
The map, the route, and the what-if
An interactive topology map draws your live IGP graph with open incidents
colored onto the exact device or link, interface counters per span, and IGP
changes as badges. A path explorer puts the correlator's inferred SPF route
beside the observed traceroute and says plainly when they disagree — honest
uncertainty as a UI state, not a footnote. And because that SPF engine is
right there, "if this link fails, what happens?" is a read-only click — which
paths shift, which lose their route, which customers are in the blast radius —
before the maintenance window, not during it. Compose a scenario — this
router and that span — and the same engine answers for the
multi-failure case, the one resilience reviews actually argue about.
A language layer over facts, never around them
Trigpoint can render an alert or a shift handoff into plain prose — but the
model renders the evidence pack; it never finds a fault, assigns a cause, or
invents a number. Detection, root cause, and impact stay deterministic and
checkable. And nothing sensitive leaves un-pseudonymized: a redaction gateway
is the single egress, mapping every customer, device, prefix and operator name
to a stable token — "8/8 paths via DEV-04, sibling DEV-05 clean" reasons
exactly as well as the real sentence, while your topology never leaves the
box. Every call is audited after redaction, so "show me exactly what left
the box" is one query, not a promise — and the same gateway drafts an RFO
from a closed incident and turns a plain-English question into a typed,
validated query that runs end-to-end: the IR executes deterministically
in the UI as editable filter chips, so a wrong translation is a visibly
wrong chip you can fix, never wrong data. Always rendering or translating
what's already true, never inventing it. Local-first by default; the cloud, if you allow one,
sees only tokens and numbers.
Notifications that respect a bad day
Named destinations and routing rules replace one firehose webhook:
PagerDuty Events v2, Slack and Teams cards carrying the narrative line and a
deep link, per-incident ServiceNow and Jira tickets (open creates, close
resolves), any JSON webhook, and a versioned stream for a SIEM or event bus —
routed by hypothesis, severity, scope, customer or POP. In front of all of it
sits storm control: per-destination dedup and rate limits, exponential
re-notify backoff, meta-incident grouping, and a global circuit breaker that
flips to summary mode before the pager melts. PagerDuty even gets the resolve
event when an alert auto-closes, and the bridge runs both ways — ack,
assign, verdict, or declare a maintenance window straight from Slack, back
through the same audited API. The monitoring must never become the incident.
Reports that stand up as evidence
Executive, operational, compliance, transit, routing, and per-customer
reports assemble from the one joined store — availability and latency
percentiles, acknowledge and resolution times, incident-by-cause, the
correlator's own hit rate. Each freezes to an immutable archive: the
computed model is stored, not the query, so re-issuing last March's SLA
report next year returns byte-identical numbers long after the raw
samples aged out. A report that isn't frozen isn't evidence — and a
closed incident drafts its own RFO from the timeline the correlator
already assembled.
Change validation, before and after
Declare a window with the device and link list from the MOP and
suppression stops being a blunt scope: an unrelated failure in the same
POP still pages, because the question becomes "does this path actually
traverse something you touched?" On declare, the what-if engine
pre-computes the blast radius — which paths shift, which lose their
route, which customers ride them. When the window closes, a path that
didn't recover and crosses a touched element is flagged a rollback
candidate; a degradation that crosses nothing you declared is called out
as correctly-alerted, not maintenance. MOP errors fail toward alerting,
never toward silence.
Built for the review, not just the demo
Viewer, operator, and admin roles map to your LDAP groups and two-factor,
enforced at a single API choke point. Every state-changing action — verdicts,
workflow, maintenance, rebaselines, logins — lands in one append-only audit
stream. Production credentials resolve from env, file, or your CyberArk vault
behind one reference syntax, so no secret rides on a command line. And the
irreplaceable data — baselines, verdicts, incidents, inventory — has a nightly
dump and a restore drill that has actually been run.
The contract with the NOC
Anatomy of an alert
An alert you can't act on is noise with a timestamp. Every Trigpoint alert
answers six questions, or it doesn't fire:
What deviated. Metric, path set, and magnitude against the learned baseline — "+18 ms vs the nearest mode, ≈60σ", not "threshold exceeded".
When it started. To the sample, including the gap between onset and confirmation.
Blast radius. POPs, services, and customers affected — from inventory, not guesswork.
Evidence checked. BGP: no change. IS-IS: no events. Interface et-0/0/1: output drops rising. Every source consulted, with its result — including the exculpatory ones.
Ranked likely cause, with confidence. Congestion, route change, hardware, service failure — or, honestly, unexplained.
The drill-down. Links straight to path forensics: raw samples around the window, traceroute history, the routing events that coincided.
And the loop closes: operators confirm or reject every hypothesis, that verdict
is recorded next to the evidence, and the correlator's hit rate is itself a
tracked metric. The system earns trust the same way an engineer does — by being
right, and by being checkable when it isn't.
One platform
Three questions every operator has to answer
Service Trust
"Is the network delivering the experience it should — POP to POP, POP to Internet?"
The continuous measurement mesh, self-learning baselines, anomaly
detection, and path forensics. Deploy a container in every POP; get a
self-baselining anomaly map of your network.
Routing Trust
"Is the control plane doing what it should?"
Observed BGP checked against RPKI validation state over RTR — an
ROV-invalid origin on your prefixes or a customer's becomes an alert.
Route-churn and flap scoring name the chronically unstable prefix or
adjacency a snapshot hides. External vantage over public RIS feeds
answers "does the Internet see us correctly?" while the internal mesh
answers "are we seeing ourselves correctly?" — and anycast catchment
validation tracks which instance each POP actually reaches.
Configuration Trust
"Are the devices configured the way we intend?"
Golden-config drift detection, compliance and security-posture checks,
and digital-twin queries — "if this link fails, what happens to these
paths?" — grounded in the same live topology the correlator already keeps.
Both halves run today. What-if: read-only failure simulation — single
element, a composed multi-failure scenario, or a planned link that
doesn't exist yet — over the live IGP graph, plus a resilience audit
naming the single points of failure. Drift: running configs snapshotted
and hashed against an operator-blessed golden per device, changes
surfacing as a red badge on the topology map with the diff attached —
on its first day in the lab it caught a test fixture leaving residue
behind that the fixture's own revert claimed to clean up. Compliance
rules and intent modeling remain the deliberately-later part.
Why another monitoring tool
Inside-out, not outside-in
Enterprise synthetics platforms watch the Internet from the outside: useful if
you're a company consuming networks, incomplete if you're the company running
one. They don't speak your IGP, they can't see your route reflectors, and they
have no idea which customers ride which links.
Trigpoint is built for the other side of the demarc. It understands
your AS, your IS-IS topology, your POPs, and
your customer inventory — so when a path degrades, the answer isn't
a red dot on a world map. It's "the ATL–CHI link changed transport path at
02:13, no IGP or BGP involvement, 14 customers affected, and here is the
evidence for each claim."
No product today — commercial or open source — combines a synthetic
measurement mesh with the operator's own routing state for root-cause
correlation. That's the gap Trigpoint exists to close.
Architecture
Small parts, sharp edges
POPs
trig-agenticmp · udp/dscp · tcp dns · http/tls · trace
trig-agentstatic binary, mTLS, signed schedules, spool
× every POPLinux box, container host, or k8s edge
resultsmTLS gRPC
Control plane
Schedulertiered mesh · adaptive escalation · signed tests
Ingest gatewayresult streams → storage
ClickHousemeasurements + events, one store, joinable
Context ingestBMP · BGP-LS · gNMI · RPKI/RTR · RIS warehouse inventory in, SLA out
Grafanawhere NOCs already live — incl. incident annotations
Storage that matches the questions
Measurement tables cluster by path and time — the shape of "this path,
this window". Event tables cluster by time — the shape of "what happened
between 02:11 and 02:14". Raw data keeps 14 days, minute rollups 90,
hourly rollups two years, because "was it always like this?" is a real
question and deserves a real answer.
Deployment without ceremony
Docker Compose for small footprints, Helm for large ones; agents as
containers or systemd units. Everything is Rust — the same properties that
make the agent trustworthy make the control plane cheap to run. Ingest for
a 50-POP tiered mesh is about 38 rows a second; sizing is not the risk.
Where it stands
Being proven the honest way
Trigpoint is in active development and validated continuously against a lab
network with thirteen scripted, labeled fault classes: latency steps, bufferbloat
congestion, clean packet loss, IS-IS reroutes, dead DNS at 03:14 with the WAN
untouched, an aggregation-uplink degrade only one attachment sees, a customer
prefix withdrawn upstream, a killed LSP, a starved QoS class, a degrading
transit provider, a hop that silently strips DSCP markings, and a resolver
that quietly stops validating DNSSEC. The detector catches every one with the
correct signature and the correct discrimination — and has held zero false
positives on a quiet mesh throughout.
The concept alert fires end-to-end, evidence and blast radius included. An
IS-IS reroute is caught as an adjacency-down event within a second, and SPF
over the live topology proves which paths crossed the failed link —
including a cross-pair set no time-window could have attributed. A customer
outage traced to a BMP withdraw event says "routing pull, not fault," naming
the prefix and the peers that saw it vanish. Mesh loss joins against interface
counters to name the dropping port. A dead LSP surfaces as the alert only a
service-plane mesh can produce: the VPN is down and the underlay beneath
it is provably clean. A starved EF class is called a class fault because
its best-effort twins stay healthy; stripped DSCP markings are caught by the
reflector echo with latency perfect; a degrading upstream is named because its
sibling transit is clean. A paired probe — one validly-signed name, one
deliberately bogus — tells a resolver that stopped validating apart from a
zone whose signatures expired, faults that are invisible to every
reachability check because resolution keeps "working". Each alert ships as a
plain-language narrative —
what happened, why we believe it, what to check — with the affected customers
attached and the raw evidence one click below.
The loop closes in the tooling that exists today: a NOC web UI with a live
mesh matrix, search over everything the platform knows (one keystroke, from
alerts to customers to path histories), one-click verdicts, and team
workflow — acknowledge, assign, escalate, annotate — so the alert list is a
shared work queue, not a read-only feed. Engineers get an interactive
topology map with read-only failure simulation, and a path explorer that
draws the inferred SPF route beside the observed traceroute and says
plainly when they disagree — honest uncertainty as a UI state, not a
footnote. Alerts reach humans through PagerDuty, Slack and Teams, open
per-incident ServiceNow and Jira tickets, and feed any JSON webhook or a
versioned SIEM/event-bus stream — routed by hypothesis, severity, scope,
customer or POP, and all of it behind a storm-control layer — dedup,
backoff, a circuit breaker — because the monitoring must never become the
incident. A trig CLI
covers bursts, triage, shift handoffs, and monthly SLA evidence reports; a
maintenance calendar with recurring windows keeps planned work from paging
anyone; and a tracked per-hypothesis hit rate means the correlator's
accuracy is itself a metric, judged by the operators it serves.
The operator plane is production-shaped too: directory (LDAP) sign-in with
PingID two-factor and viewer/operator/admin roles, every state-changing
action — verdicts, workflow, maintenance, rebaselines, logins — in one
audit stream, retention tiering with a rehearsed backup-and-restore drill
for the data that matters, and an intelligence layer that survives reboots
honestly — a restarted correlator reconciles every claim it left open, so
history never shows an alert that nobody closed. The AI layer arrives
trust-first: a redaction gateway pseudonymizes every name, prefix and
topology fact to a stable token before a byte reaches any model — and it
now runs end-to-end on a local open-weight model, on-box, chosen by a
measured nine-model bake-off rather than fashion: alert narratives,
shift handoffs, plain-English queries compiled to a validated typed IR,
watermarked RFO drafts, and "have we seen this before?" recall that
pairs deterministic history joins with local-embedding similarity —
with the entity-leak check reading zero across every call class,
embeddings included, and every outbound prompt byte-identical across
restarts.
The newest wave turns that same discipline on the control plane itself.
Routing trust now runs in the lab: announcements checked against RPKI over
RTR, so an ROV-invalid origin on your prefixes or a customer's becomes an
alert; route-churn and flap scoring that names the chronically unstable
prefix or adjacency a snapshot hides; an external vantage over public RIS
feeds asking whether the Internet still sees your prefixes; and anycast
catchment validation that catches an instance flip while the address itself
stays up. Alongside it, a reporting spine: evidence-grade SLA reports —
measured and outage-based availability, merged-percentile latency, an
incident-and-verdict ledger, a method appendix — plus executive,
operational, compliance, transit and per-customer templates, every one
frozen to an immutable archive. The computed model is stored, not the
query, so a re-issued report is byte-identical years later, after the raw
samples have aged out — because a report that isn't frozen isn't evidence.
Closed incidents draft their own RFO. And change windows now carry the
MOP's device list: suppression follows the topology, so an unrelated fault
in the same POP still pages; the blast radius is pre-computed when the
window is declared; a path that fails to recover across something you
touched is flagged a rollback candidate; and a resilience audit names the
single points of failure and single-homed customers before the change,
not after.
Where these surfaces reach outside the box, the site says so plainly. The
paging and ticketing integrations — PagerDuty, Slack, Teams, ServiceNow,
Jira — and the warehouse sync in both directions are built against each
vendor's published contract and proven against mocks, not yet pointed at a
live tenant. The real-iron measurements — STAMP against line-card
reflectors, BGP-LS from a live speaker, gNMI optics, router-reported link
delay, SR-steered probes — are built behind an already-normalized model,
waiting only on hardware to swap a lab stand-in for the real feed. And a
documented, versioned read API with a pip-installable client lets a Nornir
or Ansible playbook gate a change on "is this path healthy?" without opening
the UI. Each surface is named for exactly what it is, because the whole
argument of the product is that it tells you the truth.
That bar — a zero-false-positive week followed by caught, explained,
customer-attributed degradations — is the standard the production release has
to clear. If you run POPs and want a self-baselining anomaly map of your own
network, we'd like to talk while the roadmap is still wet.