Skip to content

Triggered capture

Set a condition on the agent, and it keeps the samples around the moment the condition fires, at full rate, without a recording running. Fetch the capture over HTTP and draw it with plot.

  • A capture is the samples that already exist, kept at full rate around the moment a condition fired. Only the agent can keep them, because only the agent has every sample. The ring already holds the last 60 s by default, so the seconds before a fire cost nothing; the seconds after cost only the wait. The default window is 5 s either side. A CAPTURE_PRE_S longer than the ring (--buffer) is a window the ring cannot supply.
  • It decides nothing about meaning. A condition is a comparison you configured. The field it compared and the value that tripped it travel in the capture’s header. The capture is the same raw delta lines /snapshot ships; nothing is turned into a percentage or a verdict.
  • It does not copy lines. The ring stores each sample as a pre-encoded, immutable line, and a capture pins the lines it needs, so a fire costs a copy of the entry headers once and nothing per tick. The design estimates that at about 3 µs for a 10 s window at 10 Hz.
  • Conditions run in the sampler’s own loop, between reading a sample and pushing it into the ring, never in a second goroutine. There is no expression language, on purpose: a parser is a dependency and an attack surface, and an operator-writable expression on the sampler’s hot path is a way to make the router slow.

The agent reads six variables from its container’s envlist. Two of them have an install flag.

Agent variable install flag Default Accepted What it sets
TRIGGERS --triggers softnet-drop,oom,kmsg<=3,reset,irq-err,flash-bad the conditions below, comma-separated Which conditions arm a capture.
CAPTURE_MB --capture-mb 4 0–256 The budget of pinned ring bytes, in MiB. 0 turns the feature off.
CAPTURE_PRE_S none 5 1–60 Seconds kept before the sample that fired.
CAPTURE_POST_S none 5 1–60 Seconds kept after it.
CAPTURE_POLICY none first first, last On a full budget: first refuses the new capture, last evicts the oldest.
TRIGGER_REFRACTORY_S none 10 0–3600 Quiet time per condition after it fires.
  • install always writes CAPTURE_MB, and writes TRIGGERS only when --triggers is given; without it the agent uses its default set. It writes none of the other four, so an installed agent runs with their defaults. mikroscope plan shows the envlist entries before anything is written.
  • --triggers goes through the agent’s own parser before the first connection. An unknown condition, a threshold out of range, a quote or a semicolon fails the command with exit status 2 and writes nothing.
  • The agent parses the list again when it starts, because an envlist can be edited by hand on the router. A value it rejects there makes it refuse to run, with one line on its standard output, which RouterOS puts in its log.

The default set is every condition without a threshold except squeeze — softnet-drop, oom, reset, irq-err and flash-bad, each firing when the kernel counts something it normally does not — plus kmsg<=3. The level conditions are not in it: their thresholds are yours to choose.

Condition Fires when field in the header Threshold In the default set
softnet-drop any softnet queue dropped a packet in the sample softnet[N].dropped none yes
squeeze any softnet queue ran out of budget (time_squeeze) in the sample softnet[N].time_squeeze none no
oom the kernel OOM-killed something (vm.oom_kill moved) vm.oom_kill none yes
reset a counter went backwards in a way that is not a 32-bit wrap resets none yes
irq-err the Err row of /proc/interrupts moved irq_err none yes
flash-bad a YAFFS partition’s bad-block count rose since the previous sample flash[<device>].bad_blocks none yes
kmsg<=N a kernel-log record at severity N or more severe (0 is emergency, 3 error) events.level 0–7 kmsg<=3
busy>=X any core’s busy ratio is at or above X cpu[N].busy_ratio 0.05–1 no
slip>=X the sample’s interval was at least X sampler periods dt_ns/period 1.1–100 no
memfall>=N MemAvailable fell by N or more in one tick mem.MemAvailable fall (MB) 1–100000 no
  • Where a condition covers several cores, queues or partitions, the header names the first one that matched.
  • memfall compares /proc/meminfo’s kB divided by 1 024, so its N is in MiB although the field calls it MB.
  • kmsg<=N needs the kernel log, which the agent reads only in a privileged container (Privileged mode); flash-bad needs a YAFFS partition. A condition whose source is absent never fires.
  • squeeze is not a default because time squeezes can be background on a healthy router (measured), and a rule that fires on any squeeze then fires constantly. The collector’s microburst rule asks for an episode of three deviations instead.
  • A capture can also be armed by hand, with POST /capture (below). Its cause is manual and its field is the reason you gave.
  1. A condition is true on sample S. If that condition fired fewer than TRIGGER_REFRACTORY_S × rate samples ago (that many seconds at the nominal rate), the fire is suppressed (reason refractory). If another capture is still collecting its window, the fire is suppressed (reason pending): one capture collects at a time, whatever condition armed it.
  2. Otherwise a capture is armed for the window from S − CAPTURE_PRE_S × rate to S + CAPTURE_POST_S × rate, and a {"trigger":{…}} line is queued for the stream.
  3. When the sample at the end of the window is in the ring, the capture pins the ring’s lines for that window. If the ring no longer holds any of them, the capture is refused (empty).
  4. If the window’s bytes alone exceed the budget, it is refused (budget). If the budget is full, first refuses it (budget) and last evicts the oldest captures until it fits.

A capture whose window holds fewer than (pre + post) × rate + 1 samples is kept with complete: false rather than silently short. That happens when the ring did not hold the whole window: a fire within CAPTURE_PRE_S of the agent starting, or a ring (BUFFER_S) shorter than the window. When the agent stops, it collects a pending capture with what the ring holds, but it cannot serve it: captures are in memory, and the HTTP server stops with the agent.

Captures live in the agent’s memory. Nothing writes them to disk, so a restart of the container, an upgrade or a reboot loses the ones not yet downloaded.

A capture’s size is its window’s sample count times the line size. A line with every source on, the PMU included, is 3 230 B. By arithmetic from that line:

Rate Default window (5 s + 5 s) Captures in the 4 MiB default
10 Hz 101 samples, about 330 kB 12
50 Hz 501 samples, about 1.6 MB 2
100 Hz 1 001 samples, about 3.2 MB 1

A bigger board — more cores, more interrupt lines — has longer lines. The bytes field of each capture is the real figure.

The budget is memory the agent holds beyond its ring: a pinned line stays alive after the ring has moved past it, so the agent checks both together at start.

  • When the agent can read the container’s memory.max and about rate × buffer × 3 456 B plus CAPTURE_MB exceeds it, the agent refuses to start, naming the three settings to lower or --memory-max to raise.
  • It warns when twice that exceeds its Go soft memory limit, MEM_LIMIT_MB (8 to 1 024 MiB). The agent’s own default for MEM_LIMIT_MB is 14. install derives it from the ring instead (rate × buffer × line, × 2.5, at least 16, at most three quarters of --memory-max), which writes 16 at the defaults, and --mem-limit-mb sets it by hand.
  • That derivation does not count CAPTURE_MB, so a large capture budget can still draw the warning. Agent cost explains why that second ratio matters.

Four endpoints on the agent. Each needs the bearer token when the agent has one, and each answers 404 with captures disabled (CAPTURE_MB=0) when the feature is off.

Request Answer
GET /captures The index, as JSON.
GET /captures/<id> One {"capture":{…}} header line, then the sample lines verbatim, as NDJSON.
DELETE /captures/<id> Frees that capture’s share of the budget; 204.
POST /capture?reason=… Arms a manual capture at the newest sample: {"id":N,"armed":true}, or 409 when a capture is pending or the manual trigger is in its refractory window. The reason defaults to operator.
Terminal window
curl -s -H "Authorization: Bearer $MIKROSCOPE_TOKEN" http://172.30.10.2:9123/captures
curl -s -H "Authorization: Bearer $MIKROSCOPE_TOKEN" http://172.30.10.2:9123/captures/3 > cap3.jsonl
mikroscope plot --in cap3.jsonl
curl -s -X DELETE -H "Authorization: Bearer $MIKROSCOPE_TOKEN" http://172.30.10.2:9123/captures/3

The index carries the policy, budget_bytes, the bytes held, the capture still pending if there is one, the configured triggers, and one entry per capture: id, cause, condition, field, value, threshold, fire_seq, fire_mono_ns, fire_wall_ns, first_seq, last_seq, samples, bytes and complete.

The sample lines of GET /captures/<id> are byte-identical to what /snapshot serves for the same samples, so a tool that reads a snapshot needs no new parser; plot skips the header line. The CLI has no command for captures: use any HTTP client.

When a capture is armed, the agent places one line before the sample it fired on, in /stream and in /snapshot?since= (not in /snapshot?seconds=):

{"trigger":{"id":3,"cause":"busy>=0.95","field":"cpu[2].busy_ratio","value":1,"threshold":0.95,"seq":48213,"wall_ns":1789000000000000000}}
  • The values above are illustrative. The line is a kind of its own, like the {"gap":…} line, not a field on the sample, so the sample schema is unchanged.
  • The agent keeps the last 64 trigger lines for pullers, so a puller more than 64 fires behind never sees the older ones. A suppressed fire produces none.
  • A manual capture fires on the newest sample already in the ring, so a puller that has already received that sample gets no trigger line for it; read /captures instead.

forward recognises the line, never mistakes it for a sample, counts it, and hands it to every sink as an annotation. The capture itself stays on the agent, under /captures/<id>.

Sink Trigger record
InfluxDB the mikroscope_trigger measurement
SQL the mikroscope_trigger table
Prometheus (the collector’s) mikroscope_collector_triggers_total{cause}
Elasticsearch a trigger document
Graphite trigger.<cause>
Loki a source="trigger" line
OTLP the mikroscope.trigger.fired{cause} sum
file the line itself

All five Grafana dashboards carry a triggers annotation, off by default in the toggle bar, beside a detections one that is on. Each reads its own store:

Dashboard Annotation query
InfluxDB, PostgreSQL the mikroscope_trigger rows
Prometheus the collector’s mikroscope_trigger_fired_total
Elasticsearch the Lucene filter kind:trigger AND host.keyword:$host, with the document’s field as text and its cause as tag (kind:detection, message and rule for detections)
Graphite aliasByNode($prefix.$host.trigger.*, 3): one marker per point titled with the cause, with no text, because Graphite holds no string (detection.* and the rule for detections)

The Elasticsearch and Graphite annotations are checked as generated JSON; they have not been run in Grafana.

record does not recognise the line: Record, mark and plot says what it does with it.

The collector’s /metrics carries the families that say how much the captures did not see, built from the agent’s GET /sampler counters, which the collector reads at start and every minute. Every condition and reason pair is rendered from the start, at 0 until it happens, so a dashboard can show “0 so far”.

Family Type Meaning
mikroscope_trigger_fired_total{condition} counter Times each condition armed a capture; condition="manual" appears once a manual capture has been armed.
mikroscope_trigger_suppressed_total{condition,reason} counter Times a condition was true and nothing was armed: refractory or pending.
mikroscope_capture_refused_total{reason} counter Captures collected and then not kept: budget or empty.
mikroscope_captures_held gauge Captures currently retained.
mikroscope_capture_bytes gauge Ring bytes the retained captures pin.
mikroscope_capture_budget_bytes gauge The budget, from CAPTURE_MB.
mikroscope_capture_bytes_served_total counter Bytes handed out over /captures/<id>.

The collector cannot recompute these from the samples; it reads them from the agent’s GET /sampler and renders them into its own exposition, so one scrape job on the collector carries them. Prometheus has the detail.

  • The capture set is a sample of events, never a census. The refractory window and the byte budget bound a trigger storm, and one capture collects at a time. mikroscope_trigger_suppressed_total and mikroscope_capture_refused_total are how much was not seen.
  • Full rate is not full detail. A capture holds samples at the sampler’s rate: at 10 Hz nothing shorter than 100 ms is reliably visible, and a busy ratio still moves in the kernel’s tick steps. Resolution limits sets that limit, not the capture.
  • A window can be short. One the ring did not wholly hold is served with complete: false. One cut off by the agent stopping is collected but never served, because the captures stop with the agent.
  • Downloading costs the router. A download runs on the same core as the sampler, and is counted in mikroscope_capture_bytes_served_total the way any puller is charged.

The cost of a fire on the device, a sustained trigger storm, captures at 50 or 100 Hz and the sizes of real captures are listed under Not tested.