Detection rules
A detection is a discrete event the collector puts on the timeline: a “look here”, never a continuous series and never a verdict. For each of the eleven rules: when it fires, what the deployment must provide for it to fire at all, and what it cannot tell you. The derived values some rules build on are on Derived values.
Detection fields
Section titled “Detection fields”Every detection has the same fields:
| Field | Meaning |
|---|---|
rule |
the rule that fired |
key |
the CPU, core, zone or port it is about; empty for a device-wide rule |
seq, wall_ns |
the sample that raised it |
value |
the quantity the rule compared |
threshold |
what it compared against: the rule’s own threshold, written into every event |
message |
the event in words |
Scroll sideways to see every column
Once per rule and key per 10 s. After a rule fires for a key, the same rule and key are suppressed for 10 s of the samples’ wall clock, so a condition that persists fires every 10 s rather than on every sample. The stage counts what it suppressed, but no sink exports that count.
The rules run in the collector process and their history lives there. A collector that restarts starts every trailing window, bin and previous value from nothing.
Destinations
Section titled “Destinations”| Sink | Form |
|---|---|
InfluxDB, Telegraf, stdout lp |
mikroscope_ with value, threshold, seq, message |
SQL (--sql, --postgres) |
a mikroscope_detection row |
file, stdout json |
a {"detection":…} line |
| Prometheus | mikroscope_, every rule at 0 from the first scrape |
| Loki | a line in the source="detection", level="warn" stream |
| OTLP | a mikroscope. delta sum of 1 |
| Graphite | detection.<rule> = 1 at the event’s second |
| Elasticsearch | a document with kind: detection |
Scroll sideways to see every column
The dashboards draw every detection as an annotation, and one of the alert
rules fires on any detection except microburst and
ipc-collapse: those two describe how a healthy router carries traffic, and they are drawn and
stored but do not page.
| Rule | Key | Fires when | Needs |
|---|---|---|---|
counter-reset |
— | the sample reports a counter that went backwards without a 32-bit wrap | any deployment |
agent-restart |
— | the sequence number went backwards | any deployment |
agent-oom |
— | the container’s own cgroup recorded an OOM kill | cgroup2 in the container |
microburst |
cpu<N> |
three burst samples on one CPU within 60 s |
softnet |
reboot |
— | the kernel’s boot id changed, or a kernel-log record’s since-boot clock is lower than the previous record’s | any deployment; the kernel-log path needs privileged=yes |
link-flap |
port | two or more link up/down records on one port within 60 s | needs privileged=yes |
conntrack-cliff |
— | nf_conntrack fell below half its previous stored value |
needs privileged=yes |
conntrack-high |
— | occupancy above 80 % of nf_conntrack_max and rising over the last 60 s |
needs privileged=yes |
thermal-high |
zone | a zone within 15 % of its own declared critical trip | a thermal zone that declares a trip |
thermal-rising |
zone | three consecutive one-minute rises of more than 1 °C each | a thermal zone |
ipc-collapse |
core<N> |
a core’s one-second IPC below half its trailing median while its cycle rate is above its median | the PMU |
Scroll sideways to see every column
counter-reset
Section titled “counter-reset”Fires when the kernel sample’s resets is above 0: the agent found a counter lower
than its previous read without a 32-bit wrap to explain it, and used the counter’s
post-reset value as that tick’s delta, a lower bound. value is the number of such
counters, threshold 0.
Needs nothing beyond a sample. The same condition marks the sample suspect, and the
per-packet derived values are withheld for it.
May not claim which counter reset, or why. Every delta in that sample is a lower bound.
agent-restart
Section titled “agent-restart”Fires when a sample’s sequence number is lower than the previous sample’s. value is
the new sequence number, threshold the previous one.
Says what the collector held from the 30 s before the restart, at the end of the
message: the CPU’s busy share over every core and the busiest core’s, MemAvailable (the last
and the lowest) against MemTotal, nf_conntrack against its limit, the softnet drops and
squeezes, any OOM kill or allocation stall, and how long no sample came after the last of
them. A reboot takes the agent’s ring and RouterOS’s own log with it, but not what the
collector had already pulled. The time without samples is left out when the router’s clock
after the restart is earlier than before it, as on a router whose clock NTP has not set yet.
When the agent reports the kernel’s boot id, the message also says whether it changed: the
same id means the router did not reboot, and the agent’s container alone restarted (an
upgrade, a stop and a start, its restart policy). It says so only with two ids to compare: the
first agent to report one, after an agent that did not, says nothing about it.
Needs the collector to have seen at least one sample before the restart. The agent’s
sequence starts again from 1 on every launch, so this is what a restart looks like from the
outside. On its next health read, one minute at most, the collector rewinds its pull cursor
to the new ring’s oldest sample and logs agent restarted: its newest sample is N and the cursor was M; resuming from K; the new agent’s low sequence against the previous one is what
this rule matches (TestResyncAfterAgentRestart covers the collector side).
May not claim why the agent restarted. The summary states readings, not a cause. A collector restarted at the same time has no previous sequence number and sees nothing.
agent-oom
Section titled “agent-oom”Fires when the container’s own cgroup records an OOM kill in the sample — a process
inside mikroscope’s container was killed by the kernel. value is the number of kills.
Needs cgroup2 readable in the container; without it the agent reports no cgroup events and this rule cannot fire.
May not claim anything about the numbers around it: every number in that window is suspect. Sizing the container’s memory to the agent’s ring is on the cost of the observer.
microburst
Section titled “microburst”Fires when a CPU’s sample carries the burst flag —
a drop, or squeezes above that CPU’s trailing 90th percentile and at least 3, while its
packet count was at or below its trailing median — and that CPU now has at least three
flagged samples within the last 60 s. value is the number of flagged samples in the
window, threshold 3; the message carries the latest sample’s squeezes, drops, packets
and the trailing median.
Needs /proc/net/softnet_stat, which every deployment reads, and ten samples of
history per CPU before its first flag. The baselines span ten seconds of wall clock at any
sampler rate.
False positives. Squeezing can be a router’s background, not an event. On the router
measured, most per-CPU samples carry no squeeze, and time_squeeze is 1 in 11.2 %,
2 in 1.2 % and 3 in 0.21 %. A trailing window of a
distribution that is mostly zeroes has a 90th percentile of 1, so “above p90” is satisfied
by any 2 — which is why the floor of 3, and not the percentile, is what the rule runs on.
Replayed over stored samples, a floor of 2 fires 77.7 /h and a floor
of 3 fires 0.5 /h. A drop flags on its own, at any squeeze count.
May not claim the size of the burst, the flow or the interface that caused it. It says the kernel ran out of budget more than it usually does while carrying fewer packets than usual — evidence of something shorter than the sample interval.
reboot
Section titled “reboot”Fires when the kernel’s boot id differs from the one the collector read before. The
agent reads it once at start and publishes it in
/healthz; the kernel draws a new one at every
boot and keeps it until the next, and a container shares the router’s kernel, so a new id
means the router rebooted, not only the agent’s container. The collector logs router rebooted: the kernel's boot id went from … to … on the health read that sees it, and the
rule fires on the first sample of the agent that came back. value and threshold are 0.
It also fires when a kernel-log record’s timestamp, microseconds since boot, is lower than
the previous record’s, the one path for an agent that does not report the id; then value
and threshold are the new and previous timestamps in seconds. A reboot the boot id has
already reported does not fire it a second time. Either way the message carries the same
summary of the 30 s before as the agent-restart that came with it.
Says, when the collector has an API tier, what RouterOS logged about the boot, quoted after
the reason: who or what rebooted it (RouterOS logged at boot: "router rebooted by ssh-cmd:admin@192.168.88.10/), or that it went down without a shutdown ("router was rebooted without proper shutdown"). The detection waits for that read, two minutes at most, since
the API connection comes back after the router (Boot log).
Without an API tier it fires on the agent’s first sample, with no such clause.
Needs a collector that keeps running across the reboot while the agent comes back, and no
RouterOS API credentials. The agent that comes back is a new process; the collector reads its
boot id in the same health read that rewinds the cursor to the new ring, within a minute (see
agent-restart). The kernel-log path needs privileged=yes, which the
kernel log requires. In the lab, a /system/reboot fired it through the boot id, and a stop
and a start of the agent’s container did not
(Tested on).
May not claim that every reboot is seen, or why the router rebooted: RouterOS’s line names who rebooted it or says it went down without a shutdown, not why the power went or the router hung. A collector started after the reboot has no earlier id to compare, and two reboots between two health reads, a minute apart, are one change of id. The kernel-log path sees only what the agent reads from the end of the log after it starts: when the first record after a reboot has a since-boot time later than the last one before it, nothing fires that way.
link-flap
Section titled “link-flap”Fires when a kernel-log record that names an interface is classified link-up or
link-down — the same classifier that puts a kind on every port record — and that port
now has two or more such records within the last 60 s. key is the port’s current
RouterOS name where the API tier’s interface inventory supplies one, the board’s default
name where the agent’s port table maps the kernel name, and the kernel name otherwise.
value is the number of records in the window, threshold 2.
Needs privileged=yes. A RouterOS name needs the board to be in the agent’s port
table, and the current name needs the API tier as well; see Port
names.
May not claim a fault. A cable pulled and reseated within a minute is a down and an up record, and fires. The rule has fired on unprovoked flaps without saying whether the cable, the device at the other end or something else caused them; no provoked flap has been captured with it running (Tested on).
conntrack-cliff
Section titled “conntrack-cliff”Fires when the nf_conntrack slab cache’s active-object count is below half its
previous stored value. value is the new count, threshold the previous one.
Needs privileged=yes, for /proc/slabinfo. The slab is read at about 6 Hz and stored
on change, so “previous” is the previous stored sample, not the previous tick.
May not claim a fault either: the message says “a flush or a reset”, and a deliberate flush of the connection table fires it.
conntrack-high
Section titled “conntrack-high”Fires when nf_conntrack active objects are above 0.8 of the kernel’s
nf_conntrack_max, and the count is higher than the oldest stored value in the last
60 s. value is the occupancy as a fraction, threshold 0.8.
Needs privileged=yes, the ceiling published by the kernel, and at least two stored
samples within the last 60 s.
May not claim when the table will be full: no time-to-full is attached, on purpose. For scale, a measured router’s table held 6 287 entries, 0.65 % of its kernel’s ceiling.
thermal-high
Section titled “thermal-high”Fires when a zone’s reading is at or above 0.85 of that zone’s own lowest declared
critical trip point. value is the reading in °C, threshold 0.85 × the trip.
Needs a thermal zone that declares a critical trip. A zone that declares none never fires; nothing is compared against a compiled number.
May not claim that cooling has failed, or anything about a zone the board does not report.
thermal-rising
Section titled “thermal-rising”Fires when a zone’s last four completed one-minute means each exceed the one before by
more than 1 °C — three consecutive rises. value is the rise from the first of the four
means to the last, threshold 3. A one-minute bin closes on the first sample at least 60 s
after it opened, and the rule is checked each time one closes, so the earliest it can fire
is after about four minutes of readings.
Needs a thermal zone. The means are over the readings the samples carried; temperature
is read at the zone’s declared polling cadence (mikroscope_).
Resolution. A sensor that quantises to about 0.42 °C, as the one measured does, resolves 1 °C per minute in 2.4 steps.
May not claim a cause, or a rise slower than 1 °C per minute.
ipc-collapse
Section titled “ipc-collapse”Fires when, on one core, a one-second bin closes with instructions per cycle below half
the median of that core’s trailing bins while its cycle rate is above the median of its
trailing rates. key is core<N>, value the bin’s IPC, threshold half the median.
Needs the PMU’s cycles and instructions per CPU (under privileged=yes), and twenty
completed one-second bins of history for that core before it can fire; the trailing
history holds up to sixty.
May not claim idleness, or anything pooled: the conjunction with the cycle rate is what separates a memory-stall regime from a core going quiet, and the rule is per core, never across cores.
Doctor checks vs rules
Section titled “Doctor checks vs rules”Standalone doctor reads the running agent’s ring and runs four checks of its own:
layer2-loop, stp-churn, link-flap and softnet-drops, described in Health
checks. They
are not detections: they run once, when doctor asks, over whatever the ring holds up to its
newest 10 000 samples, 60 s by default, and write nothing to a sink.
One name is shared and the rule is not. doctor’s link-flap counts link-downs only, two
or more on one port anywhere in the ring; the detection counts link-ups and link-downs
together, two or more within 60 s, so a cable pulled and reseated once fires the detection
and not the check. A layer-2 loop is none of the eleven rules: doctor’s layer2-loop
names it in the ring, and over history the mikroscope-l2-loop alert
rule watches the store for it.
Four of the eleven rules have fired on a real router, and only microburst has had its
behaviour measured against the samples. The other seven are exercised only by the derive
stage’s unit tests against constructed samples, and no fault has been provoked with the rules
running (Tested on).