Alert rules
mikroscope dashboards gen writes, beside each dashboard, a Grafana unified-alerting provisioning
file with the rules that follow from the dashboards’ own fault counters, from the kernel log the
agent reads, and from the collector’s detections. Every threshold is zero (a counter that should
not move), one sample (the silent-agent rule), a share of a ceiling the device itself published,
or, for one rule, a multiple of the device’s own last day. None is a number compiled in for one
router.
| File | Rules | Query language |
|---|---|---|
dashboards/ |
14 | InfluxDB 3 SQL |
dashboards/ |
15 | PromQL |
dashboards/ |
10 | PostgreSQL SQL |
Scroll sideways to see every column
The InfluxDB file has one rule fewer because “The sampler is slipping ticks” has no SQL form yet.
The counter itself does reach InfluxDB — the collector reads it from the agent’s /sampler and
writes mikroscope_ — so the rule could be written; it has not been.
The PostgreSQL file is translated from the InfluxDB one, and it also leaves out the rules the SQL
schema cannot answer in the same shape. mikroscope-l2-loop and mikroscope-port-link-down sum a
per-sample count of kernel-log port records, and the SQL sink writes one row per kernel record
instead, in mikroscope_event. mikroscope-port-errors names one column per MAC counter, and the
SQL sink writes mikroscope_api_ifcounter one row per counter. mikroscope-bridge-port-dark is
left out for the same reason: it names rx_packet, tx_unicast and tx_broadcast as columns, and
reads bridge, which that table does not have either. The translation drops those rather than
write a query that looks right and answers something else. A unit test checks every column each
PostgreSQL alert query reads against the tables the SQL sink declares; it reads the schema, not a
database (not run against PostgreSQL).
Each file is apiVersion: 1 with one rule group, mikroscope, in a folder named mikroscope,
organisation 1, evaluated every minute. The rules are provisioned rather than built into the
dashboards, so an operator who wants none copies nothing. forward --grafana,
dashboards publish and dashboards import install none of them, and
uninstall --targets dashboard removes none.
Install the rules
Section titled “Install the rules”-
Replace the datasource placeholder with your datasource’s UID (
mikroscope-<store>when the collector ordashboards publishcreated it) and write the file into Grafana’s provisioning directory. Provisioning files do not resolve a dashboard’s${DS_MIKROSCOPE}input, so the datasource is a literal placeholder,DS_UID_PLACEHOLDER:Terminal window sed 's/DS_UID_PLACEHOLDER/<uid>/g' dashboards/mikroscope-alerts-influxdb.yaml \> /etc/grafana/provisioning/alerting/mikroscope-alerts-influxdb.yamlUse the file that matches the datasource the UID belongs to.
/etc/is the path the generated file’s own header names; the directory is Grafana’s.grafana/ provisioning/ alerting/ -
Restart Grafana, or ask its Admin API to reload the provisioned files. The file format and the reload are in Grafana’s alerting file-provisioning guide.
Every rule has the same shape, the one Grafana’s own rule editor writes:
- A — the query, against your datasource, with a relative time range of the last 600 s. Every SQL query and every Prometheus counter query also bounds its own window (1, 2, 5 or 10 minutes, or 1 hour, below; the wake-up rule also reads the 24 hours before its 10 minutes, whatever A’s 600 s range says); the two Prometheus gauge rules, thermal and conntrack, read the latest value, and the egress rule takes its gauge’s maximum over the last minute.
- B — reduce A to one number per series with
last, dropping non-numeric values. - C — compare B against the threshold. C is the rule’s condition.
Each rule carries the labels severity (critical or warning) and source: mikroscope, a
summary annotation, and execErrState: Error. What Grafana then does with a rule whose query
fails is Grafana’s behaviour, set out in its own documentation of the Error and No Data
states
(not tested here).
| Rule (uid) | Fires when | Threshold (C) | Severity | for | No data means | Stores |
|---|---|---|---|---|---|---|
mikroscope-agent-silent | fewer than 1 new sample reached the store in the last 2 minutes | < 1 | critical | 2m | Alerting | all three |
mikroscope-softnet-drops | softnet_stat dropped a packet in the last 5 minutes | > 0 | critical | 0s | OK | all three |
mikroscope-oom-kill | /proc/vmstat oom_kill moved in the last 5 minutes | > 0 | critical | 0s | OK | all three |
mikroscope-detections | any detection in the last 5 minutes except microburst and ipc-collapse, which are drawn and stored but do not page | > 0 | warning | 0s | OK | all three |
mikroscope-thermal-near-critical | a zone at or above 85 % of its own critical trip | > 0 | critical | 1m | OK | all three |
mikroscope-conntrack-near-limit | nf_conntrack active objects above 0.8 of the kernel's limit | > 0.8 | warning | 5m | OK | all three |
mikroscope-agent-oom | the agent's own cgroup recorded an OOM kill in the last 5 minutes | > 0 | critical | 0s | OK | all three |
mikroscope-l2-loop | an own-address record on any port in the last 5 minutes | > 0 | critical | 0s | OK | InfluxDB 3, Prometheus |
mikroscope-port-link-down | a link-down record on any port in the last 5 minutes | > 0 | warning | 0s | OK | InfluxDB 3, Prometheus |
mikroscope-port-errors | any port's MAC counted a typed error — overflow, FCS, collision — for 5 minutes running | > 0 | warning | 5m | OK | InfluxDB 3, Prometheus |
mikroscope-bridge-port-dark | a bridge port received packets while the bridge sent it neither a unicast nor a broadcast frame, for 10 minutes | > 0 | warning | 10m | OK | InfluxDB 3, Prometheus |
mikroscope-wakeup-storm | the context-switch rate over the last 10 minutes is more than 4 times its mean over the 24 hours before, for 10 minutes | > 4 | warning | 10m | OK | all three |
mikroscope-egress-queue-drops | any port's own egress queue dropped a packet in every one of the last 10 minutes — sustained congestion, never a single burst | > 0 | warning | 10m | OK | all three |
mikroscope-ecc-failure | the NAND reported an uncorrectable ECC failure in the last hour | > 0 | critical | 0s | OK | all three |
mikroscope-ticks-slipped | the sampler slipped a tick in the last 5 minutes | > 0 | warning | 5m | OK | Prometheus |
Scroll sideways to see every column
“No data means” is the rule’s noDataState. The silent-agent rule is the one where silence is the
fault, so no data fires it; for every other rule no data is the healthy reading.
Each rule’s title, and under it its summary annotation, shortened:
- mikroscope agent stopped delivering samples. “No new samples reached the store in the last two minutes: the agent stopped, the collector stopped, or the path between them did. Every other rule is blind while this one fires.”
- Packets dropped in the kernel receive path. “softnet_stat dropped a packet: a per-CPU backlog was full. Unambiguous loss inside the router, invisible to every SNMP and RouterOS counter. Zero is the expected reading.”
- The kernel OOM-killed a process. “/proc/vmstat oom_kill moved: the kernel killed a process to get memory back. Which process is not knowable from the container (no PID namespace).”
- The collector’s derive stage flagged an event. “A detection rule fired (counter-reset,
agent-restart, agent-oom, reboot, link-flap, conntrack-cliff, conntrack-high, thermal-high,
thermal-rising). The rule, key, value and threshold are in the Detections section and on the
dashboard as an annotation.”
microburstandipc-collapseare drawn and stored like every detection but do not fire it: they describe how a healthy router carries traffic, and they were most of the detections over a measured day (measured). - A thermal zone is within 15 % of its own critical trip. “The reading is at or above 85 % of the zone’s declared critical trip point. The ceiling is the board’s own, read from /sys, not a number compiled in.”
- The connection table is above 80 % of nf_conntrack_max. “nf_conntrack active objects over the kernel’s own ceiling. Past the ceiling the router drops new connections. The limit is the sysctl the agent read, not a compiled number.”
- The sampler is slipping ticks. “Ticks finished after the next was due. The rate is not being delivered: the device is starved, the source set is too expensive for the rate, or the container’s CPU quota throttled the agent (see the observer’s throttling counter).”
- mikroscope’s own container was OOM-killed. “The kernel killed a process inside the agent’s
cgroup: the capture ring and the captures are gone, and every number in the window is suspect.
Raise
--memory-maxor lower RATE_HZ, BUFFER_S or CAPTURE_MB.” - The bridge received its own address back: a layer-2 loop signature. “The kernel log reported
received packet on <port> with own address as source address: a frame the router sent came back in, which is what a loop through a downstream switch or access point looks like. The port label says which cable. Read from /dev/kmsg by the agent, no API. The InfluxDB form needs a store that has held at least one port record classified by kind.” The layer-2 loop case study shows the signature. - A port’s link went down. “The kernel log reported a link-down on a port: a cable pulled, a peer rebooted or powered off, a renegotiation. The collector’s link-flap detection covers the repeated case; this is the single event. Read from /dev/kmsg by the agent, no API; the port, its comment and its role are on the Kernel log section’s port events.”
- A port is counting typed MAC errors. “A port’s MAC is counting typed errors: frames it could not take. The commonest on a switched LAN is rx-overflow, the receive FIFO filling faster than the chip can drain it, and it is a microburst signature rather than a load one. FCS errors and collisions mean something else: cabling, duplex, a dying port. Needs the API tier: a MAC counter is not visible from inside the container. It cannot tell a real error from a counter reset, so a router that reboots inside the window fires it once.”
- A bridge port is receiving but the bridge sends it nothing. “A port that is a member of a
bridge has received packets for ten minutes while the bridge delivered it neither a unicast nor a
broadcast frame. Whatever is behind it is transmitting and hearing nothing back: STP holds the port
discarding, or the bridge learned every host behind it on another path. RouterOS shows the port
running and error-free throughout.” A layer-2 loop reads this way on the port the bridge has
stopped delivering to (case study). It is expected to fire on a
redundant design, where an RSTP alternate port discards on purpose, and on a port with
broadcast-flood=noorhorizonset. Needs the API tier. - The kernel is switching context four times as often as over its last day. “Something started waking up very often: a busy-polling process or driver, a monitoring client, a container in a tight loop. Look at what started at the moment the rule fired (/user/active, new containers, a new integration) and at the Interrupts and softirqs section.” A monitoring integration polling the router over the API is one cause (case study). A deliberate change, such as a new container or a heavier ruleset, fires it too.
- A port’s egress queue has been dropping every minute for ten minutes. “A port’s own egress queue has dropped packets in every one of the last ten minutes. Unlike the other counter rules, this one’s counter is supposed to move: dropping is how a full queue tells a sender to slow down. What fires this is a link that is simply too small for what it is being asked to carry, or a shaper set below the traffic. Needs the API tier.”
- The NAND reported an uncorrectable ECC failure. “ecc_failures rose on an MTD partition: a read the error correction could not fix, i.e. data loss on the flash. Any increment is an incident.”
Queries
Section titled “Queries”# mikroscope-agent-silent (< 1)sum(increase(mikroscope_samples_total[2m]))# mikroscope-softnet-drops (> 0)sum(increase(mikroscope_softnet_total{kind="dropped"}[5m]))# mikroscope-oom-kill (> 0)sum(increase(mikroscope_vm_events_total{event="oom_kill"}[5m]))# mikroscope-detections (> 0)sum(increase(mikroscope_collector_detections_total{rule!~"microburst|ipc-collapse"}[5m]))# mikroscope-thermal-near-critical (> 0)count(mikroscope_thermal_celsius >= on(zone) 0.85 * mikroscope_thermal_critical_celsius)# mikroscope-conntrack-near-limit (> 0.8)max(mikroscope_slab_active_objects{cache="nf_conntrack"} / mikroscope_slab_limit_objects{cache="nf_conntrack"})# mikroscope-ticks-slipped (> 0)sum(increase(mikroscope_slipped_total[5m]))# mikroscope-agent-oom (> 0)sum(increase(mikroscope_self_oom_kills_total[5m]))# mikroscope-l2-loop (> 0)sum(increase(mikroscope_kmsg_port_records_total{kind="own-address"}[5m]))# mikroscope-port-link-down (> 0)sum(increase(mikroscope_kmsg_port_records_total{kind="link-down"}[5m]))# mikroscope-port-errors (> 0)sum(increase(mikroscope_api_interface_counter_total{counter=~"rx-overflow|rx-fcs-error|rx-fragment|rx-too-short|rx-too-long|rx-jabber|tx-fcs-error|tx-late-collision|tx-excessive-collision"}[5m]))# mikroscope-bridge-port-dark (> 0)count((sum by (interface) (increase(mikroscope_api_interface_counter_total{counter="rx-packet"}[10m])) > 0) and on (interface) (sum by (interface) (increase(mikroscope_api_interface_counter_total{counter="tx-unicast"}[10m])) == 0) and on (interface) (sum by (interface) (increase(mikroscope_api_interface_counter_total{counter="tx-broadcast"}[10m])) == 0) and on (interface) (mikroscope_api_interface_info{bridge!=""}))# mikroscope-wakeup-storm (> 4)sum(rate(mikroscope_context_switches_total[10m])) / sum(rate(mikroscope_context_switches_total[24h] offset 10m))# mikroscope-egress-queue-drops (> 0)max(max_over_time(mikroscope_api_interface{kind="tx_queue_drops"}[1m]))# mikroscope-ecc-failure (> 0)sum(increase(mikroscope_mtd_ecc_failures_total[1h]))All 15 read the collector’s /metrics, which is the
only exposition there is —
Set up in Grafana has the one
scrape job. mikroscope_slipped_total is among them: the agent owns the ticker and is the only
thing that can count a slipped tick, and it reports the number over /sampler, which the collector
reads every minute and renders. So the rule fires on a figure that is at most a minute old rather
than one scrape old, which for a counter that is 0 on a healthy device is the same alert.
-- mikroscope-agent-silent (< 1)SELECT count(1) AS value FROM mikroscope_cpu WHERE time >= now() - interval '2 minutes'-- mikroscope-softnet-drops (> 0)SELECT coalesce(sum(dropped), 0)::BIGINT AS value FROM mikroscope_softnet WHERE time >= now() - interval '5 minutes'-- mikroscope-oom-kill (> 0)SELECT coalesce(sum(oom_kill), 0)::BIGINT AS value FROM mikroscope_vm WHERE time >= now() - interval '5 minutes'-- mikroscope-detections (> 0)SELECT count(1) AS value FROM mikroscope_detection WHERE time >= now() - interval '5 minutes' AND rule NOT IN ('microburst', 'ipc-collapse')-- mikroscope-thermal-near-critical (> 0)SELECT count(1) AS value FROM (SELECT zone, max(celsius) AS c, max(critical_celsius) AS crit FROM mikroscope_thermal WHERE time >= now() - interval '2 minutes' AND critical_celsius IS NOT NULL GROUP BY zone) AS zones WHERE c >= 0.85 * crit-- mikroscope-conntrack-near-limit (> 0.8)SELECT max(active) * 1.0 / nullif(max("limit"), 0) AS value FROM mikroscope_slab WHERE time >= now() - interval '2 minutes' AND cache = 'nf_conntrack' AND "limit" IS NOT NULL-- mikroscope-agent-oom (> 0)SELECT coalesce(sum(oom_kill), 0)::BIGINT AS value FROM mikroscope_self WHERE time >= now() - interval '5 minutes' AND oom_kill IS NOT NULL-- mikroscope-l2-loop (> 0)SELECT coalesce(sum(count), 0)::BIGINT AS value FROM mikroscope_kmsg WHERE time >= now() - interval '5 minutes' AND kind = 'own-address'-- mikroscope-port-link-down (> 0)SELECT coalesce(sum(count), 0)::BIGINT AS value FROM mikroscope_kmsg WHERE time >= now() - interval '5 minutes' AND kind = 'link-down'-- mikroscope-port-errors (> 0)SELECT coalesce(sum(v), 0)::BIGINT AS value FROM (SELECT interface, greatest(max(rx_overflow) - min(rx_overflow), 0)::BIGINT + greatest(max(rx_fcs_error) - min(rx_fcs_error), 0)::BIGINT + greatest(max(rx_fragment) - min(rx_fragment), 0)::BIGINT + greatest(max(rx_too_short) - min(rx_too_short), 0)::BIGINT + greatest(max(rx_too_long) - min(rx_too_long), 0)::BIGINT + greatest(max(rx_jabber) - min(rx_jabber), 0)::BIGINT + greatest(max(tx_fcs_error) - min(tx_fcs_error), 0)::BIGINT + greatest(max(tx_late_collision) - min(tx_late_collision), 0)::BIGINT + greatest(max(tx_excessive_collision) - min(tx_excessive_collision), 0)::BIGINT AS v FROM mikroscope_api_ifcounters WHERE time >= now() - interval '5 minutes' GROUP BY interface) AS ports-- mikroscope-bridge-port-dark (> 0)SELECT count(*)::BIGINT AS value FROM (SELECT interface, max(rx_packet) - min(rx_packet) AS drx, max(tx_unicast) - min(tx_unicast) AS dtu, max(tx_broadcast) - min(tx_broadcast) AS dtb FROM mikroscope_api_ifcounters WHERE time >= now() - interval '10 minutes' AND bridge IS NOT NULL AND bridge <> '' GROUP BY interface) AS ports WHERE drx > 0 AND dtu = 0 AND dtb = 0-- mikroscope-wakeup-storm (> 4)SELECT CASE WHEN base_min >= 720 THEN (now_sum / nullif(now_min * 60.0, 0)) / nullif(base_sum / (base_min * 60.0), 0) ELSE 0.0 END AS value FROM (SELECT sum(CASE WHEN time >= now() - interval '10 minutes' THEN CAST(ctxt AS BIGINT) ELSE 0 END) AS now_sum, count(DISTINCT CASE WHEN time >= now() - interval '10 minutes' THEN date_bin(interval '1 minute', time) END) AS now_min, sum(CASE WHEN time < now() - interval '10 minutes' THEN CAST(ctxt AS BIGINT) ELSE 0 END) AS base_sum, count(DISTINCT CASE WHEN time < now() - interval '10 minutes' THEN date_bin(interval '1 minute', time) END) AS base_min FROM mikroscope_stat WHERE time >= now() - interval '1450 minutes') AS w-- mikroscope-egress-queue-drops (> 0)SELECT coalesce(max(tx_queue_drops), 0)::BIGINT AS value FROM mikroscope_api_iface WHERE time >= now() - interval '1 minute'-- mikroscope-ecc-failure (> 0)SELECT coalesce(sum(delta), 0)::BIGINT AS value FROM (SELECT max(ecc_failures) - min(ecc_failures) AS delta FROM mikroscope_mtd WHERE time >= now() - interval '1 hour' AND ecc_failures IS NOT NULL GROUP BY "partition") AS partsEvery coalesce() over an aggregate carries ::BIGINT. The InfluxDB sink writes its counters
unsigned, so coalesce(sum(count), 0) would be coalesce(UInt64, Int64), a pair Grafana’s
InfluxDB plugin cannot map: it answers HTTP 200 with no frames and no error, Grafana reads that as
NoData, and a rule declaring noDataState: OK would read OK while the thing it watches is happening
(found in Grafana). TestCoalesceIsCastInAlertSQL fails
if a rule omits the cast. The panels’ SQL pins the same fault by casting greatest() over an
aggregate to BIGINT (internal/), and
TestGreatestIsCastForTheInfluxPlugin fails if a panel omits it; a panel at least fails loudly,
with a 500. The conntrack rule reads the slab ceiling as "limit", the InfluxDB
sink’s name for it; the PostgreSQL translation turns it into limit_objs.
-- mikroscope-agent-silent (< 1)SELECT count(1) AS value FROM mikroscope_cpu WHERE time >= now() - interval '2 minutes'-- mikroscope-softnet-drops (> 0)SELECT coalesce(sum(dropped), 0)::BIGINT AS value FROM mikroscope_softnet WHERE time >= now() - interval '5 minutes'-- mikroscope-oom-kill (> 0)SELECT coalesce(sum(oom_kill), 0)::BIGINT AS value FROM mikroscope_vm WHERE time >= now() - interval '5 minutes'-- mikroscope-detections (> 0)SELECT count(1) AS value FROM mikroscope_detection WHERE time >= now() - interval '5 minutes' AND rule NOT IN ('microburst', 'ipc-collapse')-- mikroscope-thermal-near-critical (> 0)SELECT count(1) AS value FROM (SELECT zone, max(celsius) AS c, max(critical_celsius) AS crit FROM mikroscope_thermal WHERE time >= now() - interval '2 minutes' AND critical_celsius IS NOT NULL GROUP BY zone) AS zones WHERE c >= 0.85 * crit-- mikroscope-conntrack-near-limit (> 0.8)SELECT max(active_objs) * 1.0 / nullif(max("limit_objs"), 0) AS value FROM mikroscope_slab WHERE time >= now() - interval '2 minutes' AND cache = 'nf_conntrack' AND "limit_objs" IS NOT NULL-- mikroscope-agent-oom (> 0)SELECT coalesce(sum(oom_kill), 0)::BIGINT AS value FROM mikroscope_self WHERE time >= now() - interval '5 minutes' AND oom_kill IS NOT NULL-- mikroscope-wakeup-storm (> 4)SELECT CASE WHEN base_min >= 720 THEN (now_sum / nullif(now_min * 60.0, 0)) / nullif(base_sum / (base_min * 60.0), 0) ELSE 0.0 END AS value FROM (SELECT sum(CASE WHEN time >= now() - interval '10 minutes' THEN CAST(ctxt AS BIGINT) ELSE 0 END) AS now_sum, count(DISTINCT CASE WHEN time >= now() - interval '10 minutes' THEN date_bin(interval '1 minute', time, TIMESTAMPTZ 'epoch') END) AS now_min, sum(CASE WHEN time < now() - interval '10 minutes' THEN CAST(ctxt AS BIGINT) ELSE 0 END) AS base_sum, count(DISTINCT CASE WHEN time < now() - interval '10 minutes' THEN date_bin(interval '1 minute', time, TIMESTAMPTZ 'epoch') END) AS base_min FROM mikroscope_stat WHERE time >= now() - interval '1450 minutes') AS w-- mikroscope-egress-queue-drops (> 0)SELECT coalesce(max(tx_queue_drops), 0)::BIGINT AS value FROM mikroscope_api_iface WHERE time >= now() - interval '1 minute'-- mikroscope-ecc-failure (> 0)SELECT coalesce(sum(delta), 0)::BIGINT AS value FROM (SELECT max(ecc_failures) - min(ecc_failures) AS delta FROM mikroscope_mtd WHERE time >= now() - interval '1 hour' AND ecc_failures IS NOT NULL GROUP BY "partition") AS partsThresholds
Section titled “Thresholds”- Zero, for a counter that should not move. Kernel RX drops, kernel OOM kills, detections,
slipped ticks, the agent’s own OOM kills, uncorrectable ECC failures, typed MAC errors on any
port, and the kernel-log port records whose kind is
own-addressorlink-down. A healthy device reads zero on each. - Zero dark bridge ports, for ten minutes.
mikroscope-bridge-port-darkcounts bridge ports that received packets while the bridge sent them neither a unicast nor a broadcast frame; the count is 0 on a healthy bridge, and the 10-minute pending period is what stops a quiet moment from firing it. - One sample, for the silent agent. Fewer than one sample in two minutes is none at all, at any configured rate.
- A share of the board’s own thermal trip. 85 % of the lowest critical trip point each zone
declares, which the agent reads from
/sys/class/thermaland ships beside every reading. A zone that declares no critical trip is left out of the query rather than compared against a made-up ceiling. - A share of the kernel’s own connection limit. 0.8 of the
nf_conntracklimit the agent read. The occupancy comes from/proc/slabinfo, which needs a privileged container. - Zero again, for a counter that is supposed to move — with the judgement in the duration
instead. The egress queue is the one case here where the healthy reading is not exactly zero on
every device: dropping is how a full queue tells a sender to slow down, so a link that is briefly
saturated drops a few packets and is working as designed. A packets-per-second threshold would be
a number the device did not publish, so
mikroscope-egress-queue-dropskeeps the zero and asks only the last minute with a ten-minute pending period: it takes ten consecutive minutes of dropping to fire, and no single burst can do it however large (a measured burst). This is the shape to copy for any future rule whose healthy reading is not zero. - A multiple of the device’s own last day. A context-switch rate has no healthy value that holds
across boards, so
mikroscope-wakeup-stormcompares the router with itself: the mean rate over the last ten minutes against the mean over the 24 hours before, firing above 4 (healthy and storm ratios). Each side is divided by the minutes it actually covers, and the rule waits for 12 hours of baseline, so a store younger than a day does not read as a storm. The rule announces a change of regime and clears as the new rate becomes the baseline.
The rules do not alert on softnet squeezes: on the router measured, about 11.2 %
of per-CPU samples carry one squeeze, so “squeeze > 0” would page forever. What
reaches the detections alert instead is the microburst detection: three burst samples on one CPU
within 60 s. A burst sample is one in which a softnet queue dropped a packet, or ran out of budget
more often than that CPU’s trailing 90th percentile and at least three times, while the sample’s
packet count was at or below its trailing median. See
Detection rules.
Limitations
Section titled “Limitations”- Which process. The OOM rule says the kernel killed something. From inside the container there is no PID namespace to say what.
- Which port. Every port rule reduces to one number: the two kernel-log rules and the MAC-error rule sum over ports, the egress rule takes the maximum, and the dark-bridge-port rule counts ports. So a firing rule says a loop signature, a link-down or a dark port happened, not on which cable. The port is on the Kernel log section’s two port-event panels, with its comment and its interface lists, or in the Interface traffic section.
- An unprivileged agent’s blind spots. The connection-table and ECC rules read sources that need a privileged container. Without one those measurements never reach the store, and on Prometheus a query over a missing metric returns no data — which these rules read as OK.
- A collector without the API tier. The port-error, dark-bridge-port and egress-queue rules read
RouterOS counters that only the API tier polls. Under
--api-mode offnone of that reaches the store: on Prometheus the three rules read no data, and on PostgreSQL the egress rule, the only one of the three in that file, reads 0; both are OK. On an InfluxDB store that has never held those tables the queries fail andexecErrState: Errorapplies instead. - A table the store has never held. InfluxDB 3 refuses a query naming a missing table or column
when it plans it, so on a store that has never held a detection, or
mikroscope_mtdon an unprivileged agent, that rule’s query is expected to fail andexecErrState: Errorto apply instead of OK (not observed). - Anything while the silent-agent rule fires. Every rule reads the same stream — including the slipped-ticks rule, whose only query is against the collector, not the agent — so with no samples arriving they read zero or no data, and both are OK.
Which forms of these rules have been loaded into Grafana and watched evaluate, and what has not been tried, is on Tested on.