Diagnose faults
Match what the data shows against the signatures below, then open the case study for the full reading and the commands that reproduce it on your own device. Each signature is a change against the idle shape: a kernel log that is normally silent carrying a steady stream, one core pinned while the device total stays modest, one interrupt line rising several-fold.
Before you test
Section titled “Before you test”Nothing the provoked case studies do can break the router’s uplink or cut an administrator’s path to it. Keep that property on your own device:
-
Find the port you are connected through:
/interface/bridge/host/print where mac-address="<your machine's MAC>" -
Leave that port, the WAN port and any port carrying a service alone.
Fault signatures
Section titled “Fault signatures”Read the idle shape first. Without it, every other signature looks like an anomaly.
| Fault | Where it shows | Signature |
|---|---|---|
| Idle baseline | per-core busy, time_squeeze, events |
a low, steady busy floor on every core, a squeeze that is never zero, no kernel events |
| Layer-2 loop | events (the kmsg source) |
a steady stream of events at idle, spaced at the STP hello interval, while RouterOS’s log and port monitor stay clean |
| Port names | events |
the kernel names a port by its own netdev name, which can differ from the RouterOS name |
| CPU-bound core | per-core busy, temperature, frequency | one core near 100 % while the device total sits near 100 % divided by the core count; the load hops between cores before it settles |
| Wake-up storm | context switches, timer interrupts | a context-switch rate several times its mean over the previous day, timer interrupts in bursts on one core at a time, busy columns flat |
| Packet flood | interrupts, softirqs, time_squeeze |
one interrupt line rising several-fold with a sharp onset and offset, its cost on the one core it is pinned to |
| Flash wear | the yaffs source, MTD ECC counters |
pw and er deltas at idle that you did not cause (look first at logging actions set to disk); corrected_bits rising or any ecc_failures |
| Conntrack without the API | the nf_conntrack slab cache |
the router’s real connection count, where the container’s own namespace reports 0 |
| Port losing frames | the API tier’s per-port MAC counters | rx overflow on one port, in intervals that carry a small fraction of the link’s capacity: a burst, not a load |
Scroll sideways to see every column
The wake-up storm has no case study of its own: its reading is part of the CPU-bound core study,
and the rule that watches for it is mikroscope-wakeup-storm in
Alert rules.
Check the data first
Section titled “Check the data first”- The sampler was not starved.
mikroscope_slipped_totalshould be 0 over the window you are reading. A slipped tick is one whose read finished after the next tick was due, and the sampler’s own accounting is then the first thing to distrust. - The kernel log was kept whole.
mikroscope_counts loss events, not records: one per tick that hit the agent’s cap of 64 records, and one per kernel ring overrun, which can stand for many records. While it is non-zero, the per-level counts inkmsg_ dropped_ total mikroscope_are a lower bound, and so is an event rate taken from them.kmsg_ records_ total
To measure what the agent costs while you run a scenario, read the collector’s /metrics, not a
snapshot: Agent cost has the procedure.