Skip to content

Diagnose faults

Match what the data shows against the signatures below, then open the case study for the full reading and the commands that reproduce it on your own device. Each signature is a change against the idle shape: a kernel log that is normally silent carrying a steady stream, one core pinned while the device total stays modest, one interrupt line rising several-fold.

Nothing the provoked case studies do can break the router’s uplink or cut an administrator’s path to it. Keep that property on your own device:

  1. Find the port you are connected through:

    /interface/bridge/host/print where mac-address="<your machine's MAC>"
  2. Leave that port, the WAN port and any port carrying a service alone.

Read the idle shape first. Without it, every other signature looks like an anomaly.

Fault Where it shows Signature
Idle baseline per-core busy, time_squeeze, events a low, steady busy floor on every core, a squeeze that is never zero, no kernel events
Layer-2 loop events (the kmsg source) a steady stream of events at idle, spaced at the STP hello interval, while RouterOS’s log and port monitor stay clean
Port names events the kernel names a port by its own netdev name, which can differ from the RouterOS name
CPU-bound core per-core busy, temperature, frequency one core near 100 % while the device total sits near 100 % divided by the core count; the load hops between cores before it settles
Wake-up storm context switches, timer interrupts a context-switch rate several times its mean over the previous day, timer interrupts in bursts on one core at a time, busy columns flat
Packet flood interrupts, softirqs, time_squeeze one interrupt line rising several-fold with a sharp onset and offset, its cost on the one core it is pinned to
Flash wear the yaffs source, MTD ECC counters pw and er deltas at idle that you did not cause (look first at logging actions set to disk); corrected_bits rising or any ecc_failures
Conntrack without the API the nf_conntrack slab cache the router’s real connection count, where the container’s own namespace reports 0
Port losing frames the API tier’s per-port MAC counters rx overflow on one port, in intervals that carry a small fraction of the link’s capacity: a burst, not a load

The wake-up storm has no case study of its own: its reading is part of the CPU-bound core study, and the rule that watches for it is mikroscope-wakeup-storm in Alert rules.

  • The sampler was not starved. mikroscope_slipped_total should be 0 over the window you are reading. A slipped tick is one whose read finished after the next tick was due, and the sampler’s own accounting is then the first thing to distrust.
  • The kernel log was kept whole. mikroscope_kmsg_dropped_total counts loss events, not records: one per tick that hit the agent’s cap of 64 records, and one per kernel ring overrun, which can stand for many records. While it is non-zero, the per-level counts in mikroscope_kmsg_records_total are a lower bound, and so is an event rate taken from them.

To measure what the agent costs while you run a scenario, read the collector’s /metrics, not a snapshot: Agent cost has the procedure.