Skip to content

Resolution limits

Every CPU number mikroscope produces is counted in kernel ticks, so the smallest change it can show is set by the kernel, not by the agent. One source reads beneath the tick, the CPU’s own counters, and the agent’s ring sets how far back any reading goes.

A tick lasts 10 ms. /proc/stat does not count time, it counts USER_HZ ticks, 100 per second. A 100 ms sample can therefore hold 10 ticks per core, so the busy ratio of one core resolves to 10 % steps, and the average of four cores to 2.5 %. Over 1 s the resolution of one core is 1 %.

That is arithmetic from the tick, not a property of the agent, and no rate or setting changes it. The agent ships the raw ticks and the real interval of every sample (dt_ns), never a percentage, so the window you divide over is yours to choose.

Tick accounting is also quantised at the edges: a 100.3 ms interval can carry 11 ticks, which would make a busy ratio above 1. The derived busy ratio is capped at 1; the ticks themselves stay raw.

Sampling faster does not refine this. At 100 Hz a sample holds 0 or 1 busy tick, so the per-sample busy ratio has two possible values; above roughly 20 Hz the tick counters are an occupancy indicator rather than a percentage. What a higher rate does buy is on Rate ceiling.

One source reads below the tick: the CPU’s performance monitoring unit (PMU), through perf_event_open, which the agent collects as the perf source when the container is privileged. It is the only source in mikroscope that does not come from a file.

A sample in which /proc/stat shows zero busy ticks on every core can still hold millions of cycles and instructions retired (measured). The tick rounds that work away; the counter does not.

The number worth watching is instructions per cycle (IPC), and the agent does not compute it: it ships the raw counts and you divide. IPC separates a core doing work from a core stalled on memory, which no tick counter can express. Compare it per core: the idle IPC of two cores on one board can differ by more than a factor of two (measured).

  • The agent asks for seven counters: cycles, instructions, cache-references, cache-misses, branch-instructions, branch-misses and bus-cycles.
  • Which of them open depends on the CPU. A counter that cannot be opened is absent, never zero, so read the counter label of mikroscope_perf_events_total{counter,cpu} rather than assuming a set (CPUs not tried).
  • The generic stalled-frontend and stalled-backend events are not asked for. They return ENOENT where they were tried and would need raw, CPU-specific event codes.
  • Every count carries enabled_ns and running_ns. On a CPU that multiplexes its counters, the raw count covers only the part of the interval the counter ran, not the interval.
  • Without privileged=yes the whole family is absent: the counters are opened system-wide, which the unprivileged container cannot do. See Privileged mode.

The collector uses the same two counters for one of its detections, ipc-collapse: per core, once at least 20 s of history exists, the one-second instructions-per-cycle falling below half its trailing 60 s median while the cycle rate is above its own median. That is described with the other rules on Detection rules.

Do not expect a finer clock from PSI, whose stall totals are in microseconds, or from schedstat, whose run and wait times are in nanoseconds. The RouterOS kernel measured so far has neither /proc/pressure nor /proc/schedstat (measured). The agent reads both where a kernel has them (not yet on a router).

On a kernel whose /proc/stat irq column stays at 0, as on the measured one, hard-IRQ time is counted inside system.

There is no tracing path either. The measured kernel has no eBPF, kprobes or ftrace, and the router has no kernel modules directory, headers or compiler to add them (probed). A kernel feature that is missing cannot be added from a container.

The agent detects what a kernel has at start and reports it on /capabilities: the sources map says which sources this deployment actually reads. Fields for an absent source are absent from every sample and every sink, never zero.

The other hard bound is depth, not resolution. The agent keeps its samples in a ring of --buffer seconds, 60 s by default and 10–3600 s allowed, which holds rate × buffer samples. Nothing older exists anywhere on the router.

That depth is also the whole window of doctor’s health section, which reads the ring once (the whole ring when it holds no more than 10 000 samples, otherwise the newest 10 000, so the shorter of --buffer and 10 000 / rate seconds) and prints how many seconds it covered. A clean doctor says there was no loop, STP churn, link flap or softnet drop in the last minute at the defaults, not today; a fault that happens once a day belongs to the dashboards and the alert rules.

A collector or recorder outage shorter than the ring is backfilled on reconnect: it asks for since=<seq> and receives every sample it missed. An outage longer than the ring is reported as a gap of known length, never papered over. The agent answers with a {"gap":{"from":…,"to":…}} line naming the sequence numbers that are gone, before the samples it still holds; record writes it as a marker reading samples N..M lost, and forward counts it in mikroscope_collector_gaps_total and hands it to every sink.

A longer ring costs memory in the agent, and the agent refuses one that cannot fit. At start it charges each line 3 456 B: the allocator size class that a line with every source on (3 230 B) is served from, because that is what the heap pays. A board with more cores or interrupt lines costs more per line, and one with no PMU less. The agent adds the triggered-capture budget and exits with an error if the total exceeds the container’s memory.max. If the total is more than half the Go soft memory limit it starts but logs a warning, because a heap that tight keeps the garbage collector running. How to size both limits is on Agent cost.