Resolution limits
Every CPU number mikroscope produces is counted in kernel ticks, so the smallest change it can show is set by the kernel, not by the agent. One source reads beneath the tick, the CPU’s own counters, and the agent’s ring sets how far back any reading goes.
Tick resolution
Section titled “Tick resolution”A tick lasts 10 ms. /proc/stat does not count time, it counts
USER_HZ ticks, 100 per second. A 100 ms sample can therefore hold 10 ticks per core, so the busy ratio of
one core resolves to 10 % steps, and the average of four cores to 2.5 %. Over 1 s the resolution of one core is 1 %.
That is arithmetic from the tick, not a property of the agent, and no rate or setting
changes it. The agent ships the raw ticks and the real interval of every sample (dt_ns),
never a percentage, so the window you divide over is yours to choose.
Tick accounting is also quantised at the edges: a 100.3 ms interval can carry 11 ticks, which would make a busy ratio above 1. The derived busy ratio is capped at 1; the ticks themselves stay raw.
Sampling faster does not refine this. At 100 Hz a sample holds 0 or 1 busy tick, so the per-sample busy ratio has two possible values; above roughly 20 Hz the tick counters are an occupancy indicator rather than a percentage. What a higher rate does buy is on Rate ceiling.
PMU counters
Section titled “PMU counters”One source reads below the tick: the CPU’s performance monitoring unit (PMU), through
perf_event_open, which the agent collects as the perf source when the container is
privileged. It is the only source in mikroscope that does not come from a file.
A sample in which /proc/stat shows zero busy ticks on every core can still hold millions of
cycles and instructions retired (measured). The tick
rounds that work away; the counter does not.
The number worth watching is instructions per cycle (IPC), and the agent does not compute it: it ships the raw counts and you divide. IPC separates a core doing work from a core stalled on memory, which no tick counter can express. Compare it per core: the idle IPC of two cores on one board can differ by more than a factor of two (measured).
- The agent asks for seven counters:
cycles,instructions,cache-references,cache-misses,branch-instructions,branch-missesandbus-cycles. - Which of them open depends on the CPU. A counter that cannot be opened is absent, never zero,
so read the
counterlabel ofmikroscope_rather than assuming a set (CPUs not tried).perf_ events_ total{counter, cpu} - The generic
stalled-frontendandstalled-backendevents are not asked for. They returnENOENTwhere they were tried and would need raw, CPU-specific event codes. - Every count carries
enabled_nsandrunning_ns. On a CPU that multiplexes its counters, the raw count covers only the part of the interval the counter ran, not the interval. - Without
privileged=yesthe whole family is absent: the counters are opened system-wide, which the unprivileged container cannot do. See Privileged mode.
The collector uses the same two counters for one of its detections, ipc-collapse: per core,
once at least 20 s of history exists, the one-second instructions-per-cycle falling below
half its trailing 60 s median while the cycle rate is above its own median. That is described with the other rules on
Detection rules.
PSI, schedstat and tracing
Section titled “PSI, schedstat and tracing”Do not expect a finer clock from PSI, whose stall totals are in microseconds, or from
schedstat, whose run and wait times are in nanoseconds. The RouterOS kernel
measured so far has neither /proc/pressure nor /proc/schedstat (measured).
The agent reads both where a kernel has them (not yet on a router).
On a kernel whose /proc/stat irq column stays at 0, as on the measured one, hard-IRQ time is
counted inside system.
There is no tracing path either. The measured kernel has no eBPF, kprobes or ftrace, and the router has no kernel modules directory, headers or compiler to add them (probed). A kernel feature that is missing cannot be added from a container.
The agent detects what a kernel has at start and reports it on /capabilities: the
sources map says which sources this deployment actually reads. Fields for an absent
source are absent from every sample and every sink, never zero.
Ring depth
Section titled “Ring depth”The other hard bound is depth, not resolution. The agent keeps its samples in a ring of
--buffer seconds, 60 s by default and 10–3600 s allowed, which holds rate × buffer
samples. Nothing older exists anywhere on the router.
That depth is also the whole window of doctor’s health section, which reads the ring once
(the whole ring when it holds no more than 10 000 samples, otherwise the newest 10 000, so the
shorter of --buffer and 10 000 / rate seconds) and prints how
many seconds it covered. A clean doctor says there was no loop, STP churn, link flap or softnet
drop in the last minute at the defaults, not today; a fault that happens once a day belongs to the
dashboards and the alert rules.
A collector or recorder outage shorter than the ring is backfilled on reconnect: it asks
for since=<seq> and receives every sample it missed. An outage longer than the ring is
reported as a gap of known length, never papered over. The agent answers with a
{"gap":{"from":…, line naming the sequence numbers that are gone, before the
samples it still holds; record writes it as a marker reading samples N..M lost, and
forward counts it in mikroscope_ and hands it to every sink.
A longer ring costs memory in the agent, and the agent refuses one that cannot fit. At start it
charges each line 3 456 B: the allocator size class that a line
with every source on (3 230 B) is served from, because that is what the
heap pays. A board with more cores or interrupt lines costs more per line, and one with no PMU
less. The agent adds the triggered-capture budget and exits with an error if the total exceeds
the container’s memory.max. If the total is more than half the Go soft memory limit it starts
but logs a warning, because a heap that tight keeps the garbage collector running. How to size
both limits is on Agent cost.