Rate ceiling
The agent lost no sample on the RB5009 at any rate up to the CLI’s own 100 Hz cap, with every source read on every tick in the worst case. That is not an extrapolation from the 10 Hz figure: it is six runs.
Every figure is from the agent’s own cgroup, carried in every sample and read back from the store, with the full source set. Each row is one window with the ring already full:
Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · 300 s windows at steady state (ring full), full source set, the shipped configuration — a 60 s ring, the memory limit derived from it and the default 64M container cap — with the collector forwarding to InfluxDB 3
| rate | floors | CPU of one core | µs/ | RSS | slipped ticks | gaps / |
|---|---|---|---|---|---|---|
| 10 Hz (default) | default | 2.69 % | 2 685 | 13.2 MiB | 0 | 0 / 0 |
| 20 Hz | default | 4.61 % | 2 303 | 15.4 MiB | 0 | 0 / 0 |
| 50 Hz | default | 9.63 % | 1 926 | 23.3 MiB | 0 | 0 / 0 |
| 100 Hz | default | 16.83 % | 1 684 | 45.7 MiB | 5 (0.02 %) | 0 / 0 |
| 50 Hz | FLOOR_HZ=50 | 22.56 % | 4 511 | 25.1 MiB | 4 (0.03 %) | 0 / 0 |
| 100 Hz | FLOOR_HZ=100 | 42.70 % | 4 270 | 49.5 MiB | 178 (0.59 %) | 0 / 0 |
Scroll sideways to see every column
FLOOR_HZ means every source is read on every tick, with no per-source floor at
all — the worst case the agent can be asked for.
Every row ran the shipped configuration and passed no memory flag at all:
the 60 s ring is the default and the limit is derived from it — 16 MiB at 10 and
20 Hz, 25 at 50, 48 at 100 — under the default 64M container cap. The campaign
before this one could not do that. It needed --memory-max 96M at 50 Hz and
128M at 100, and its 10 Hz row held a 300 s ring. Same device, same sources:
38 to 59 % less memory per row, and the CPU unmoved: 2.85 % of one core at
10 Hz on 2026-09-15 against 2.69 % on 2026-09-18.
Memory still differs by row because the ring does: it holds 60 s of samples whatever the rate, so ten times the rate is ten times the ring.
Reading the table
Section titled “Reading the table”Per-sample cost
Section titled “Per-sample cost”A sample costs 2 685 µs at 10 Hz against 1 684 µs at 100 Hz. That is not a
paradox, it is the floors working: the expensive sources are amortised over more samples.
/proc/slabinfo, one of the expensive sources (13.8 kB), is read every 2nd tick at 10 Hz and every
17th at 100 Hz, so an average sample costs less while the rate of slabinfo reads
stays near its 6 Hz floor either way (5 Hz at 10 Hz, about 5.9 Hz at 100 Hz). So ten times the
data costs about 6.3 times the CPU, not ten: 16.83 % at 100 Hz
against 2.69 % at 10.
With FLOOR_HZ there is no amortisation to be had and the per-sample cost is
flat — 4 511 µs at 50 Hz and 4 270 µs at
100 Hz — so CPU scales linearly with the rate: 22.56 %, then 42.70 %.
Slipped ticks
Section titled “Slipped ticks”The sample is still produced and still delivered, carrying its real dt_ns, so
any rate computed from it stays correct. It is a smear, not a hole, and
mikroscope_ is where you see it. At 100 Hz with
everything on every tick, 29 777 of 29 994 intervals — 99.3 % — still landed
within 11 ms of a 10 ms period; the 99.5th percentile was 11.3 ms and the worst
single interval 21.7 ms.
Read time vs CPU
Section titled “Read time vs CPU”At the default floors a whole tick’s sources are read in under 2 ms for 97.7 % of
samples at 100 Hz, comfortably inside a 10 ms period. FLOOR_HZ pushes 1.45 %
of reads past 5 ms — and not one read of 29 996 came in under 2 ms — and those
are the ticks that slip: 0.593 % of them at 100 Hz with every
source on every tick, against 0.017 % at the default floors. The CPU headroom is
larger than the timing headroom, which is why the ceiling is a statement about
I/O rather than about the A72. Even at the default floors at 100 Hz the mean read was 1.2 ms but the
worst 17.9 ms, and a tick whose read outlasts its period is late by definition. That, and not a
CPU wall at 16.83 % of one core, is where the few slips there come from.
Memory per rate
Section titled “Memory per rate”The ring is the memory. Every other figure — the Go runtime, the collector’s headroom, the allocator’s fragmentation — follows from how many entries it holds, because the soft memory limit is derived from exactly that. So the question “can this board sample at 100 Hz” is mostly “can it hold 100 × your buffer seconds of samples”, and the answer for the whole table above is yes inside the default 64M cap.
Buy the room in buffer seconds, not in cleverness. At 100 Hz a 20 s buffer holds 6.59 MiB and derives a 17 MiB limit — the same relief that compressing the ring 4.7× would give, and compression was measured on this device at +1 331 µs a sample, which at 100 Hz is +74 % CPU and six times the slipped ticks. The seconds are free; the compression is not.
Higher rates
Section titled “Higher rates”Not CPU-percent resolution. The jiffie is 10 ms, so at 100 Hz a sample holds either 0 or 1 busy tick and the per-sample busy ratio has two possible values. Above roughly 20 Hz the tick counters stop being a percentage and become an occupancy indicator; the PMU is the resolution from there.
What a higher rate does buy is everything that is not jiffie-quantised — softnet packet counts, interrupt deltas, PMU counters, the kernel log’s own timestamps — and a tighter bound on how long a burst can hide between two samples.