Skip to content

Source cadences

Not every source changes as fast as the sampler ticks, and storing a value more often than the hardware refreshes it is storage for no information. The agent reads every counter on every tick except the MTD ECC counters (every 10 s), slows a level source only for a reason it names, and publishes each cadence so a consumer never has to infer it. FLOOR_HZ turns every floor off at once.

CPU ticks, interrupts, softirqs, the vmstat counters, the PMU and the memory levels (which change many times a second, measured) are read and shipped on every tick. For a counter a delta of zero is real information, the core was idle, so there is nothing to skip.

Three more are read every tick with their own rule:

  • NAND wear (/proc/yaffs) and block I/O (/proc/diskstats) are counters and cheap to read. A device’s row is stored only when its delta is non-zero: a router can list many idle devices, nbd among them, and a row per tick for each would be payload and nothing else.
  • The kernel log (/dev/kmsg) is drained every tick and never slowed. An event’s value is its timestamp, and a marker delayed by a second no longer lines up with the spike it explains.

A level source (a temperature, a clock frequency, a cache population) may be read or stored less often than every tick, but only for a reason the agent names. It is never slowed because it was seen changing slowly, since that would measure one night on one device. Each level source reports its cadence and reason on /capabilities under cadences, on /metrics as mikroscope_source_cadence_hz{source,reason}, and to every sink in the device-info stream, so a consumer reads the true cadence of a field instead of inferring it from the data.

The cadence reasons a source can report
reasonWhat it means
rateread at the full sampler rate; nothing the device declares justifies less
declaredthe device publishes its own refresh cadence, and reading faster returns the same value with new dither
policya setting says the value cannot move on its own: a userspace cpufreq governor
budgeta measured parse cost
changeread every tick, stored only when it moves
overrideFLOOR_HZ is set, and every level source is on its one cadence

rate describes the read cadence only, so /proc/buddyinfo reports rate although it is stored on change. The MTD counters report budget although no parse cost was measured for them. Both are how the code labels them, and neither matches the table above exactly.

Source Read Stored reason
thermal zones at the zone’s declared polling_delay, else every tick every read declared, or rate where the board declares none or the sampler rate is no faster than the declared cadence
scaling_cur_freq every tick on change, or heartbeat change, or policy under a userspace governor
/proc/slabinfo at about 6 Hz on change, or heartbeat budget
/proc/buddyinfo every tick on change, or heartbeat rate
MTD ECC counters every 10 s on change, or heartbeat budget

A slower cadence is a whole number of ticks: the sampler rate divided by the floor, rounded to the nearest integer and never less than one. So the reported rate is what actually happens, not the nominal floor.

Temperature. A thermal zone declares its own polling_delay. At 1 000 ms the kernel itself re-reads the sensor at 1 Hz, and at 10 Hz the agent reads every 10th tick. The sensor quantises to steps of about 0.42 °C, and the raw reading dithers across a step boundary tens of times a second. Storing on change would store that dither as signal; sampling and holding at the declared cadence captures the real curve and drops the dither. On a board that declares no cadence the zones are read every tick.

CPU frequency. A frequency step is a real DVFS transition, not noise, so the frequency is read every tick and stored when any core moves. On a device with a pinned clock that stores nothing after the first reading but the heartbeat; on a scaling device it catches every move. Under a userspace cpufreq governor the clock cannot move without a write, so the reason reads policy.

/proc/slabinfo is the expensive one, the largest per-tick parse by an order of magnitude, and that cost is the reason it is slowed at all. The 6 Hz it is slowed to is the rate its fastest cache, nf_conntrack, was measured changing. It is stored on change. At 10 Hz that is every 2nd tick (5 Hz); at 100 Hz, every 17th (about 5.9 Hz). How that amortisation shows in the per-sample cost is on Rate ceiling.

/proc/buddyinfo is about 100 bytes, among the cheapest files the agent reads, and has no floor because none has been measured. The free lists churn with every allocation, so on a busy router it will be stored on most ticks, and that is the measurement rather than noise.

MTD ECC counters are read every 10 seconds. That is not a measured floor: the counters move on the scale of a device’s life, and each read is six small sysfs files per partition, so ten seconds is an arbitrary but generous choice. The code still labels that cadence budget, although no parse cost was measured for it. They need privileged=yes.

The 6 Hz of /proc/slabinfo is the one measured floor, and it comes from one device over one idle night (Tested on). A router with a scaling clock, a busy dirty-page workload or a different board has different floors, so re-measure with FLOOR_HZ equal to the sampler rate before trusting them on yours. The floors are named constants in internal/agent/source.go.

Every source stored on change is also re-emitted about once every 60 seconds (on the first read due after 60 s), so a value that sits still for an hour still has a recent row in any store. The container’s own memory.max, a constant, is re-emitted in the samples once per heartbeat (every tick under FLOOR_HZ); the limits on /capabilities and the device-info stream carry it from start.

On /metrics, a floored gauge between emissions holds the last reading, because a level’s value between readings is the last one read, not nothing. mikroscope_source_age_seconds{source} says how old that held reading is.

The change filter is re-armed after the sampler’s baseline read, so the first sample carries every source stored on change, and the floored gauge families reach /metrics with it rather than a heartbeat later.

Every floor is overridable at once, without a rebuild:

Terminal window
mikroscope install --rate 10 --floor-hz 10 # writes FLOOR_HZ=10 into the agent's envlist

FLOOR_HZ is the agent’s environment variable; --floor-hz is the deployment flag (install, upgrade, and plan for the listing) that writes it into the container’s envlist, only when it is above zero. It is one global setting in hertz, from 0 to 1000.

  • 0, the default, keeps the per-source floors above.
  • Any N > 0 puts the thermal zones, /proc/slabinfo and the MTD counters on one cadence of N Hz, and turns off the store-on-change filter for every level source, so nothing is held back and a capture sees every read. Every level source then reports the reason override.
  • N equal to or above --rate reads and emits everything every tick. That is the configuration the floors were measured from, and the one to re-run before trusting any number on this page on your device.

Two things FLOOR_HZ does not change. The per-device row filter on /proc/yaffs and /proc/diskstats stays: a device whose counters did not move still has no row, which loses nothing because the delta was zero. And scaling_cur_freq and /proc/buddyinfo are read every tick with or without it; with N below --rate they are emitted every tick, but their reported cadence is the override’s, not the sampler rate.

What it costs to read everything every tick, at 50 and 100 Hz, is on Rate ceiling.