Skip to content

CPU-bound core

In short: one saturated core on the four-core RB5009 reads as about 29 % of the device, so it is the per-core row, not the total, that shows a single-threaded bottleneck: here one core at 99.8 % while three idled. At 10 Hz the scheduler moving the load between cores is visible for about 20 s before it settles; sampled once a second, that is a vague plateau. The load was provoked on purpose.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · agent at 10 Hz in an ephemeral privileged container

A console loop that burns one core and ends on its own, with no configuration change and no external dependency:

:local i 0; :while ($i < 4000000) do={ :set i ($i + 1) }

It ran 58.8 s.

In 5-second buckets. The agent ships raw tick deltas; the percentages below were computed from them afterwards. There is no chart: the run predates the history of the reference InfluxDB store, which begins on 2026-09-19.

t busy% per-core busy% ctxt/s temp MHz
0 22.3 31.8 7.4 40.4 9.6 9192 36.5 1400
1 30.9 11.7 11.4 89.7 10.6 11325 37.0 1400
2 29.9 14.3 9.4 19.6 76.0 9476 37.0 1400
3 29.1 7.4 9.4 17.6 81.6 10805 37.0 1400
4 29.3 3.1 9.4 4.9 99.8 9257 37.2 1400
8 28.4 2.2 7.0 4.6 99.8 11947 37.2 1400
11 33.0 11.8 11.2 9.0 99.7 12222 37.3 1400
  • Total ~29 % is one core of four. A device total is a trap; always look at the per-core row. “29 % CPU” here means “one core is saturated and three are idle”, which for a single-threaded bottleneck is the whole story.
  • The scheduler took ~20 s to settle. Buckets 0–3 show the work moving between cores (40 % → 90 % → 76 % → 82 %) before pinning on core 3 at 99.8 %. If you sample at 1 s or slower you see a vague plateau; at 10 Hz you see the migration.
  • Temperature followed: 36.5 → 37.3 °C. One saturated core is worth ~0.8 °C on this passively cooled board. Small, but it tracks, and it is read from /sys/class/thermal with no API call.
  • Frequency stayed flat at 1400 MHz. On this device that is a configuration, not an observation: the owner has pinned the clock to maximum, which /system/routerboard/settings reports as Warning: cpu not running at default frequency. On a device that scales, the freq_khz field is where you would see a busy tick that was earned slowly. The agent reads the frequency every tick and stores freq_khz only when it changes, plus a heartbeat once every 60 s, so on a pinned clock you get one row a minute, and on a scaling clock you get every step that lasts at least one tick.

Before you conclude anything from a CPU number, confirm mikroscope_slipped_total is 0. A slipped tick is one whose read finished after the next tick was due, and the sampler’s own accounting is then the first thing you should distrust.

The opposite shape: the cores are not busy, but the kernel is. It reached the reference RB5009 on its own, on RouterOS 7.24.4, at 2026-09-23 12:08:50 UTC, when a Home Assistant MikroTik integration started a wake-up storm:

  • Timer interrupts went from about 2 500 to about 35 000 a second, in bursts of 20–90 s on one core at a time, while user time and traffic stayed flat.
  • It barely showed in RouterOS’s own profile, and the busy columns above would not have caught it. The context-switch rate did: on that router it moved with the timer interrupts (r = 1.0).
  • Disabling the integration on 2026-09-24 brought the timer back to about 2 500 a second within 30 s.

It is one event on one router, read from the InfluxDB reference store, and it has no case study of its own yet. The rule that watches for it is mikroscope-wakeup-storm in Alert rules, which compares the last ten minutes’ context-switch rate with the router’s own previous 24 hours.

Provoked on purpose ·

  • A device total near 100 % divided by the core count — here ~29 % on four cores — with one per-core column at 99.7–99.8 %.
  • Before it settles, the load visibly hopping between cores for tens of seconds.
  • A temperature rise of under a degree that tracks the load.