Skip to content

Agent cost

At the install default — 10 Hz, the default per-source floors, a 60 s ring — the agent costs a MikroTik RB5009 2.69 % of one core and 13.2 MiB RSS, read from its own cgroup:

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · 300 s windows at steady state (ring full), full source set, the shipped configuration — a 60 s ring, the memory limit derived from it and the default 64M container cap — with the collector forwarding to InfluxDB 3

The agent reports its own cost on every sample, and CI asserts the image-size budget.

The budget is ≤ 2 % of one core, ≤ 16 MiB RSS, ≤ 8 MiB image. The published 1.2.1 arm64 image is 6.38 MiB (6 690 304 B, measured on the release asset on 2026-09-24); the 1.0.0 image was 6.1 MiB, the figure that release’s notes give. The other two depend on the rate and on how much you ask the agent to read, so the answer is a table rather than a number: Rate ceiling is that table.

The memory is inside the budget and the CPU is not. 13.2 MiB against the 16 MiB the budget asks for, and 2.69 % against the 2 %, with every source read: the perf timings, buddyinfo, the MTD ECC counters, slabinfo, the kernel log and the cgroup events among them. Per-interface traffic is not in that set: the container cannot see it, and it comes from the API tier.

The 2026-09-15 campaign was over on both. Its ring held 300 s rather than 60 and its memory limit was a flat 40 MiB that never bound, which put the same agent at 31.3 MiB; neither was a statement about the rate. Nothing about the sampling changed between the two campaigns: the CPU moved from 2.85 % to 2.69 %, which is the noise between two windows.

The budget says what the agent is allowed to cost. The comparison a reader usually wants is against the obvious way: a busybox shell loop. On 2026-09-11, on the same router, a loop reading the agent’s file set of that date at 10 Hz cost 2.40–2.48 % of one core, while the reads themselves took about 0.77 ms per sample.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · a busybox shell loop reading the agent's file set of that date, seven files, at 10 Hz, one fork per iteration, in a container on the router; two 60 s runs

So most of the shell loop’s cost was not the reading. It pays a fork per iteration and the agent pays none: one process starts, opens its files once and keeps them open, which is why the agent is a binary rather than a script. The two figures cannot be set side by side, though. The agent reads far more today — perf, buddyinfo, the MTD ECC counters, slabinfo, the kernel log — and the loop has not been re-measured against that set, so it cannot be set against 2.69 %.

Pullers add to it. Besides the collector and record, a standalone doctor reads the ring once, in one request: the whole ring when it holds no more than 10 000 samples, otherwise the newest 10 000. At the defaults that is the whole 60 s ring, about 1.9 MB at 10 Hz, and that request is served on the sampler’s core; the cost of that read on the RB5009 has not been measured.

The single most expensive mistake available here is giving the process less memory than its ring needs. The Go garbage collector responds to a tight soft limit by running more often, and the CPU cost jumps for a reason that has nothing to do with the sampling rate. The ring holds pre-encoded lines rather than structs for the same reason. If you raise --buffer or the rate, raise --mem-limit-mb and the container’s --memory-max with it; the flags each run used are listed beside it on Rate ceiling.

install derives MEM_LIMIT_MB from the ring: rate × buffer × the line size, times 2.5, floored at 16 MiB and capped at three quarters of memory-max, so the default install writes 16 for a 60 s ring at 10 Hz. The factor is measured. On the RB5009 (RouterOS 7.24.2, 10 Hz, a 300 s ring, every source on) on 2026-09-17, four limits over four windows of about 12 000 samples each:

limit ring multiple RSS CPU per sample What it cost
40 MiB 4.0× 32.9 MiB 2 657 µs the limit never binds
24 MiB 2.4× 26.3 MiB 2 780 µs no measurable cost
21 MiB 2.1× 23.5 MiB 3 250 µs +22 %, and climbing
18 MiB 1.8× 20.4 MiB 14 800 µs +457 %, worst tick 52 ms

The rows are the ones the comment on memLimitRingFactor in internal/router/options.go records. The 2.5× factor would be 25 MiB at 300 s; the window measured nearest it was 24 MiB, and none of the four windows was written down with its spread. The same cliff was measured from the other side on 2026-09-12: 9.38 % of one core at 14 MiB against 1.39 % with room, both with the ring full.

The parse is not where the time goes. Parsing the seven global /proc files plus one delta took 27 µs and 239 allocations per sample on the amd64 development host (Go 1.27.1, three runs, 26.7–27.5 µs, 2026-09-11). On the RB5009’s Cortex-A72 a whole tick — timers, JSON and the garbage collector included — costs 2 685 µs at 10 Hz with the default floors, and the six runs give it at every rate. The parse was not measured on the A72 on its own.

Each SSH connect costs the RB5009 20–27 % CPU for its duration.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.1 ·  · one ssh connect, for its duration

That reading predates the agent and was taken on RouterOS 7.24.1. On 7.24.2, on 2026-09-11, the same cost showed in /tool profile as 17–33 % in one or two snapshots, not a measured window. So the CLI batches every read into one connect — doctor is one, status is one, the state questions of install are one — and each write takes one more, plus the scp upload. SSH is never a data path: record and forward reach the agent over HTTP or the RouterOS API.

Measured with RouterOS’s own profiler, once with the collector stopped and once with it running a round every second:

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · /tool profile duration=60s cpu=total, once with the collector stopped and once with it running --interfaces bridge,ether1,PPPoE_DIGI --counters-every 10s --api-every 1s; the first five one-second snapshots of each profile dropped, because they hold the SSH connect that asked for it

Profile row Collector stopped Collector running
total 5.93 % 6.04 %
interface-mgmt 0.40 % 0.87 %
config-db about 0 % 0.15 %

The total moves by 0.11 points, inside the traffic noise of a minute; the two rows that answer the tier’s questions move by about half a point between them. No api process row appears in either profile: the API process relays, and the work lands on the subsystem that answers. So a tier running a round every second costs the router about 0.5 % of its total CPU. That table is one 60 s profile per condition, with no spread.

A larger configuration was measured as an A/B on the same router on 2026-09-19, the day of the RouterOS upgrade that left it on 7.24.4; which version the A/B itself ran on was not recorded. Every interface but lo in --interfaces (16) and --conntrack-every 10s, 8 min without against 14 min with: mean cpu-load went from 7.12 % to 7.91 % and mean kernel busy from 7.70 % to 8.46 %, about +0.8 points. That is a lower bound, since bridge traffic fell from 48.0 to 38.9 Mbit/s between the windows. The p50 of cpu-load stayed at 6 and the p95 went from 16 to 18.

--conntrack-every asks /ip/firewall/connection/print count-only: a call took 1.3 ms at 6 212 entries, the median of ten calls, the fastest 1.1 ms and the first 71 ms, cold.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · the connection table counted over the RouterOS binary API, /ip/firewall/connection/print count-only, ten calls

Nothing here transfers to a board that is not an RB5009: a different core count, clock, kernel and flash all move it. The procedure is three commands and takes a minute:

Terminal window
curl -s http://<collector host>:9124/metrics | grep -E 'self_cpu_usec_total|self_rss_bytes'
sleep 60
curl -s http://<collector host>:9124/metrics | grep -E 'self_cpu_usec_total|self_rss_bytes'

The difference in self_cpu_usec_total divided by 60 000 000 is the share of one core. Or let the shell do the division, over three minutes:

Terminal window
U=http://<collector host>:9124/metrics
get() { curl -s "$U" | awk -v k="$1" '$1==k{print $2}'; }
c0=$(get mikroscope_self_cpu_usec_total); t0=$(date +%s)
sleep 180
c1=$(get mikroscope_self_cpu_usec_total); t1=$(date +%s)
echo "$c0 $c1 $t0 $t1" | awk '{printf "%.2f %% of one core\n", 100*($2-$1)/1e6/($4-$3)}'

Three things to expect:

  • Cost and memory rise until the ring fills. With the default 60 s ring at 10 Hz the agent holds 600 pre-encoded samples; a figure taken in the first minute after install is measured on a nearly empty heap and reads low. Wait out BUFFER_S before quoting a steady-state figure.
  • A snapshot overstates it. A 60 s /snapshot hands over about 600 lines, about 1.9 MB, and self.cpu_us inside those samples includes the cost of serving them; reading the cost out of repeated snapshots overstates it more.
  • mikroscope_slipped_total is the number that matters. A sampler that costs a little more but never slips is telling you the truth; one that slips is not.

The address is the collector’s, not the agent’s: the agent serves no /metrics, and these two counters reach the collector in every sample. With no Prometheus sink configured, the same two numbers are cpu_us and rss in mikroscope_self on InfluxDB, Telegraf, stdout and the SQL stores, self.cpu_us and self.rss in the sample documents on Elasticsearch, self.cpu_us and self.rss_bytes on Graphite, and mikroscope.self.cpu.time / mikroscope.self.memory{kind="rss"} on OTLP. Measured on a board other than the RB5009, with the rate, the date and what the router was doing, the figure belongs in a board report.