Benchmarking
The benchmark command measures RouterOS API operations against a real router. It is intended for development and capacity planning, not for normal deployments.
Command
Section titled “Command”Build and run the benchmark:
go run ./cmd/benchmarkThe config path is optional and defaults to config/test.yaml:
go run ./cmd/benchmark config/test.yamlThe benchmark reads the MikroTik section from the selected YAML file, connects to RouterOS, then runs API-level operations.
What it measures
Section titled “What it measures”- Single IPv4 and IPv6 address-list adds/removes.
- Address-list find and list calls.
- Firewall rule create, find, and remove calls.
- Sequential IPv4 batch add/list/find/remove timings.
The benchmark uses documentation-prefix test addresses such as 198.51.100.0/24 and 2001:db8::/32, but it still writes to the configured RouterOS address lists and firewall menus.
Example output
Section titled “Example output”Connected to: lab-router
=== SINGLE OPERATION BENCHMARKS (RouterOS API) === Add single IPv4 35ms OK Find IPv4 (1 entry) 12ms OK Remove IPv4 by .id 18ms OK
=== BATCH ADD BENCHMARKS (sequential via API) === Add 100 IPv4 (sequential) 4s (40ms/ip, failures=0)Use the numbers to compare RouterOS versions, hardware models, TLS vs plaintext API, and network latency.
For real bouncer startup performance, prefer the functional CAPI tests because reconciliation uses script-based bulk adds and pool-based removals, not only sequential API calls.
Real-world results
Section titled “Real-world results”These figures come from the functional CAPI test suite (tests/functional/) run against a MikroTik RB5009, not from the single-operation cmd/benchmark tool above. See Architecture: Connection pool and Script-based bulk add for the mechanisms behind these numbers.
| Scenario | Result |
|---|---|
| Cache-first optimistic add (address not yet on the router) | ~1–3 ms per operation |
| Check-first add (list entries before adding, avoided by cache) | ~400 ms per IP |
| Script-based bulk add vs. sequential API calls | ~97× faster for large batches |
| Full RB5009 bulk-add, ~22,000 CAPI-origin entries | ~29 s bulk, ~34 s from start |
The gap between cache-first (~1-3 ms) and check-first (~400 ms) is why the in-memory address cache described in Architecture exists: it lets the bouncer skip the RouterOS list-then-add round trip for addresses it already knows about.
For CAPI-scale deployments (tens of thousands of entries), also see CAPI Blocklists and Performance Tuning for pool_size sizing guidance.
The full lifecycle, sampled every 100 ms
Section titled “The full lifecycle, sampled every 100 ms”The chart below is a single continuous 70-second capture from the production
RB5009, taken with a 10 Hz sampler reading /proc/stat on the router itself
(the section after it names the instrument, and what has replaced it). It
covers the three states an operator actually meets: the router idle with the
bouncer stopped and the address list empty, the first reconciliation
importing ~22,000 entries, and the quiet state that follows with the full list
in place. CPU reads on the left axis, RAM on the right, and the dashed markers
are taken from the bouncer’s own log timestamps. Each point is the mean of five
100 ms samples; the sub-second peaks reach 57% for single 100 ms slices during
the import.
The import is the plateau: roughly 31% sustained for ~26 seconds, spread across all four cores with the connection pool rotating single-core peaks of 50–100%. RAM is the quiet part of the story, with one surprise in the timing: the ~19 MiB step lands at connection time, before a single entry is imported — and 22,000 entries later it has not moved again. Holding the list costs the router almost nothing; memory is not where this workload bites, CPU during list reads is. The idle stretch also shows the router’s own background blips — a real device is never a flat line.
Where this capture came from
Section titled “Where this capture came from”The series was taken by cmd/perfmon, a one-off instrument this repository
carried from 1.5.0 until 1.6.0 removed it: it cross-built a /proc/stat
sampler, side-loaded it onto the router as a container, pushed batches to Loki,
pulled a window back as CSV and drew this SVG pair. That code is gone —
mikroscope does the same job as a
supported tool, and the next section is how to use it. What the instrument
measured stays: docs/src/data/perf-lifecycle.csv (141 rows, 100 ms apart) and
perf-lifecycle-markers.csv (three markers, read off the bouncer’s own log)
are still committed beside the figure, so every number above can be checked
against the data it came from.
What that costs, stated here rather than discovered later: the two SVGs can
no longer be regenerated from this repository. They were rendered with the
site’s own palette read out of theme.css, so a palette change will not reach
them any more. They are a frozen artefact of 2026-08-26, not a build output —
treat a new palette, or a new question, as a reason to recapture rather than to
recolour. mikroscope’s own plot is not a substitute for that generator either: it
draws three panels on one time axis from a recording, in its own palette with no dark
variant, so it produces a different figure rather than this one.
The method, as it was measured: CPU from /proc/stat at 10 Hz in a container
on the router, RAM from /proc/meminfo once per second, since memory did not
move at 100 ms on this workload. RouterOS’s own cpu-load updates only once
per second, so a script cannot see below that, which is the whole reason a
sampler existed. The capture protocol: stop the bouncer, delete both address
lists, let the router settle, then start the bouncer and record through the
first reconciliation into steady state.
Two things the dataset does not carry, and neither did this page before: the
RouterOS version of that particular capture — the other reference numbers here
were taken on 7.22.1 — and how ram_used_mib was derived. The instrument’s own
capture verb emitted mem_avail_kib and mem_free_kib, so the column
plotted as RAM went through one conversion step that never lived in the
repository. The CPU column is the instrument’s own output; the RAM column is
one arithmetic step away from it.
Measuring the router: use mikroscope
Section titled “Measuring the router: use mikroscope”cmd/perfmon is gone because it grew into its own project.
mikroscope is that project: a static
Go agent that runs on the router in a scratch container and reads the
shared kernel’s /proc at 1 to 100 Hz, 10 by default, plus a CLI on your
machine that installs it, records a window with markers, draws the chart, or
runs as a collector into eleven sinks. It carries the deployment steps, the
dockerless image builder and the RouterOS API client that started here, with
attribution in each package — so what follows is the same mechanism, now
tested, released and documented on its own.
What the observer costs the router is measured rather than promised: 2.69% of one core and 13.2 MiB resident at the 10 Hz install default, read from the agent’s own cgroup on an RB5009UG+S+ (4 × 1.4 GHz Cortex-A72, RouterOS 7.24.2), over one 300-second window at steady state with the collector forwarding to InfluxDB 3, on 2026-09-18. That is the same board this bouncer’s own numbers are taken on, which is what makes the two comparable; it transfers to no other board. What it costs says how the figure is taken and how to take it on your own device.
Putting the agent on the router
Section titled “Putting the agent on the router”curl -fsSL https://raw.githubusercontent.com/jmrplens/mikroscope/main/install.sh | bash
export MIKROSCOPE_ROUTER=admin@192.168.88.1mikroscope doctor # read-only preflightmikroscope install --ephemeral # lists every command it will run, asks, then writesmikroscope status # what is installed, and whether it answers--ephemeral keeps the one habit worth carrying over from the old instrument:
it forces the container root onto tmpfs and registers it with
start-on-boot=no, so an instrument leaves no trace on a device it was only
borrowing. Leave it off for a deployment meant to survive a reboot. With no Go
toolchain on your machine, add --remote-image jmrplens/mikroscope-agent:latest
and the router pulls the image itself — nothing is built and nothing is
uploaded.
The two default-firewall traps that used to eat this repository’s sampler
traffic are the same two, and mikroscope’s installer handles both: the raw rule
drop the rest (in-interface-list=!LAN) drops every packet from a veth outside
the LAN interface list, and drop local if not from default IP range matches
a source outside the LANs address list. The signature is worth recognising
because nothing logs it: ICMP to the container answers while its outbound TCP
produces no conntrack entry and no masquerade counter, and the sniffer shows
nothing. The two firewall
traps is the long version.
Recording a window and drawing it
Section titled “Recording a window and drawing it”mikroscope record --for 3m --out first-reconcile # a line typed here becomes a marker# in another shell: systemctl start cs-routeros-bouncermikroscope plot --in first-reconcile # first-reconcile.svg, deterministicrecord writes first-reconcile.jsonl, .csv and .markers.csv, and
mikroscope mark --out first-reconcile "reconciliation complete" adds a marker
to a recording that already exists. It starts at the agent’s newest sample, so
start it before the bouncer — or pass --from-start to pull whatever the
agent’s ring still holds first, 60 seconds of it at the install default.
The capture protocol from the 2026-08 run is the part that decides whether the numbers mean anything, and none of it has changed: stop the bouncer, delete both address lists, let the router settle, then start recording, and only then start the bouncer. One thing has improved. Memory is now read and shipped on every tick rather than once a second — on the reference device the kernel’s memory levels moved about 24 times a second, so a floor there would have thrown away real movement — which means the RAM series comes back at the sampler rate instead of at 1 Hz.
Reading the 10 Hz series without touching the router
Section titled “Reading the 10 Hz series without touching the router”The point of a collector is that the window you want is already stored.
mikroscope forward --influx http://influx:8181 --influx-db mikroscope writes
the kernel tier into InfluxDB 3 continuously — the token comes from
MIKROSCOPE_INFLUX_TOKEN, never from a flag — and from then on a CPU question
is a query, not a deployment. The token reaches curl on standard input, as a
--config file, because an argument is visible to any local user through ps
and -H "Authorization: …" is an argument like any other:
printf 'header = "Authorization: Bearer %s"\n' "$MIKROSCOPE_INFLUX_TOKEN" | curl -s --config - --get \ http://influx:8181/api/v3/query_sql \ --data-urlencode "db=mikroscope" \ --data-urlencode "format=json" \ --data-urlencode "q=SELECT time, avg(busy_ratio) * 100 AS cpu_pct FROM mikroscope_cpu WHERE host = 'router' AND time >= '2026-09-24T18:00:00Z' AND time < '2026-09-24T18:02:00Z' GROUP BY time ORDER BY time"host is whatever the collector’s --host-tag sets — router by default, and
rb5009 on the collector behind this project’s own numbers. busy_ratio is a per-core
fraction between 0 and 1, computed by the sink and clamped at 1, so averaging it over the
cpu tag at one timestamp is the all-core percentage the retired dataset called
cpu_pct. The raw tick deltas
sit on the same row, with the sample’s real interval dt_ns, if you would
rather do that arithmetic yourself — and a rate taken over a nominal period
instead of the real one is exactly the kind of error the next section is about.
| Column of the retired dataset | Where the same quantity comes from now |
|---|---|
cpu_pct | avg(busy_ratio) * 100 over the cpu tag of mikroscope_cpu at one time |
ram_used_mib | (total_kb - available_kb) / 1024.0 from mikroscope_mem, one row per tick |
perf-lifecycle-markers.csv | a line typed into record, or mikroscope mark afterwards |
scripts/router-perf-query.sh in this repository wraps that curl for the
handful of questions this project actually asks — per-core CPU, memory,
softnet drops, the agent’s own cost, and the continuity of the window — and
takes MIKROSCOPE_INFLUX_URL and MIKROSCOPE_INFLUX_TOKEN from the
environment, the same variable names the collector uses. Its header documents
the two traps worth knowing: the softnet and observer counters are per-sample
deltas, so they must be summed rather than differenced, and a partial
date_bin bin at the edge of a window averages over less than a second.
InfluxDB and SQL
measurements is the
column list for every table, and the InfluxDB
sink is what writes them.
Grafana reading the same database is the shortest path to a chart; this
repository’s own grafana/dashboard.json is Prometheus-only and covers the
bouncer’s metrics, not the router’s kernel.
Measuring against a live router
Section titled “Measuring against a live router”Two habits separate a number you can act on from one that misleads, and both were learned the expensive way while tuning this bouncer.
Interleave the arms. A router is not a quiet machine: it routes traffic, runs its own background work, and drifts over minutes. Measuring variant A for a while and then variant B for a while attributes that drift to the change. Alternate them instead — A, B, A, B — inside the same connection and the same minute, and compare paired samples. The first in-situ comparison of the proplist reduction ran as two separate windows and reported −2%; the interleaved run reported −11.2%, with every one of 48 pairs favouring the same side. Same code, same router; only the schedule differed.
Distrust a curve that is not monotonic. Sweeping the bulk chunk size once per value, in sequence, produced this:
chunk 50 6.926schunk 100 4.679schunk 250 9.760s <- twice its neighbours on both sideschunk 500 5.395schunk 1000 6.194sNo cost model produces that shape.
Report a spread, not a single number. The obvious repair — interleave the sizes across rounds and take the best time each — is still wrong, just less obviously. A minimum is a defensible estimator when noise is one-sided, but it throws away exactly the information that decides whether the comparison means anything. Re-run keeping every sample, the same sweep says:
| chunk | median | mean | IQR |
|---|---|---|---|
| 50 | 5.78 s | 6.86 s | 1.76 s |
| 100 | 5.06 s | 5.31 s | 0.22 s |
| 250 | 5.31 s | 6.39 s | 2.47 s |
| 500 | 5.50 s | 5.77 s | 0.84 s |
| 1000 | 4.36 s | 5.75 s | 2.55 s |
The spread within a single chunk size reaches 2.55 s, while the gap between the best and worst size is 1.43 s. The noise on one arm is larger than the difference between arms, so the sweep resolves nothing — and sure enough, which size “wins” depends on the estimator: median and minimum pick 1,000, the mean picks 100. Reporting only the best of three rounds would have hidden all of that behind a clean-looking table.
The conclusion survived, because it never rested on the ranking: inserting a row dominates, so no chunk size in this range moves the import measurably. But the first write-up of it stated a two-term cost model to two significant figures that the data does not support.
A useful reflex: before believing a result, ask what else changed between the two measurements. Usually it is time.
Safety checklist
Section titled “Safety checklist”- Confirm the target is a lab router.
- Use dedicated test address lists when possible.
- Keep
command_timeouthigh enough for slow devices. - Check the router after interruption and remove any
benchmark-*comments if a run is stopped midway.