Tested on
Where each part of mikroscope was run, measured or checked: the device, the RouterOS version, the
date and the conditions, what each run found, and what has not been tested. The guides state what
the tool does and link here for the proof. The feature verdicts were checked against the code at
1.2.0 and 1.2.2 on 2026-09-24, the sections that came from the guides against the code of 1.3.1 on
2026-09-26, what this page says of the code after 1.3.1, which 1.4.0 ships, against that code on
2026-09-27, and its sections on publishing to Grafana, uninstall and the test suites against 1.5.0’s
code on the same day; the current release is 1.6.1.
Reference hardware
Section titled “Reference hardware”Every figure measured on router hardware in this documentation comes from one MikroTik RB5009UG+S+: arm64, 4 × 1.4 GHz Cortex-A72 (r0p1), 1 GiB of RAM, Linux 5.6.3. It is the production router of the project’s maintainer, José Manuel Requena Plens, so nothing that needs a reboot is done on it outside a maintenance window, and no conntrack storm is provoked on it, which risks locking out the path being worked through. There is no second physical device. The only other RouterOS the project runs is the virtual lab. Which RouterOS it ran when is under Devices and versions.
Kernel and PMU
Section titled “Kernel and PMU”- Accounting files (campaign
kernel-2026-09-11): no/proc/pressureand no/proc/schedstat, and theirqcolumn of/proc/statalways 0, so hard-IRQ time is counted insidesystem.USER_HZis 100, one tick of 10 ms. - Tracing: a privileged discovery container on 2026-09-12 found no eBPF, kprobes or ftrace: no
BTF, no debugfs, no tracefs, and the mount points do not exist. A re-probe on 2026-09-14 found
/sys/fs/bpfpresent and empty and/proc/moduleslisting 238 loaded modules, but no/lib/modules, no kernel headers and no compiler on the device. - PMU, from inside a privileged container on 2026-09-12 (RouterOS 7.24.2): the probe opened
cycles,instructions,cache-misses,branch-missesandbus-cyclessystem-wide on 4 of 4 CPUs. In the agent’s own data over the 24 h ending 2026-09-12, six of the seven counters it asks for reported on 4 of 4 cores,cache-referencesamong them, andbranch-instructionsproduced no rows at all. The genericstalled-frontendandstalled-backendevents returnENOENTon the A72. - Work the tick misses: in a 100.4 ms sample where
/proc/statreported zero busy ticks on all four cores, the PMU counted 2.2–4.6 million cycles and 0.8–2.1 million instructions retired, at a 3.8–5.7 % cache-miss rate (recorded 2026-09-12). Over 2 s the same day, instructions per cycle ranged from 0.381 on cpu0 to 0.992 on cpu1. - Privileges:
perf_event_paranoidreads 2 and does not block a privileged container, which keepsCAP_SYS_ADMINeffective against the host (measured 2026-09-12). An unprivileged container already holds all 38 capabilities, and its root maps to host uid 32768, so/proc/slabinfoand/dev/kmsgcome backEACCES(RouterOS 7.24.2; the date of these two readings is not recorded). Withprivileged=yesthe kernel log,/proc/slabinfo, the MTD ECC counters and the PMU were found readable on 2026-09-12 (RouterOS 7.24.2, kernel 5.6.3).
Sensors and clocks
Section titled “Sensors and clocks”/sys/class/hwmonis empty, even privileged. The two thermal zones,cpu-thermalandsoc-thermal, are the whole sensor set of the base model, and both are readable without privileged. RouterOS’s/system/healthon this board reports exactly one sensor,cpu-temperature, which is thesoc-thermalzone truncated to whole degrees: a mean offset of +0.55 °C over the 8 minutes both series existed (2026-09-14).- Both zones declare a
polling_delayof 1 000 ms and apolling-delay-passiveof 250 ms (read 2026-09-14), so the kernel re-reads the sensor at 1 Hz. The declared critical trip is 105 °C, read from/sys(date not recorded). The sensor quantises to about 0.42 °C (floors-overnight). - The cpufreq governor read
userspaceon 2026-09-14, and the cpufreq clusters, read fromrelated_cpusandaffected_cpusinside the container the same day (RouterOS 7.24.2), are{0,1}and{2,3}.
Container view
Section titled “Container view”- Host-wide inside the container (established 2026-09-11, RouterOS 7.24.2, kernel 5.6.3): the
CPU, interrupt, memory and block-device files, and
/proc/net/softnet_stat. An ordinary container also read these as the router’s on 2026-09-12: the two thermal zones,scaling_cur_freqper core,/proc/yaffsand/proc/buddyinfo./proc/device-tree/modelreadsRB5009even unprivileged. - The container’s own (campaign
netns-2026-09-12):/proc/net/dev,/proc/net/snmp,/proc/net/netstatandnf_conntrack_countdescribed the container’s veth, 4 packets while the router forwarded millions, andprivileged=yesdid not change that. The rest of that boundary:
| What | What the container sees | Read |
|---|---|---|
/sys/class/net |
only lo and the veth; no device for any front-panel port, privileged or not |
2026-09-14 |
/sys/class/mdio_bus, /sys/class/phy |
mdio_bus holds only fixed-0; phy is empty |
2026-09-14 |
netdev_budget, netdev_max_backlog and the other global net.core entries |
absent; per-namespace entries such as somaxconn are present |
2026-09-14 |
/proc/net/* files created by modules |
the container’s own: fib_trie shows only the veth’s /30, snmp6 counts its packets |
2026-09-15 |
nf_conntrack_max |
the router’s real ceiling, 966 656, the same as RouterOS reports as max-entries |
2026-09-14 |
| conntrack timeouts | Linux defaults (tcp_timeout_established 432 000 s against RouterOS’s 1d) |
2026-09-14 |
Scroll sideways to see every column
- Host mounts do not cross the boundary: see the verified fact of 2026-09-15.
- The device-info record carries the board’s own ceilings, never a number from elsewhere: the 966 656 conntrack ceiling and the container’s 64 MiB memory cap.
Storage
Section titled “Storage”nbd0–nbd15are present and idle andmtdblock0–2read all-zero (2026-09-12, RouterOS 7.24.2, kernel 5.6.3;internal/, the disk in-flight panel). A device’sdashboards/ panels.go /proc/yaffsor/proc/diskstatsrow is stored only when its delta is non-zero, which here is a few times a minute. The reference InfluxDB store still held no block-device table, and no PSI table, on 2026-09-25, on RouterOS 7.24.4.- The MTD ECC counters are all zero after years of service (date of the reading not recorded).
/proc/slabinfois 13 833 bytes and 129 lines per read (2026-09-14), the largest per-tick parse by an order of magnitude.- The router log held 66 217 rows when
recordread it, because itsdnstopic logs to disk (date not recorded). - It has a tmpfs disk, which
--ephemeraluses. Its raw firewall carries MikroTik’s twodefconf:drop rules in their list form,in-interface-list=!LANandsrc-address-list=!LANs(Firewall lists).
Ports and interfaces
Section titled “Ports and interfaces”- Kernel names against RouterOS names (Port names):
eth1=ether2from the layer-2 loop of the case study (2026-09-13).eth5=ether6andeth6=ether7on 2026-09-15: both ports are commentedUnusedand neither wasRUNNING, so each was disabled and enabled again over the API while the agent read/dev/kmsg. The kernel loggedbr0: port 7(eth5) entered disabled stateinsideether6’s window andport 7(eth6)insideether7’s, nine seconds apart, which rules out reading one flap twice. No traffic was interrupted and both ports were back within four seconds. The six inferred pairs and the switch rest on a read-only/interface/of 2026-09-13: consecutive MAC addresses,ethernet/ print …:55forether1through…:5Dforsfp-sfpplus1, in RouterOS’s enumeration order, and all nine ports reportingswitch=switch1against the kernel’s oneswitch0. The device tree, parsed on 2026-09-14, shows the SoC’sethernet@0with three MACs, onlyeth0enabled (the 10 G uplink) andeth1/eth2disabled; the nine front-panel ports are netdevs the switch driver creates at runtime. The table is this board’s, observed on RouterOS 7.24.2. - Kernel-log record shapes for each port event kind were seen between 2026-09-12 and 2026-09-15 (kernel 5.6.3).
- Inventory (2026-09-16, RouterOS 7.24.2): the three configuration reads return 17 interfaces.
ether1is anetherin theLANlist, a port ofbridge, labelled “TrueNAS - High Performance Storage”, MTU 9000;ether5is anetherinWANand in no bridge, labelled “DIGI ONT”;PPPoE_DIGIispppoe-out,WAN, MTU 1480;VLAN_DIGIisvlan,WAN;wg_devicesandwg_trasteroarewginLAN,VPN;ether6andether7carry the comment “Unused”. - Counters: 9 ports × 45 counters, and 15 on each bridge, VLAN, PPPoE, WireGuard, veth or
loopback interface, counted from the collector’s own
mikroscope_rows on 2026-09-24. The RouterOS version was not read with them; the router ran 7.24.4 that day.api_ ifcounters
Virtual lab
Section titled “Virtual lab”MikroTik’s Cloud Hosted Router under QEMU, in a Docker container, on the machine the lab was built
on (x86_64, 12 cores, 27 GB RAM, Docker 29.8, QEMU 10.0.13), provisioned into a clean snapshot
with the container package and device-mode container=yes. Measured on 2026-09-26 with RouterOS
7.24.4 and the published 1.3.1 CLI and agent image:
- CHR x86_64 under KVM: 2 cores, kernel
5.6.3-64. - CHR arm64 under QEMU’s TCG emulation,
cortex-a72, 2 vCPU, 1 GiB: kernel5.6.3, the RB5009’s architecture, kernel version and CPU model, at the speed of the host.
What it shows: correctness. The RouterOS commands mikroscope sends and what RouterOS answers,
doctor’s reading of the router, RouterOS choosing and pulling the image for its architecture, the
container’s start and stop, the agent’s capability detection, its sample format and its ring. The
install routes that ran there are under Install routes tested.
How the lab reaches its agents: through a static route. The lab’s LAN side, 192.168.88.10/24
inside the lab’s container, has its default route through Docker’s bridge, not through the router,
and reaches every agent through 172.30.0.0/16 via 192.168.88.1, the default of
LAB_AGENT_ROUTES: the wider prefix of Static route.
On 2026-09-27 the default 172.30.10.0/30 and a 172.30.11.0/30 install both answered /healthz
through it.
What it cannot show:
- A board. There is no switch chip, flash, sensor or device-tree model, so no port map. The agent’s
/capabilitiesreportscpufreq,mtd,psi,schedstatandthermalabsent on both,perfandkmsgpresent on both, andyaffsabsent on x86_64 and present on arm64, a/proc/yaffswith no flash behind it. - Rates above 10 Hz. The free CHR licence caps what the router sends at 1 Mbit/s per interface:
2 MiB copied off the router over ether2 took 15.6 s, about 1.07 Mbit/s, against 0.2 s onto it.
/streammeasured 2 526 B a line on the lab, 0.21 Mbit/s at 10 Hz, so 50 Hz is at the cap and 100 Hz over it. Neither was measured. - Any cost on arm64. Under TCG the guest’s clock follows the host’s, so no duration, no CPU figure
(
cpu-load,self.cpu_us,read_ns,wake_ns, thedt_nsspread,slipped), no interrupt, softirq or context-switch rate and no PMU count from the arm64 lab measures a Cortex-A72. The agent kept up at 10 Hz there, 3 003 ticks in 300.1 s with 1 slipped and 1 202 in 120.1 s with none, and its resident size of 14–16 MiB (32 MiB charged to its cgroup) is an indication only.
A routing loop in the lab itself, since fixed, made the emulated router crawl whenever something
probed an agent address with no veth behind it: before the fix, arm64 install took 38.4, 11.8 and
34.1 s, and the first and third exited 1 on a complete install. The arm64 figures here are from after
the fix; the x86_64 ones were taken before it. What 1.3.1 met there is under Known issues.
Lab timings
Section titled “Lab timings”On the machine above, from the clean snapshot. How each suite is run is on Test suites.
- A boot from the snapshot until ssh answers, 2026-09-26: 7 s on x86_64 (6.9, 7.2, 7.4 and 7.4 s) and 26 to 28 s on arm64 (25.7, 26.6, 27.0 and 27.5 s).
make roundtrip, 2026-09-26:doctor,install,status,upgradeanduninstall, every verb with--ephemeral, and the router’s/exporthashed before and after. 28 to 34 s on x86_64 over three runs and 40 to 45 s on arm64 over four, the export byte-identical every time. With 1.4.0’s code, which makes oneuninstallattempt, on 2026-09-27: 20 s on x86_64 and 35 s on arm64, one run each with make’s build steps included, the export byte-identical. With 1.5.0’s code on the same day: 21 s on x86_64 and 30 s on arm64, measured the same way, the export byte-identical.make test-labwith the scenarios of 1.3.1, 2026-09-26: 7 min 21 s to 9 min 49 s on x86_64 over six runs and 12 min 12 s to 16 min 48 s on arm64 over three, the slowest of each with both suites running side by side on a busy host. No install failed.make test-labwith the install options’ scenarios, with the code after 1.3.1, 2026-09-27: 28 min 32 s on x86_64 alone, then 28 min 17 s on x86_64 and 57 min 29 s on arm64 side by side. After the fixes the review asked for, 29 min 2 s and 57 min 13 s side by side, every test passing; S9’s twenty uninstalls with a client on/stream, ten per architecture, were each clean at the first attempt, in 8.0 to 12.6 s. The same day S17 ran on a CHR x86_64 lab with RouterOS 7.23.7:doctorwas MISSINGRouterOS 7.24 or laterand read every other check.- RouterOS 7.24.4 changes its own
/export: it adds and drops a/system keymat-provider … name=defaultline by itself, so the suite’s comparison leaves that line out (2026-09-26).
Power cut after install
Section titled “Power cut after install”On 2026-09-26, a power cut made as soon as a fresh install answered brought back no agent within
90 s in three of four tries on the arm64 lab. The two containers looked at could not start
(Exec format error, Segmentation fault), most likely because RouterOS had not yet written the
install to its disk; that was not examined. On x86_64 three of three came back. The start-on-boot
scenario, S7, waits 45 s between the install and its cut, and asks only whether start-on-boot works.
S6 cuts the power under an --ephemeral install on the lab’s tmpfs disk: afterwards the container
is configured and stopped, its root, its image and the manifest are gone with the disk’s contents,
the disk is there and empty, and nothing answers; uninstall --ephemeral then leaves nothing at
all. It passed on both architectures in the runs of 2026-09-27.
Reboot detections
Section titled “Reboot detections”On 2026-10-05, on the x86_64 lab (CHR, RouterOS 7.24.4, KVM), the collector of the code after 1.5.0
ran forward --api-mode off --stdout json across a reboot made with RouterOS’s own
/system/reboot, the 1.5.0 agent installed by the Docker Hub pull with start-on-boot. The agent
answered again about 20 s after the reboot; the collector logged agent restarted: its newest sample is 439 and the cursor was 1057; resuming from 1, and agent-restart fired with its summary
of the 30 s before: 300 samples, CPU 0 % busy, MemAvailable 767.5 MiB of 863.0 MiB, nf_conntrack 24
of 843 776, no softnet drop, then no sample for 16.2 s. reboot did not fire: no kernel-log record
after the agent came back had a since-boot time below the last one before the reboot. One run, on an
idle router.
Later the same day, on the same lab, the agent and the collector of the code after 1.5.0 that reads
the kernel’s boot id (the agent from the branch’s tar, with start-on-boot) ran seven minutes of
forward --api-mode off --stdout json across a stop and a start of the agent’s container and then a
/system/reboot. The agent inside the RouterOS container read the id and served it in /healthz.
After /container/stop and /container/start (20.6 s without a sample) the id was the same, and
agent-restart said so: the kernel's boot id did not change, so the router did not reboot; no
reboot. After /system/reboot (16.5 s without a sample) the id was new: the collector logged
router rebooted: the kernel's boot id went from 30b9831d-… to 8d4e19f2-…, and agent-restart and
reboot fired on the agent’s first sample with the same summary (300 samples, CPU 0 % busy,
nf_conntrack 24 of 843 776, no softnet drop). The new boot’s kernel log fired no second reboot.
One run of each, on an idle router; the arm64 lab was not run.
What RouterOS logs about a boot was read the same evening over the API, from /log/print, after
each way of taking the x86_64 lab down. After /system/reboot from an SSH session its memory buffer
held router rebooted by ssh-cmd:admin@192.168.88.10/ (topics system,info); from a script,
…/script:rb/reboot; after /system/shutdown and a start, …/shutdown; after the power was pulled
(power-cycle), router was rebooted without proper shutdown (system,error,critical), and the
same after QEMU’s system_reset (reset-button). With a logging action that also sent the system
topic to disk, a plain /log/print returned the previous boot’s lines as well, and /log/print ?buffer=memory returned this boot’s alone.
Then the collector of the code that reads it (forward --api-mode slow --stdout json, the API
tier given the lab’s credentials through LAB_CLI_API=lab) ran eight minutes across a
/system/reboot and a reset-button. Both times the read succeeded at the health read that noticed
the new boot id, before the API tier’s own round had reconnected (the read reopened the
connection), so reboot fired on the agent’s first sample, with RouterOS logged at boot: "router rebooted by ssh-cmd:admin@192.168.88.10/ the first time and "router was rebooted without proper shutdown" the second, 19.4 s and 22.1 s without a sample. One run of each; the two-minute
wait for an API that is not back was not exercised on a router, only in the unit tests.
On the RB5009 (RouterOS 7.24.4), the 1.6.0 agent installed on 2026-10-06 served the kernel’s boot
id in /healthz from inside its container. The agent-restart of that upgrade said the kernel's boot id did not change, so the router did not reboot, which was true but not known: the 1.5.0 agent
before it reported no id. In the code after 1.6.0 the first agent to report an id after one that
did not makes no claim about the kernel.
A power-cycle of the lab is no test of the collector: it restarts the lab’s container, and a
collector started before it with mikroscope-lab cli keeps the old container’s network namespace and
reached nothing after the cut (no route to host until it stopped, the same day). The reboot from
inside keeps the lab’s network.
Doctor exposure advisories
Section titled “Doctor exposure advisories”On 2026-10-05, on the x86_64 lab (CHR, RouterOS 7.24.4, KVM), doctor of the code after 1.5.0 that
adds the two advisories ran against the lab router set up three ways. With allow-remote-requests=yes
and no firewall rule, it warned the router does not answer DNS from its uplink (…, uplink ether1: no rule drops a query that comes in on it). With the default configuration’s two input rules (accept
established,, drop in-interface-list=!LAN) added ten seconds before, it passed,
naming the !LAN drop; with in-interface=ether1 protocol=udp dst-port=53 action=drop alone it
passed too. With a filter rule whose interface list had been removed and a raw rule whose veth had
been removed, it named both under no firewall rule doctor reads is invalid or names a deleted list.
What RouterOS does with a rule whose reference goes away was read over SSH the same day. An
interface list removed under a rule left the rule valid, reading in-interface-list=!*2000010;
creating a list of the same name did not change it, and the rule counted packets as a catch-all
beside it did. A !LAN drop left that way on the input chain cut the lab’s own SSH from the LAN
until the lab was reset. A veth removed under a rule made the rule invalid (in-interface=*4,
about=vprobe not ready), and it counted no packet. A rule read right after it was added was
invalid with no about, and valid five seconds later. RouterOS refuses to add a rule that names a
list that does not exist (input does not match any value of interface-list).
For the fix’s placement: place-before=0 failed from a one-command SSH session (no such item),
and a place-before of the input chain’s first rule failed on an empty chain and worked with a rule
there. Not tested: IPv6, and a DNS query sent from outside the lab, whose uplink is QEMU’s user
network; doctor reads the rules and sends no packet.
On the RB5009 (RouterOS 7.24.4, 2026-10-06), 1.6.0’s doctor read 80 enabled rules in raw
prerouting and filter forward and input, none invalid and none naming a deleted list, and said the
DNS check had no uplink to judge: the router’s active default route is over PPPoE, and its
immediate-gw is the interface alone (PPPoE_DIGI), which the uplink read did not handle; the
lab’s DHCP uplink reads 10.0.2.2%ether1. With that read fixed, the code after 1.6.0 found
PPPoE_DIGI in the WAN list and judged the queries a maybe, because the default configuration’s
accept to local loopback (for CAPsMAN) rule, dst-address=127.0.0.1, stands before its drop all not coming from LAN and read as a maybe accept. With a loopback destination taken as no match for a
packet from the Internet, it named defconf: drop all not coming from LAN as dropping the queries.
Each of the three runs was one read-only connection.
RouterOS x86 from the ISO
Section titled “RouterOS x86 from the ISO”The opt-in lab that installs RouterOS x86 from MikroTik’s ISO ran on 2026-09-26 with RouterOS
7.24.4. The router reported the board x86 QEMU Standard PC (Q35 + ICH9, 2009) and a trial licence
with no level and 24 hours to run. The 1.3.1 CLI and the agent of the lab’s branch behaved as on
CHR x86_64: doctor missed the same two lists, a tar install answered at once, status recognised
every object and uninstall verified the router clean, leaving the same empty mikroscope
directory. The agent’s /capabilities were CHR x86_64’s.
Lab runs in CI
Section titled “Lab runs in CI”The workflow first ran on GitHub’s runners on 2026-09-26, as the lab job of CI on pull request #69:
four x86_64 runs between 21:06 UTC that day and 00:45 UTC on 2026-09-27, each passing the ten tests
the suite then had (471.7 to 485.8 s) in a job of 10 min 30 s to 13 min 30 s. Of the three runs
below, the first two had an empty cache, so make lab-up downloaded RouterOS and provisioned it; the
third took the downloads from the cache and provisioned a new snapshot.
| Run | Commit | Runner | make lab-up |
Suite | Job |
|---|---|---|---|---|---|
| arm64, a dispatch on main | e7efbbe, the lab as merged |
ubuntu-latest: 4 CPUs, 15 989 MB, /dev/kvm present |
4 min 29 s: provisioning 83 s, a boot of the snapshot 24 s | 638.6 s, the ten tests of that commit passing | 16 min 24 s |
| x86_64, pull request #70 | 3c34d9a, the install options |
ubuntu-latest: 4 CPUs, 15 989 MB, /dev/kvm used by KVM |
4 min 28 s: provisioning 46 s, a boot of the snapshot 8 s | 1 591.4 s (26 min 31 s), every test passing; S17 skipped, as it needs a RouterOS below 7.24 | 32 min 42 s |
| x86_64, pull request #70 | 1826aa5, with the registry credential |
ubuntu-latest, /dev/kvm used by KVM |
1 min 25 s: the downloads from the cache, provisioning 76 s | 1 633.1 s (27 min 13 s), every test passing; S17 skipped. The router pulled as the repository’s Docker Hub account, S2 included; S1, S18 and S5’s Docker Hub and GHCR scripts booted without it | 29 min 52 s |
Scroll sideways to see every column
In the first two of those runs one download from MikroTik was cut (connection reset by peer) and
resumed at the first retry. The probe for KVM on GitHub’s arm64 runner found ubuntu-24.04-arm with
4 CPUs and no /dev/kvm, so the arm64 lab stays emulated on an x86_64 runner.
The two runs on pull request #70 held it for 30 and 33 min, which is why lab.yml now runs weekly and on
dispatch only, never on a pull request or before a release.
On 1.4.0’s release commit (79b7c2f, a dispatch on 2026-09-27, run 36329638066) the whole suite
passed on both architectures: 1 517.2 s on x86_64 in a job of 31 min 38 s, and 3 042.6 s on arm64,
emulated, in a job of 54 min 52 s. The agent image 1.4.0 was not published yet, so the thirteen
golden cases that pull their image pulled 1.3.1’s, as S5 does before a tag and says in its log.
On 1.5.0’s release commit (e61ffb6, a dispatch on 2026-09-27, run 36344788569) the whole suite
passed again on both architectures: 1 573.2 s on x86_64 in a job of 31 min 38 s, and 3 065.8 s on
arm64, emulated, in a job of 54 min 55 s. The thirteen golden cases that pull their image pulled
1.4.0’s, 1.5.0’s being unpublished, as S5 logged.
Devices and versions
Section titled “Devices and versions”| Device | Kind | RouterOS | When | What ran there |
|---|---|---|---|---|
| RB5009UG+S+ | hardware, arm64 | 7.24.1 | until the upgrade of 2026-09-10 | the SSH connect cost, 2026-08-26 |
| RB5009UG+S+ | hardware, arm64 | 7.24.2 | 2026-09-10 to about 2026-09-18 22:43 UTC | most campaigns, 2026-09-11 to 2026-09-18 |
| RB5009UG+S+ | hardware, arm64 | 7.24.4 | from about 2026-09-18 22:43 UTC | the port-errors campaign, the /stream timings, the charts from the reference store, and the deployment commands from 2026-09-21; upgrade with 1.5.0 and 1.6.0, doctor with 1.6.0 |
| CHR x86_64, virtual lab | virtual (KVM), amd64 | 7.24.4 | 2026-09-26 and 2026-09-27 | doctor, plan, three install routes, --expose, uninstall with 1.3.1; the whole lab suite with the code after it and on 1.4.0’s commit |
| CHR arm64, virtual lab | emulated (TCG), arm64 | 7.24.4 | 2026-09-26 and 2026-09-27 | doctor, plan, three install routes, uninstall with 1.3.1; the whole lab suite with the code after it and on 1.4.0’s commit |
| CHR x86_64, virtual lab | virtual (KVM), amd64 | 7.23.7 | 2026-09-27 | doctor, which refuses a RouterOS below 7.24 (S17) |
| RouterOS x86 from the ISO, virtual lab | virtual (KVM), amd64 | 7.24.4 | 2026-09-26 | doctor, a tar install, status and uninstall with 1.3.1 |
Scroll sideways to see every column
When the RB5009 moved to 7.24.4. The upgrade itself was not written down, but it is bounded: 7.24.4’s build time is 2026-09-16 11:32:21; on 2026-09-24 the router reported 7.24.4 with an uptime of 5d15h49m12s, a boot at about 2026-09-18 22:43 UTC; and a RouterOS upgrade rebooted it at 00:43:30 CEST on 2026-09-19, 22:43:30 UTC the day before, after which the kernel tier resumed at 00:45:03. The API tier’s uptime in the reference store grows at clock rate from 2026-09-19 11:13:35 UTC. So the router has run 7.24.4 since that boot at the latest. Campaigns dated 2026-09-16 to 2026-09-18 are recorded as 7.24.2; the bound neither confirms nor refutes that. None of the 7.24.2 measurements has been repeated on 7.24.4.
What ran on 7.24.4. doctor, plan, install, status and uninstall on 2026-09-21, with
the installed agent running continuously since 2026-09-19. On 2026-09-23 the 1.2.x form of
doctor’s registry warning was reproduced read-only, and its token warning with a throwaway install
that was exposed, upgraded without a token and removed, so upgrade has run there too. The same
day doctor read 600 samples from the running agent’s ring and found nothing. On 2026-09-24
RouterOS was given the agent’s reference with its registry host inside remote-image=
(two verified facts). The API tier has run continuously on 7.24.4 since
2026-09-19; its inventory and loss-key readings and its cost profile are from 7.24.2 (2026-09-15
and 16), and its reconnection test and 16-interface A/B from 2026-09-19, after that day’s upgrade,
with the version not recorded beside them.
The container settings install writes were verified on the RB5009 on RouterOS 7.24.2 and
re-exercised on 7.24.4; no other board was tried.
The agent on the RB5009 was upgraded to 1.0.9 on 2026-09-19, to 1.2.0 at 08:08:54 UTC on
2026-09-24 and to 1.2.1 and 1.2.2 later that day, to 1.5.0 on 2026-09-27 and to 1.6.0 at 07:15 UTC
on 2026-10-06, each with upgrade --remote-image; the last took 11.8 s from the command to the
probe’s answer. The /stream timings of 2026-09-21 are from agent 1.0.9.
Platforms
Section titled “Platforms”Of the four platforms a release publishes, arm64 has run on hardware, the RB5009, and under
emulation in the lab. amd64 has run only in the lab, on CHR x86_64 and on RouterOS x86 from the
ISO, both under KVM, never on x86 hardware. arm/v7 and arm/v5 have never run on RouterOS: they
are cross-built, and CI starts each image with -version under QEMU user-mode emulation
(make agent-smoke), which is not RouterOS. What each release publishes, and which version numbers
never became a release, is on
Releases.
Feature status
Section titled “Feature status”Everything below ran against the RB5009 unless it says the container suite or the lab: on RouterOS 7.24.2 when dated up to 2026-09-18, and on 7.24.4 when dated from 2026-09-19 (Devices and versions). Each part opens with its verdict and the date behind it.
The agent works on the RB5009: the installed one was running there on 2026-09-23, when doctor
read 600 samples from its ring. It reads the shared kernel’s /proc, /sys, /dev/kmsg and
perf_event_open counters on a fixed ticker at 1 to 100 Hz (10 Hz by default; 10, 20, 50 and
100 Hz measured), keeps the samples in a ring, and serves them:
/healthz, /capabilities, /snapshot, /stream, /sampler, and the triggered-capture
endpoints /captures and /capture. It serves no /metrics; the exposition is the collector’s.
Deployment commands
Section titled “Deployment commands”All six deployment commands have run on the RB5009 on RouterOS 7.24.4, all six by 2026-09-23,
each with the CLI of its date. Of 1.4.0 and later, only upgrade (1.5.0 on 2026-09-27, 1.6.0 on
2026-10-06, the second reading the install manifest the first wrote) and doctor (1.6.0, 2026-10-06)
have run there; the rest of the code after 1.3.1, which 1.4.0 ships, has run only in the virtual lab
(2026-09-27). Read from the code, 1.5.0 changes none of
the six commands’ router steps; its uninstall refuses --grafana-dry-run and lists the Grafana
and store targets before it touches the router, both checked only in unit tests
(uninstall --targets). doctor, plan, install, status,
upgrade and uninstall install, upgrade and remove the agent, with every write listed before it
happens and every removal verified by ownership counts.
- Round trip, 2026-09-12:
doctor→install→status→upgrade→uninstallleft the router’s/exportbyte-identical (verified). The agent answered 3 s after install, at a 5–7 ms round trip: two probes, 7 ms afterinstalland 5 ms afterupgrade, the first printingdirect transport ok: agent 4857d0a-dirty, 10 Hz, seq 29, 0 slipped, 7ms round trip. That is one install on one network, not a figure for another. The run passed--ephemeraltodoctor,installandupgrade, and it predates 1.1.0, from whichuninstallremoves only with--yes.scripts/roundtrip.shnow passes--yesto its uninstall (through 1.2.0 it did not, so there it only listed and its export check failed) and--ephemeralto every verb. In that form it runs in the lab asmake roundtrip(Lab timings), and it has not been run against the RB5009, where it ismake roundtrip-device. - Install routes: four routes end to end on 2026-09-17, and the lab’s on 2026-09-26, under Install routes tested.
- Standalone
doctoralso pulls the running agent’s ring once and names four faults RouterOS’s own tools do not show: a layer-2 loop, STP churn, a link flap and softnet drops. It reports an agent that does not answer within 3 s and skips it, and its findings never change the exit status. On 2026-09-23 it read 600 samples and found nothing; the four faults are covered by tests that replay the real record text and the 2.0 s cadence of the RB5009’s loop, not by a live fault. Its STP-churn finding rests on how a healthy link-up looks: on the RB5009’s ether7, five link-ups on 2026-09-21 each logged three moves to learning at once and reached forwarding 2.1 to 2.8 s later, and across 30 days of that router’s store every healthy link-up left learning minus forwarding at 0. - Its
WARNchecks, which change neither the exit status nor whetherinstallproceeds. Two are older than 1.4.0: with--remote-image, a/container/configusername set whileregistry-urlis empty or names a host other than the one the image is pulled from; and an install of the same--namepublished on the LAN with noTOKENin its environment. It reads whether a username is set and how manyTOKENentries exist, never a value. The 1.2.x form of the registry warning was reproduced read-only on 7.24.4 on 2026-09-23; the host comparison that replaced it has warned only against fake router answers in the tests. 1.4.0 adds aWARNfor a pull on a 32-bit ARM router, for a pull with less than 16 MiB free beyond the container’smemory-max, for start-on-boot with the root on a tmpfs disk, for a firewall rule that may drop the agent’s replies, for a--lan-addresson the uplink, for objects tagged for the install that the flags do not select, and for a route table or firewall it could not read. Of these, the tmpfs and the tagged-objects warnings fired in the lab on CHR x86_64, RouterOS 7.24.4, on 2026-09-27; none has run on the RB5009. - Tar extraction, 7.24.2, 2026-09-11: a 1.8 MiB tar, a build of that date from before 1.0.0,
was extracted within the same second as its
/container/add. The CLI of that time deleted the tar after a fixed wait; 1.4.0 waits for the container to readstopped, bounded by--extract-timeout. - An exposed install upgraded without its token, 7.24.4, 2026-09-21: the upgrade passed its
check, left both rules in place and wrote an envlist with no
TOKEN, and/snapshotthrough the router’s LAN address went from401to200. From 1.2.0doctorreports that state, and the code after 1.3.1 refuses such anupgrade(lab, 2026-09-27). - A plain
uninstallof an exposed install, 7.24.4, 2026-09-21: it removed everything else, printedverified: nothing mikroscope created remains on the router, and left the dst-nat pointing at an address that no longer existed. The code after 1.3.1 reads the install’s shape from the router, and a plainuninstallremoved both rules in the lab (2026-09-27).
record, mark and plot
Section titled “record, mark and plot”record, mark and plot work end to end, shown by a recording made on the RB5009 on
2026-09-12: 60 s at 10 Hz, exactly 600 samples, 0 gaps and −7 ms of clock skew. What it showed is
under its campaign.
In the lab on 2026-09-27 (CHR x86_64, RouterOS 7.24.4, the published 1.3.1 agent image),
record --for 60s wrote 600 samples, seq 11 to 610, with 0 gaps. Three mark notes from a second
shell landed in the same .markers.csv while record ran, and record’s summary said
0 marker(s): it counts only the notes typed into it. mark --log-markers over the lab’s API added
7 markers from the router’s log in the window, and plot drew the 600 samples and 10 markers.
forward
Section titled “forward”forward has run from the RB5009 into three of its eleven sinks, file, Prometheus and InfluxDB 3,
all three at once on 2026-09-15; the other eight have been read back only from real products in
containers, since 2026-09-16. forward merges the kernel tier with the RouterOS API tier and
writes to eleven sinks. Loki, OTLP, Graphite, Elasticsearch, SQL, PostgreSQL, Telegraf and stdout
have not had router samples pushed through them; the container suite writes into each real product
and reads it back through its own API (Test suites).
- 2026-09-12, eight minutes: 4 800 kernel and 479 API samples forwarded with 0 gaps and 0 drops,
running
mikroscope forward --for 8m --prom :9124 --influx … --interfaces bridge,.ether1, PPPoE_ DIGI --conntrack-every 10s - 2026-09-15, five rate runs at 10, 50 and 100 Hz into a file, a Prometheus exposition and
InfluxDB 3 at once: every sink reported 0 gaps and 0 drops
(
rates-2026-09-15). - Prometheus: a Prometheus 3.14 scraped the collector’s exposition every 5 s with the RB5009 feeding the collector, on 2026-09-12 and again on 2026-09-15. No other scrape interval and no other Prometheus version is recorded.
- InfluxDB 3: the first real target, on 2026-09-12, answered
422: would exceed limit of 5 databases, the node limit InfluxData documents for Core. The measured runs used an InfluxDB 3 Core instance of their own. On 2026-09-19 the collector was moved onto InfluxDB 3 Enterprise 3.11.4 and wrote 44 820 rows across 36 tables in its first 19 minutes, 0 dropped and 0 errors. Everything before that date was measured against Core. - The collector against a faster agent, 2026-09-21: with the agent at 100 Hz and again at
50 Hz,
/healthzreadrate_hz100 and 50 and the collector’s log named the same rate back, whilemikroscope_info{rate_hz}read10at both;mikroscope_carried the bucketscpu_ busy_ ticks le="0"throughle="11"plus+Infat both; andwindow="1s",window="10s"andwindow="60s"were all present at both. - Detections in the reference store, 2026-09-19 11:13 UTC to 2026-09-24 08:26 UTC: four of the
eleven rules fired,
microburst116 times,ipc-collapse41 (on all four cores),link-flap11 andagent-restartonce. Onlymicrobursthas had its behaviour measured against the samples (microburst-replay-2026-09-16); the other three were counted, not checked against what the router was doing. The 11link-flaprows are onether1,ether2,ether4,ether6andether7, each at 2 to 6 link records in 60 s, all unprovoked; which was a cable, a device at the other end or something else was not checked. - Restarts: the RouterOS upgrade reboot of 2026-09-19 exercised the collector’s resync for real
(the kernel tier resumed at 00:45:03 CEST). Whether a
mikroscope_detectionrow foragent-restartwas written then was not checked, and the store that would hold it was replaced later that day; the current one starts at 11:13 UTC. It holds oneagent-restartrow, at 08:08:54 UTC on 2026-09-24,value1 againstthreshold4 338 037, when the agent was upgraded to 1.2.0. - Device info: sent once at start, the four device panels read “No data” over every window after it: on 2026-09-17 the last device row in the reference deployment was 26 hours old and those panels had been empty as long. Repeated every five minutes, it costs twelve rows an emission on the RB5009.
- Delivery over 24 hours, 2026-09-17, in the reference deployment: the largest interruption was 114.5 s, and it was self-inflicted, a container swap plus the minute the collector takes to notice a restarted agent. In ordinary running the collector never fell behind. That is why the ring holds 60 s by default: it covers a restart of either side on a LAN.
RouterOS API tier
Section titled “RouterOS API tier”The API tier has run on one board, the RB5009, continuously on 7.24.4 since 2026-09-19.
- Reconnection, 2026-09-19, after that day’s upgrade, without touching the router: a collector
running
monitor-trafficon 16 interfaces (all butlo) at 1 Hz had its API socket destroyed from the host withss -K, which is what the router’s side of a reboot looks like to it. The tier reopened the connection and retried inside the same round: 0 failed commands, 0 dropped rounds, 16 interfaces in every one of the 108 seconds either side of the kill. - What RouterOS returns: loss keys,
an
rx-overflowcount only the port counters carry, the two counting planes of a switch port and its bridge, and log times. cpu-loadis a trailing mean of about one second, fitted against/proc/stat(cpu-load-window-2026-09-15).- What it costs the router: API tier cost.
Dashboards
Section titled “Dashboards”Every panel of all five dashboards answered without error in a real Grafana over the container
suite’s stores (2026-09-17), and the InfluxDB and Prometheus dashboards passed dashboards check
over the RB5009’s own data (2026-09-16). There is one per store, InfluxDB 3, Prometheus,
PostgreSQL, Graphite and Elasticsearch, generated from one panel list. Graphite and Elasticsearch
carry fewer panels on purpose (39
and 28
against 177): Graphite has no labels and Elasticsearch no
nested documents. The PostgreSQL panels are the InfluxDB ones rewritten, and the container suite has
a real PostgreSQL plan every one of their queries. The Grafana versions used are 12.3.2 on
2026-09-12; 13.2.1 for the browser passes of 2026-09-12 and 2026-09-14, check and render on
2026-09-15, check on 2026-09-16 and the container suite; 12.3.0 on 2026-09-21 and in the renders
of 2026-09-25; and 13.2.2 on 2026-09-24 and 2026-09-25. No other version has been tried.
- 2026-09-12. On Grafana 12.3.2, the InfluxDB dashboard against an isolated InfluxDB 3 Core fed
by
forwardfrom the RB5009: every panel returned rows, 158–316 per panel over 10 minutes. The Prometheus dashboard against a Prometheus 3.14 scraping the collector every 5 s: every panel returned rows, 228–2 052 over 5 minutes. Without the datasource’stokenfield, panels failed withflightsql: Unauthenticated. The same daydashboards check --store influxdb --window 12hpassed on all 125 panels it walked while about 90 were unreadable in a browser: legends reading “value core 0”, two xycharts stuck on “Loading plugin panel…”, a continuity lane that stayed green over 2 170 missing ticks. An eight-panel Overview rendered 2 188 px tall in an 844 px phone viewport, and 900 data points is the width the browser sent for the graphs. - 2026-09-14. An eleven-panel Overview measured 1 052 px tall in a 1 080 px browser viewport,
one desktop screen; a render walk against
a 10.5 h capture from the RB5009 rendered 140 InfluxDB panels with 0 error badges and 0 “No
data”; the InfluxDB datasource escaped
$__interval_msin five panels, in the browser, into SQL InfluxDB 3 could not parse; and a query naming a field the store had never received failed at planning,No field named limit. Valid fields are …, through the datasource proxy. - 2026-09-15, Grafana 13.2.1:
check, and a headless row-by-row walk of both dashboards in Chromium at 1600x1000, 0 error badges and 0 “No data” over 168 InfluxDB and 130 Prometheus panels. It has not been repeated for the three panels it did not cover: the two port-event panels and “What each interface is: type, role, bridge and label”. The same walk saw the fast-path share swing 0–100 % between polls on interfaces moving a few packets. Regenerating that day reproduced the committed files byte for byte. - 2026-09-16, Grafana 13.2.1, against the agent on the RB5009 (RouterOS 7.24.2, privileged, the
default triggers),
forward --prom :9124 --influx … --interfaces bridge,for 30 minutes into an isolated InfluxDB 3 Core, and a Prometheus 3.14 scraping the collector every 5 s plus the agent directly for the families the collector could not recompute then:ether1, PPPoE_ DIGI --counters-every 10s
| Store | Window | Panels | Failing | Known-empty tolerated |
|---|---|---|---|---|
| InfluxDB 3 | 30 minutes | 171 | 0 | 10 (the two port-event panels, the opt-in conntrack poll, the two trigger panels, PSI, the four idle block devices) |
| Prometheus | 30 minutes | 133 | 0 | 9 (the two port-event panels, the two conntrack API panels, PSI, the four block devices) |
Scroll sideways to see every column
The 10 and 9 are that deployment over that window. The two port-event panels’ SQL was validated the
same day against a synthetic table in the same InfluxDB 3, because the live store had no kind
column until the first port record classified by kind was written; it has one since 2026-09-19,
holding all eight kinds. The dashboards page of that date also recorded that the agent’s own
exposition carried the kind label only when the agent itself classified, which the one on the
RB5009 then did not; agents serve no exposition now.
- 2026-09-17, container suite: all five dashboards imported into a real Grafana over the stores the suite filled, and every panel asked: 57, 98, 136, 35 and 28 panels returned data and none failed.
- 2026-09-21, Grafana 12.3.0:
importagainst the live InfluxDB 3 store, thencheckover a 15 min window: 167 panels returned rows, the 9 known-empty ones were empty, and the Overview’s detections tile was empty because the store held no detection in that window, which is the one casecheckreports and a browser reads as the healthy “none”. - 2026-09-24, the production Grafana, whose health endpoint reported 13.2.2 later that day: the InfluxDB dashboard imported and captured at 390x844 and 1600x1000 to check the 32 px stat text and the memory time series. With Grafana’s automatic size a single “0” was drawn about 60 px tall at 390x844 and each stat tile took about a third of the screen; with the fixed size the numbers read at a normal size at both. No row-by-row badge count was taken, the Overview’s height was not measured, and the other four dashboards were not captured.
- 2026-09-25, Grafana 12.3.0, 13.2.1 and 13.2.2 over a throwaway InfluxDB 3.11.2 Core:
- The not-available row as committed in 1.3.0 sent five queries and painted five
table … not foundbadges on 12.3.0 and 13.2.1; with its queries shipped hidden it sent none and painted none, and 13.2.2 did the same in a second render. - A detections query against a missing
mikroscope_detectiontable cannot be written to avoid InfluxDB 3’s planning error:WHERE false, aUNION ALLand anEXISTSguard overinformation_schemaall failed at planning. On 13.2.1 the failing layer showed nothing on the dashboard and wrote one error-levelPartial data response errorline to Grafana’s log per load. 12.3.0 never sent that layer’s query, with or without the table (known issue). - A probed
importfrom 1.3.0 removed a missing panel’s queries, and Grafana 12.3.0 gives a panel with no query a default one: the InfluxDB plugin answered it withNo SQL statements were provided in the query string, as a red badge. 13.2.1 sent nothing for the same panel. - Elasticsearch 9.5.3: the Host variable’s lookup, written up to 1.3.0 as a JSON object, went out
from 13.2.1 and 13.2.2 as an empty query answered 400
invalid query, missing metrics and aggregations, and 12.3.0 did not send it; opened without?var-host=, six Overview panels failed withFailed to parse query [host.keyword:]. Written as the JSON string the datasource parses, Host filled from the index and no panel showed a badge on any of the three. - An InfluxDB panel that bins with
$__dateBinunder a sub-second step drew a 0-second bin and came back empty, incheckand in a browser zoomed to a few minutes, on 13.2.1 and 13.2.2.
- The not-available row as committed in 1.3.0 sent five queries and painted five
- 2026-09-25, the production Grafana 13.2.2 against the reference InfluxDB 3 store, with a
build that hides the not-available queries:
dashboards check --window 1h, with the probe and with--no-probe, gave 164 panels with rows, 12 known-empty tolerated and 1 failing, “Detections in the window”, because the store held no detection that hour. Without the probe the five not-available panels readnone … rows=0 frames=0with no error text, where 1.3.0 printed400 … table … not found; the two trigger panels answeredmikroscope_, because that store has no trigger table. The detections annotation’s SQL returned 44 rows over 6 hours. Run again later that day, after the other four stores’ unprobed files got their not-available queries back, it gave the same counts, and the detections layer 32 rows over 6 hours; the InfluxDB file checked was byte-identical to the first run’s.trigger not found - 2026-09-25, container suite (Grafana 13.2.1, Prometheus 3.14.0, PostgreSQL 18.6, graphite-statsd 1.1.10-5, Elasticsearch 9.5.3): the committed files’ not-available queries, against stores that held nothing for them, all answered 200 with no error, and on Elasticsearch every point was 0 whether or not another router in the same index had mapped the field. On PostgreSQL the detections layer’s query answered 200 with no rows over a window before any detection and 10 rows over the run’s own. Both were queries through Grafana’s API, not renders.
- Continuity is derived from every sample’s sequence number, and it found 4 493 missing ticks and one restart in the 2026-09-11/12 capture that the collector’s gap record said nothing about.
- The default range is 3 hours because
now-15mopened on 108 empty panels with the agent stopped, and with it running drew unreadable walls of noise from 100 ms samples (date not recorded).
Alert rules
Section titled “Alert rules”Only the InfluxDB 3 form of the rules has been loaded into Grafana and watched evaluate, on 2026-09-21 against the maintainer’s store; the Prometheus and PostgreSQL forms never have. The rules number 14 for InfluxDB 3, 15 for Prometheus and 10 for PostgreSQL.
- The load of 2026-09-21, into Grafana 13.2.1 against the live InfluxDB 3 store, held the twelve
InfluxDB rules of that date and found two defects, both fixed in 1.1.0.
coalesce()over the unsigned aggregate the InfluxDB sink writes returned HTTP 200 with no frames, which Grafana reads as NoData and these rules as OK: seven of the twelve could never fire,mikroscope-l2-loopamong them, while the same SQL over/api/v3/query_sqlreturned 109 on the live loop signature. And the conntrack rule namedlimit_objs, the SQL sink’s column, where the InfluxDB sink writeslimit; the alert rules page had predicted that one before it was confirmed. After the fix all twelve returned a value, andmikroscope-l2-loopwent to Alerting at 16:58:50Z on the liveown-addresssignature; Grafana’s instance list for it still carriedNormal (NoData)at 16:56:50Z from the run before the fix. - Queries run by hand against the live store on 2026-09-21:
mikroscope-l2-loopread 109 over its own five-minute window, theown-addresssignature onether2having run continuously at about 0.5 records/s since 2026-09-19;mikroscope-port-link-downread 3 over the window holding a realether7link-down at 16:33:49 and 4 over the cluster of four onether4at 09:19 on 2026-09-20. Both thresholds are 0 withgt, so both conditions were met. - The two rules 1.2.0 added were not in that load, and no Grafana has evaluated them since.
Their InfluxDB SQL was backtested on the reference store:
mikroscope-bridge-port-darkreturned 1 on three windows inside the loop and 0 after it, and markedsfp-sfpplus1in 204 ten-minute bins andether2in 376 between 2026-09-19 11:13 and 2026-09-23 22:44 UTC, the two phases of that week’s layer-2 loop, and no other port.mikroscope-wakeup-stormreturned 0.94–0.95 on two healthy windows, never exceeded 1.81 over 551 ten-minute windows with a full day behind them (2026-09-20 to 23), and opened at 35.7 on the wake-up storm of 2026-09-23; it cleared as the new rate became the baseline, after about four hours. Their PromQL was checked for syntax only. The wake-up rule’s PostgreSQL form has no recorded run, and the dark-port rule has no PostgreSQL form: 1.2.1 dropped it, because the SQL sink’s interface-counter table does not hold the columns it reads. - The detections rule no longer fires on
microburstoripc-collapse, which describe how a healthy router carries traffic: from 2026-09-23 10:30 to 2026-09-24 10:30 UTC on the RB5009 they were 64 of 71 detections, and left in they fired the rule in 43 of 288 five-minute windows; without them it fired in 6. - The egress rule: a 1 GbE port to a server lost 3 337 packets in six one-second bursts over 6.5 h, peaking at 436 packets/s (2026-09-19), visible on the egress queue panel and correctly not an alert.
- The thermal rule reads the zone’s own critical trip, 105 °C on the RB5009.
forward --grafana
Section titled “forward --grafana”forward --grafana has run against the maintainer’s Grafana (2026-09-19, InfluxDB 3) and against
the container suite’s for all five stores (2026-09-20, and with 1.5.0’s code on 2026-09-27).
dashboards publish has run only there, on 2026-09-27, over what forward --grafana had just
made. On 2026-09-20 the collector was pointed at a real Grafana and each of the five stores in
turn, allowed to create the datasource, and then dashboards check ran every panel’s query through
Grafana’s API against the datasource the collector had built. All five answered with no datasource
error as the test of that date read it: it looked for flightsql: Unauthenticated, for the TLS
error named below and for err with a space on either side, which is no mark check prints, so an
error worded any other way passed it. The same run found a defect no unit test had: a collector
writing to a live PostgreSQL through --postgres and nothing else published nothing and reported
“there is nothing to publish”, because the store list only knew --sql. Against a plain-HTTP
InfluxDB store, a datasource without insecureGrpc answered every panel
tls: first record does not look like a TLS handshake while the store was fine (measured
2026-09-19).
On 2026-09-27 the test ran with 1.5.0’s code (commit ffb934e, whose Go code is the release’s;
e2e.yml run 36344009521) against Grafana 13.2.1, InfluxDB 3.11.2 Core, Elasticsearch 9.5.3,
PostgreSQL 18.6, Prometheus 3.14.0 and graphite-statsd 1.1.10-5, now failing on any FAIL line of
check that carries an error. For each store in turn, forward --grafana created the datasource
and published the dashboard, check ran every panel against that datasource over a 15-minute
window, and dashboards publish, given the same store and Grafana flags, exited 0, reported the
datasource unchanged and printed the dashboard’s address. No panel that check counts as failing
carried an error; a panel expected to be empty is marked none when it answers no rows or an
error, and the test reads those lines only for the two InfluxDB errors above. The Elasticsearch
datasource was the collector’s, with @timestamp as its time field (1.4.0’s named time, a field
no document carries), and its check exited 0; on InfluxDB, PostgreSQL, Prometheus and Graphite 5,
20, 33 and 4 panels returned no rows in the window, with no error. Not exercised in the suite: the
Authorization header the Elasticsearch datasource sends, because its Elasticsearch runs without
authentication; a datasource that carries a token or a password, which no store there has; and a
run that publishes more than one store, since each run published one.
uninstall --targets
Section titled “uninstall --targets”uninstall --targets removed only what this project wrote in the container suite (2026-09-20,
and --targets data again with 1.5.0’s code on 2026-09-27), and has not been run against the
maintainer’s production store. The suite asserts that a table this project did not write, in the
same database and schema, is neither listed nor removed. It runs the verb with --targets data
alone, against PostgreSQL and InfluxDB 3, with no router and no Grafana. What 1.5.0 changed in the
verb beyond that path has only unit tests: it refuses --grafana-dry-run, lists the Grafana and
store targets before it touches the router, stops on a Grafana it could not read, which 1.4.0
counted as holding nothing to remove, and reads Grafana’s address from GRAFANA_URL when
MIKROSCOPE_GRAFANA_URL is unset.
Test suites
Section titled “Test suites”Four suites need no router: the unit and end-to-end suites run in CI on Linux, macOS and Windows against captured trees and fakes, one runs against nine real stores in docker compose, first in full on 2026-09-16, and one runs the CLI and the agent against a virtual RouterOS, first green on 2026-09-26 and first run on GitHub’s runners the same day. What each proves and how to run it is on Test suites.
- End to end: both binaries against a captured
/proctree of the RB5009 and a fake agent, with one receiver per sink protocol asserting the bytes. It needs no router, no Grafana and no network, and runs in CI on macOS and Windows as well as Linux, where a difference in the executable’s name, in the files the file and SQL sinks write, or in how a child process is stopped would surface. Every sink has been tested against a local receiver since 2026-09-12, on the amd64 development host. The suite used to scrape both the agent and the collector, and now checks the collector alone. - Stores (
make test-e2e-docker): the collector, against the same canned agent and with the API tier off, into Loki 3, the OpenTelemetry Collector, graphite-statsd, Elasticsearch 9, Telegraf 1.39 over HTTP, PostgreSQL 18 (through--sql, and through--postgressince that sink landed on 2026-09-21), InfluxDB 3 and Prometheus, with the file sink as the oracle the others are compared against. Each store is read back through its own API, then all five dashboards are imported into Grafana and every panel’s query runs through Grafana’s API; the suite hasforward --grafanapublish all five datasources and checks their dashboards against them, then, since 2026-09-27, runsdashboards publishwith the same store and Grafana flags, which has to exit 0 and name each store’s datasource and dashboard, and empties the stores again withuninstall --targets data. It needs Docker and no router: the samples are canned, so the run is reproducible anywhere. The first full run, on 2026-09-16, found a panel that named two columns the store has only when the API tier ran: the port-event table nameslabelandrole, which the sink writes onto a kernel-log row only from the API tier’s inventory, and on InfluxDB 3 a column that is not in the table failed the query withSchema error: No field named label, so the panel could not render at all. The sink pages record the collector running into these stores since 2026-09-17. On the development machine on 2026-09-16 the stack came up in 55 to 81 s with the images already pulled, and the whole suite took 75 to 100 s from nothing. The suite joined the pull-request path on 2026-09-18. Until then it ran weekly, on dispatch and before a release, and that is how 1.0.5 reachedmainwith two tests in it still reading the agent’s/metrics, which that release had removed: every pull-request check was green, and the release gate found it. On the release run that failed it took 2 min 37 s from job start to result, containers included; on pull request #70, on 2026-09-27, 4 min 9 s. The OpenTelemetry Collector, which takes JSON, is the only real OTLP receiver the sink has run against. On 2026-09-20 the SQL and PostgreSQL sinks, run side by side into two databases, held 34 tables matching byte for byte, every column of every row hashed per row and summed, plus identicalinformation_for the whole schema.schema. columns - Virtual RouterOS lab (
make test-lab): the CLI’s deploy verbs and the agent against MikroTik’s Cloud Hosted Router 7.24.4 in QEMU, x86_64 and emulated arm64: install by both image routes, upgrade, uninstall,--ephemeraland start-on-boot through a power cut, each scenario ending with the router’s/exportcompared with its start. It tests the installer’s correctness; it is not a board, and no figure under Agent cost or Campaigns comes from it. The lab, its timings and its first runs on GitHub’s runners are under Virtual lab, its scenarios on Test suites. - Sink fixtures, not a router: the SQL, OTLP, Graphite, Elasticsearch and Telegraf sizes on
Other sinks come from two-core test fixtures of 2026-09-12. The SQL
header for all forty-three tables, rendered by the sink’s own
header()on 2026-09-19, is 10 482 B, 13 966 B with the TimescaleDB hypertable statements. Twelve renders of an 8-name map gave 7 orders (2026-09-12), recorded ininternal/.sinks/ telegraf.go - The development host (amd64, kernel 6.12.107):
TestParsePressureReadsTheRunningKernelandTestParseSchedstatReadsTheRunningKernelparse the host’s own/proc/andpressure/ {cpu, memory, io} /proc/schedstatand compare them with a second reading of the same bytes, skipping where the files are absent, which is the RB5009’s case. On that host they pass. Its/proc/pressure/cpucarries afullline, all zeros when read again on 2026-09-24: the kernel’s PSI documentation says CPUfullis undefined system-wide and has been reported since 5.13 as zero, so an older kernel has no such line and the parser accepts either. Its PMU, recorded on 2026-09-21, has six counters,cycles,instructions,cache-references,cache-misses,branch-instructionsandbranch-misses, multiplexed, each running about 84 % of the time it was enabled. A 100.3 ms interval there carried 11 ticks (2026-09-11/12);sample_test.goasserts the busy ratio is capped at 1. - The change filter’s re-arm, against the fixture tree on 2026-09-15: the floored gauge families
and
mikroscope_were present in 6 of 6 scrapes from 5 s after start. The comment onslab_ limit_ objects ProcSource.ReArmininternal/agent/source.gorecords the same check as four scrapes from t+5 s, 0 before the fix and 4 of 4 after; which run each count is from is not recorded. - The script generator, 2026-09-27:
pnpm run rsc:checkrenders the 20 golden cases of the steps spec in JavaScript and matches each script, each step’s commands, the removal order, the values and the command line with the Go output byte for byte; a one-line change to the renderer made it report 72 failures.pnpm test:generatormade 168 checks in headless Chromium (Playwright 1.63) against the built site: the 20 cases set through the form’s own controls, the errors, the token, copy, keyboard-only use, no network request, and the.rscdownloads, which work under the site’s meta CSP. A one-off comparison with the Go code over 304 option sets agreed on 301; the other three are inputs the page refuses and Go accepts (a/030prefix, abusy>=.5threshold and an IPv4-mapped IPv6 address). - The site, 2026-09-25: headless Chromium 153 requested only
favicon.svg, and resolved a manifestidof./to the origin’s root; in headless Chromium and WebKit, in both schemes, after a pick, a stored pick and a reload, the phone’s toggle and with JavaScript off, thetheme-colortag held the header’s colour every time, where the pair split byprefers-color-schemeit replaced held the other theme’s colour after each pick.
Install routes tested
Section titled “Install routes tested”| Route | RB5009UG+S+ | CHR x86_64, lab | CHR arm64, lab |
|---|---|---|---|
Docker Hub pull, --remote-image |
2026-09-17, 7.24.2: /healthz at 2 ms; the 1.0.1 image, sent without its host and pulled through registry-url. 2026-09-24, 7.24.4: the reference with its host, by a hand-written /container/add, never started; no whole install |
2026-09-26: install in 9.6 s, pulled anonymously with the factory /container/config; 2026-09-27: S2, S4 and S5 |
2026-09-26: install in 10.3 and 10.2 s, pulled anonymously; 2026-09-27: S2, S4 and S5 |
GHCR pull, --remote-image |
2026-09-21, 7.24.4: failed, auth error |
2026-09-26: pulled with no registry credential and the host in remote-image=, /healthz 5 s after install began; 2026-09-27: S18 and S5 |
2026-09-27: S18 and S5 |
plan --rsc, /imported |
2026-09-17, 7.24.2: /healthz at 15 ms on its first samples, no CLI in the install |
2026-09-26: /healthz 14 s after the upload began; 2026-09-27: S4, and the fifteen golden scripts the lab can run (S5); the default script pasted at the ] > prompt, and a tar script without its tar, which stopped before any write |
2026-09-26: /healthz 9.1 and 9.5 s after the upload began; the script was the x86_64 one; 2026-09-27: S4 and S5 |
Image tar, --agent-tar |
2026-09-17, 7.24.2: the published arm64 tar, /healthz at 2 ms |
2026-09-26: the release’s amd64 tar, 6 s; 2026-09-27: the branch’s tar, S3 and ten installs in a row; the release’s amd64 tar checked with sha256sum and cosign verify-blob, then installed and upgraded |
2026-09-26: the release’s arm64 tar, checksum as published, 9.6 and 9.2 s; 2026-09-27: the branch’s tar, S3 and three installs in a row |
| Script generator | not run | 2026-09-27: the page’s default script, equal to plan --rsc’s, and a tar script with other settings, each downloaded from the built page and pasted at the ] > prompt; the agent answered each time, and the CLI’s uninstall and the page’s uninstall script each left /export as it was |
not run |
| Manual install, terminal | not run | 2026-09-27: the page’s commands one at a time, registry pull and tar; /healthz answered, status recognised the install from its manifest, and the page’s removal left nothing; again with --expose’s two rules, and with both memberships skipped |
not run |
| Manual install, WebFig | not run | 2026-09-27: both image routes, every form submitted; /healthz 200 each time; the pull removed through WebFig, the tar by uninstall reading the manifest WebFig wrote; /export as it was after each |
not run |
| From a checkout | 2026-09-17, 7.24.2: /healthz at 2 ms; the round trip of 2026-09-12 |
not run | not run |
--expose |
2026-09-11: the two rules verified; 2026-09-23, 7.24.4: a throwaway exposed install, upgraded without a token and removed | 2026-09-26: /healthz 200, /capabilities 401 without the token and 200 with it; 2026-09-27: S8 and S14 |
2026-09-27: S8 and S14 |
uninstall |
2026-09-17: verified by ownership count after each route; /export after all four byte-identical to the one before |
2026-09-26, 1.3.1: 5 of 15 first attempts failed; a second run cleaned up every time. 2026-09-27: every first attempt clean, after every route (F4) | 2026-09-26, 1.3.1: 8 of 8 first attempts cleaned up; with a client on /stream it failed. 2026-09-27: every first attempt clean, after every route (F4) |
Scroll sideways to see every column
- In the lab on 2026-09-27, with the code after 1.3.1 (the changelog’s 1.4.0 section): the
whole lab suite on both architectures, every test passing (Lab timings). After
every route, the CLI’s pull and tar routes, a tar install then an upgrade, an imported
plan --rscscript, every golden script the lab can run and installs made by the released 1.3.1 CLI,uninstallleft/exportequal to the one taken before the install and no mikroscope path on/file. The GHCR pull of 2026-09-26 loggedregistry=ghcr.io, one 3 093 207-byte layer anddownload/extract done2 s later, with/container/configholding noregistry-urland no username. - The install pages’ runs, 2026-09-27, on the x86_64 lab, with the code after 1.3.1 and the
published 1.3.1 agent image and tar:
- The generator’s default script answered 8.8 s after the paste began. Its tar script (name
gen2,172.30.11.0/30, port 9200, both listsnone, rate 20, triggersbusy>=0.9,oom, a token generated in the page,--exposeon 192.168.88.1,--restart-max-count 3,--start-on-boot no) was byte-identical toplan --rscfor the page’s command line, answered on172.30.11.2:9200and deleted its tar after extraction. plan --rsc’s default script, byte-identical topull-dockerhub.rsc, installed a running agent both pasted and uploaded then/imported (Script file loaded and executed successfully). The tar script without the tar stopped withmikroscope: upload mikroscope.tar firstand wrote nothing.- The manual terminal page’s commands, each sent on its own: both routes answered, the tar
route’s wait deleted the tar after extraction, and a manual tar install was also removed by
uninstall --yesalone. A find-and-replace of the placeholders that also rewrote the envlist’s keys made the removal leave the whole envlist behind, silently, which is why the page says to replace values only. - The same page again, once it wrote a manifest per image source and gave
--exposea section: a registry pull exposed on192.168.88.1with a token, a tar install, and a registry pull with both memberships skipped. Each manifest its commands wrote was byte-identical to the CLI’s for the same settings (pull-dockerhub,expose-token,default-tar,lists-none);statusread each from its manifest, the exposed one with its two rules as the install’s;/capabilitiesanswered 401 without the token and 200 with it; and the page’s removal, the two rules first, left the leftover count at 0 and/export, the residue and/fileequal to the baseline. - A WebFig install left an
/exportequal to the one the golden script’s/importleft, on both routes, apart from the veth’s two MAC addresses, which RouterOS draws at random;statuslisted every object of it as the install’s. - The offline page’s checks on the published 1.3.1 files:
sha256sum --ignore-missing -c checksums.txtOK andcosign verify-blob(cosign v3.1.3)Verified OK. Theninstallandupgrade --agent-tarwith the amd64 tar, 7 053 KiB uploaded, and a cleanuninstall. - Installs made by the released 1.3.1 CLI:
statusread their shape from the tags,upgrade --remote-imagewrote a manifest, anduninstall --yesremoved everything, 1.3.1’smikroscope/directory included. upgradewith neither--remote-imagenor--agent-tarlooked for Go to build the agent: it does not take the image from the manifest. A secondinstallon an installed router created nothing (install done: 0 step(s) created).
- The generator’s default script answered 8.8 s after the paste began. Its tar script (name
- The Configure pages’ runs, 2026-09-27, x86_64 lab:
doctoron a stock CHR was MISSING only the interface listLAN, with a fix offering--iface-list none, and with both listsnoneevery check passed. Built-in lists were refused before any connect. The trap check against theadvanced-firewallandadvanced-firewall-rangeprofiles named the rule that drops the agent’s replies, and its fix. An address list that did not exist was created by the install’s entry and was gone afteruninstall. The probe’s first round trip read 1.02 to 1.03 s on three of six installs and 2 ms on the other three, andstatusa second later 1 to 2 ms. - On the RB5009, 2026-09-17, the four routes ran one after another, each under its own name,
veth and
/30so that nothing already on the device was touched, and each was removed before the next./healthzwas read from the collector host. - GHCR, 2026-09-21, a second install beside the running one under its own
--name,--veth,--subnetand--port. The CLI of that date put only the rest of the reference intoremote-image=, soregistry-urlwas pointed athttps://ghcr.iofor the run and put back afterwards;doctorrefused first withMISSING registry-url is https://ghcr.io. With the registry pointed at GHCR the router created the container and failed the pull:download/extract error: fetch manifest failed: getting https://.ghcr.io/ v2/ jmrplens/ mikroscope-agent/ manifests/1.0.9 failed: auth error /container/configholds one username and password for every registry, and the RB5009’s is a Docker Hub login. From the operator host the same day, against the same public package, no credential returned 200 once the anonymous token was fetched and a foreign credential 403 at the token endpoint. The run also found that a--remote-imageinstall identified its container by the registry reference, the same string for every install, so a second install on one router refused as though the first were someone else’s; it is identified by its veth now. - Docker Hub’s limits, read on 2026-09-24: 100 anonymous pulls per 6 hours per IPv4 address or
IPv6
/64, and the anonymous token carries the same 6 h (pull_limit_interval21600), while the registry’sratelimit-limitheader on aHEADof the agent’s manifest read100;w=3600. Which one Docker enforces was not measured. The one-pull-per-install figure comes from Docker’s counting rule and a pull made withcurlfrom the operator host, not from a pull by RouterOS. - MikroTik’s sources on the default registry, read on 2026-09-24: the 7.18 changelog adds
registry-url=https://, the 7.21.2 one says “changed default container registry to docker.io”, no changelog from 7.21.3 to 7.24.4 mentions the registry, and the container documentation still giveslscr.io https://lscr.io/. The lab’s CHR on 7.24.4 had noregistry-urlor username and reportedassumed-registry-url: docker.io.
Agent cost
Section titled “Agent cost”At the install default, 10 Hz, default per-source floors and a 60 s ring, the agent costs 2.69 % of one core and 13.2 MiB RSS, read from its own cgroup at steady state with the ring full:
Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · 300 s windows at steady state (ring full), full source set, the shipped configuration — a 60 s ring, the memory limit derived from it and the default 64M container cap — with the collector forwarding to InfluxDB 3
| rate | floors | CPU of one core | µs/ | RSS | slipped ticks | gaps / |
|---|---|---|---|---|---|---|
| 10 Hz (default) | default | 2.69 % | 2 685 | 13.2 MiB | 0 | 0 / 0 |
Scroll sideways to see every column
That is above the budget of 2 % of one core and inside
the 16 MiB RSS one. The 2026-09-15 campaign was over on both, with a 300 s ring and a
flat memory limit, before 1.0.6 brought the 60 s ring and the derived limit. The image budget
is 8 MiB; the published 1.2.1 image for arm64
is 6.38 MiB (image). The budget is guidance rather than a
contract: cost scales with the device, the source set and the ring size.
Two rules keep a cost figure honest. Wait out the ring (BUFFER_S) first: on the RB5009, at a
14 MiB soft memory limit, a reading taken in the first minute after install came back at 1.47 % of
one core against a 9.38 % steady state. And read cost from the collector’s
/metrics, or from cpu_us and rss in mikroscope_self in the store you write to, rather than
from a large /snapshot, whose ~1.9 MB response the agent must serialise. The procedure is on
Agent cost.
Cost by configuration
Section titled “Cost by configuration”Every row was measured on the same RB5009 at 10 Hz, and each is one window with no spread recorded.
| Configuration | CPU of one core | RSS | Measured | Note |
|---|---|---|---|---|
| Every source read every tick | 2.43 % | not recorded | 2026-09-12 | above the budget |
Ring full, MEM_LIMIT_MB 14 |
9.38 % (9 374 µs/sample) | not recorded | 2026-09-12 | a 300 s ring of lines of about 2.4 kB, the line of that date, holds ~7.3 MB; the Go GC runs without pause |
Ring full, --mem-limit-mb 40, --memory-max 64M |
1.39 % (1 388 µs/sample) | 25.13 MiB | 2026-09-12 | 0 slipped ticks |
PMU counters on, a live forward writing to InfluxDB |
1.72 % | not recorded | 2026-09-12 | 0 slipped ticks; 2 400 samples forwarded, 0 gaps, 0 drops |
| The install default | 2.69 % | 13.2 MiB | 2026-09-18 | the figure above |
Scroll sideways to see every column
Campaigns
Section titled “Campaigns”Each campaign in one entry: the device, the RouterOS version, the date and the conditions, the figures the guides quote from it, and what it found. A figure on a guide links its campaign here.
Cost and rate
Section titled “Cost and rate”Rates at the shipped configuration
Section titled “Rates at the shipped configuration”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · 300 s windows at steady state (ring full), full source set, the shipped configuration — a 60 s ring, the memory limit derived from it and the default 64M container cap — with the collector forwarding to InfluxDB 3
read.under2ms97.7 %read.floorOver5ms1.45 %read.100hz.meanMs1.2 msread.100hz.worstMs17.9 msload.ordinary19 Mbit/srun.10hz.cpu2.69 %run.10hz.rss13.2 MiBrun.10hz.usPerSample2 685 µsrun.10hz.slippedPct0.000 %run.20hz.cpu4.61 %run.20hz.rss15.4 MiBrun.20hz.usPerSample2 303 µsrun.20hz.slippedPct0.000 %run.50hz.cpu9.63 %run.50hz.rss23.3 MiBrun.50hz.usPerSample1 926 µsrun.50hz.slippedPct0.000 %run.100hz.cpu16.83 %run.100hz.rss45.7 MiBrun.100hz.usPerSample1 684 µsrun.100hz.slippedPct0.017 %run.50hz-floor.cpu22.56 %run.50hz-floor.rss25.1 MiBrun.50hz-floor.usPerSample4 511 µsrun.50hz-floor.slippedPct0.027 %run.100hz-floor.cpu42.70 %run.100hz-floor.rss49.5 MiBrun.100hz-floor.usPerSample4 270 µsrun.100hz-floor.slippedPct0.593 %
Six windows of 300 s, at 10, 20, 50 and 100 Hz and at 50 and 100 Hz with FLOOR_HZ, none of which
lost a sample: Rate ceiling has the table. The WAN carried
about 19 Mbit/s on average over the campaign’s 44 minutes, with peaks to 933 Mbit/s.
Rates into three sinks
Section titled “Rates into three sinks”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · 60 s windows at steady state (ring full), full source set, collector forwarding to a file, a Prometheus exposition and InfluxDB 3 at once
Five rate runs at 10, 50 and 100 Hz with the collector writing to a file, a Prometheus exposition
and InfluxDB 3 at once; every sink reported 0 gaps and 0 drops. The 10 Hz install default of that
date cost 2.85 % of one core and 31.3 MiB RSS, with a 300 s ring and a flat 40 MiB memory limit;
50 Hz needed --memory-max 96M and 100 Hz 128M. The WAN carried about 30 Mbit/s, in the evening.
The install-default figure the guides quote is from rates-2026-09-18, not from this campaign.
Memory limit and garbage collector
Section titled “Memory limit and garbage collector”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · 10 Hz, 300 s ring full, the same sources in both runs; only the memory limits differ
gc.tight.cpu9.38 %gc.tight.us9 374 µsgc.roomy.cpu1.39 %gc.roomy.us1 388 µs
The same agent, ring and sources under a 14 MiB soft limit and then with room: Cost by configuration has both rows.
Ring line with every source
Section titled “Ring line with every source”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · one ring line with every source on, privileged, 10 Hz, 4 cores, IRQ_TOP_K=8 — including the PMU and the sampler's own timing, which the 2026-09-12 measurement predates
ring.lineKB3.5 kBring.lineBytes3 230 Bring.lineChargedBytes3 456 Brelay.maxBatch13
3 230 B measured, charged from Go’s 3 456 B size class. The 2 439 B of 2026-09-12 understated the ring by 35 %, and the agent ran at 32.9 MiB of RSS where 24.3 was available. A line with no PMU costs 35 % less. Line sizes at today’s default per-source floors have not been measured.
First ring line measurement
Section titled “First ring line measurement”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · the mean pre-encoded ring line with every source of that date, the slow sources refreshed at 1 Hz, privileged, 10 Hz, 4 cores, IRQ_TOP_K=8
The mean line before the PMU and the sampler’s own timing were in it, rounded to about 2.4 kB on the guides of that date.
Busybox shell loop
Section titled “Busybox shell loop”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 · · a busybox shell loop reading the agent's file set of that date, seven files, at 10 Hz, one fork per iteration, in a container on the router; two 60 s runs
busybox.cpu2.40–2.48 %procread.ms0.77 ms
Two 60 s runs, 2.40 and 2.48 % of one core, from the cgroup’s cpu.stat over /proc/uptime; the
reads themselves took about 0.77 ms per sample. Not re-measured against the agent’s source set of
today: Agent cost.
SSH connect cost
Section titled “SSH connect cost”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.1 · · one ssh connect, for its duration
ssh.connectCpu20–27 %
Measured before the agent existed, on RouterOS 7.24.1. On 7.24.2, on 2026-09-11, the same cost
showed in /tool profile as 17–33 % in one or two snapshots, not a measured window:
SSH cost.
API tier router profile
Section titled “API tier router profile”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · /tool profile duration=60s cpu=total, once with the collector stopped and once with it running --interfaces bridge,; the first five one-second snapshots of each profile dropped, because they hold the SSH connect that asked for it
One 60 s profile per condition, no spread: API tier cost.
Published image size
Section titled “Published image size”· arm64 image tar of the published v1.2.1 release, the file the router loads
image.size6.38 MiB
Measured with ls -l on the v1.2.1 release asset, its checksum verified against the signed
checksums.txt: 6 690 304 B, of which the agent binary is 6 684 832 B. The armv5 and armv7 tars are
7 214 592 B and the amd64 one 7 222 784 B. CI’s agent-size job holds the agent binary, which is
all the image carries, under 8 MiB on arm64, armv7 and amd64.
Image size at 1.0.0
Section titled “Image size at 1.0.0”date not recorded · agent image size of the 1.0.0 release
image.size.v1006.1 MiB
The figure the 1.0.0 release notes give; v1.0.0 was tagged on 2026-09-16. The images between it and 1.2.1 have no measurement of their own.
Kernel and sources
Section titled “Kernel and sources”Kernel accounting files
Section titled “Kernel accounting files”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 · · /proc/pressure and /proc/schedstat absent
The irq column of /proc/stat is always 0 on this kernel, so hard-IRQ time is inside system.
USER_HZ is 100, so a tick is 10 ms and the busy-time steps on the guides are arithmetic from it.
The files that are the router’s inside the container were established the same day
(Container view).
Privileged discovery container
Section titled “Privileged discovery container”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 · · a privileged discovery container reading the host's /proc once, before the agent read slabinfo
conntrack.slabDiscovery6 582
The round that found no tracing path (Kernel and PMU), and whose slabinfo is
the one in testdata/proc/rb5009: nf_conntrack 6 582 active of 8 075.
Network namespace under privileged
Section titled “Network namespace under privileged”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · privileged=yes does not change the network namespace
Inside the container /proc/net/dev counted the veth, 4 packets while the router forwarded
millions, with and without privileged=yes.
Overnight source cadences
Section titled “Overnight source cadences”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 · · a 10.5 h capture at 50 Hz over one idle night, clock pinned
thermal.quantumC0.42 °C
The capture that set the 6 Hz floor /proc/slabinfo is slowed to, the rate its fastest cache,
nf_conntrack, changed. The memory levels moved about 24 times a second. It is one board, one idle
night, its clock pinned at the maintainer’s setting: “cpufreq never changed” means it did not
change that night.
`cpu-load` averaging window
Section titled “`cpu-load` averaging window”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 · · /system/resource polled at 1 Hz against the agent's per-core busy ratio at 10 Hz, two separate hours
cpuLoad.window11.0 scpuLoad.delay10.6 scpuLoad.r10.9825cpuLoad.samples13 499cpuLoad.window21.1 scpuLoad.delay20.1 scpuLoad.r20.9734cpuLoad.samples23 594
The best fit in each hour: a trailing mean of 1.0 s that reached the
API 0.6 s late (r = 0.9825
over 3 499 API samples), and a mean
of 1.1 s that reached it 0.1 s late (r = 0.9734
over 3 594). /system/resource reports cpu-load as an integer percent, and
the API series was correlated against the agent’s per-core busy ratio, the same /proc/stat
jiffies, for a range of window lengths and delays. Widening the window only made the fit worse:
1.5 s gave 0.955, 2 s 0.919, 5 s 0.822, 8 s 0.791. A sixty-second average is ruled out twice: its
correlation is 0.238, and at the sharpest load step of the day the kernel went from 5 % to 27 % in
one second and cpu-load from 5 to 26 in that same second, then from 22 % to 6 % on the way down
just as fast. MikroTik’s
/system/resource documentation
defines cpu-load as the percentage of used CPU resources, all CPUs combined, and names no window
(read on 2026-09-24).
Recording and transport
Section titled “Recording and transport”Recording of a script loop
Section titled “Recording of a script loop”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · a 60 s record at 10 Hz, 600 samples over 59.9 s, a RouterOS script loop started over ssh and the router log added as markers
Exactly 600 samples, 0 gaps and a clock skew of −7 ms. A scripted RouterOS loop
(:for … 400 000) showed as one core’s worth of load at 100 % from t = 21.0 to 25.8 s, with the
sample at 25.8 s reading about 60 %, onset and offset resolved to 100 ms. The router’s log markers
explained a 2 s plateau at 15–17 s that nobody had caused: the ensure-ipv6-nd-prefix scheduler.
Log times came over the API as full dates (verified).
Recording at rest
Section titled “Recording at rest”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · a 70 s record at 10 Hz, 700 samples over 69.9 s, the router otherwise at rest, three notes typed into record's terminal
The recording the chart on Idle baseline is drawn from.
`/stream` connection lifetime
Section titled “`/stream` connection lifetime”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.4 · · three GET /stream connections against the installed agent at 10 Hz, opened from the operator host on the LAN and held until the server closed them
stream.lifetime30.01–30.06 sstream.lines304–306
Agent 1.0.9 at 10 Hz. Each connection ended at the server’s 30 s write timeout, whatever it still
had to send, so a /stream consumer reconnects with the last seq it saw.
Relay fetch limit
Section titled “Relay fetch limit”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 · · /tool fetch output=user called over the binary API; each call took either about 3 ms or about 1 s, and about half took 1 s
relay.replyMaxBytes64 512 B
/tool fetch output=user truncated the body silently at 64 512 B for 64 K, 256 K, 1 M and 4 M
bodies.
Line-protocol sample size
Section titled “Line-protocol sample size”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 · · one kernel sample at 10 Hz rendered to InfluxDB line protocol, with the agent's source set of that date
About 1.2 KiB a sample, the figure the InfluxDB and Telegraf sinks size their byte budgets from. Not measured above 10 Hz, and not re-measured against today’s source set.
Network and faults
Section titled “Network and faults”Case-study readings
Section titled “Case-study readings”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 · · agent at 10 Hz in an ephemeral privileged container
loop.eventsBefore1.49 /sloop.eventsAfter0.03 /sloop.helloGap2.00–2.01 sconntrack.slab6 287
The readings the case studies are built on, from an aarch64 kernel
on a board with 1 GB of RAM; where a fault was provoked, the case study says how. Three case
studies also carry a chart from the reference InfluxDB store, whose history begins on 2026-09-19:
the loop’s return, the conntrack cross-check and the port losing frames. The port losing frames is
a separate campaign a week later, read from the API tier’s port counters rather than from the
agent. time_squeeze was nonzero even at idle. While the layer-2 reflection was live the kernel log
ran at 1.49 /s, and a “blocking state” then “learning state” pair arrived
inside one 100 ms tick at the same level, so their order within the tick is the signal. The
conntrack table held 6 287 entries, 0.65 % of the 966 656 ceiling the
kernel reported on 2026-09-14.
Softnet squeeze distribution
Section titled “Softnet squeeze distribution”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · 3 738 704 per-CPU samples over 24 h at 10 Hz, softnet time_squeeze per sample, 0 drops in the whole window
squeeze.backgroundOne11.2 %squeeze.backgroundTwo1.2 %squeeze.exactlyThree0.21 %
The 24 h ended at 07:00 UTC on 2026-09-16. time_squeeze was 0 in 87.3 % of per-CPU samples, and
a trailing window of that distribution has a 90th percentile of 1, so “above p90” is met by any 2.
On 2026-09-15 a trigger condition that fires on any squeeze fired 92 times in twenty minutes
(RouterOS 7.24.2), which is why squeeze is not a default trigger.
Microburst rule replay
Section titled “Microburst rule replay”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · the rule replayed over 6 h of stored samples, 863 944 rows, 4 CPUs
microburst.firesPerHourAtTwo77.7 /hmicroburst.firesPerHourAtThree0.5 /h
At a floor of 3 the rule still flagged 88 samples for the burst counter and for derived.burst.
Conntrack count over the API
Section titled “Conntrack count over the API”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 · · the connection table counted over the RouterOS binary API, /ip/, ten calls
conntrack.api6 212api.conntrackMs1.3 ms
The count was 6 212, the day before the agent’s slab read 6 287: the same order of magnitude, which says nothing about whether the two track. The ten calls took a median of 1.3 ms, the fastest 1.1 ms and the first 71 ms, cold; the time of day is not recorded.
Port and fast-path counters
Section titled “Port and fast-path counters”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · · RouterOS port and interface counters read over the API, cumulative since boot or since each port's last counter reset
ether1, 2.5 GbE to a NAS, had received 255.8 GB on the wire (rx-bytes) and handed 29.7 GB of it
to the CPU (driver-rx-byte) since the port’s last counter reset; the switch chip forwarded the rest
in hardware. Every switch port read a fast-path share of about 100 %, fp-rx-byte equal to
driver-rx-byte within a few kB. bridge had fast-pathed 211.9 GB of the 663.0 GB it took to the
CPU since boot (32 %), and PPPoE_DIGI 99.97 %. fp-tx-byte stood at 0 on every interface after
hundreds of GB transmitted (verified).
Port overflow before and after
Section titled “Port overflow before and after”Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.4 · · ether1, the 2.5 GbE port to the NAS, over 10 s counter intervals; the before figures are the three hours preceding the fix and the after figures the 39 minutes following it, at the same load
overflow.shareBefore0.502 %overflow.perHourBefore5 399overflow.occupancy0.36 %overflow.burstRate8.96 Mbit/s
The case study is Port losing frames.
Measured with no campaign of its own: on 2026-09-15 a burst peaked at 736 packets in one 20 ms sample, a 50 Hz sample, against a median of 28.
Verified RouterOS behaviour
Section titled “Verified RouterOS behaviour”Yes-or-no facts about RouterOS the guides rely on, each checked once, with where and when.
Read user sees envlists
Section titled “Read user sees envlists”Over the binary API, /container/print returned every property of every container, cmd and envlist included, to a user with only read,api, the user mikroscope; only the property names were printed, and reading the values in /container/envs as that user was not checked (RB5009UG+S+, RouterOS 7.24.2, ).
Test policy for fetch and profile
Section titled “Test policy for fetch and profile”/tool fetch and /tool profile both require the test policy (RB5009UG+S+, RouterOS 7.24.2, ).
Quoted values in find
Section titled “Quoted values in find”In a RouterOS find, address and port attributes match only when their values are quoted, and a bare word is read as a variable name: over the same 15 dstnat rules (RouterOS 7.24.4, 2026-09-21), protocol=tcp found 0 and protocol="tcp" found 10 (RB5009UG+S+, RouterOS 7.24.2, ).
Expose rules reach the agent
Section titled “Expose rules reach the agent”A dst-nat rule plus a forward accept rule reach the agent from the LAN, and both are removed by their tag (RB5009UG+S+, RouterOS 7.24.2, ).
LAN reaches the veth
Section titled “LAN reaches the veth”The LAN reaches the veth once the veth joins the interface list LAN and the /30 joins the address list LANs (RB5009UG+S+, RouterOS 7.24.2, ).
Container removal returns early
Section titled “Container removal returns early”/container/remove returns before the container is gone, and a /file/remove issued meanwhile does nothing, silently (RB5009UG+S+, RouterOS 7.24.2, ).
Removing the tar re-extracts
Section titled “Removing the tar re-extracts”With ignore-remote-image-change=no, removing the image tar makes RouterOS stop and remove the container and extract it again minutes later (RB5009UG+S+, RouterOS 7.24.2, ).
Round trip leaves /export unchanged
Section titled “Round trip leaves /export unchanged”doctor, install, status, upgrade and uninstall, run in that order, left the router's /export byte-identical, its # header lines aside, compared by hash in memory and never written to disk (RB5009UG+S+, RouterOS 7.24.2, ).
tmpfs install writes no NAND
Section titled “tmpfs install writes no NAND”A tmpfs disk holds the image tar and the container root with no NAND writes: write-sect-since-reboot stayed at 58 279 across install, run and removal (RB5009UG+S+, RouterOS 7.24.2, ).
Privileged keeps network and PID
Section titled “Privileged keeps network and PID”privileged=yes drops the container's user namespace but not its network or PID namespace (RB5009UG+S+, RouterOS 7.24.2, ).
Registry host in the reference
Section titled “Registry host in the reference”A registry host inside remote-image= overrides /container/, and docker.io is pulled as registry-1.docker.io: with registry-url=https://, a reference on registry.invalid was logged as registry=registry. and failed with resolving error, while the agent's 1.2.2 image named as docker.io/… and as registry-1.docker.io/… was logged as registry=registry-1. and ended in download/extract done, its arm64 layer 2 799 648 bytes; /container/config read the same afterwards. The containers were created in a temporary veth, never started and removed again (RB5009UG+S+, RouterOS 7.24.4, ).
Docker Hub pull without login
Section titled “Docker Hub pull without login”With the /container/config username and password cleared, RouterOS pulled the agent's 1.2.2 image from Docker Hub anonymously, both as remote-image=registry-1. and without a host through registry-url=https://, each ending in download/extract done 5 s after the add; the configuration was then restored and verified identical (RB5009UG+S+, RouterOS 7.24.4, ).
Log times over the API
Section titled “Log times over the API”Over the API, /log/print returns an entry's time as a full date and time with no zone, 2026-09-12 02:21:24, and accepts that format in a ?>time= query (RB5009UG+S+, RouterOS 7.24.2, ).
Loss keys from monitor-traffic
Section titled “Loss keys from monitor-traffic”/interface/ returns rx-drops, tx-drops and tx-queue-drops per second and no error keys at all (RB5009UG+S+, RouterOS 7.24.2, ).
rx-overflow only in port counters
Section titled “rx-overflow only in port counters”ether1 had counted 652 364 rx-overflow events, and growing, in its port counters while monitor-traffic returned no error key for that port: that count reaches a consumer only through the port counters (RB5009UG+S+, RouterOS 7.24.2, ).
Switch port and bridge planes
Section titled “Switch port and bridge planes”An ether port in a bridge counts its wire, frames the switch chip forwarded in hardware included, and the bridge counts its CPU side, so neither is a subset of the other: since the port's last counter reset ether1 had received 255.8 GB on the wire (rx-bytes) and handed 29.7 GB of it to the CPU (driver-rx-byte) (RB5009UG+S+, RouterOS 7.24.2, ).
fp-tx-byte stays at zero
Section titled “fp-tx-byte stays at zero”fp-tx-byte read 0 on every interface after hundreds of GB transmitted, while fp-rx-byte counted; why RouterOS leaves it at 0 is not established (RB5009UG+S+, RouterOS 7.24.2, ).
Host paths in a privileged container
Section titled “Host paths in a privileged container”A privileged container given the host's /proc, /sys and / as bind mounts read zero PIDs in the host's /proc and found no class/net in its /sys, so the namespaces held; the host's / did mount, and it exposed the RouterOS flash filesystem, configuration and files, secrets included. Run with the maintainer's consent (RB5009UG+S+, RouterOS 7.24.2, ).
Credential sent when the host matches
Section titled “Credential sent when the host matches”RouterOS presents the /container/config username and password for a reference whose host is registry-url as written: with a deliberately wrong credential, the 1.3.1 agent image named as registry-1. failed with auth error under registry-url=registry-1., and was pulled anonymously under https://, the value MikroTik's examples use, and under the same with a trailing slash. A router given a credential did not fall back to an anonymous pull (CHR x86_64, RouterOS 7.24.4, ).
No WebFig field for ignore-remote-image-change
Section titled “No WebFig field for ignore-remote-image-change”WebFig's New Container and container edit forms show no field for ignore-remote-image-change, with File or Remote Image set or not; the terminal sets it (CHR x86_64, RouterOS 7.24.4, ).
A container set resets restart-policy
Section titled “A container set resets restart-policy”A /container/set that leaves restart-policy out (of ignore-remote-image-change, of comment, of logging) put it from on-failure back to always and changed nothing else /container/print detail shows; a set that names it kept it, and so did Apply or OK in WebFig's edit form. The CLI never runs /container/set (CHR x86_64, RouterOS 7.24.4, ).
WebFig shows env values
Section titled “WebFig shows env values”WebFig shows a /container/envs entry's value in clear, in the Envs list and in its form, for the key TOKEN too (CHR x86_64, RouterOS 7.24.4, ).
No Device Mode page in WebFig
Section titled “No Device Mode page in WebFig”WebFig's System menu has no Device Mode page, so device mode is set from the terminal (CHR x86_64, RouterOS 7.24.4, ).
fetch as-value returns the body
Section titled “fetch as-value returns the body”:put ([/tool/fetch url="http://172.30.10.2:9123/, run on the router, printed the agent's /healthz JSON, in an ssh session, in WebFig's terminal and at the end of a pasted script (CHR x86_64, RouterOS 7.24.4, ).
Licence question takes the first lines
Section titled “Licence question takes the first lines”The first interactive login after a reset asks Do you want to see the software license? [Y/n]: before the ] > prompt. A script pasted into that question lost its first lines: one or two header comment lines, then a fragment run as a command (bad command name . or syntax error); the install script's { … } block still ran and installed, and the uninstall script still removed everything (CHR x86_64, RouterOS 7.24.4, ).
Not tested
Section titled “Not tested”What no run, measurement or check here covers. A guide that states one of these as expected behaviour links this list.
Hardware and versions
Section titled “Hardware and versions”- A second board of any kind. The hEX S (2025), 32-bit RouterOS on an ARM64 chip, which is what the
agent’s
linux/arm/v5build is for, has not arrived, so the 32-bit counter-wrap path and neither 32-bit ARM image have run on RouterOS. A board report from another board is what changes this. - x86 RouterOS on hardware: the amd64 image has run only in QEMU, on the lab’s CHR x86_64 and on RouterOS x86 installed from MikroTik’s ISO.
- Any RouterOS before 7.24 beyond
doctor’s version check, which S17 ran on a 7.23.7 CHR, and any of the 7.24.2 figures re-measured on 7.24.4. - Anything that needs a reboot of the RB5009, which waits for a maintenance window. So a persistent
install surviving a reboot with
start-on-boot=yesis untested on hardware. On the lab’s CHR it answered again within 90 s of a power cut (S7, on both architectures). - Which files are namespaced, on another RouterOS version or another board: every row of Container view was read on one RB5009 on 7.24.2.
- A board with hwmon sensors, or a later RouterOS that builds containers differently: the sensor set, the 38 capabilities and the uid mapping were read on one RB5009 on 7.24.2.
- The port mapping on another board: the shift by one, the position of the SFP+ cage and the reuse
of
port 7were observed on one RB5009UG+S+ on 7.24.2, and the agent does not apply them to any other board. - The PSI and
schedstatpaths on a router: the RB5009’s kernel has neither. The PMU counter set is known on the RB5009’s Cortex-A72 and the amd64 development host only. - Every source cadence on another device or workload: the cadences are the RB5009’s, on 7.24.2.
Re-measure with
FLOOR_HZequal to the sampler rate. - Which loss keys and counters another RouterOS version or board returns.
- What
--goarm 7saves against the ARMv5 build: no ARM hardware has run either. - Envlist entries on a RouterOS before 7.24: mikroscope writes
key=, becausename=failed on the RB5009 on 7.24.2, and how an earlier 7.x takes either was not tried. - Traffic heavier than the RB5009’s ordinary load: about 19 Mbit/s on the WAN during the 2026-09-18 campaign. No rate is claimed for any other board.
Install and upgrade
Section titled “Install and upgrade”- On the RB5009, a pull from GHCR with no registry credential set, and any GHCR pull with the host
inside
remote-image=: its one GHCR run took the host fromregistry-url. The lab’s CHR did both on 7.24.4 (Install routes tested). - Which credential RouterOS presents to a host named only in
remote-image=whenregistry-urlnames another. The one run with a host that differed fromregistry-url’s, on 2026-09-24, namedregistry.invalid, which never resolved. With the same host, the lab showed RouterOS presents it only whenregistry-urlis written as the bare host (verified). - On the RB5009: a whole
installorupgradethat sends the full reference, aregistry-urlat its factory default, and the full reference on any RouterOS but 7.24.4. The lab’s CHR ran both with 1.3.1 on 7.24.4. - A pull through a mirror or pull-through cache named in the reference.
install,statusanduninstallof 1.4.0 and later on any hardware, and the guarded script: they have run only on the lab’s CHRs. On the RB5009, 1.6.0’supgraderead the install manifest and the architecture (--arch auto) from the router, and itsdoctorran every check.--ssh-optionandMIKROSCOPE_SSH_OPTIONSagainst any router: the lab’s CLI connects through the lab’s ownssh_config, and only unit tests pass the option.--extract-timeoutrunning out, which keeps the tar, on any router.- Doctor’s registry host comparison against a real router: it has run only against fake router answers in the tests.
- Whether a
read,apiuser can read the envlist values in/container/envs. - The byte-identical round trip on hardware: on any board but the RB5009, on the RB5009 with any
RouterOS but 7.24.2, with
installandupgraderun without--ephemeral, or with the current container settings (privileged=yes,memory-max=64M, the envlist entriesMEM_LIMIT_MB,CAPTURE_MB,TRIGGERSandFLOOR_HZ): on the RB5009 it ran withmemory-max=32M.install,statusanduninstallhave run there on 7.24.4 since. On the lab’s CHR 7.24.4,make roundtripruns it with--ephemeral(2026-09-26); S2 and S3 runinstall,upgradeanduninstallwithout it, withprivileged=yes,memory-max=64M,MEM_LIMIT_MBandCAPTURE_MB, and S5 imports a script that also writesTRIGGERSandFLOOR_HZ; each ends with/exportequal to its start, RouterOS’skeymat-providerline aside (both architectures, 2026-09-27). - Firewalls other than the RB5009’s, where on 2026-09-11 the two list memberships were enough, and
the lab’s profiles. A firewall with other drop rules in
raw,inputorforwardmay drop the container’s traffic elsewhere, andinstalladds nothing for that beyond the two memberships. Whether a router’s factory default configuration carries the two raw rules was not checked. --exposefrom outside the LAN: only a LAN host reaching the router’s LAN address was tested. Whether anything outside the LAN reaches that address depends on the rest of the firewall, and no such path was tried.- Winbox: the WebFig install page says Winbox has the same menus and fields, and no Winbox ran. The
--exposerules through WebFig’s NAT and Filter Rules forms, a whole script pasted into WebFig’s or Winbox’s terminal, and WebFig on any RouterOS but 7.24.4. - The script generator’s scripts, the manual terminal install and the WebFig install on the arm64 lab or on the RB5009.
Recording and capture
Section titled “Recording and capture”- A recording above 10 Hz: the lossless 20, 50 and 100 Hz runs were the collector’s, which uses the same batch sizing. A recording through the relay, at any rate.
- The relay’s throughput. From
/tool fetchround trips of about 1 s for half the calls and about 3 ms for the rest (measured on the RB5009 on 7.24.2), arithmetic gives on the order of 26 samples a second, less when the slow calls cluster. The relay cap and the start-up warning are read from the code (2026-09-15), not measured against a device. The relay’s cap and its 1 s share on any RouterOS but 7.24.2. - The cost of a trigger fire on the device; 3 µs for a 10 s window at 10 Hz is the design’s estimate. How the agent behaves under a sustained trigger storm. Any capture at 50 or 100 Hz. The capture sizes on Triggered capture are arithmetic from the line size, not sizes of captures taken on the RB5009.
Collector and sinks
Section titled “Collector and sinks”- Sinks fed from the RB5009 into a running backend: only file, Prometheus and InfluxDB 3. Loki, an OTLP receiver, carbon, Elasticsearch, OpenSearch, Telegraf, PostgreSQL and TimescaleDB have had no router samples in a recorded run. OpenSearch, TimescaleDB’s hypertables and standard output have never run against a real store; only the byte-contract tests cover them.
- An Elasticsearch or OpenSearch that requires authentication. The container suite’s Elasticsearch
runs with security off, so the
Authorizationheader built fromMIKROSCOPE_ELASTIC_AUTH, basic auth foruser:passwordandApiKeyotherwise, has no recorded run against one: neither from the sink nor from the Elasticsearch datasource that--grafanaanddashboards publishbuild, which sends the same header from 1.5.0. Unit tests pin both forms, and the end-to-end suite the sink’sApiKeyagainst a fake receiver. Nor has that datasource run against an OpenSearch, with authentication or without. - A live TimescaleDB: the
create_hypertablecalls follow TimescaleDB 2.x’s documented signature and are not verified. - The SQL sink’s size on a router. The fixture figures (1 375 B of SQL for a kernel event against
716 B of line protocol, 1 138 B for an API event against 608 B, about 14 KiB/s at 10 Hz plus the
1 Hz API tier after a 5.6 KiB header, 2 749 B with the privileged sources) cover eighteen of the
forty-three tables the sink writes: eleven were added after the fixture, and
load,stat,buddy,mtd,api_ifcounter,api_ifinfo,trigger,derived,derived_iface,detectionand the fourdevicetables were not exercised. They predate theportandkindcolumns ofmikroscope_event. --sql - | psqlwith apsqlthat falls behind a 10 Hz agent, which blocks the pull loop.- The standard-output budget for
json, whose lines are larger than line protocol’s. - An OTLP render of a four-core sample with a real interrupt top-K, a four-core sample in the Elasticsearch format, and the Graphite budget above 10 Hz.
- A write to InfluxDB 2’s
/api/v2/write. The line-protocol size above 10 Hz or with today’s source set. Whether turning off the HTTP client’s own resend avoids the duplicates of 2026-09-13. - Loki delivery during a kernel-log storm.
- Any Prometheus scrape interval but 5 s, and any Prometheus but 3.14.
- The collector’s constant 10 Hz against a faster agent: the burst baseline and the top-K pruning are read from the code; the 1/600 EWMA weight and the 36 000-sample eviction were never watched happening, since no run made an interrupt line fall out of the top-K and stay out.
- The zero deltas written beside an absent fast-path share (read
from
internal/), and a ceiling or cadence change without a hash change reaching the sinks at the next five-minute repeat (read fromderive/ derive.go capsHash). - The per-packet PMU cost compared across two router configurations. Whether the burst baseline’s ten seconds suit another board or traffic mix: it was tuned against one RB5009 on one day. A tx fast-path share from live counters.
- Detections provoked with the rules running: no OOM kill in the container, reboot, link flap,
conntrack flush or storm, thermal excursion or IPC collapse. Whether
rebootfired on the 2026-09-19 reboot. A provoked flap withlink-flaprunning; the flaps of 2026-09-15 were not. - Whether RouterOS computes
cpu-load’s one-second window on a wall clock or on jiffies. Whether the API’s conntrack count and the slab count track each other.
Other tools
Section titled “Other tools”- None of the six tools on Compared with alternatives was run for
that page. Every cell in their rows is what their own documentation or source code says, read on
2026-09-24: mktxp at commit
1e4412a(dated 2026-09-21), mikrotik-exporter at428dbfc(dated 2026-07-03), and the MIB file MikroTik publishes for RouterOS 7.24.4. What SNMP polling, The Dude, Graphing, the Profiler, mktxp or mikrotik-exporter costs a router, how often each can usefully poll one, and what RouterOS puts inhrProcessorLoadwere not measured.
Dashboards and alerts
Section titled “Dashboards and alerts”- Any Grafana but 12.3.0, 12.3.2, 13.2.1 and 13.2.2. On 12.3.0 the only renders are those of 2026-09-25; the rest of the dashboard was checked there by query only.
dashboards publishagainst any Grafana but 13.2.1, and over anything but whatforward --grafanahad just made there: in the container suite on 2026-09-27 it ran once per store, after the collector, with no credential in any datasource, and found the folder and each datasource as the collector had left them, reportedunchanged. Withdashboards publish, a run of more than one store, a store that fails while the rest carry on, a datasource that carries a token or a password (written again at every run and reportedupdated), and--grafana-dry-runhave run only in the unit tests, against fake Grafanas.- The height of the thirteen-panel Overview in a browser; the committed JSON makes it 31 grid units.
- Byte-for-byte regeneration of the dashboard files added after 2026-09-15.
- The Elasticsearch and Graphite annotations and detection layers in Grafana: only their generated JSON is checked, by unit tests.
- A render pass over the two port-event panels and the interface inventory table.
- A rule resolving: the loop on the RB5009 ended on 2026-09-23, but
mikroscope-l2-loopreturning to Normal has not been checked in Grafana’s state history. Neithermikroscope-l2-loopnormikroscope-port-link-downhas been watched going pending, firing and resolving. - The two rules 1.2.0 added inside any Grafana, their PromQL beyond syntax, and the wake-up rule’s PostgreSQL SQL.
- The Prometheus and PostgreSQL alert files in any Grafana, and any PostgreSQL alert query against a real PostgreSQL: a unit test reads the schema, not a database.
- Notification delivery: no contact point was configured, so what was watched is each rule’s state.
- A store missing a table or column a rule reads: the query is expected to fail and
execErrState: Errorto apply instead of OK. What Grafana then does is its Error and No Data behaviour, not tested here.
Known issues
Section titled “Known issues”Checked against the code on 2026-09-24, and the end-of-run summary’s stream again at 1.3.1 on 2026-09-26:
forward’s end-of-run summary goes to stdout, the same stream the--stdoutsink writes records to, soforward --stdout=lp | telegrafends every run with lines the consumer cannot parse.- A wrong
--tokendoes not fail the run. Each pull is logged as401 Unauthorized, and the run then ends at its--fordeadline with exit status 0 andforwarded 0 kernel samples. The health checkforwardmakes before it starts pulling reads/healthz, which needs no token, so it passes. The message is right; the exit status tells a unit file nothing. - Duplicate sequence numbers in the 2026-09-13 overnight run. The 50 Hz run wrote repeated
seqvalues into InfluxDB; the dashboards name those rows “duplicate sample (seq repeated)”. The leading explanation, which is unverified, is Go’s HTTP client replaying a POST on a dead pooled connection after the server had committed it. The InfluxDB, Loki, OTLP, Elasticsearch and Telegraf sinks refuse that replay, so a dead connection is an error the sink retries and counts. No run since that night shows whether the duplicates are gone. - A negated tag inside an OR returns nothing, silently, on InfluxDB 3. On 2026-09-23, on the
reference store (InfluxDB 3 Enterprise since 2026-09-19; the version was not recorded with the
query),
NOT (port='ether2' AND kind IN (…))overmikroscope_kmsgreturned 0 rows where 14 matched, and so did its De Morgan form and the explicit OR. It returned 14 onceport='ether2'was also filtered outside the OR. The shipped panels and rules do not use the pattern. - Grafana 12.3.0 never sends the InfluxDB detections layer’s query, with or without the table: its toggle kept a loading indicator, switching it off and on and refreshing sent nothing, and no marker was drawn, on the full dashboard and on a one-panel copy (2026-09-25). Grafana 13.2.1 sent it at once over the same store. Why, and whether 12.3.2 does the same, has not been looked into.
Found in the virtual lab
Section titled “Found in the virtual lab”What the published 1.3.1 CLI and agent image met on the lab’s CHR, RouterOS 7.24.4, on 2026-09-26. Each item ends with what the code after 1.3.1 does, checked in the lab on 2026-09-27; 1.4.0 ships that code, and the changelog’s 1.4.0 section lists each change.
doctorwith its defaults fails three checks on a CHR.--archdefaults to arm64, which fails on x86_64 only; the interface listLANdoes not exist; and the address listLANs“has entries” fails on an empty or missing list, although its own fix says an empty list is fine if no such rule exists. A router without that rule passes only by giving the list an entry, or with--no-doctor, which skips every other check. After 1.3.1:--archdefaults toautoand reads the router’s architecture, an empty address list is no longer a failure, and a missing interface list, still MISSING with the defaults, has a fix that offers--iface-list none(S1).- A built-in interface list such as
static,allordynamicpasses doctor’s existence check, and RouterOS refuses to add a member to it. After 1.3.1:--iface-listrefusesall,dynamicandstatic; RouterOS answerscannot add to builtin list. --ephemeralneeds a tmpfs disk, and a CHR lists no disk. Doctor’s fix,/disk/add type=tmpfs tmpfs-max-size=64M slot=tmpfs, worked, and the install went totmpfs/withmikroscope/ mikroscope start-on-boot=no. After 1.3.1: the same, anddoctoralso warns when start-on-boot is forced on a root on a tmpfs disk, which a reboot empties.uninstallraces the container’s stop. It stops the container, waits a fixed 4 s and removes it, and the agent took 0 s to stop six times, 4 s four times and 5 s once (RouterOS log, 1 s resolution): 5 of 15 first attempts on x86_64 failed withfailure: cannot remove running. With a client on/streamit failed every time on both architectures, because the agent’s HTTP shutdown waits up to 5 s for it. The steps after it still ran, and a seconduninstallcleaned up every time. After 1.3.1: the removal waits up to 30 s while the container isrunningorstopping, and every first attempt was clean (Lab timings).- An empty
mikroscopedirectory stays in/fileafter everyuninstall: the parent of the container’s root-dir, which the ownership count does not include. After 1.3.1: every install route writes an install manifest, anduninstallremoves what it lists, then the manifest and themikroscope/directory when nothing else is in it; no mikroscope path was left on/fileafter any route (F4). - The probe after
installmisreads a running container on 7.24.4. It asks/container/find … status="running", andstatusis not a property of/containerthere (/container/find status=runninganswersbad parameter status), while the flagrunningreads 1. So when the probe failed, the CLI said “the container is not running on the router” with the agent running and about to answer. Seen twice on arm64. After 1.3.1: the probe reads therunningflag. - The plan listing named a tar with
--remote-image:image=mikroscope.tar, a fileinstallnever uploads, and it printed the upload or the pull as a step of its own under the container step’s number, so two steps shared a number. After 1.3.1:image=names the reference the router pulls, and the pull or upload is a line of the container step. - The agent token travelled on ssh’s command line, where the host’s process table showed it
while the command ran: the 1.3.1 code showed it on two command lines. After 1.3.1 a command
that carries it goes to ssh on standard input; reading the process table every 2 ms through an
install and an upgrade with
--exposefound it on none, on both architectures. - RouterOS’s words were lost when its ssh exited 1: the error read
ssh "<script>": exit status 1, and RouterOS’s message, on the lines after it, was dropped by uninstall’s skip line. RouterOS’s ssh exited 1 or 0 on the same failure, about half and half. After 1.3.1 the error starts with what RouterOS printed.
Uncollected sources
Section titled “Uncollected sources”Two readable files on the RB5009, read there on 2026-09-15, are not collected: /proc/cmdline,
which carries board=5009 ver=7.24.1, a second device identity without the API, and the hardware
watchdog at /sys/.