Skip to content

Tested on

Where each part of mikroscope was run, measured or checked: the device, the RouterOS version, the date and the conditions, what each run found, and what has not been tested. The guides state what the tool does and link here for the proof. The feature verdicts were checked against the code at 1.2.0 and 1.2.2 on 2026-09-24, the sections that came from the guides against the code of 1.3.1 on 2026-09-26, what this page says of the code after 1.3.1, which 1.4.0 ships, against that code on 2026-09-27, and its sections on publishing to Grafana, uninstall and the test suites against 1.5.0’s code on the same day; the current release is 1.6.1.

Every figure measured on router hardware in this documentation comes from one MikroTik RB5009UG+S+: arm64, 4 × 1.4 GHz Cortex-A72 (r0p1), 1 GiB of RAM, Linux 5.6.3. It is the production router of the project’s maintainer, José Manuel Requena Plens, so nothing that needs a reboot is done on it outside a maintenance window, and no conntrack storm is provoked on it, which risks locking out the path being worked through. There is no second physical device. The only other RouterOS the project runs is the virtual lab. Which RouterOS it ran when is under Devices and versions.

  • Accounting files (campaign kernel-2026-09-11): no /proc/pressure and no /proc/schedstat, and the irq column of /proc/stat always 0, so hard-IRQ time is counted inside system. USER_HZ is 100, one tick of 10 ms.
  • Tracing: a privileged discovery container on 2026-09-12 found no eBPF, kprobes or ftrace: no BTF, no debugfs, no tracefs, and the mount points do not exist. A re-probe on 2026-09-14 found /sys/fs/bpf present and empty and /proc/modules listing 238 loaded modules, but no /lib/modules, no kernel headers and no compiler on the device.
  • PMU, from inside a privileged container on 2026-09-12 (RouterOS 7.24.2): the probe opened cycles, instructions, cache-misses, branch-misses and bus-cycles system-wide on 4 of 4 CPUs. In the agent’s own data over the 24 h ending 2026-09-12, six of the seven counters it asks for reported on 4 of 4 cores, cache-references among them, and branch-instructions produced no rows at all. The generic stalled-frontend and stalled-backend events return ENOENT on the A72.
  • Work the tick misses: in a 100.4 ms sample where /proc/stat reported zero busy ticks on all four cores, the PMU counted 2.2–4.6 million cycles and 0.8–2.1 million instructions retired, at a 3.8–5.7 % cache-miss rate (recorded 2026-09-12). Over 2 s the same day, instructions per cycle ranged from 0.381 on cpu0 to 0.992 on cpu1.
  • Privileges: perf_event_paranoid reads 2 and does not block a privileged container, which keeps CAP_SYS_ADMIN effective against the host (measured 2026-09-12). An unprivileged container already holds all 38 capabilities, and its root maps to host uid 32768, so /proc/slabinfo and /dev/kmsg come back EACCES (RouterOS 7.24.2; the date of these two readings is not recorded). With privileged=yes the kernel log, /proc/slabinfo, the MTD ECC counters and the PMU were found readable on 2026-09-12 (RouterOS 7.24.2, kernel 5.6.3).
  • /sys/class/hwmon is empty, even privileged. The two thermal zones, cpu-thermal and soc-thermal, are the whole sensor set of the base model, and both are readable without privileged. RouterOS’s /system/health on this board reports exactly one sensor, cpu-temperature, which is the soc-thermal zone truncated to whole degrees: a mean offset of +0.55 °C over the 8 minutes both series existed (2026-09-14).
  • Both zones declare a polling_delay of 1 000 ms and a polling-delay-passive of 250 ms (read 2026-09-14), so the kernel re-reads the sensor at 1 Hz. The declared critical trip is 105 °C, read from /sys (date not recorded). The sensor quantises to about 0.42 °C (floors-overnight).
  • The cpufreq governor read userspace on 2026-09-14, and the cpufreq clusters, read from related_cpus and affected_cpus inside the container the same day (RouterOS 7.24.2), are {0,1} and {2,3}.
  • Host-wide inside the container (established 2026-09-11, RouterOS 7.24.2, kernel 5.6.3): the CPU, interrupt, memory and block-device files, and /proc/net/softnet_stat. An ordinary container also read these as the router’s on 2026-09-12: the two thermal zones, scaling_cur_freq per core, /proc/yaffs and /proc/buddyinfo. /proc/device-tree/model reads RB5009 even unprivileged.
  • The container’s own (campaign netns-2026-09-12): /proc/net/dev, /proc/net/snmp, /proc/net/netstat and nf_conntrack_count described the container’s veth, 4 packets while the router forwarded millions, and privileged=yes did not change that. The rest of that boundary:
What What the container sees Read
/sys/class/net only lo and the veth; no device for any front-panel port, privileged or not 2026-09-14
/sys/class/mdio_bus, /sys/class/phy mdio_bus holds only fixed-0; phy is empty 2026-09-14
netdev_budget, netdev_max_backlog and the other global net.core entries absent; per-namespace entries such as somaxconn are present 2026-09-14
/proc/net/* files created by modules the container’s own: fib_trie shows only the veth’s /30, snmp6 counts its packets 2026-09-15
nf_conntrack_max the router’s real ceiling, 966 656, the same as RouterOS reports as max-entries 2026-09-14
conntrack timeouts Linux defaults (tcp_timeout_established 432 000 s against RouterOS’s 1d) 2026-09-14
  • Host mounts do not cross the boundary: see the verified fact of 2026-09-15.
  • The device-info record carries the board’s own ceilings, never a number from elsewhere: the 966 656 conntrack ceiling and the container’s 64 MiB memory cap.
  • nbd0–nbd15 are present and idle and mtdblock0–2 read all-zero (2026-09-12, RouterOS 7.24.2, kernel 5.6.3; internal/dashboards/panels.go, the disk in-flight panel). A device’s /proc/yaffs or /proc/diskstats row is stored only when its delta is non-zero, which here is a few times a minute. The reference InfluxDB store still held no block-device table, and no PSI table, on 2026-09-25, on RouterOS 7.24.4.
  • The MTD ECC counters are all zero after years of service (date of the reading not recorded).
  • /proc/slabinfo is 13 833 bytes and 129 lines per read (2026-09-14), the largest per-tick parse by an order of magnitude.
  • The router log held 66 217 rows when record read it, because its dns topic logs to disk (date not recorded).
  • It has a tmpfs disk, which --ephemeral uses. Its raw firewall carries MikroTik’s two defconf: drop rules in their list form, in-interface-list=!LAN and src-address-list=!LANs (Firewall lists).
  • Kernel names against RouterOS names (Port names): eth1 = ether2 from the layer-2 loop of the case study (2026-09-13). eth5 = ether6 and eth6 = ether7 on 2026-09-15: both ports are commented Unused and neither was RUNNING, so each was disabled and enabled again over the API while the agent read /dev/kmsg. The kernel logged br0: port 7(eth5) entered disabled state inside ether6’s window and port 7(eth6) inside ether7’s, nine seconds apart, which rules out reading one flap twice. No traffic was interrupted and both ports were back within four seconds. The six inferred pairs and the switch rest on a read-only /interface/ethernet/print of 2026-09-13: consecutive MAC addresses, …:55 for ether1 through …:5D for sfp-sfpplus1, in RouterOS’s enumeration order, and all nine ports reporting switch=switch1 against the kernel’s one switch0. The device tree, parsed on 2026-09-14, shows the SoC’s ethernet@0 with three MACs, only eth0 enabled (the 10 G uplink) and eth1/eth2 disabled; the nine front-panel ports are netdevs the switch driver creates at runtime. The table is this board’s, observed on RouterOS 7.24.2.
  • Kernel-log record shapes for each port event kind were seen between 2026-09-12 and 2026-09-15 (kernel 5.6.3).
  • Inventory (2026-09-16, RouterOS 7.24.2): the three configuration reads return 17 interfaces. ether1 is an ether in the LAN list, a port of bridge, labelled “TrueNAS - High Performance Storage”, MTU 9000; ether5 is an ether in WAN and in no bridge, labelled “DIGI ONT”; PPPoE_DIGI is pppoe-out, WAN, MTU 1480; VLAN_DIGI is vlan, WAN; wg_devices and wg_trastero are wg in LAN,VPN; ether6 and ether7 carry the comment “Unused”.
  • Counters: 9 ports × 45 counters, and 15 on each bridge, VLAN, PPPoE, WireGuard, veth or loopback interface, counted from the collector’s own mikroscope_api_ifcounters rows on 2026-09-24. The RouterOS version was not read with them; the router ran 7.24.4 that day.

MikroTik’s Cloud Hosted Router under QEMU, in a Docker container, on the machine the lab was built on (x86_64, 12 cores, 27 GB RAM, Docker 29.8, QEMU 10.0.13), provisioned into a clean snapshot with the container package and device-mode container=yes. Measured on 2026-09-26 with RouterOS 7.24.4 and the published 1.3.1 CLI and agent image:

  • CHR x86_64 under KVM: 2 cores, kernel 5.6.3-64.
  • CHR arm64 under QEMU’s TCG emulation, cortex-a72, 2 vCPU, 1 GiB: kernel 5.6.3, the RB5009’s architecture, kernel version and CPU model, at the speed of the host.

What it shows: correctness. The RouterOS commands mikroscope sends and what RouterOS answers, doctor’s reading of the router, RouterOS choosing and pulling the image for its architecture, the container’s start and stop, the agent’s capability detection, its sample format and its ring. The install routes that ran there are under Install routes tested.

How the lab reaches its agents: through a static route. The lab’s LAN side, 192.168.88.10/24 inside the lab’s container, has its default route through Docker’s bridge, not through the router, and reaches every agent through 172.30.0.0/16 via 192.168.88.1, the default of LAB_AGENT_ROUTES: the wider prefix of Static route. On 2026-09-27 the default 172.30.10.0/30 and a 172.30.11.0/30 install both answered /healthz through it.

What it cannot show:

  • A board. There is no switch chip, flash, sensor or device-tree model, so no port map. The agent’s /capabilities reports cpufreq, mtd, psi, schedstat and thermal absent on both, perf and kmsg present on both, and yaffs absent on x86_64 and present on arm64, a /proc/yaffs with no flash behind it.
  • Rates above 10 Hz. The free CHR licence caps what the router sends at 1 Mbit/s per interface: 2 MiB copied off the router over ether2 took 15.6 s, about 1.07 Mbit/s, against 0.2 s onto it. /stream measured 2 526 B a line on the lab, 0.21 Mbit/s at 10 Hz, so 50 Hz is at the cap and 100 Hz over it. Neither was measured.
  • Any cost on arm64. Under TCG the guest’s clock follows the host’s, so no duration, no CPU figure (cpu-load, self.cpu_us, read_ns, wake_ns, the dt_ns spread, slipped), no interrupt, softirq or context-switch rate and no PMU count from the arm64 lab measures a Cortex-A72. The agent kept up at 10 Hz there, 3 003 ticks in 300.1 s with 1 slipped and 1 202 in 120.1 s with none, and its resident size of 14–16 MiB (32 MiB charged to its cgroup) is an indication only.

A routing loop in the lab itself, since fixed, made the emulated router crawl whenever something probed an agent address with no veth behind it: before the fix, arm64 install took 38.4, 11.8 and 34.1 s, and the first and third exited 1 on a complete install. The arm64 figures here are from after the fix; the x86_64 ones were taken before it. What 1.3.1 met there is under Known issues.

On the machine above, from the clean snapshot. How each suite is run is on Test suites.

  • A boot from the snapshot until ssh answers, 2026-09-26: 7 s on x86_64 (6.9, 7.2, 7.4 and 7.4 s) and 26 to 28 s on arm64 (25.7, 26.6, 27.0 and 27.5 s).
  • make roundtrip, 2026-09-26: doctor, install, status, upgrade and uninstall, every verb with --ephemeral, and the router’s /export hashed before and after. 28 to 34 s on x86_64 over three runs and 40 to 45 s on arm64 over four, the export byte-identical every time. With 1.4.0’s code, which makes one uninstall attempt, on 2026-09-27: 20 s on x86_64 and 35 s on arm64, one run each with make’s build steps included, the export byte-identical. With 1.5.0’s code on the same day: 21 s on x86_64 and 30 s on arm64, measured the same way, the export byte-identical.
  • make test-lab with the scenarios of 1.3.1, 2026-09-26: 7 min 21 s to 9 min 49 s on x86_64 over six runs and 12 min 12 s to 16 min 48 s on arm64 over three, the slowest of each with both suites running side by side on a busy host. No install failed.
  • make test-lab with the install options’ scenarios, with the code after 1.3.1, 2026-09-27: 28 min 32 s on x86_64 alone, then 28 min 17 s on x86_64 and 57 min 29 s on arm64 side by side. After the fixes the review asked for, 29 min 2 s and 57 min 13 s side by side, every test passing; S9’s twenty uninstalls with a client on /stream, ten per architecture, were each clean at the first attempt, in 8.0 to 12.6 s. The same day S17 ran on a CHR x86_64 lab with RouterOS 7.23.7: doctor was MISSING RouterOS 7.24 or later and read every other check.
  • RouterOS 7.24.4 changes its own /export: it adds and drops a /system keymat-provider … name=default line by itself, so the suite’s comparison leaves that line out (2026-09-26).

On 2026-09-26, a power cut made as soon as a fresh install answered brought back no agent within 90 s in three of four tries on the arm64 lab. The two containers looked at could not start (Exec format error, Segmentation fault), most likely because RouterOS had not yet written the install to its disk; that was not examined. On x86_64 three of three came back. The start-on-boot scenario, S7, waits 45 s between the install and its cut, and asks only whether start-on-boot works. S6 cuts the power under an --ephemeral install on the lab’s tmpfs disk: afterwards the container is configured and stopped, its root, its image and the manifest are gone with the disk’s contents, the disk is there and empty, and nothing answers; uninstall --ephemeral then leaves nothing at all. It passed on both architectures in the runs of 2026-09-27.

On 2026-10-05, on the x86_64 lab (CHR, RouterOS 7.24.4, KVM), the collector of the code after 1.5.0 ran forward --api-mode off --stdout json across a reboot made with RouterOS’s own /system/reboot, the 1.5.0 agent installed by the Docker Hub pull with start-on-boot. The agent answered again about 20 s after the reboot; the collector logged agent restarted: its newest sample is 439 and the cursor was 1057; resuming from 1, and agent-restart fired with its summary of the 30 s before: 300 samples, CPU 0 % busy, MemAvailable 767.5 MiB of 863.0 MiB, nf_conntrack 24 of 843 776, no softnet drop, then no sample for 16.2 s. reboot did not fire: no kernel-log record after the agent came back had a since-boot time below the last one before the reboot. One run, on an idle router.

Later the same day, on the same lab, the agent and the collector of the code after 1.5.0 that reads the kernel’s boot id (the agent from the branch’s tar, with start-on-boot) ran seven minutes of forward --api-mode off --stdout json across a stop and a start of the agent’s container and then a /system/reboot. The agent inside the RouterOS container read the id and served it in /healthz. After /container/stop and /container/start (20.6 s without a sample) the id was the same, and agent-restart said so: the kernel's boot id did not change, so the router did not reboot; no reboot. After /system/reboot (16.5 s without a sample) the id was new: the collector logged router rebooted: the kernel's boot id went from 30b9831d-… to 8d4e19f2-…, and agent-restart and reboot fired on the agent’s first sample with the same summary (300 samples, CPU 0 % busy, nf_conntrack 24 of 843 776, no softnet drop). The new boot’s kernel log fired no second reboot. One run of each, on an idle router; the arm64 lab was not run.

What RouterOS logs about a boot was read the same evening over the API, from /log/print, after each way of taking the x86_64 lab down. After /system/reboot from an SSH session its memory buffer held router rebooted by ssh-cmd:admin@192.168.88.10/reboot (topics system,info); from a script, …/script:rb/reboot; after /system/shutdown and a start, …/shutdown; after the power was pulled (power-cycle), router was rebooted without proper shutdown (system,error,critical), and the same after QEMU’s system_reset (reset-button). With a logging action that also sent the system topic to disk, a plain /log/print returned the previous boot’s lines as well, and /log/print ?buffer=memory returned this boot’s alone.

Then the collector of the code that reads it (forward --api-mode slow --stdout json, the API tier given the lab’s credentials through LAB_CLI_API=lab) ran eight minutes across a /system/reboot and a reset-button. Both times the read succeeded at the health read that noticed the new boot id, before the API tier’s own round had reconnected (the read reopened the connection), so reboot fired on the agent’s first sample, with RouterOS logged at boot: "router rebooted by ssh-cmd:admin@192.168.88.10/reboot" the first time and "router was rebooted without proper shutdown" the second, 19.4 s and 22.1 s without a sample. One run of each; the two-minute wait for an API that is not back was not exercised on a router, only in the unit tests.

On the RB5009 (RouterOS 7.24.4), the 1.6.0 agent installed on 2026-10-06 served the kernel’s boot id in /healthz from inside its container. The agent-restart of that upgrade said the kernel's boot id did not change, so the router did not reboot, which was true but not known: the 1.5.0 agent before it reported no id. In the code after 1.6.0 the first agent to report an id after one that did not makes no claim about the kernel.

A power-cycle of the lab is no test of the collector: it restarts the lab’s container, and a collector started before it with mikroscope-lab cli keeps the old container’s network namespace and reached nothing after the cut (no route to host until it stopped, the same day). The reboot from inside keeps the lab’s network.

On 2026-10-05, on the x86_64 lab (CHR, RouterOS 7.24.4, KVM), doctor of the code after 1.5.0 that adds the two advisories ran against the lab router set up three ways. With allow-remote-requests=yes and no firewall rule, it warned the router does not answer DNS from its uplink (…, uplink ether1: no rule drops a query that comes in on it). With the default configuration’s two input rules (accept established,related,untracked, drop in-interface-list=!LAN) added ten seconds before, it passed, naming the !LAN drop; with in-interface=ether1 protocol=udp dst-port=53 action=drop alone it passed too. With a filter rule whose interface list had been removed and a raw rule whose veth had been removed, it named both under no firewall rule doctor reads is invalid or names a deleted list.

What RouterOS does with a rule whose reference goes away was read over SSH the same day. An interface list removed under a rule left the rule valid, reading in-interface-list=!*2000010; creating a list of the same name did not change it, and the rule counted packets as a catch-all beside it did. A !LAN drop left that way on the input chain cut the lab’s own SSH from the LAN until the lab was reset. A veth removed under a rule made the rule invalid (in-interface=*4, about=vprobe not ready), and it counted no packet. A rule read right after it was added was invalid with no about, and valid five seconds later. RouterOS refuses to add a rule that names a list that does not exist (input does not match any value of interface-list).

For the fix’s placement: place-before=0 failed from a one-command SSH session (no such item), and a place-before of the input chain’s first rule failed on an empty chain and worked with a rule there. Not tested: IPv6, and a DNS query sent from outside the lab, whose uplink is QEMU’s user network; doctor reads the rules and sends no packet.

On the RB5009 (RouterOS 7.24.4, 2026-10-06), 1.6.0’s doctor read 80 enabled rules in raw prerouting and filter forward and input, none invalid and none naming a deleted list, and said the DNS check had no uplink to judge: the router’s active default route is over PPPoE, and its immediate-gw is the interface alone (PPPoE_DIGI), which the uplink read did not handle; the lab’s DHCP uplink reads 10.0.2.2%ether1. With that read fixed, the code after 1.6.0 found PPPoE_DIGI in the WAN list and judged the queries a maybe, because the default configuration’s accept to local loopback (for CAPsMAN) rule, dst-address=127.0.0.1, stands before its drop all not coming from LAN and read as a maybe accept. With a loopback destination taken as no match for a packet from the Internet, it named defconf: drop all not coming from LAN as dropping the queries. Each of the three runs was one read-only connection.

The opt-in lab that installs RouterOS x86 from MikroTik’s ISO ran on 2026-09-26 with RouterOS 7.24.4. The router reported the board x86 QEMU Standard PC (Q35 + ICH9, 2009) and a trial licence with no level and 24 hours to run. The 1.3.1 CLI and the agent of the lab’s branch behaved as on CHR x86_64: doctor missed the same two lists, a tar install answered at once, status recognised every object and uninstall verified the router clean, leaving the same empty mikroscope directory. The agent’s /capabilities were CHR x86_64’s.

The workflow first ran on GitHub’s runners on 2026-09-26, as the lab job of CI on pull request #69: four x86_64 runs between 21:06 UTC that day and 00:45 UTC on 2026-09-27, each passing the ten tests the suite then had (471.7 to 485.8 s) in a job of 10 min 30 s to 13 min 30 s. Of the three runs below, the first two had an empty cache, so make lab-up downloaded RouterOS and provisioned it; the third took the downloads from the cache and provisioned a new snapshot.

Run Commit Runner make lab-up Suite Job
arm64, a dispatch on main e7efbbe, the lab as merged ubuntu-latest: 4 CPUs, 15 989 MB, /dev/kvm present 4 min 29 s: provisioning 83 s, a boot of the snapshot 24 s 638.6 s, the ten tests of that commit passing 16 min 24 s
x86_64, pull request #70 3c34d9a, the install options ubuntu-latest: 4 CPUs, 15 989 MB, /dev/kvm used by KVM 4 min 28 s: provisioning 46 s, a boot of the snapshot 8 s 1 591.4 s (26 min 31 s), every test passing; S17 skipped, as it needs a RouterOS below 7.24 32 min 42 s
x86_64, pull request #70 1826aa5, with the registry credential ubuntu-latest, /dev/kvm used by KVM 1 min 25 s: the downloads from the cache, provisioning 76 s 1 633.1 s (27 min 13 s), every test passing; S17 skipped. The router pulled as the repository’s Docker Hub account, S2 included; S1, S18 and S5’s Docker Hub and GHCR scripts booted without it 29 min 52 s

In the first two of those runs one download from MikroTik was cut (connection reset by peer) and resumed at the first retry. The probe for KVM on GitHub’s arm64 runner found ubuntu-24.04-arm with 4 CPUs and no /dev/kvm, so the arm64 lab stays emulated on an x86_64 runner.

The two runs on pull request #70 held it for 30 and 33 min, which is why lab.yml now runs weekly and on dispatch only, never on a pull request or before a release.

On 1.4.0’s release commit (79b7c2f, a dispatch on 2026-09-27, run 36329638066) the whole suite passed on both architectures: 1 517.2 s on x86_64 in a job of 31 min 38 s, and 3 042.6 s on arm64, emulated, in a job of 54 min 52 s. The agent image 1.4.0 was not published yet, so the thirteen golden cases that pull their image pulled 1.3.1’s, as S5 does before a tag and says in its log.

On 1.5.0’s release commit (e61ffb6, a dispatch on 2026-09-27, run 36344788569) the whole suite passed again on both architectures: 1 573.2 s on x86_64 in a job of 31 min 38 s, and 3 065.8 s on arm64, emulated, in a job of 54 min 55 s. The thirteen golden cases that pull their image pulled 1.4.0’s, 1.5.0’s being unpublished, as S5 logged.

Device Kind RouterOS When What ran there
RB5009UG+S+ hardware, arm64 7.24.1 until the upgrade of 2026-09-10 the SSH connect cost, 2026-08-26
RB5009UG+S+ hardware, arm64 7.24.2 2026-09-10 to about 2026-09-18 22:43 UTC most campaigns, 2026-09-11 to 2026-09-18
RB5009UG+S+ hardware, arm64 7.24.4 from about 2026-09-18 22:43 UTC the port-errors campaign, the /stream timings, the charts from the reference store, and the deployment commands from 2026-09-21; upgrade with 1.5.0 and 1.6.0, doctor with 1.6.0
CHR x86_64, virtual lab virtual (KVM), amd64 7.24.4 2026-09-26 and 2026-09-27 doctor, plan, three install routes, --expose, uninstall with 1.3.1; the whole lab suite with the code after it and on 1.4.0’s commit
CHR arm64, virtual lab emulated (TCG), arm64 7.24.4 2026-09-26 and 2026-09-27 doctor, plan, three install routes, uninstall with 1.3.1; the whole lab suite with the code after it and on 1.4.0’s commit
CHR x86_64, virtual lab virtual (KVM), amd64 7.23.7 2026-09-27 doctor, which refuses a RouterOS below 7.24 (S17)
RouterOS x86 from the ISO, virtual lab virtual (KVM), amd64 7.24.4 2026-09-26 doctor, a tar install, status and uninstall with 1.3.1

When the RB5009 moved to 7.24.4. The upgrade itself was not written down, but it is bounded: 7.24.4’s build time is 2026-09-16 11:32:21; on 2026-09-24 the router reported 7.24.4 with an uptime of 5d15h49m12s, a boot at about 2026-09-18 22:43 UTC; and a RouterOS upgrade rebooted it at 00:43:30 CEST on 2026-09-19, 22:43:30 UTC the day before, after which the kernel tier resumed at 00:45:03. The API tier’s uptime in the reference store grows at clock rate from 2026-09-19 11:13:35 UTC. So the router has run 7.24.4 since that boot at the latest. Campaigns dated 2026-09-16 to 2026-09-18 are recorded as 7.24.2; the bound neither confirms nor refutes that. None of the 7.24.2 measurements has been repeated on 7.24.4.

What ran on 7.24.4. doctor, plan, install, status and uninstall on 2026-09-21, with the installed agent running continuously since 2026-09-19. On 2026-09-23 the 1.2.x form of doctor’s registry warning was reproduced read-only, and its token warning with a throwaway install that was exposed, upgraded without a token and removed, so upgrade has run there too. The same day doctor read 600 samples from the running agent’s ring and found nothing. On 2026-09-24 RouterOS was given the agent’s reference with its registry host inside remote-image= (two verified facts). The API tier has run continuously on 7.24.4 since 2026-09-19; its inventory and loss-key readings and its cost profile are from 7.24.2 (2026-09-15 and 16), and its reconnection test and 16-interface A/B from 2026-09-19, after that day’s upgrade, with the version not recorded beside them.

The container settings install writes were verified on the RB5009 on RouterOS 7.24.2 and re-exercised on 7.24.4; no other board was tried.

The agent on the RB5009 was upgraded to 1.0.9 on 2026-09-19, to 1.2.0 at 08:08:54 UTC on 2026-09-24 and to 1.2.1 and 1.2.2 later that day, to 1.5.0 on 2026-09-27 and to 1.6.0 at 07:15 UTC on 2026-10-06, each with upgrade --remote-image; the last took 11.8 s from the command to the probe’s answer. The /stream timings of 2026-09-21 are from agent 1.0.9.

Of the four platforms a release publishes, arm64 has run on hardware, the RB5009, and under emulation in the lab. amd64 has run only in the lab, on CHR x86_64 and on RouterOS x86 from the ISO, both under KVM, never on x86 hardware. arm/v7 and arm/v5 have never run on RouterOS: they are cross-built, and CI starts each image with -version under QEMU user-mode emulation (make agent-smoke), which is not RouterOS. What each release publishes, and which version numbers never became a release, is on Releases.

Everything below ran against the RB5009 unless it says the container suite or the lab: on RouterOS 7.24.2 when dated up to 2026-09-18, and on 7.24.4 when dated from 2026-09-19 (Devices and versions). Each part opens with its verdict and the date behind it.

The agent works on the RB5009: the installed one was running there on 2026-09-23, when doctor read 600 samples from its ring. It reads the shared kernel’s /proc, /sys, /dev/kmsg and perf_event_open counters on a fixed ticker at 1 to 100 Hz (10 Hz by default; 10, 20, 50 and 100 Hz measured), keeps the samples in a ring, and serves them: /healthz, /capabilities, /snapshot, /stream, /sampler, and the triggered-capture endpoints /captures and /capture. It serves no /metrics; the exposition is the collector’s.

All six deployment commands have run on the RB5009 on RouterOS 7.24.4, all six by 2026-09-23, each with the CLI of its date. Of 1.4.0 and later, only upgrade (1.5.0 on 2026-09-27, 1.6.0 on 2026-10-06, the second reading the install manifest the first wrote) and doctor (1.6.0, 2026-10-06) have run there; the rest of the code after 1.3.1, which 1.4.0 ships, has run only in the virtual lab (2026-09-27). Read from the code, 1.5.0 changes none of the six commands’ router steps; its uninstall refuses --grafana-dry-run and lists the Grafana and store targets before it touches the router, both checked only in unit tests (uninstall --targets). doctor, plan, install, status, upgrade and uninstall install, upgrade and remove the agent, with every write listed before it happens and every removal verified by ownership counts.

  • Round trip, 2026-09-12: doctor → install → status → upgrade → uninstall left the router’s /export byte-identical (verified). The agent answered 3 s after install, at a 5–7 ms round trip: two probes, 7 ms after install and 5 ms after upgrade, the first printing direct transport ok: agent 4857d0a-dirty, 10 Hz, seq 29, 0 slipped, 7ms round trip. That is one install on one network, not a figure for another. The run passed --ephemeral to doctor, install and upgrade, and it predates 1.1.0, from which uninstall removes only with --yes. scripts/roundtrip.sh now passes --yes to its uninstall (through 1.2.0 it did not, so there it only listed and its export check failed) and --ephemeral to every verb. In that form it runs in the lab as make roundtrip (Lab timings), and it has not been run against the RB5009, where it is make roundtrip-device.
  • Install routes: four routes end to end on 2026-09-17, and the lab’s on 2026-09-26, under Install routes tested.
  • Standalone doctor also pulls the running agent’s ring once and names four faults RouterOS’s own tools do not show: a layer-2 loop, STP churn, a link flap and softnet drops. It reports an agent that does not answer within 3 s and skips it, and its findings never change the exit status. On 2026-09-23 it read 600 samples and found nothing; the four faults are covered by tests that replay the real record text and the 2.0 s cadence of the RB5009’s loop, not by a live fault. Its STP-churn finding rests on how a healthy link-up looks: on the RB5009’s ether7, five link-ups on 2026-09-21 each logged three moves to learning at once and reached forwarding 2.1 to 2.8 s later, and across 30 days of that router’s store every healthy link-up left learning minus forwarding at 0.
  • Its WARN checks, which change neither the exit status nor whether install proceeds. Two are older than 1.4.0: with --remote-image, a /container/config username set while registry-url is empty or names a host other than the one the image is pulled from; and an install of the same --name published on the LAN with no TOKEN in its environment. It reads whether a username is set and how many TOKEN entries exist, never a value. The 1.2.x form of the registry warning was reproduced read-only on 7.24.4 on 2026-09-23; the host comparison that replaced it has warned only against fake router answers in the tests. 1.4.0 adds a WARN for a pull on a 32-bit ARM router, for a pull with less than 16 MiB free beyond the container’s memory-max, for start-on-boot with the root on a tmpfs disk, for a firewall rule that may drop the agent’s replies, for a --lan-address on the uplink, for objects tagged for the install that the flags do not select, and for a route table or firewall it could not read. Of these, the tmpfs and the tagged-objects warnings fired in the lab on CHR x86_64, RouterOS 7.24.4, on 2026-09-27; none has run on the RB5009.
  • Tar extraction, 7.24.2, 2026-09-11: a 1.8 MiB tar, a build of that date from before 1.0.0, was extracted within the same second as its /container/add. The CLI of that time deleted the tar after a fixed wait; 1.4.0 waits for the container to read stopped, bounded by --extract-timeout.
  • An exposed install upgraded without its token, 7.24.4, 2026-09-21: the upgrade passed its check, left both rules in place and wrote an envlist with no TOKEN, and /snapshot through the router’s LAN address went from 401 to 200. From 1.2.0 doctor reports that state, and the code after 1.3.1 refuses such an upgrade (lab, 2026-09-27).
  • A plain uninstall of an exposed install, 7.24.4, 2026-09-21: it removed everything else, printed verified: nothing mikroscope created remains on the router, and left the dst-nat pointing at an address that no longer existed. The code after 1.3.1 reads the install’s shape from the router, and a plain uninstall removed both rules in the lab (2026-09-27).

record, mark and plot work end to end, shown by a recording made on the RB5009 on 2026-09-12: 60 s at 10 Hz, exactly 600 samples, 0 gaps and −7 ms of clock skew. What it showed is under its campaign.

In the lab on 2026-09-27 (CHR x86_64, RouterOS 7.24.4, the published 1.3.1 agent image), record --for 60s wrote 600 samples, seq 11 to 610, with 0 gaps. Three mark notes from a second shell landed in the same .markers.csv while record ran, and record’s summary said 0 marker(s): it counts only the notes typed into it. mark --log-markers over the lab’s API added 7 markers from the router’s log in the window, and plot drew the 600 samples and 10 markers.

forward has run from the RB5009 into three of its eleven sinks, file, Prometheus and InfluxDB 3, all three at once on 2026-09-15; the other eight have been read back only from real products in containers, since 2026-09-16. forward merges the kernel tier with the RouterOS API tier and writes to eleven sinks. Loki, OTLP, Graphite, Elasticsearch, SQL, PostgreSQL, Telegraf and stdout have not had router samples pushed through them; the container suite writes into each real product and reads it back through its own API (Test suites).

  • 2026-09-12, eight minutes: 4 800 kernel and 479 API samples forwarded with 0 gaps and 0 drops, running mikroscope forward --for 8m --prom :9124 --influx … --interfaces bridge,ether1,PPPoE_DIGI --conntrack-every 10s.
  • 2026-09-15, five rate runs at 10, 50 and 100 Hz into a file, a Prometheus exposition and InfluxDB 3 at once: every sink reported 0 gaps and 0 drops (rates-2026-09-15).
  • Prometheus: a Prometheus 3.14 scraped the collector’s exposition every 5 s with the RB5009 feeding the collector, on 2026-09-12 and again on 2026-09-15. No other scrape interval and no other Prometheus version is recorded.
  • InfluxDB 3: the first real target, on 2026-09-12, answered 422: would exceed limit of 5 databases, the node limit InfluxData documents for Core. The measured runs used an InfluxDB 3 Core instance of their own. On 2026-09-19 the collector was moved onto InfluxDB 3 Enterprise 3.11.4 and wrote 44 820 rows across 36 tables in its first 19 minutes, 0 dropped and 0 errors. Everything before that date was measured against Core.
  • The collector against a faster agent, 2026-09-21: with the agent at 100 Hz and again at 50 Hz, /healthz read rate_hz 100 and 50 and the collector’s log named the same rate back, while mikroscope_info{rate_hz} read 10 at both; mikroscope_cpu_busy_ticks carried the buckets le="0" through le="11" plus +Inf at both; and window="1s", window="10s" and window="60s" were all present at both.
  • Detections in the reference store, 2026-09-19 11:13 UTC to 2026-09-24 08:26 UTC: four of the eleven rules fired, microburst 116 times, ipc-collapse 41 (on all four cores), link-flap 11 and agent-restart once. Only microburst has had its behaviour measured against the samples (microburst-replay-2026-09-16); the other three were counted, not checked against what the router was doing. The 11 link-flap rows are on ether1, ether2, ether4, ether6 and ether7, each at 2 to 6 link records in 60 s, all unprovoked; which was a cable, a device at the other end or something else was not checked.
  • Restarts: the RouterOS upgrade reboot of 2026-09-19 exercised the collector’s resync for real (the kernel tier resumed at 00:45:03 CEST). Whether a mikroscope_detection row for agent-restart was written then was not checked, and the store that would hold it was replaced later that day; the current one starts at 11:13 UTC. It holds one agent-restart row, at 08:08:54 UTC on 2026-09-24, value 1 against threshold 4 338 037, when the agent was upgraded to 1.2.0.
  • Device info: sent once at start, the four device panels read “No data” over every window after it: on 2026-09-17 the last device row in the reference deployment was 26 hours old and those panels had been empty as long. Repeated every five minutes, it costs twelve rows an emission on the RB5009.
  • Delivery over 24 hours, 2026-09-17, in the reference deployment: the largest interruption was 114.5 s, and it was self-inflicted, a container swap plus the minute the collector takes to notice a restarted agent. In ordinary running the collector never fell behind. That is why the ring holds 60 s by default: it covers a restart of either side on a LAN.

The API tier has run on one board, the RB5009, continuously on 7.24.4 since 2026-09-19.

  • Reconnection, 2026-09-19, after that day’s upgrade, without touching the router: a collector running monitor-traffic on 16 interfaces (all but lo) at 1 Hz had its API socket destroyed from the host with ss -K, which is what the router’s side of a reboot looks like to it. The tier reopened the connection and retried inside the same round: 0 failed commands, 0 dropped rounds, 16 interfaces in every one of the 108 seconds either side of the kill.
  • What RouterOS returns: loss keys, an rx-overflow count only the port counters carry, the two counting planes of a switch port and its bridge, and log times.
  • cpu-load is a trailing mean of about one second, fitted against /proc/stat (cpu-load-window-2026-09-15).
  • What it costs the router: API tier cost.

Every panel of all five dashboards answered without error in a real Grafana over the container suite’s stores (2026-09-17), and the InfluxDB and Prometheus dashboards passed dashboards check over the RB5009’s own data (2026-09-16). There is one per store, InfluxDB 3, Prometheus, PostgreSQL, Graphite and Elasticsearch, generated from one panel list. Graphite and Elasticsearch carry fewer panels on purpose (39 and 28 against 177): Graphite has no labels and Elasticsearch no nested documents. The PostgreSQL panels are the InfluxDB ones rewritten, and the container suite has a real PostgreSQL plan every one of their queries. The Grafana versions used are 12.3.2 on 2026-09-12; 13.2.1 for the browser passes of 2026-09-12 and 2026-09-14, check and render on 2026-09-15, check on 2026-09-16 and the container suite; 12.3.0 on 2026-09-21 and in the renders of 2026-09-25; and 13.2.2 on 2026-09-24 and 2026-09-25. No other version has been tried.

  • 2026-09-12. On Grafana 12.3.2, the InfluxDB dashboard against an isolated InfluxDB 3 Core fed by forward from the RB5009: every panel returned rows, 158–316 per panel over 10 minutes. The Prometheus dashboard against a Prometheus 3.14 scraping the collector every 5 s: every panel returned rows, 228–2 052 over 5 minutes. Without the datasource’s token field, panels failed with flightsql: Unauthenticated. The same day dashboards check --store influxdb --window 12h passed on all 125 panels it walked while about 90 were unreadable in a browser: legends reading “value core 0”, two xycharts stuck on “Loading plugin panel…”, a continuity lane that stayed green over 2 170 missing ticks. An eight-panel Overview rendered 2 188 px tall in an 844 px phone viewport, and 900 data points is the width the browser sent for the graphs.
  • 2026-09-14. An eleven-panel Overview measured 1 052 px tall in a 1 080 px browser viewport, one desktop screen; a render walk against a 10.5 h capture from the RB5009 rendered 140 InfluxDB panels with 0 error badges and 0 “No data”; the InfluxDB datasource escaped $__interval_ms in five panels, in the browser, into SQL InfluxDB 3 could not parse; and a query naming a field the store had never received failed at planning, No field named limit. Valid fields are …, through the datasource proxy.
  • 2026-09-15, Grafana 13.2.1: check, and a headless row-by-row walk of both dashboards in Chromium at 1600x1000, 0 error badges and 0 “No data” over 168 InfluxDB and 130 Prometheus panels. It has not been repeated for the three panels it did not cover: the two port-event panels and “What each interface is: type, role, bridge and label”. The same walk saw the fast-path share swing 0–100 % between polls on interfaces moving a few packets. Regenerating that day reproduced the committed files byte for byte.
  • 2026-09-16, Grafana 13.2.1, against the agent on the RB5009 (RouterOS 7.24.2, privileged, the default triggers), forward --prom :9124 --influx … --interfaces bridge,ether1,PPPoE_DIGI --counters-every 10s for 30 minutes into an isolated InfluxDB 3 Core, and a Prometheus 3.14 scraping the collector every 5 s plus the agent directly for the families the collector could not recompute then:
Store Window Panels Failing Known-empty tolerated
InfluxDB 3 30 minutes 171 0 10 (the two port-event panels, the opt-in conntrack poll, the two trigger panels, PSI, the four idle block devices)
Prometheus 30 minutes 133 0 9 (the two port-event panels, the two conntrack API panels, PSI, the four block devices)

The 10 and 9 are that deployment over that window. The two port-event panels’ SQL was validated the same day against a synthetic table in the same InfluxDB 3, because the live store had no kind column until the first port record classified by kind was written; it has one since 2026-09-19, holding all eight kinds. The dashboards page of that date also recorded that the agent’s own exposition carried the kind label only when the agent itself classified, which the one on the RB5009 then did not; agents serve no exposition now.

  • 2026-09-17, container suite: all five dashboards imported into a real Grafana over the stores the suite filled, and every panel asked: 57, 98, 136, 35 and 28 panels returned data and none failed.
  • 2026-09-21, Grafana 12.3.0: import against the live InfluxDB 3 store, then check over a 15 min window: 167 panels returned rows, the 9 known-empty ones were empty, and the Overview’s detections tile was empty because the store held no detection in that window, which is the one case check reports and a browser reads as the healthy “none”.
  • 2026-09-24, the production Grafana, whose health endpoint reported 13.2.2 later that day: the InfluxDB dashboard imported and captured at 390x844 and 1600x1000 to check the 32 px stat text and the memory time series. With Grafana’s automatic size a single “0” was drawn about 60 px tall at 390x844 and each stat tile took about a third of the screen; with the fixed size the numbers read at a normal size at both. No row-by-row badge count was taken, the Overview’s height was not measured, and the other four dashboards were not captured.
  • 2026-09-25, Grafana 12.3.0, 13.2.1 and 13.2.2 over a throwaway InfluxDB 3.11.2 Core:
    • The not-available row as committed in 1.3.0 sent five queries and painted five table … not found badges on 12.3.0 and 13.2.1; with its queries shipped hidden it sent none and painted none, and 13.2.2 did the same in a second render.
    • A detections query against a missing mikroscope_detection table cannot be written to avoid InfluxDB 3’s planning error: WHERE false, a UNION ALL and an EXISTS guard over information_schema all failed at planning. On 13.2.1 the failing layer showed nothing on the dashboard and wrote one error-level Partial data response error line to Grafana’s log per load. 12.3.0 never sent that layer’s query, with or without the table (known issue).
    • A probed import from 1.3.0 removed a missing panel’s queries, and Grafana 12.3.0 gives a panel with no query a default one: the InfluxDB plugin answered it with No SQL statements were provided in the query string, as a red badge. 13.2.1 sent nothing for the same panel.
    • Elasticsearch 9.5.3: the Host variable’s lookup, written up to 1.3.0 as a JSON object, went out from 13.2.1 and 13.2.2 as an empty query answered 400 invalid query, missing metrics and aggregations, and 12.3.0 did not send it; opened without ?var-host=, six Overview panels failed with Failed to parse query [host.keyword:]. Written as the JSON string the datasource parses, Host filled from the index and no panel showed a badge on any of the three.
    • An InfluxDB panel that bins with $__dateBin under a sub-second step drew a 0-second bin and came back empty, in check and in a browser zoomed to a few minutes, on 13.2.1 and 13.2.2.
  • 2026-09-25, the production Grafana 13.2.2 against the reference InfluxDB 3 store, with a build that hides the not-available queries: dashboards check --window 1h, with the probe and with --no-probe, gave 164 panels with rows, 12 known-empty tolerated and 1 failing, “Detections in the window”, because the store held no detection that hour. Without the probe the five not-available panels read none … rows=0 frames=0 with no error text, where 1.3.0 printed 400 … table … not found; the two trigger panels answered mikroscope_trigger not found, because that store has no trigger table. The detections annotation’s SQL returned 44 rows over 6 hours. Run again later that day, after the other four stores’ unprobed files got their not-available queries back, it gave the same counts, and the detections layer 32 rows over 6 hours; the InfluxDB file checked was byte-identical to the first run’s.
  • 2026-09-25, container suite (Grafana 13.2.1, Prometheus 3.14.0, PostgreSQL 18.6, graphite-statsd 1.1.10-5, Elasticsearch 9.5.3): the committed files’ not-available queries, against stores that held nothing for them, all answered 200 with no error, and on Elasticsearch every point was 0 whether or not another router in the same index had mapped the field. On PostgreSQL the detections layer’s query answered 200 with no rows over a window before any detection and 10 rows over the run’s own. Both were queries through Grafana’s API, not renders.
  • Continuity is derived from every sample’s sequence number, and it found 4 493 missing ticks and one restart in the 2026-09-11/12 capture that the collector’s gap record said nothing about.
  • The default range is 3 hours because now-15m opened on 108 empty panels with the agent stopped, and with it running drew unreadable walls of noise from 100 ms samples (date not recorded).

Only the InfluxDB 3 form of the rules has been loaded into Grafana and watched evaluate, on 2026-09-21 against the maintainer’s store; the Prometheus and PostgreSQL forms never have. The rules number 14 for InfluxDB 3, 15 for Prometheus and 10 for PostgreSQL.

  • The load of 2026-09-21, into Grafana 13.2.1 against the live InfluxDB 3 store, held the twelve InfluxDB rules of that date and found two defects, both fixed in 1.1.0. coalesce() over the unsigned aggregate the InfluxDB sink writes returned HTTP 200 with no frames, which Grafana reads as NoData and these rules as OK: seven of the twelve could never fire, mikroscope-l2-loop among them, while the same SQL over /api/v3/query_sql returned 109 on the live loop signature. And the conntrack rule named limit_objs, the SQL sink’s column, where the InfluxDB sink writes limit; the alert rules page had predicted that one before it was confirmed. After the fix all twelve returned a value, and mikroscope-l2-loop went to Alerting at 16:58:50Z on the live own-address signature; Grafana’s instance list for it still carried Normal (NoData) at 16:56:50Z from the run before the fix.
  • Queries run by hand against the live store on 2026-09-21: mikroscope-l2-loop read 109 over its own five-minute window, the own-address signature on ether2 having run continuously at about 0.5 records/s since 2026-09-19; mikroscope-port-link-down read 3 over the window holding a real ether7 link-down at 16:33:49 and 4 over the cluster of four on ether4 at 09:19 on 2026-09-20. Both thresholds are 0 with gt, so both conditions were met.
  • The two rules 1.2.0 added were not in that load, and no Grafana has evaluated them since. Their InfluxDB SQL was backtested on the reference store: mikroscope-bridge-port-dark returned 1 on three windows inside the loop and 0 after it, and marked sfp-sfpplus1 in 204 ten-minute bins and ether2 in 376 between 2026-09-19 11:13 and 2026-09-23 22:44 UTC, the two phases of that week’s layer-2 loop, and no other port. mikroscope-wakeup-storm returned 0.94–0.95 on two healthy windows, never exceeded 1.81 over 551 ten-minute windows with a full day behind them (2026-09-20 to 23), and opened at 35.7 on the wake-up storm of 2026-09-23; it cleared as the new rate became the baseline, after about four hours. Their PromQL was checked for syntax only. The wake-up rule’s PostgreSQL form has no recorded run, and the dark-port rule has no PostgreSQL form: 1.2.1 dropped it, because the SQL sink’s interface-counter table does not hold the columns it reads.
  • The detections rule no longer fires on microburst or ipc-collapse, which describe how a healthy router carries traffic: from 2026-09-23 10:30 to 2026-09-24 10:30 UTC on the RB5009 they were 64 of 71 detections, and left in they fired the rule in 43 of 288 five-minute windows; without them it fired in 6.
  • The egress rule: a 1 GbE port to a server lost 3 337 packets in six one-second bursts over 6.5 h, peaking at 436 packets/s (2026-09-19), visible on the egress queue panel and correctly not an alert.
  • The thermal rule reads the zone’s own critical trip, 105 °C on the RB5009.

forward --grafana has run against the maintainer’s Grafana (2026-09-19, InfluxDB 3) and against the container suite’s for all five stores (2026-09-20, and with 1.5.0’s code on 2026-09-27). dashboards publish has run only there, on 2026-09-27, over what forward --grafana had just made. On 2026-09-20 the collector was pointed at a real Grafana and each of the five stores in turn, allowed to create the datasource, and then dashboards check ran every panel’s query through Grafana’s API against the datasource the collector had built. All five answered with no datasource error as the test of that date read it: it looked for flightsql: Unauthenticated, for the TLS error named below and for err with a space on either side, which is no mark check prints, so an error worded any other way passed it. The same run found a defect no unit test had: a collector writing to a live PostgreSQL through --postgres and nothing else published nothing and reported “there is nothing to publish”, because the store list only knew --sql. Against a plain-HTTP InfluxDB store, a datasource without insecureGrpc answered every panel tls: first record does not look like a TLS handshake while the store was fine (measured 2026-09-19).

On 2026-09-27 the test ran with 1.5.0’s code (commit ffb934e, whose Go code is the release’s; e2e.yml run 36344009521) against Grafana 13.2.1, InfluxDB 3.11.2 Core, Elasticsearch 9.5.3, PostgreSQL 18.6, Prometheus 3.14.0 and graphite-statsd 1.1.10-5, now failing on any FAIL line of check that carries an error. For each store in turn, forward --grafana created the datasource and published the dashboard, check ran every panel against that datasource over a 15-minute window, and dashboards publish, given the same store and Grafana flags, exited 0, reported the datasource unchanged and printed the dashboard’s address. No panel that check counts as failing carried an error; a panel expected to be empty is marked none when it answers no rows or an error, and the test reads those lines only for the two InfluxDB errors above. The Elasticsearch datasource was the collector’s, with @timestamp as its time field (1.4.0’s named time, a field no document carries), and its check exited 0; on InfluxDB, PostgreSQL, Prometheus and Graphite 5, 20, 33 and 4 panels returned no rows in the window, with no error. Not exercised in the suite: the Authorization header the Elasticsearch datasource sends, because its Elasticsearch runs without authentication; a datasource that carries a token or a password, which no store there has; and a run that publishes more than one store, since each run published one.

uninstall --targets removed only what this project wrote in the container suite (2026-09-20, and --targets data again with 1.5.0’s code on 2026-09-27), and has not been run against the maintainer’s production store. The suite asserts that a table this project did not write, in the same database and schema, is neither listed nor removed. It runs the verb with --targets data alone, against PostgreSQL and InfluxDB 3, with no router and no Grafana. What 1.5.0 changed in the verb beyond that path has only unit tests: it refuses --grafana-dry-run, lists the Grafana and store targets before it touches the router, stops on a Grafana it could not read, which 1.4.0 counted as holding nothing to remove, and reads Grafana’s address from GRAFANA_URL when MIKROSCOPE_GRAFANA_URL is unset.

Four suites need no router: the unit and end-to-end suites run in CI on Linux, macOS and Windows against captured trees and fakes, one runs against nine real stores in docker compose, first in full on 2026-09-16, and one runs the CLI and the agent against a virtual RouterOS, first green on 2026-09-26 and first run on GitHub’s runners the same day. What each proves and how to run it is on Test suites.

  • End to end: both binaries against a captured /proc tree of the RB5009 and a fake agent, with one receiver per sink protocol asserting the bytes. It needs no router, no Grafana and no network, and runs in CI on macOS and Windows as well as Linux, where a difference in the executable’s name, in the files the file and SQL sinks write, or in how a child process is stopped would surface. Every sink has been tested against a local receiver since 2026-09-12, on the amd64 development host. The suite used to scrape both the agent and the collector, and now checks the collector alone.
  • Stores (make test-e2e-docker): the collector, against the same canned agent and with the API tier off, into Loki 3, the OpenTelemetry Collector, graphite-statsd, Elasticsearch 9, Telegraf 1.39 over HTTP, PostgreSQL 18 (through --sql, and through --postgres since that sink landed on 2026-09-21), InfluxDB 3 and Prometheus, with the file sink as the oracle the others are compared against. Each store is read back through its own API, then all five dashboards are imported into Grafana and every panel’s query runs through Grafana’s API; the suite has forward --grafana publish all five datasources and checks their dashboards against them, then, since 2026-09-27, runs dashboards publish with the same store and Grafana flags, which has to exit 0 and name each store’s datasource and dashboard, and empties the stores again with uninstall --targets data. It needs Docker and no router: the samples are canned, so the run is reproducible anywhere. The first full run, on 2026-09-16, found a panel that named two columns the store has only when the API tier ran: the port-event table names label and role, which the sink writes onto a kernel-log row only from the API tier’s inventory, and on InfluxDB 3 a column that is not in the table failed the query with Schema error: No field named label, so the panel could not render at all. The sink pages record the collector running into these stores since 2026-09-17. On the development machine on 2026-09-16 the stack came up in 55 to 81 s with the images already pulled, and the whole suite took 75 to 100 s from nothing. The suite joined the pull-request path on 2026-09-18. Until then it ran weekly, on dispatch and before a release, and that is how 1.0.5 reached main with two tests in it still reading the agent’s /metrics, which that release had removed: every pull-request check was green, and the release gate found it. On the release run that failed it took 2 min 37 s from job start to result, containers included; on pull request #70, on 2026-09-27, 4 min 9 s. The OpenTelemetry Collector, which takes JSON, is the only real OTLP receiver the sink has run against. On 2026-09-20 the SQL and PostgreSQL sinks, run side by side into two databases, held 34 tables matching byte for byte, every column of every row hashed per row and summed, plus identical information_schema.columns for the whole schema.
  • Virtual RouterOS lab (make test-lab): the CLI’s deploy verbs and the agent against MikroTik’s Cloud Hosted Router 7.24.4 in QEMU, x86_64 and emulated arm64: install by both image routes, upgrade, uninstall, --ephemeral and start-on-boot through a power cut, each scenario ending with the router’s /export compared with its start. It tests the installer’s correctness; it is not a board, and no figure under Agent cost or Campaigns comes from it. The lab, its timings and its first runs on GitHub’s runners are under Virtual lab, its scenarios on Test suites.
  • Sink fixtures, not a router: the SQL, OTLP, Graphite, Elasticsearch and Telegraf sizes on Other sinks come from two-core test fixtures of 2026-09-12. The SQL header for all forty-three tables, rendered by the sink’s own header() on 2026-09-19, is 10 482 B, 13 966 B with the TimescaleDB hypertable statements. Twelve renders of an 8-name map gave 7 orders (2026-09-12), recorded in internal/sinks/telegraf.go.
  • The development host (amd64, kernel 6.12.107): TestParsePressureReadsTheRunningKernel and TestParseSchedstatReadsTheRunningKernel parse the host’s own /proc/pressure/{cpu,memory,io} and /proc/schedstat and compare them with a second reading of the same bytes, skipping where the files are absent, which is the RB5009’s case. On that host they pass. Its /proc/pressure/cpu carries a full line, all zeros when read again on 2026-09-24: the kernel’s PSI documentation says CPU full is undefined system-wide and has been reported since 5.13 as zero, so an older kernel has no such line and the parser accepts either. Its PMU, recorded on 2026-09-21, has six counters, cycles, instructions, cache-references, cache-misses, branch-instructions and branch-misses, multiplexed, each running about 84 % of the time it was enabled. A 100.3 ms interval there carried 11 ticks (2026-09-11/12); sample_test.go asserts the busy ratio is capped at 1.
  • The change filter’s re-arm, against the fixture tree on 2026-09-15: the floored gauge families and mikroscope_slab_limit_objects were present in 6 of 6 scrapes from 5 s after start. The comment on ProcSource.ReArm in internal/agent/source.go records the same check as four scrapes from t+5 s, 0 before the fix and 4 of 4 after; which run each count is from is not recorded.
  • The script generator, 2026-09-27: pnpm run rsc:check renders the 20 golden cases of the steps spec in JavaScript and matches each script, each step’s commands, the removal order, the values and the command line with the Go output byte for byte; a one-line change to the renderer made it report 72 failures. pnpm test:generator made 168 checks in headless Chromium (Playwright 1.63) against the built site: the 20 cases set through the form’s own controls, the errors, the token, copy, keyboard-only use, no network request, and the .rsc downloads, which work under the site’s meta CSP. A one-off comparison with the Go code over 304 option sets agreed on 301; the other three are inputs the page refuses and Go accepts (a /030 prefix, a busy>=.5 threshold and an IPv4-mapped IPv6 address).
  • The site, 2026-09-25: headless Chromium 153 requested only favicon.svg, and resolved a manifest id of ./ to the origin’s root; in headless Chromium and WebKit, in both schemes, after a pick, a stored pick and a reload, the phone’s toggle and with JavaScript off, the theme-color tag held the header’s colour every time, where the pair split by prefers-color-scheme it replaced held the other theme’s colour after each pick.
Route RB5009UG+S+ CHR x86_64, lab CHR arm64, lab
Docker Hub pull, --remote-image 2026-09-17, 7.24.2: /healthz at 2 ms; the 1.0.1 image, sent without its host and pulled through registry-url. 2026-09-24, 7.24.4: the reference with its host, by a hand-written /container/add, never started; no whole install 2026-09-26: install in 9.6 s, pulled anonymously with the factory /container/config; 2026-09-27: S2, S4 and S5 2026-09-26: install in 10.3 and 10.2 s, pulled anonymously; 2026-09-27: S2, S4 and S5
GHCR pull, --remote-image 2026-09-21, 7.24.4: failed, auth error 2026-09-26: pulled with no registry credential and the host in remote-image=, /healthz 5 s after install began; 2026-09-27: S18 and S5 2026-09-27: S18 and S5
plan --rsc, /imported 2026-09-17, 7.24.2: /healthz at 15 ms on its first samples, no CLI in the install 2026-09-26: /healthz 14 s after the upload began; 2026-09-27: S4, and the fifteen golden scripts the lab can run (S5); the default script pasted at the ] > prompt, and a tar script without its tar, which stopped before any write 2026-09-26: /healthz 9.1 and 9.5 s after the upload began; the script was the x86_64 one; 2026-09-27: S4 and S5
Image tar, --agent-tar 2026-09-17, 7.24.2: the published arm64 tar, /healthz at 2 ms 2026-09-26: the release’s amd64 tar, 6 s; 2026-09-27: the branch’s tar, S3 and ten installs in a row; the release’s amd64 tar checked with sha256sum and cosign verify-blob, then installed and upgraded 2026-09-26: the release’s arm64 tar, checksum as published, 9.6 and 9.2 s; 2026-09-27: the branch’s tar, S3 and three installs in a row
Script generator not run 2026-09-27: the page’s default script, equal to plan --rsc’s, and a tar script with other settings, each downloaded from the built page and pasted at the ] > prompt; the agent answered each time, and the CLI’s uninstall and the page’s uninstall script each left /export as it was not run
Manual install, terminal not run 2026-09-27: the page’s commands one at a time, registry pull and tar; /healthz answered, status recognised the install from its manifest, and the page’s removal left nothing; again with --expose’s two rules, and with both memberships skipped not run
Manual install, WebFig not run 2026-09-27: both image routes, every form submitted; /healthz 200 each time; the pull removed through WebFig, the tar by uninstall reading the manifest WebFig wrote; /export as it was after each not run
From a checkout 2026-09-17, 7.24.2: /healthz at 2 ms; the round trip of 2026-09-12 not run not run
--expose 2026-09-11: the two rules verified; 2026-09-23, 7.24.4: a throwaway exposed install, upgraded without a token and removed 2026-09-26: /healthz 200, /capabilities 401 without the token and 200 with it; 2026-09-27: S8 and S14 2026-09-27: S8 and S14
uninstall 2026-09-17: verified by ownership count after each route; /export after all four byte-identical to the one before 2026-09-26, 1.3.1: 5 of 15 first attempts failed; a second run cleaned up every time. 2026-09-27: every first attempt clean, after every route (F4) 2026-09-26, 1.3.1: 8 of 8 first attempts cleaned up; with a client on /stream it failed. 2026-09-27: every first attempt clean, after every route (F4)
  • In the lab on 2026-09-27, with the code after 1.3.1 (the changelog’s 1.4.0 section): the whole lab suite on both architectures, every test passing (Lab timings). After every route, the CLI’s pull and tar routes, a tar install then an upgrade, an imported plan --rsc script, every golden script the lab can run and installs made by the released 1.3.1 CLI, uninstall left /export equal to the one taken before the install and no mikroscope path on /file. The GHCR pull of 2026-09-26 logged registry=ghcr.io, one 3 093 207-byte layer and download/extract done 2 s later, with /container/config holding no registry-url and no username.
  • The install pages’ runs, 2026-09-27, on the x86_64 lab, with the code after 1.3.1 and the published 1.3.1 agent image and tar:
    • The generator’s default script answered 8.8 s after the paste began. Its tar script (name gen2, 172.30.11.0/30, port 9200, both lists none, rate 20, triggers busy>=0.9,oom, a token generated in the page, --expose on 192.168.88.1, --restart-max-count 3, --start-on-boot no) was byte-identical to plan --rsc for the page’s command line, answered on 172.30.11.2:9200 and deleted its tar after extraction.
    • plan --rsc’s default script, byte-identical to pull-dockerhub.rsc, installed a running agent both pasted and uploaded then /imported (Script file loaded and executed successfully). The tar script without the tar stopped with mikroscope: upload mikroscope.tar first and wrote nothing.
    • The manual terminal page’s commands, each sent on its own: both routes answered, the tar route’s wait deleted the tar after extraction, and a manual tar install was also removed by uninstall --yes alone. A find-and-replace of the placeholders that also rewrote the envlist’s keys made the removal leave the whole envlist behind, silently, which is why the page says to replace values only.
    • The same page again, once it wrote a manifest per image source and gave --expose a section: a registry pull exposed on 192.168.88.1 with a token, a tar install, and a registry pull with both memberships skipped. Each manifest its commands wrote was byte-identical to the CLI’s for the same settings (pull-dockerhub, expose-token, default-tar, lists-none); status read each from its manifest, the exposed one with its two rules as the install’s; /capabilities answered 401 without the token and 200 with it; and the page’s removal, the two rules first, left the leftover count at 0 and /export, the residue and /file equal to the baseline.
    • A WebFig install left an /export equal to the one the golden script’s /import left, on both routes, apart from the veth’s two MAC addresses, which RouterOS draws at random; status listed every object of it as the install’s.
    • The offline page’s checks on the published 1.3.1 files: sha256sum --ignore-missing -c checksums.txt OK and cosign verify-blob (cosign v3.1.3) Verified OK. Then install and upgrade --agent-tar with the amd64 tar, 7 053 KiB uploaded, and a clean uninstall.
    • Installs made by the released 1.3.1 CLI: status read their shape from the tags, upgrade --remote-image wrote a manifest, and uninstall --yes removed everything, 1.3.1’s mikroscope/ directory included.
    • upgrade with neither --remote-image nor --agent-tar looked for Go to build the agent: it does not take the image from the manifest. A second install on an installed router created nothing (install done: 0 step(s) created).
  • The Configure pages’ runs, 2026-09-27, x86_64 lab: doctor on a stock CHR was MISSING only the interface list LAN, with a fix offering --iface-list none, and with both lists none every check passed. Built-in lists were refused before any connect. The trap check against the advanced-firewall and advanced-firewall-range profiles named the rule that drops the agent’s replies, and its fix. An address list that did not exist was created by the install’s entry and was gone after uninstall. The probe’s first round trip read 1.02 to 1.03 s on three of six installs and 2 ms on the other three, and status a second later 1 to 2 ms.
  • On the RB5009, 2026-09-17, the four routes ran one after another, each under its own name, veth and /30 so that nothing already on the device was touched, and each was removed before the next. /healthz was read from the collector host.
  • GHCR, 2026-09-21, a second install beside the running one under its own --name, --veth, --subnet and --port. The CLI of that date put only the rest of the reference into remote-image=, so registry-url was pointed at https://ghcr.io for the run and put back afterwards; doctor refused first with MISSING registry-url is https://ghcr.io. With the registry pointed at GHCR the router created the container and failed the pull: download/extract error: fetch manifest failed: getting https://ghcr.io/v2/jmrplens/mikroscope-agent/manifests/1.0.9 failed: auth error. /container/config holds one username and password for every registry, and the RB5009’s is a Docker Hub login. From the operator host the same day, against the same public package, no credential returned 200 once the anonymous token was fetched and a foreign credential 403 at the token endpoint. The run also found that a --remote-image install identified its container by the registry reference, the same string for every install, so a second install on one router refused as though the first were someone else’s; it is identified by its veth now.
  • Docker Hub’s limits, read on 2026-09-24: 100 anonymous pulls per 6 hours per IPv4 address or IPv6 /64, and the anonymous token carries the same 6 h (pull_limit_interval 21600), while the registry’s ratelimit-limit header on a HEAD of the agent’s manifest read 100;w=3600. Which one Docker enforces was not measured. The one-pull-per-install figure comes from Docker’s counting rule and a pull made with curl from the operator host, not from a pull by RouterOS.
  • MikroTik’s sources on the default registry, read on 2026-09-24: the 7.18 changelog adds registry-url=https://lscr.io, the 7.21.2 one says “changed default container registry to docker.io”, no changelog from 7.21.3 to 7.24.4 mentions the registry, and the container documentation still gives https://lscr.io/. The lab’s CHR on 7.24.4 had no registry-url or username and reported assumed-registry-url: docker.io.

At the install default, 10 Hz, default per-source floors and a 60 s ring, the agent costs 2.69 % of one core and 13.2 MiB RSS, read from its own cgroup at steady state with the ring full:

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · 300 s windows at steady state (ring full), full source set, the shipped configuration — a 60 s ring, the memory limit derived from it and the default 64M container cap — with the collector forwarding to InfluxDB 3

The measured runs
ratefloorsCPU of one coreµs/sampleRSSslipped ticksgaps / drops
10 Hz (default)default2.69 %2 68513.2 MiB00 / 0

That is above the budget of 2 % of one core and inside the 16 MiB RSS one. The 2026-09-15 campaign was over on both, with a 300 s ring and a flat memory limit, before 1.0.6 brought the 60 s ring and the derived limit. The image budget is 8 MiB; the published 1.2.1 image for arm64 is 6.38 MiB (image). The budget is guidance rather than a contract: cost scales with the device, the source set and the ring size.

Two rules keep a cost figure honest. Wait out the ring (BUFFER_S) first: on the RB5009, at a 14 MiB soft memory limit, a reading taken in the first minute after install came back at 1.47 % of one core against a 9.38 % steady state. And read cost from the collector’s /metrics, or from cpu_us and rss in mikroscope_self in the store you write to, rather than from a large /snapshot, whose ~1.9 MB response the agent must serialise. The procedure is on Agent cost.

Every row was measured on the same RB5009 at 10 Hz, and each is one window with no spread recorded.

Configuration CPU of one core RSS Measured Note
Every source read every tick 2.43 % not recorded 2026-09-12 above the budget
Ring full, MEM_LIMIT_MB 14 9.38 % (9 374 µs/sample) not recorded 2026-09-12 a 300 s ring of lines of about 2.4 kB, the line of that date, holds ~7.3 MB; the Go GC runs without pause
Ring full, --mem-limit-mb 40, --memory-max 64M 1.39 % (1 388 µs/sample) 25.13 MiB 2026-09-12 0 slipped ticks
PMU counters on, a live forward writing to InfluxDB 1.72 % not recorded 2026-09-12 0 slipped ticks; 2 400 samples forwarded, 0 gaps, 0 drops
The install default 2.69 % 13.2 MiB 2026-09-18 the figure above

Each campaign in one entry: the device, the RouterOS version, the date and the conditions, the figures the guides quote from it, and what it found. A figure on a guide links its campaign here.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · 300 s windows at steady state (ring full), full source set, the shipped configuration — a 60 s ring, the memory limit derived from it and the default 64M container cap — with the collector forwarding to InfluxDB 3

  • read.under2ms 97.7 %
  • read.floorOver5ms 1.45 %
  • read.100hz.meanMs 1.2 ms
  • read.100hz.worstMs 17.9 ms
  • load.ordinary 19 Mbit/s
  • run.10hz.cpu 2.69 %
  • run.10hz.rss 13.2 MiB
  • run.10hz.usPerSample 2 685 µs
  • run.10hz.slippedPct 0.000 %
  • run.20hz.cpu 4.61 %
  • run.20hz.rss 15.4 MiB
  • run.20hz.usPerSample 2 303 µs
  • run.20hz.slippedPct 0.000 %
  • run.50hz.cpu 9.63 %
  • run.50hz.rss 23.3 MiB
  • run.50hz.usPerSample 1 926 µs
  • run.50hz.slippedPct 0.000 %
  • run.100hz.cpu 16.83 %
  • run.100hz.rss 45.7 MiB
  • run.100hz.usPerSample 1 684 µs
  • run.100hz.slippedPct 0.017 %
  • run.50hz-floor.cpu 22.56 %
  • run.50hz-floor.rss 25.1 MiB
  • run.50hz-floor.usPerSample 4 511 µs
  • run.50hz-floor.slippedPct 0.027 %
  • run.100hz-floor.cpu 42.70 %
  • run.100hz-floor.rss 49.5 MiB
  • run.100hz-floor.usPerSample 4 270 µs
  • run.100hz-floor.slippedPct 0.593 %

Six windows of 300 s, at 10, 20, 50 and 100 Hz and at 50 and 100 Hz with FLOOR_HZ, none of which lost a sample: Rate ceiling has the table. The WAN carried about 19 Mbit/s on average over the campaign’s 44 minutes, with peaks to 933 Mbit/s.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · 60 s windows at steady state (ring full), full source set, collector forwarding to a file, a Prometheus exposition and InfluxDB 3 at once

Five rate runs at 10, 50 and 100 Hz with the collector writing to a file, a Prometheus exposition and InfluxDB 3 at once; every sink reported 0 gaps and 0 drops. The 10 Hz install default of that date cost 2.85 % of one core and 31.3 MiB RSS, with a 300 s ring and a flat 40 MiB memory limit; 50 Hz needed --memory-max 96M and 100 Hz 128M. The WAN carried about 30 Mbit/s, in the evening. The install-default figure the guides quote is from rates-2026-09-18, not from this campaign.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · 10 Hz, 300 s ring full, the same sources in both runs; only the memory limits differ

  • gc.tight.cpu 9.38 %
  • gc.tight.us 9 374 µs
  • gc.roomy.cpu 1.39 %
  • gc.roomy.us 1 388 µs

The same agent, ring and sources under a 14 MiB soft limit and then with room: Cost by configuration has both rows.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · one ring line with every source on, privileged, 10 Hz, 4 cores, IRQ_TOP_K=8 — including the PMU and the sampler's own timing, which the 2026-09-12 measurement predates

  • ring.lineKB 3.5 kB
  • ring.lineBytes 3 230 B
  • ring.lineChargedBytes 3 456 B
  • relay.maxBatch 13

3 230 B measured, charged from Go’s 3 456 B size class. The 2 439 B of 2026-09-12 understated the ring by 35 %, and the agent ran at 32.9 MiB of RSS where 24.3 was available. A line with no PMU costs 35 % less. Line sizes at today’s default per-source floors have not been measured.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · the mean pre-encoded ring line with every source of that date, the slow sources refreshed at 1 Hz, privileged, 10 Hz, 4 cores, IRQ_TOP_K=8

The mean line before the PMU and the sampler’s own timing were in it, rounded to about 2.4 kB on the guides of that date.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · a busybox shell loop reading the agent's file set of that date, seven files, at 10 Hz, one fork per iteration, in a container on the router; two 60 s runs

  • busybox.cpu 2.40–2.48 %
  • procread.ms 0.77 ms

Two 60 s runs, 2.40 and 2.48 % of one core, from the cgroup’s cpu.stat over /proc/uptime; the reads themselves took about 0.77 ms per sample. Not re-measured against the agent’s source set of today: Agent cost.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.1 ·  · one ssh connect, for its duration

  • ssh.connectCpu 20–27 %

Measured before the agent existed, on RouterOS 7.24.1. On 7.24.2, on 2026-09-11, the same cost showed in /tool profile as 17–33 % in one or two snapshots, not a measured window: SSH cost.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · /tool profile duration=60s cpu=total, once with the collector stopped and once with it running --interfaces bridge,ether1,PPPoE_DIGI --counters-every 10s --api-every 1s; the first five one-second snapshots of each profile dropped, because they hold the SSH connect that asked for it

One 60 s profile per condition, no spread: API tier cost.

· arm64 image tar of the published v1.2.1 release, the file the router loads

  • image.size 6.38 MiB

Measured with ls -l on the v1.2.1 release asset, its checksum verified against the signed checksums.txt: 6 690 304 B, of which the agent binary is 6 684 832 B. The armv5 and armv7 tars are 7 214 592 B and the amd64 one 7 222 784 B. CI’s agent-size job holds the agent binary, which is all the image carries, under 8 MiB on arm64, armv7 and amd64.

date not recorded · agent image size of the 1.0.0 release

  • image.size.v100 6.1 MiB

The figure the 1.0.0 release notes give; v1.0.0 was tagged on 2026-09-16. The images between it and 1.2.1 have no measurement of their own.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · /proc/pressure and /proc/schedstat absent

The irq column of /proc/stat is always 0 on this kernel, so hard-IRQ time is inside system. USER_HZ is 100, so a tick is 10 ms and the busy-time steps on the guides are arithmetic from it. The files that are the router’s inside the container were established the same day (Container view).

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · a privileged discovery container reading the host's /proc once, before the agent read slabinfo

  • conntrack.slabDiscovery 6 582

The round that found no tracing path (Kernel and PMU), and whose slabinfo is the one in testdata/proc/rb5009: nf_conntrack 6 582 active of 8 075.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · privileged=yes does not change the network namespace

Inside the container /proc/net/dev counted the veth, 4 packets while the router forwarded millions, with and without privileged=yes.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · a 10.5 h capture at 50 Hz over one idle night, clock pinned

  • thermal.quantumC 0.42 °C

The capture that set the 6 Hz floor /proc/slabinfo is slowed to, the rate its fastest cache, nf_conntrack, changed. The memory levels moved about 24 times a second. It is one board, one idle night, its clock pinned at the maintainer’s setting: “cpufreq never changed” means it did not change that night.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · /system/resource polled at 1 Hz against the agent's per-core busy ratio at 10 Hz, two separate hours

  • cpuLoad.window1 1.0 s
  • cpuLoad.delay1 0.6 s
  • cpuLoad.r1 0.9825
  • cpuLoad.samples1 3 499
  • cpuLoad.window2 1.1 s
  • cpuLoad.delay2 0.1 s
  • cpuLoad.r2 0.9734
  • cpuLoad.samples2 3 594

The best fit in each hour: a trailing mean of 1.0 s that reached the API 0.6 s late (r = 0.9825 over 3 499 API samples), and a mean of 1.1 s that reached it 0.1 s late (r = 0.9734 over 3 594). /system/resource reports cpu-load as an integer percent, and the API series was correlated against the agent’s per-core busy ratio, the same /proc/stat jiffies, for a range of window lengths and delays. Widening the window only made the fit worse: 1.5 s gave 0.955, 2 s 0.919, 5 s 0.822, 8 s 0.791. A sixty-second average is ruled out twice: its correlation is 0.238, and at the sharpest load step of the day the kernel went from 5 % to 27 % in one second and cpu-load from 5 to 26 in that same second, then from 22 % to 6 % on the way down just as fast. MikroTik’s /system/resource documentation defines cpu-load as the percentage of used CPU resources, all CPUs combined, and names no window (read on 2026-09-24).

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · a 60 s record at 10 Hz, 600 samples over 59.9 s, a RouterOS script loop started over ssh and the router log added as markers

Exactly 600 samples, 0 gaps and a clock skew of −7 ms. A scripted RouterOS loop (:for … 400 000) showed as one core’s worth of load at 100 % from t = 21.0 to 25.8 s, with the sample at 25.8 s reading about 60 %, onset and offset resolved to 100 ms. The router’s log markers explained a 2 s plateau at 15–17 s that nobody had caused: the ensure-ipv6-nd-prefix scheduler. Log times came over the API as full dates (verified).

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · a 70 s record at 10 Hz, 700 samples over 69.9 s, the router otherwise at rest, three notes typed into record's terminal

The recording the chart on Idle baseline is drawn from.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.4 ·  · three GET /stream connections against the installed agent at 10 Hz, opened from the operator host on the LAN and held until the server closed them

  • stream.lifetime 30.01–30.06 s
  • stream.lines 304–306

Agent 1.0.9 at 10 Hz. Each connection ended at the server’s 30 s write timeout, whatever it still had to send, so a /stream consumer reconnects with the last seq it saw.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · /tool fetch output=user called over the binary API; each call took either about 3 ms or about 1 s, and about half took 1 s

  • relay.replyMaxBytes 64 512 B

/tool fetch output=user truncated the body silently at 64 512 B for 64 K, 256 K, 1 M and 4 M bodies.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · one kernel sample at 10 Hz rendered to InfluxDB line protocol, with the agent's source set of that date

About 1.2 KiB a sample, the figure the InfluxDB and Telegraf sinks size their byte budgets from. Not measured above 10 Hz, and not re-measured against today’s source set.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · agent at 10 Hz in an ephemeral privileged container

  • loop.eventsBefore 1.49 /s
  • loop.eventsAfter 0.03 /s
  • loop.helloGap 2.00–2.01 s
  • conntrack.slab 6 287

The readings the case studies are built on, from an aarch64 kernel on a board with 1 GB of RAM; where a fault was provoked, the case study says how. Three case studies also carry a chart from the reference InfluxDB store, whose history begins on 2026-09-19: the loop’s return, the conntrack cross-check and the port losing frames. The port losing frames is a separate campaign a week later, read from the API tier’s port counters rather than from the agent. time_squeeze was nonzero even at idle. While the layer-2 reflection was live the kernel log ran at 1.49 /s, and a “blocking state” then “learning state” pair arrived inside one 100 ms tick at the same level, so their order within the tick is the signal. The conntrack table held 6 287 entries, 0.65 % of the 966 656 ceiling the kernel reported on 2026-09-14.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · 3 738 704 per-CPU samples over 24 h at 10 Hz, softnet time_squeeze per sample, 0 drops in the whole window

  • squeeze.backgroundOne 11.2 %
  • squeeze.backgroundTwo 1.2 %
  • squeeze.exactlyThree 0.21 %

The 24 h ended at 07:00 UTC on 2026-09-16. time_squeeze was 0 in 87.3 % of per-CPU samples, and a trailing window of that distribution has a 90th percentile of 1, so “above p90” is met by any 2. On 2026-09-15 a trigger condition that fires on any squeeze fired 92 times in twenty minutes (RouterOS 7.24.2), which is why squeeze is not a default trigger.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · the rule replayed over 6 h of stored samples, 863 944 rows, 4 CPUs

  • microburst.firesPerHourAtTwo 77.7 /h
  • microburst.firesPerHourAtThree 0.5 /h

At a floor of 3 the rule still flagged 88 samples for the burst counter and for derived.burst.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · the connection table counted over the RouterOS binary API, /ip/firewall/connection/print count-only, ten calls

  • conntrack.api 6 212
  • api.conntrackMs 1.3 ms

The count was 6 212, the day before the agent’s slab read 6 287: the same order of magnitude, which says nothing about whether the two track. The ten calls took a median of 1.3 ms, the fastest 1.1 ms and the first 71 ms, cold; the time of day is not recorded.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · RouterOS port and interface counters read over the API, cumulative since boot or since each port's last counter reset

ether1, 2.5 GbE to a NAS, had received 255.8 GB on the wire (rx-bytes) and handed 29.7 GB of it to the CPU (driver-rx-byte) since the port’s last counter reset; the switch chip forwarded the rest in hardware. Every switch port read a fast-path share of about 100 %, fp-rx-byte equal to driver-rx-byte within a few kB. bridge had fast-pathed 211.9 GB of the 663.0 GB it took to the CPU since boot (32 %), and PPPoE_DIGI 99.97 %. fp-tx-byte stood at 0 on every interface after hundreds of GB transmitted (verified).

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.4 ·  · ether1, the 2.5 GbE port to the NAS, over 10 s counter intervals; the before figures are the three hours preceding the fix and the after figures the 39 minutes following it, at the same load

  • overflow.shareBefore 0.502 %
  • overflow.perHourBefore 5 399
  • overflow.occupancy 0.36 %
  • overflow.burstRate 8.96 Mbit/s

The case study is Port losing frames.

Measured with no campaign of its own: on 2026-09-15 a burst peaked at 736 packets in one 20 ms sample, a 50 Hz sample, against a median of 28.

Yes-or-no facts about RouterOS the guides rely on, each checked once, with where and when.

Over the binary API, /container/print returned every property of every container, cmd and envlist included, to a user with only read,api, the user mikroscope; only the property names were printed, and reading the values in /container/envs as that user was not checked (RB5009UG+S+, RouterOS 7.24.2, ).

/tool fetch and /tool profile both require the test policy (RB5009UG+S+, RouterOS 7.24.2, ).

In a RouterOS find, address and port attributes match only when their values are quoted, and a bare word is read as a variable name: over the same 15 dstnat rules (RouterOS 7.24.4, 2026-09-21), protocol=tcp found 0 and protocol="tcp" found 10 (RB5009UG+S+, RouterOS 7.24.2, ).

A dst-nat rule plus a forward accept rule reach the agent from the LAN, and both are removed by their tag (RB5009UG+S+, RouterOS 7.24.2, ).

The LAN reaches the veth once the veth joins the interface list LAN and the /30 joins the address list LANs (RB5009UG+S+, RouterOS 7.24.2, ).

/container/remove returns before the container is gone, and a /file/remove issued meanwhile does nothing, silently (RB5009UG+S+, RouterOS 7.24.2, ).

With ignore-remote-image-change=no, removing the image tar makes RouterOS stop and remove the container and extract it again minutes later (RB5009UG+S+, RouterOS 7.24.2, ).

doctor, install, status, upgrade and uninstall, run in that order, left the router's /export byte-identical, its # header lines aside, compared by hash in memory and never written to disk (RB5009UG+S+, RouterOS 7.24.2, ).

A tmpfs disk holds the image tar and the container root with no NAND writes: write-sect-since-reboot stayed at 58 279 across install, run and removal (RB5009UG+S+, RouterOS 7.24.2, ).

privileged=yes drops the container's user namespace but not its network or PID namespace (RB5009UG+S+, RouterOS 7.24.2, ).

A registry host inside remote-image= overrides /container/config registry-url, and docker.io is pulled as registry-1.docker.io: with registry-url=https://registry-1.docker.io, a reference on registry.invalid was logged as registry=registry.invalid and failed with resolving error, while the agent's 1.2.2 image named as docker.io/… and as registry-1.docker.io/… was logged as registry=registry-1.docker.io and ended in download/extract done, its arm64 layer 2 799 648 bytes; /container/config read the same afterwards. The containers were created in a temporary veth, never started and removed again (RB5009UG+S+, RouterOS 7.24.4, ).

With the /container/config username and password cleared, RouterOS pulled the agent's 1.2.2 image from Docker Hub anonymously, both as remote-image=registry-1.docker.io/… and without a host through registry-url=https://registry-1.docker.io, each ending in download/extract done 5 s after the add; the configuration was then restored and verified identical (RB5009UG+S+, RouterOS 7.24.4, ).

Over the API, /log/print returns an entry's time as a full date and time with no zone, 2026-09-12 02:21:24, and accepts that format in a ?>time= query (RB5009UG+S+, RouterOS 7.24.2, ).

/interface/monitor-traffic returns rx-drops, tx-drops and tx-queue-drops per second and no error keys at all (RB5009UG+S+, RouterOS 7.24.2, ).

ether1 had counted 652 364 rx-overflow events, and growing, in its port counters while monitor-traffic returned no error key for that port: that count reaches a consumer only through the port counters (RB5009UG+S+, RouterOS 7.24.2, ).

An ether port in a bridge counts its wire, frames the switch chip forwarded in hardware included, and the bridge counts its CPU side, so neither is a subset of the other: since the port's last counter reset ether1 had received 255.8 GB on the wire (rx-bytes) and handed 29.7 GB of it to the CPU (driver-rx-byte) (RB5009UG+S+, RouterOS 7.24.2, ).

fp-tx-byte read 0 on every interface after hundreds of GB transmitted, while fp-rx-byte counted; why RouterOS leaves it at 0 is not established (RB5009UG+S+, RouterOS 7.24.2, ).

A privileged container given the host's /proc, /sys and / as bind mounts read zero PIDs in the host's /proc and found no class/net in its /sys, so the namespaces held; the host's / did mount, and it exposed the RouterOS flash filesystem, configuration and files, secrets included. Run with the maintainer's consent (RB5009UG+S+, RouterOS 7.24.2, ).

RouterOS presents the /container/config username and password for a reference whose host is registry-url as written: with a deliberately wrong credential, the 1.3.1 agent image named as registry-1.docker.io/jmrplens/mikroscope-agent failed with auth error under registry-url=registry-1.docker.io, and was pulled anonymously under https://registry-1.docker.io, the value MikroTik's examples use, and under the same with a trailing slash. A router given a credential did not fall back to an anonymous pull (CHR x86_64, RouterOS 7.24.4, ).

No WebFig field for ignore-remote-image-change

Section titled “No WebFig field for ignore-remote-image-change”

WebFig's New Container and container edit forms show no field for ignore-remote-image-change, with File or Remote Image set or not; the terminal sets it (CHR x86_64, RouterOS 7.24.4, ).

A /container/set that leaves restart-policy out (of ignore-remote-image-change, of comment, of logging) put it from on-failure back to always and changed nothing else /container/print detail shows; a set that names it kept it, and so did Apply or OK in WebFig's edit form. The CLI never runs /container/set (CHR x86_64, RouterOS 7.24.4, ).

WebFig shows a /container/envs entry's value in clear, in the Envs list and in its form, for the key TOKEN too (CHR x86_64, RouterOS 7.24.4, ).

WebFig's System menu has no Device Mode page, so device mode is set from the terminal (CHR x86_64, RouterOS 7.24.4, ).

:put ([/tool/fetch url="http://172.30.10.2:9123/healthz" output=user as-value]->"data"), run on the router, printed the agent's /healthz JSON, in an ssh session, in WebFig's terminal and at the end of a pasted script (CHR x86_64, RouterOS 7.24.4, ).

The first interactive login after a reset asks Do you want to see the software license? [Y/n]: before the ] > prompt. A script pasted into that question lost its first lines: one or two header comment lines, then a fragment run as a command (bad command name . or syntax error); the install script's { … } block still ran and installed, and the uninstall script still removed everything (CHR x86_64, RouterOS 7.24.4, ).

What no run, measurement or check here covers. A guide that states one of these as expected behaviour links this list.

  • A second board of any kind. The hEX S (2025), 32-bit RouterOS on an ARM64 chip, which is what the agent’s linux/arm/v5 build is for, has not arrived, so the 32-bit counter-wrap path and neither 32-bit ARM image have run on RouterOS. A board report from another board is what changes this.
  • x86 RouterOS on hardware: the amd64 image has run only in QEMU, on the lab’s CHR x86_64 and on RouterOS x86 installed from MikroTik’s ISO.
  • Any RouterOS before 7.24 beyond doctor’s version check, which S17 ran on a 7.23.7 CHR, and any of the 7.24.2 figures re-measured on 7.24.4.
  • Anything that needs a reboot of the RB5009, which waits for a maintenance window. So a persistent install surviving a reboot with start-on-boot=yes is untested on hardware. On the lab’s CHR it answered again within 90 s of a power cut (S7, on both architectures).
  • Which files are namespaced, on another RouterOS version or another board: every row of Container view was read on one RB5009 on 7.24.2.
  • A board with hwmon sensors, or a later RouterOS that builds containers differently: the sensor set, the 38 capabilities and the uid mapping were read on one RB5009 on 7.24.2.
  • The port mapping on another board: the shift by one, the position of the SFP+ cage and the reuse of port 7 were observed on one RB5009UG+S+ on 7.24.2, and the agent does not apply them to any other board.
  • The PSI and schedstat paths on a router: the RB5009’s kernel has neither. The PMU counter set is known on the RB5009’s Cortex-A72 and the amd64 development host only.
  • Every source cadence on another device or workload: the cadences are the RB5009’s, on 7.24.2. Re-measure with FLOOR_HZ equal to the sampler rate.
  • Which loss keys and counters another RouterOS version or board returns.
  • What --goarm 7 saves against the ARMv5 build: no ARM hardware has run either.
  • Envlist entries on a RouterOS before 7.24: mikroscope writes key=, because name= failed on the RB5009 on 7.24.2, and how an earlier 7.x takes either was not tried.
  • Traffic heavier than the RB5009’s ordinary load: about 19 Mbit/s on the WAN during the 2026-09-18 campaign. No rate is claimed for any other board.
  • On the RB5009, a pull from GHCR with no registry credential set, and any GHCR pull with the host inside remote-image=: its one GHCR run took the host from registry-url. The lab’s CHR did both on 7.24.4 (Install routes tested).
  • Which credential RouterOS presents to a host named only in remote-image= when registry-url names another. The one run with a host that differed from registry-url’s, on 2026-09-24, named registry.invalid, which never resolved. With the same host, the lab showed RouterOS presents it only when registry-url is written as the bare host (verified).
  • On the RB5009: a whole install or upgrade that sends the full reference, a registry-url at its factory default, and the full reference on any RouterOS but 7.24.4. The lab’s CHR ran both with 1.3.1 on 7.24.4.
  • A pull through a mirror or pull-through cache named in the reference.
  • install, status and uninstall of 1.4.0 and later on any hardware, and the guarded script: they have run only on the lab’s CHRs. On the RB5009, 1.6.0’s upgrade read the install manifest and the architecture (--arch auto) from the router, and its doctor ran every check.
  • --ssh-option and MIKROSCOPE_SSH_OPTIONS against any router: the lab’s CLI connects through the lab’s own ssh_config, and only unit tests pass the option.
  • --extract-timeout running out, which keeps the tar, on any router.
  • Doctor’s registry host comparison against a real router: it has run only against fake router answers in the tests.
  • Whether a read,api user can read the envlist values in /container/envs.
  • The byte-identical round trip on hardware: on any board but the RB5009, on the RB5009 with any RouterOS but 7.24.2, with install and upgrade run without --ephemeral, or with the current container settings (privileged=yes, memory-max=64M, the envlist entries MEM_LIMIT_MB, CAPTURE_MB, TRIGGERS and FLOOR_HZ): on the RB5009 it ran with memory-max=32M. install, status and uninstall have run there on 7.24.4 since. On the lab’s CHR 7.24.4, make roundtrip runs it with --ephemeral (2026-09-26); S2 and S3 run install, upgrade and uninstall without it, with privileged=yes, memory-max=64M, MEM_LIMIT_MB and CAPTURE_MB, and S5 imports a script that also writes TRIGGERS and FLOOR_HZ; each ends with /export equal to its start, RouterOS’s keymat-provider line aside (both architectures, 2026-09-27).
  • Firewalls other than the RB5009’s, where on 2026-09-11 the two list memberships were enough, and the lab’s profiles. A firewall with other drop rules in raw, input or forward may drop the container’s traffic elsewhere, and install adds nothing for that beyond the two memberships. Whether a router’s factory default configuration carries the two raw rules was not checked.
  • --expose from outside the LAN: only a LAN host reaching the router’s LAN address was tested. Whether anything outside the LAN reaches that address depends on the rest of the firewall, and no such path was tried.
  • Winbox: the WebFig install page says Winbox has the same menus and fields, and no Winbox ran. The --expose rules through WebFig’s NAT and Filter Rules forms, a whole script pasted into WebFig’s or Winbox’s terminal, and WebFig on any RouterOS but 7.24.4.
  • The script generator’s scripts, the manual terminal install and the WebFig install on the arm64 lab or on the RB5009.
  • A recording above 10 Hz: the lossless 20, 50 and 100 Hz runs were the collector’s, which uses the same batch sizing. A recording through the relay, at any rate.
  • The relay’s throughput. From /tool fetch round trips of about 1 s for half the calls and about 3 ms for the rest (measured on the RB5009 on 7.24.2), arithmetic gives on the order of 26 samples a second, less when the slow calls cluster. The relay cap and the start-up warning are read from the code (2026-09-15), not measured against a device. The relay’s cap and its 1 s share on any RouterOS but 7.24.2.
  • The cost of a trigger fire on the device; 3 µs for a 10 s window at 10 Hz is the design’s estimate. How the agent behaves under a sustained trigger storm. Any capture at 50 or 100 Hz. The capture sizes on Triggered capture are arithmetic from the line size, not sizes of captures taken on the RB5009.
  • Sinks fed from the RB5009 into a running backend: only file, Prometheus and InfluxDB 3. Loki, an OTLP receiver, carbon, Elasticsearch, OpenSearch, Telegraf, PostgreSQL and TimescaleDB have had no router samples in a recorded run. OpenSearch, TimescaleDB’s hypertables and standard output have never run against a real store; only the byte-contract tests cover them.
  • An Elasticsearch or OpenSearch that requires authentication. The container suite’s Elasticsearch runs with security off, so the Authorization header built from MIKROSCOPE_ELASTIC_AUTH, basic auth for user:password and ApiKey otherwise, has no recorded run against one: neither from the sink nor from the Elasticsearch datasource that --grafana and dashboards publish build, which sends the same header from 1.5.0. Unit tests pin both forms, and the end-to-end suite the sink’s ApiKey against a fake receiver. Nor has that datasource run against an OpenSearch, with authentication or without.
  • A live TimescaleDB: the create_hypertable calls follow TimescaleDB 2.x’s documented signature and are not verified.
  • The SQL sink’s size on a router. The fixture figures (1 375 B of SQL for a kernel event against 716 B of line protocol, 1 138 B for an API event against 608 B, about 14 KiB/s at 10 Hz plus the 1 Hz API tier after a 5.6 KiB header, 2 749 B with the privileged sources) cover eighteen of the forty-three tables the sink writes: eleven were added after the fixture, and load, stat, buddy, mtd, api_ifcounter, api_ifinfo, trigger, derived, derived_iface, detection and the four device tables were not exercised. They predate the port and kind columns of mikroscope_event.
  • --sql - | psql with a psql that falls behind a 10 Hz agent, which blocks the pull loop.
  • The standard-output budget for json, whose lines are larger than line protocol’s.
  • An OTLP render of a four-core sample with a real interrupt top-K, a four-core sample in the Elasticsearch format, and the Graphite budget above 10 Hz.
  • A write to InfluxDB 2’s /api/v2/write. The line-protocol size above 10 Hz or with today’s source set. Whether turning off the HTTP client’s own resend avoids the duplicates of 2026-09-13.
  • Loki delivery during a kernel-log storm.
  • Any Prometheus scrape interval but 5 s, and any Prometheus but 3.14.
  • The collector’s constant 10 Hz against a faster agent: the burst baseline and the top-K pruning are read from the code; the 1/600 EWMA weight and the 36 000-sample eviction were never watched happening, since no run made an interrupt line fall out of the top-K and stay out.
  • The zero deltas written beside an absent fast-path share (read from internal/derive/derive.go), and a ceiling or cadence change without a hash change reaching the sinks at the next five-minute repeat (read from capsHash).
  • The per-packet PMU cost compared across two router configurations. Whether the burst baseline’s ten seconds suit another board or traffic mix: it was tuned against one RB5009 on one day. A tx fast-path share from live counters.
  • Detections provoked with the rules running: no OOM kill in the container, reboot, link flap, conntrack flush or storm, thermal excursion or IPC collapse. Whether reboot fired on the 2026-09-19 reboot. A provoked flap with link-flap running; the flaps of 2026-09-15 were not.
  • Whether RouterOS computes cpu-load’s one-second window on a wall clock or on jiffies. Whether the API’s conntrack count and the slab count track each other.
  • None of the six tools on Compared with alternatives was run for that page. Every cell in their rows is what their own documentation or source code says, read on 2026-09-24: mktxp at commit 1e4412a (dated 2026-09-21), mikrotik-exporter at 428dbfc (dated 2026-07-03), and the MIB file MikroTik publishes for RouterOS 7.24.4. What SNMP polling, The Dude, Graphing, the Profiler, mktxp or mikrotik-exporter costs a router, how often each can usefully poll one, and what RouterOS puts in hrProcessorLoad were not measured.
  • Any Grafana but 12.3.0, 12.3.2, 13.2.1 and 13.2.2. On 12.3.0 the only renders are those of 2026-09-25; the rest of the dashboard was checked there by query only.
  • dashboards publish against any Grafana but 13.2.1, and over anything but what forward --grafana had just made there: in the container suite on 2026-09-27 it ran once per store, after the collector, with no credential in any datasource, and found the folder and each datasource as the collector had left them, reported unchanged. With dashboards publish, a run of more than one store, a store that fails while the rest carry on, a datasource that carries a token or a password (written again at every run and reported updated), and --grafana-dry-run have run only in the unit tests, against fake Grafanas.
  • The height of the thirteen-panel Overview in a browser; the committed JSON makes it 31 grid units.
  • Byte-for-byte regeneration of the dashboard files added after 2026-09-15.
  • The Elasticsearch and Graphite annotations and detection layers in Grafana: only their generated JSON is checked, by unit tests.
  • A render pass over the two port-event panels and the interface inventory table.
  • A rule resolving: the loop on the RB5009 ended on 2026-09-23, but mikroscope-l2-loop returning to Normal has not been checked in Grafana’s state history. Neither mikroscope-l2-loop nor mikroscope-port-link-down has been watched going pending, firing and resolving.
  • The two rules 1.2.0 added inside any Grafana, their PromQL beyond syntax, and the wake-up rule’s PostgreSQL SQL.
  • The Prometheus and PostgreSQL alert files in any Grafana, and any PostgreSQL alert query against a real PostgreSQL: a unit test reads the schema, not a database.
  • Notification delivery: no contact point was configured, so what was watched is each rule’s state.
  • A store missing a table or column a rule reads: the query is expected to fail and execErrState: Error to apply instead of OK. What Grafana then does is its Error and No Data behaviour, not tested here.

Checked against the code on 2026-09-24, and the end-of-run summary’s stream again at 1.3.1 on 2026-09-26:

  • forward’s end-of-run summary goes to stdout, the same stream the --stdout sink writes records to, so forward --stdout=lp | telegraf ends every run with lines the consumer cannot parse.
  • A wrong --token does not fail the run. Each pull is logged as 401 Unauthorized, and the run then ends at its --for deadline with exit status 0 and forwarded 0 kernel samples. The health check forward makes before it starts pulling reads /healthz, which needs no token, so it passes. The message is right; the exit status tells a unit file nothing.
  • Duplicate sequence numbers in the 2026-09-13 overnight run. The 50 Hz run wrote repeated seq values into InfluxDB; the dashboards name those rows “duplicate sample (seq repeated)”. The leading explanation, which is unverified, is Go’s HTTP client replaying a POST on a dead pooled connection after the server had committed it. The InfluxDB, Loki, OTLP, Elasticsearch and Telegraf sinks refuse that replay, so a dead connection is an error the sink retries and counts. No run since that night shows whether the duplicates are gone.
  • A negated tag inside an OR returns nothing, silently, on InfluxDB 3. On 2026-09-23, on the reference store (InfluxDB 3 Enterprise since 2026-09-19; the version was not recorded with the query), NOT (port='ether2' AND kind IN (…)) over mikroscope_kmsg returned 0 rows where 14 matched, and so did its De Morgan form and the explicit OR. It returned 14 once port='ether2' was also filtered outside the OR. The shipped panels and rules do not use the pattern.
  • Grafana 12.3.0 never sends the InfluxDB detections layer’s query, with or without the table: its toggle kept a loading indicator, switching it off and on and refreshing sent nothing, and no marker was drawn, on the full dashboard and on a one-panel copy (2026-09-25). Grafana 13.2.1 sent it at once over the same store. Why, and whether 12.3.2 does the same, has not been looked into.

What the published 1.3.1 CLI and agent image met on the lab’s CHR, RouterOS 7.24.4, on 2026-09-26. Each item ends with what the code after 1.3.1 does, checked in the lab on 2026-09-27; 1.4.0 ships that code, and the changelog’s 1.4.0 section lists each change.

  • doctor with its defaults fails three checks on a CHR. --arch defaults to arm64, which fails on x86_64 only; the interface list LAN does not exist; and the address list LANs “has entries” fails on an empty or missing list, although its own fix says an empty list is fine if no such rule exists. A router without that rule passes only by giving the list an entry, or with --no-doctor, which skips every other check. After 1.3.1: --arch defaults to auto and reads the router’s architecture, an empty address list is no longer a failure, and a missing interface list, still MISSING with the defaults, has a fix that offers --iface-list none (S1).
  • A built-in interface list such as static, all or dynamic passes doctor’s existence check, and RouterOS refuses to add a member to it. After 1.3.1: --iface-list refuses all, dynamic and static; RouterOS answers cannot add to builtin list.
  • --ephemeral needs a tmpfs disk, and a CHR lists no disk. Doctor’s fix, /disk/add type=tmpfs tmpfs-max-size=64M slot=tmpfs, worked, and the install went to tmpfs/mikroscope/mikroscope with start-on-boot=no. After 1.3.1: the same, and doctor also warns when start-on-boot is forced on a root on a tmpfs disk, which a reboot empties.
  • uninstall races the container’s stop. It stops the container, waits a fixed 4 s and removes it, and the agent took 0 s to stop six times, 4 s four times and 5 s once (RouterOS log, 1 s resolution): 5 of 15 first attempts on x86_64 failed with failure: cannot remove running. With a client on /stream it failed every time on both architectures, because the agent’s HTTP shutdown waits up to 5 s for it. The steps after it still ran, and a second uninstall cleaned up every time. After 1.3.1: the removal waits up to 30 s while the container is running or stopping, and every first attempt was clean (Lab timings).
  • An empty mikroscope directory stays in /file after every uninstall: the parent of the container’s root-dir, which the ownership count does not include. After 1.3.1: every install route writes an install manifest, and uninstall removes what it lists, then the manifest and the mikroscope/ directory when nothing else is in it; no mikroscope path was left on /file after any route (F4).
  • The probe after install misreads a running container on 7.24.4. It asks /container/find … status="running", and status is not a property of /container there (/container/find status=running answers bad parameter status), while the flag running reads 1. So when the probe failed, the CLI said “the container is not running on the router” with the agent running and about to answer. Seen twice on arm64. After 1.3.1: the probe reads the running flag.
  • The plan listing named a tar with --remote-image: image=mikroscope.tar, a file install never uploads, and it printed the upload or the pull as a step of its own under the container step’s number, so two steps shared a number. After 1.3.1: image= names the reference the router pulls, and the pull or upload is a line of the container step.
  • The agent token travelled on ssh’s command line, where the host’s process table showed it while the command ran: the 1.3.1 code showed it on two command lines. After 1.3.1 a command that carries it goes to ssh on standard input; reading the process table every 2 ms through an install and an upgrade with --expose found it on none, on both architectures.
  • RouterOS’s words were lost when its ssh exited 1: the error read ssh "<script>": exit status 1, and RouterOS’s message, on the lines after it, was dropped by uninstall’s skip line. RouterOS’s ssh exited 1 or 0 on the same failure, about half and half. After 1.3.1 the error starts with what RouterOS printed.

Two readable files on the RB5009, read there on 2026-09-15, are not collected: /proc/cmdline, which carries board=5009 ver=7.24.1, a second device identity without the API, and the hardware watchdog at /sys/class/watchdog/watchdog0.