Skip to content

RouterOS API tier

forward can hold one persistent session to the RouterOS binary API (a second one when the kernel tier itself comes through the relay) to read what the container cannot see: per-interface traffic and counters. Choose how much it asks with --api-mode.

Most of what the API tier can fetch, the agent reads itself:

  • /system/resource and /system/resource/cpu are RouterOS’s one-second average of the same /proc/stat jiffies the agent differences at 10 Hz.
  • /system/health’s temperature is /sys/class/thermal, readable from the container.
  • The conntrack count is the global nf_conntrack slab cache, which the agent reads under privileged=yes; see Conntrack without the API.

What is left is per-interface bytes and packets. They live in the router’s network namespace and stay out of reach whatever the container is given; see Container visibility.

RouterOS cpu-load is a trailing mean of about one second of the kernel’s /proc/stat busy time, and it reaches the API a fraction of a second late (measured). MikroTik’s /system/resource documentation defines cpu-load as the percentage of used CPU resources, all CPUs combined, and names no window. It cannot resolve anything shorter than its own second, which is why mikroscope reads the kernel.

--api-mode picks a preset. It is full unless you say otherwise, and an explicit --api-every, --no-health or --conntrack-every still wins over it.

Mode What it does Use it when
off no API tier: --api-every 0 production, when per-interface traffic already comes from somewhere else
slow one round every 10 s, without /system/health and without the conntrack count; /system/resource, monitor-traffic and the port counters still run on each round production, when you want interface rates too
full one round every second, with /system/health; port counters every 10 s; the conntrack count only if --conntrack-every asks for it experiments and short forward runs

off gives per-interface traffic up; slow keeps it cheaply. Under slow two dashboard panels are blank by configuration rather than by device — /system/health and the RouterOS connection count — and they name the flag that fills them.

Every API-tier command runs on the tier’s one session, one at a time, with a 15 s timeout each. The slower cadences are checked on each round, so none of them runs more often than --api-every.

Command Asks for Runs Off when
/system/resource/print cpu-load, free-memory, total-memory, free-hdd-space, uptime, version every round the tier is off
/system/resource/cpu/print cpu, load, irq, disk every round the tier is off
/system/health/print name, value every round --no-health, or slow
/interface/monitor-traffic with interface=<list> and once the four traffic rates and whichever loss rates the router returns every round, one call for all interfaces --interfaces is empty
/interface/print with .proplist=name,default-name,type,comment,actual-mtu what each interface is: its current name, the board’s name for it, its type, its comment and its MTU at start, before the first kernel pull, then every --labels-every (5 min) the tier is off
/interface/list/member/print with .proplist=list,interface which interface lists name each interface, which is its role with the read above the tier is off
/interface/bridge/port/print with .proplist=interface,bridge which bridge each port belongs to with the read above the tier is off
/interface/ethernet/print stats and /interface/print stats-detail every numeric counter of every interface every --counters-every (10 s) --counters-every 0
/ip/firewall/connection/print count-only the connection count every --conntrack-every --conntrack-every 0, the default
/log/print with .proplist=topics,message and ?buffer=memory the lines RouterOS wrote about the boot it is in (Boot log) once per reboot, when the agent’s kernel boot id changed the tier is off
  • A command that fails leaves its part of the sample empty and records why; the rest of the round still stands. The failures are logged on standard error as api tier: …, written as lines to Loki and rows to SQL, and counted on OTLP — never turned into a value.
  • An /interface read that fails keeps the inventory already held, so a transient error does not blank every panel’s label, and is reported as inventory: …. The list and bridge reads are best effort; without them the inventory still carries names, types and comments.
  • The API-tier sample is stamped with the collector’s clock plus the skew measured against the agent, so it lands on the agent’s timeline.

The three inventory reads say what the numbers belong to. Per interface they give:

Field Source Example
name /interface/print the current name
default name /interface/print the factory ether5 of a physical port; empty for a bridge, VLAN or tunnel
type /interface/print ether, bridge, vlan, pppoe-out, wg, veth, loopback
comment /interface/print the label a panel shows
MTU /interface/print (actual-mtu) 9000
role /interface/list/member/print the interface lists, sorted and comma-joined: WAN, LAN,VPN
bridge /interface/bridge/port/print the bridge it is a port of
  • A bridge member that is in no list of its own takes its bridge’s lists, because that is how a RouterOS firewall rule matches it. A bridge port that names an interface list rather than an interface is not labelled.
  • This is configuration, not telemetry, so it is read once at start — before the first kernel pull, so a kernel-log record is labelled from the first line — and then on the slow --labels-every cadence, never per poll. An edited comment, or a port moved between lists, reaches the dashboard within minutes, and a comment removed in RouterOS disappears here too: the read replaces the inventory wholesale. None of the five properties it asks for can carry a secret.
  • Every monitor-traffic rate row and every per-port counter row carries label (the comment), type, role and bridge, so a panel says what is plugged in rather than a port number, and which numbers may be compared.

Prometheus is the exception by design: a comment is edited by a human, and a changing label would spawn a new series for every edit, so the collector’s /metrics carries one info series per interface — mikroscope_api_interface_info{interface,label,type,role,bridge,default_name} 1, for every interface whether or not it has a comment — and a query joins it:

mikroscope_api_interface_counter_total * on(interface) group_left(label, type, role) mikroscope_api_interface_info

The inventory also gives the kernel log its names. A kernel record names a port as eth1, the board’s port table maps that to the RouterOS default name, and the inventory maps the default name to the current one, so an operator who renamed ether5 to WAN reads WAN on the panel, with the port’s label and role beside it. Without the API tier a kernel record keeps the board’s default name and gets no label.

Flag Default Meaning
--api host:port MIKROSCOPE_API_ADDR the RouterOS binary API, e.g. 192.168.88.1:8728
--api-user MIKROSCOPE_API_USER the API user; its password only from MIKROSCOPE_API_PASSWORD
--api-mode full off, slow or full, as above
--api-every 1s the round’s cadence; 0 disables the tier
--interfaces MIKROSCOPE_INTERFACES, empty comma-separated interfaces for monitor-traffic, e.g. bridge,ether1
--counters-every 10s how often to read every port’s cumulative counters; 0 never
--labels-every 5m how often to re-read what each interface is — comment, type, interface lists, bridge; 0 means the default
--conntrack-every 0 how often to ask the connection count; 0 never, because it is a table scan
--no-health off skip /system/health

Without --api, --api-user and MIKROSCOPE_API_PASSWORD, forward runs the kernel tier alone. When they are given but the session cannot be opened within 10 s it logs api tier: not connected yet, will keep trying: … and attaches the reader anyway: a failed first dial is a warning, not a tier disabled for the life of the process.

In any mode the API tier’s commands are reads: a user with the read and api policies is enough. test is needed only for /tool fetch, which the relay transport and nothing else in the API tier uses. The group, the address restriction and why the credential never lives on the router are on API user.

The API tier keeps exactly what the router returned and invents nothing for what it did not.

  • Loss rates. monitor-traffic may return rx-drops, tx-drops, tx-queue-drops, rx-errors and tx-errors per second, and a sink writes each one only when it came back. A router can return the three drop rates and no error keys at all (verified).
  • Port counters. The per-port counters are a map of RouterOS’s own field names, not a fixed set of fields, and only plain unsigned integers that count something are kept. mtu, actual-mtu, l2mtu, max-l2mtu and sfp-shutdown-temperature parse as integers but are sizes and configuration, and every sink renders this map as a counter family, so they are dropped rather than left for each consumer to know about; the MTU travels in the inventory instead.
  • Two commands, merged per interface. /interface/ethernet/print stats brings the MAC’s typed errors, the collision family, the frame-size buckets and the driver counters; /interface/print stats-detail brings the fast-path fp-* counters, link-downs, tx-queue-drop and the kernel-side totals. A board without collision counters produces no collision entries rather than a row of zeroes that reads as “no collisions”.
  • Errors only the counters carry. A port can count rx-overflow while monitor-traffic returns no error key for it, so that count reaches a consumer only through the port counters (verified).
  • Every interface. The counters cover every interface the router lists, not only --interfaces: the port a fault lives on is often one nobody thought to monitor.
  • Cumulative. The counters run since boot or the port’s last reset, so a consumer differences them. They separate what a switch port received on the wire (rx-bytes) from what reached the CPU (driver-rx-byte); the rest the switch chip forwarded in hardware, and no counter inside the container has a number for it. The collector turns the fast-path counters beside them into a share of the traffic each interface hands the CPU.
  • A switch port and a bridge count different things, which is why the type travels with every row. An ether in a bridge counts its wire, including the frames the switch chip forwarded without the CPU; the bridge counts its own CPU side; a VLAN or a PPPoE link counts what the CPU sent and received. A port and its bridge are two planes, neither a subset of the other: drawn side by side without their type they read as peers, and summed they double-count. Never add them.
  • Interface labels. What each row is labelled with — the comment, the type, the role and the bridge — is the inventory above, and an interface the inventory does not list carries none of them rather than an invented blank identity.

The kernel tier cannot see rx-overflow by construction. A frame the switch chip forwards in hardware never reaches the CPU, so no /proc file on the router has a number for it; it takes the per-port counters, and only the API has those. Port losing frames works through one such port, from the dashboard tile to the fix.

The API process relays each command to the subsystem that answers it (interface-mgmt, config-db), and that is where its cost shows. The profiles and the A/B of a round every second are on Agent cost.

The API is still the costly path. Per-port data comes from the container wherever the container can see it, and configuration is read at start and on the slow labels cadence, never per poll.

The tier holds one socket, so anything that takes the router away takes the tier with it: a reboot, a RouterOS upgrade, an operator restarting the API service. The tier reopens the connection itself.

Failure Reopens the connection Why
transport error: EOF, a broken pipe, a connection reset, a command timeout yes, and retries that one command on the new connection the socket is finished
!trap no the router is alive and refusing the command; asking again only spends its CPU
!fatal yes RouterOS sends it as it closes the session, as MikroTik’s API documentation states
  • Attempts are spaced at five seconds. One round issues four or five commands, so a router that is down would otherwise be dialed several times a second, and every dial carries a login.
  • The inventory is dropped and re-read after a reconnection: an upgrade is exactly when an interface can change its name, type or bridge, and stale labels on fresh rates would be worse than a moment’s gap.
  • A collector that starts while the router is down is the same case seen from the other end: it connects on the first round the router answers.

A socket killed under a running collector is reopened and the command retried inside the same round (tested). The collector says so once per event rather than once per failed command:

api tier: reconnected (1 since start)
api tier: recovered after 137 failed round(s)

and the minute report carries api: N failed round(s), N reconnect(s) while either is nonzero. api counts rounds attempted, so those two counts are what shows a tier in which every command fails. Troubleshooting covers what that looks like on the dashboards.

When the agent’s kernel boot id changes, the router rebooted, and the collector asks RouterOS once what it logged about the boot it is in. It reads the memory buffer alone, which RouterOS empties at every boot: on a router that also logs to disk, /log/print returns the earlier boots’ lines too. It keeps the system lines that begin router rebooted or router was rebooted, and any line that mentions the previous boot.

How it went down The line Topics
/system/reboot from a session router rebooted by ssh-cmd:admin@192.168.88.10/reboot system,info
/system/reboot from a script router rebooted by ssh-cmd:admin@…/script:rb/reboot system,info
/system/shutdown, then power on router rebooted by ssh-cmd:admin@…/shutdown system,info
the power pulled, or a reset button router was rebooted without proper shutdown system,error,critical

The line names the session and the user, or the script, that rebooted the router, or says that it went down without a shutdown (Tested on). The reboot detection carries it, and the collector logs it:

router's boot log: "router was rebooted without proper shutdown"

The connection died with the router, so the read is tried when the reboot is noticed and again after each round until the tier is back, for up to two minutes; the reboot detection waits for it that long. A router that answers with an error is not asked again: the detection then says RouterOS's log could not be read (…). When the memory buffer holds no such line, because it wrapped or no logging rule sends the system topic to memory, it says RouterOS's memory log holds no line about the boot.

--conntrack-every 10s asks /ip/firewall/connection/print count-only at that cadence; one call took 1.3 ms at 6 212 entries. It is off by default because it is a table scan over an API session, and under privileged=yes the agent’s nf_conntrack slab count is the same population read from a file at the sampler’s rate. The two will not match exactly — they are sampled at different instants, and the slab counts objects the allocator still holds. Whether they track is not measured: run the count and read the slab in the same minute on your own device to compare them.

The boards and RouterOS versions the tier has run on are under Devices and versions. Which loss keys and which counters another version or another board returns is that router’s to say; the sinks carry whatever comes back and nothing else.