Skip to content

Port names

The kernel log names netdevs (eth0, eth5), RouterOS names interfaces (ether1, sfp-sfpplus1), and they do not agree: eth1 in the log is not ether1. Map a kernel-log record to the cable it is about, as the loop case study had to.

RouterOS exposes no mapping, and the container cannot read one. Network devices are namespaced, so /sys/class/net inside the container shows only lo and the veth, and /sys/class/mdio_bus and /sys/class/phy name no port either. privileged=yes does not change that (verified). RouterOS’s names live in RouterOS’s configuration, not in the kernel.

Toggle a port that carries nothing and read the name the kernel prints. It is one safe step: a dead port interrupts no traffic.

  1. Find a genuinely dead port, with zero packets in both directions for its whole life:

    /interface/print stats where name="ether6" or name="ether7"
  2. With the agent running, disable and enable it:

    /interface/ethernet/disable [find name="ether6"]
    :delay 4s
    /interface/ethernet/enable [find name="ether6"]
  3. Read the kernel’s records for that moment. They name the netdev:

    [6] br0: port 7(eth5) entered blocking state
    [4] eth5: set isolation from 0 to 1
    [4] eth5: set isolation from 1 to 0

    Here RouterOS ether6 is kernel eth5.

  4. Repeat on a second dead port. ether7 gave eth6: a −1 offset, seen at two points.

The agent’s table is keyed by the exact /proc/device-tree/model string. On the board below the mapping is a shift by one — RouterOS numbers ports from 1 and the kernel from 0, the switch chip included (switch0 is the switch=switch1 every port reports):

Board (/proc/device-tree/model) RouterOS kernel how known
RB5009 ether1 eth0 inferred
RB5009 ether2 eth1 measured — the case-study loop
RB5009 ether3 … ether5 eth2 … eth4 inferred
RB5009 ether6 eth5 measured — flapped
RB5009 ether7 eth6 measured — flapped
RB5009 ether8 eth7 inferred
RB5009 sfp-sfpplus1 eth8 inferred, last in the enumeration
RB5009 switch1 switch0 inferred

Three pairs were measured, all on the same shift by one; the six inferred ports and the switch rest on it (how each pair was measured). The inferred pairs rest on two observations from a read-only /interface/ethernet/print: the nine ports carry consecutive MAC addresses, …:55 for ether1 through …:5D for sfp-sfpplus1, in RouterOS’s own enumeration order; and all nine report switch=switch1 against the kernel’s one switch0.

A mapping holds for its own board only. The shift by one, the position of the SFP+ cage and the reuse of port 7 are not claimed for another board, and the agent does not apply them to one (not tested elsewhere).

  • The name is reliable; the bridge port number is not. Both flaps reported port 7, because the kernel reuses port slots when a port leaves and rejoins the bridge. Match on eth5, never on port 7.
  • The offset is a property of this model’s driver, not a rule. Re-measure on a different device rather than assuming, and be careful with the intuition that the SFP+ must be eth0 — here it is last, not first. The device tree does not help either: it shows the SoC’s ethernet@0 with three MACs, only eth0 enabled (the 10 G uplink) and eth1/eth2 disabled, while the nine front-panel ports are netdevs the switch driver creates at runtime. The kernel-log names are the runtime ones.

The agent reads the board model from /proc/device-tree/model at start, which works even unprivileged, and, on a board in its table, names ports without asking RouterOS:

  • every kernel-log record whose text names a port carries both names and what happened to that port: iface, ros_iface and kind on the event, port and kind tags on the InfluxDB rows, and iface=eth1 ros_iface=ether2 port_event=own-address on a Loki line;
  • the collector’s /metrics carries mikroscope_kmsg_port_records_total{port,kind,level}, built from those records, with the RouterOS name as port on a board in the table, the kernel name on a board that is not, and no port series at all on a device whose device tree reports no model. It is a subset of mikroscope_kmsg_records_total, not a partition of it: records naming no port are absent from it;
  • the collector’s link-flap detection is keyed by the port name and reads the same classification: a flap is link-up and link-down records on one port, counted, through the same classifier rather than by re-reading the text;
  • mikroscope status prints the board and whether it has a map, for example eth1 (ether2); /healthz and /capabilities carry the board, and /capabilities and mikroscope_device_info carry the table’s evidence string as ports_from, so nobody has to take the mapping on trust.

mikroscope doctor, which is the CLI and not the agent, reads the same records from the running agent’s ring: the whole ring when it holds no more than 10 000 samples, otherwise the newest 10 000; 60 s by default (the agent’s BUFFER_S). It classifies each kernel-log record again from its text, with the board the agent reports on /healthz, rather than trusting the agent’s kind: an agent that ships records without one would otherwise make every port check pass by saying nothing. On a device whose device tree reports no model it still classifies, and names the kernel’s port. Per RouterOS port it reports layer2-loop at three or more own-address records in the window, stp-churn when the port entered learning at least three more times than it reached forwarding, and link-flap at two or more link-down records; softnet-drops is reported per CPU, not per port.

An unknown board gets no port names, not guessed ones. The shift by one is not applied to a board nobody has measured, because a wrong port name points at the wrong cable. On such a board status says so and asks for the pair: bring one port down, see which ethN the log names, and send that pair with the board string in a board report.

A port name alone does not say what happened to the port, so every record that names one is classified as well, by procfs.KmsgKind over the record’s text. The kinds, and the records a RouterOS kernel prints for each (where they were seen):

kind the record
link-up eth8: link up, 1Gbps, full-duplex, eth1: Link is Up - 1Gbps/Full, eth1: phy link up
link-down eth1: link down
stp-<state> br0: port 2(eth1) entered blocking state — blocking, listening, learning, forwarding, disabled
own-address br0: received packet on eth1 with own address as source address (addr:…, vlan:0) — the layer-2 loop signature
other anything else that names a port, including eth1: link becomes ready, which is IPv6 address configuration noticing the carrier and not a transition of its own

The agent classifies at read time and ships kind in the record. A collector in front of an agent whose records carry none classifies them itself with the same function.

Four records are not four faults. A normal link-up is followed by stp-blocking, stp-learning and stp-forwarding as the bridge walks the port back into service.

The table maps to RouterOS’s default names. A port renamed on the router (ether5 → WAN) has a name the table cannot know, so the collector asks the API tier instead. Its interface inventory — read once before the first kernel pull and again every --labels-every, five minutes by default — holds each interface’s factory default name, its current name, its comment and its interface lists, and the collector uses it on every record that names a port to:

  • replace the default name with the current one, so an operator who renamed ether5 to WAN reads WAN on the event, on the InfluxDB port tag, on the Loki line and in the link-flap key;
  • attach label, the port’s comment, and role, its interface lists — so the eth5 flap arrives as ether6, label Unused, rather than as a netdev number somebody has to look up.

Without an API tier the record keeps the board’s default name and gets no label, which is what it can support.

Two panels in the dashboards’ Kernel log row are where the name, the kind and the comment meet:

  • Port events from the kernel log, per port and kind — bars per bin, one series per port and kind, from sum by (port, kind) (increase(mikroscope_kmsg_port_records_total[$__interval])) on Prometheus and from the port/kind tags on InfluxDB;
  • Port events in the window, per port — a table with one row per port that the log named: the port, its label and role, and a column per kind (link down, link up, own address (loop), STP blocking, STP disabled, STP learning, STP forwarding, other).

Both are known-empty panels — a quiet set of ports is the healthy state — and the Prometheus form of each reads the collector’s /metrics, whose copy of the family carries a kind on every port record however the agent shipped it. In an InfluxDB store the kind column exists only once a first port record classified by kind has been written; before that, both InfluxDB queries fail at planning time.

Two alert rules ship beside them, in the InfluxDB and Prometheus provisioning files; the PostgreSQL file has neither, because that store keeps each kernel-log record as a row rather than a count. Both read /dev/kmsg through the agent and ask RouterOS nothing:

  • mikroscope-l2-loop, critical: any own-address record in five minutes. That signature can run for hours while every RouterOS counter looks healthy, as in the loop case study.
  • mikroscope-port-link-down, warning: any link-down record in five minutes — the single event, where link-flap covers the repeats.

The InfluxDB form of each needs a store that has already held one port record classified by kind; before that the query fails at planning time.

Both thresholds are 0 with gt. Both rules’ InfluxDB queries return a firing value on records a router really wrote (verified). A rule going pending, firing and resolving in Grafana’s own evaluation is not tested (Tested on), the gap Alert rules states.