Skip to content

Layer-2 loop

In short: a layer-2 loop behind a port of a MikroTik bridge shows in the Linux kernel’s log even while RouterOS’s own log and port monitor stay clean, as they did here: received packet on <port> with own address as source address, the reflected frames 2.00–2.01 s apart, which is the STP hello interval. The port in that line is where the reflected frame came in, not always the port the loop has cut off, so compare each bridge port’s rx with its tx before you pull a cable. Here the stream fell from 1.49 /s to 0.03 /s after a firmware update on the mesh access points, and a week later the loop came back.

A real fault, not provoked. It was found by accident on 2026-09-12, while measuring something else, on the owner’s production RB5009UG+S+ (RouterOS 7.24.2, kernel 5.6.3), and was diagnosed and gone the same day. A week later it came back. For its whole life, RouterOS’s log and port monitor reported a healthy device.

Steps 1 to 6 are the diagnosis in the order it ran. Follow them in the same order on your own router.

/snapshot carried a steady stream of events — the agent’s kmsg source — at 1.49 /s:

[6] br0: port 2(eth1) entered blocking state
[4] br0: received packet on eth1 with own address as source address (addr:00:00:5e:00:53:5d, vlan:0)
[6] br0: port 2(eth1) entered learning state

The address in the second line is the router’s own: the MAC of sfp-sfpplus1, coming back into the bridge. It is shown here as 00:00:5e:00:53:5d, from the block RFC 9542 reserves for documentation (it replaced RFC 7042 in 2024) — on your device it is your bridge’s address, and that is how you recognise the line.

Count events per second over a window long enough to be stable. A single log line is an anecdote; a rate is a measurement, and it tells you later whether a fix worked:

Terminal window
curl -s "http://172.30.10.2:9123/snapshot?seconds=120" \
| python3 -c '
import sys, json, collections
rows=[json.loads(l) for l in sys.stdin if l.strip()]
secs=sum(r["dt_ns"] for r in rows)/1e9
ev=[e for r in rows for e in r.get("events",[])]
print(f"{len(ev)/secs:.2f} events/s over {secs:.0f}s")
c=collections.Counter(e["msg"].split("(")[0][:60] for e in ev)
for k,v in c.most_common(): print(f" {v:5d} {k}")'

?seconds= takes 1 to 3600, and the agent can only return what its ring still holds — 60 s at the default BUFFER_S. The window’s length comes from summing each sample’s own dt_ns, not from the number you asked for.

Take the gaps between consecutive occurrences of one message:

Terminal window
# gaps between consecutive occurrences of one message
... | python3 -c '
import sys, json
ts=sorted(e["us"]/1e6 for l in sys.stdin if l.strip()
for e in json.loads(l).get("events",[]) if "own address" in e["msg"])
print([round(ts[i+1]-ts[i],2) for i in range(len(ts)-1)][:12])'

us is the kernel’s own timestamp for the record, in microseconds since boot on the monotonic clock, not the time the agent read it, so the gaps are the kernel’s, not the sampler’s.

The gaps between the reflected frames were 2.00–2.01 s, every time. 2 s is the STP hello interval, RouterOS’s default HelloTime. So the router was sending a BPDU and receiving its own BPDU back: a loop, not a misbehaving client.

mikroscope doctor, run on its own, does this counting for you (internal/health/health.go). It reads the running agent’s ring once and reports:

  • WARN layer2-loop, with the port named, when three or more frames in the window came back carrying the bridge’s own address;
  • stp-churn for a port that STP moved to learning at least three more times than it let forward, unless that port already carries the own-address finding.

Two alert rules in internal/dashboards/alerts.go watch the same fault over time:

  • mikroscope-l2-loop fires on the first own-address record, from the agent’s kernel log;
  • mikroscope-bridge-port-dark fires on a bridge port that has received for ten minutes while the bridge sent it nothing. It needs the API tier.

See Health checks and Alert rules.

Check RouterOS’s log and port monitor explicitly; the answer decides where you look next:

/log/print where topics~"bridge" or topics~"stp" or topics~"interface"
/interface/bridge/port/monitor [find] once

Both came back clean: zero log rows on those topics, every port designated-port / in-bridge. RouterOS was not hiding the fault; it does not surface this class of kernel event.

That is the log and the port monitor, not every counter the API offers. When the loop returned on 2026-09-19..23, the per-port rx/tx counters the API tier reads did show it, as a bridge port that receives while the bridge sends it nothing. That is the mikroscope-bridge-port-dark rule, backtested over those four days. It was not checked against the 2026-09-12 episode.

The line the interface topic would have carried is the one people search for when they meet this fault. RouterOS has printed it in its own log, under the topics interface,warning, in this form:

<port>: bridge port received packet with own address as source address (<MAC>), probably loop

That form is taken from a user’s RouterOS log posted on MikroTik’s forum on 2017-12-10, with the port and the address replaced; which RouterOS versions still print it is not established here. On the reference router, on RouterOS 7.24.2, no such line appeared, while the kernel’s own received packet on eth1 with own address arrived every 2 s.

Besides STP, RouterOS’s guard against a loop is Loop Protect, which sends its own packets out of an interface, disables the interface when one of them comes back, and logs that. MikroTik’s page recommends (R/M)STP over it on bridge ports. It played no part in this diagnosis, and it is not claimed that it would have caught this loop.

Look for a second, independent observation outside the kernel log. The MAC in the message belonged to sfp-sfpplus1, which pointed at the SFP+ segment; the bridge’s own tables pointed elsewhere. Count the hosts learned on each port:

# hosts learned per port
:foreach p in=[/interface/bridge/port/find] do={ \
:local n [/interface/bridge/port/get $p interface]; \
:put ($n . " hosts=" . [:len [/interface/bridge/host/find interface=$n]]) }
Port Learned MACs Traffic RSTP edge
sfp-sfpplus1 60 23.8 GB rx true
ether2 0 9.15 GB rx / 22.4 M packets false
others 0–2 — true

ether2 was passing 22 million packets on a healthy 1 Gbps link and the bridge had learned nothing behind it, which is what a bridge does on a port where it keeps seeing its own addresses. It was also the only port receiving BPDUs: both anomalies on the same port.

The kernel log had said eth1, not ether2: the kernel and RouterOS name the same port differently. Port names maps one to the other in one safe step, and says why the agent does it for you on this board.

The hypothesis: the device on ether2 has a second path to the router through the other APs, which hang off the switch. Two changes on the switch came first, and neither moved the rate:

Change Event rate
baseline 1.49 /s
switch change 1 1.67 /s
switch change 2 1.47 /s
(noise band) ±0.2 /s

One further reading in the same series, 1.50 /s, was also inside that band.

That ruled out the switch with evidence. Re-measure after every change, even one you expect to work, and know your noise band before you read a difference.

The firmware of the mesh APs (Deco units) was updated; the cabling was not touched. Then:

events 9 -> 0.03/s (baseline 1.49/s)

and of those nine, none was the loop message: they were the APs coming back (eth1: phy link up, eth1: set isolation from 0 to 1). The loop message, which had appeared every 2 s (30 times a minute, with twice as many blocking/learning records around it), appeared zero times in 300 s.

Confirm with a second, independent measurement. Here the bridge’s view became coherent:

Before After
MACs on ether2 0 24
MACs on sfp-sfpplus1 60 36
ether2 edge false true

The same 60 devices, redistributed 36/24 instead of 60/0. While the loop existed, the router was learning nearly every device on the wrong port; the Zigbee coordinator 00:4B:12:96:80:33, for instance, moved from the SFP+ to ether2, where it lives. A second unrelated measurement falling into place is what makes a fix more than a coincidence of timing. It confirms the fault was gone when it was measured, not that it stays gone: a week later it came back.

The same signature returned on 2026-09-19, on RouterOS 7.24.4, again a layer-2 loop through a second path via a mesh access point, and lasted until that path was broken on 2026-09-23. For its first 34 h the port the bridge had stopped delivering to was sfp-sfpplus1, then ether2 for 62 h. Between 12:07 and 18:55 UTC on 2026-09-21 and 22, ether2 received 12 and 13 packets a second and sent 1; over the same hours on 2026-09-23, after the second path went at 12:06, it received 192 and sent 178.

Two things were different the second time:

  • For most of the first phase the own-address records in the kernel log named ether2, while the port receiving and hearing nothing back was sfp-sfpplus1: the message names the port the reflected frame came in on, not the port the loop had cut off.
  • RouterOS’s per-port counters did show it, read as receive without transmit. The backtest of mikroscope-bridge-port-dark over 2026-09-19 11:13 to 2026-09-23 22:44 UTC marks those two ports in those two stretches and no other port.

The reference InfluxDB store kept both halves, and the chart below is drawn from it, hour by hour: the kernel’s own-address records on top, then what each of the two ports received and was sent. While sfp-sfpplus1 was the port cut off, the records came a few an hour; once ether2 was, they came every 2 s, as on 2026-09-12.

The loop's return, hour by hour Hourly, 2026-09-19 11:00 to 2026-09-24 11:00 UTC, from the reference InfluxDB store (RB5009UG+S+, RouterOS 7.24.4, Linux 5.6.3). Until 2026-09-20 21:00 the kernel logged at most 12 own-address records an hour, while the bridge sent sfp-sfpplus1 nothing in any hour and received 6.5 a second from it. From then until the last record, at 2026-09-23 12:06:46, it logged a median of 1 798 an hour, one every 2.0 s, and the bridge sent ether2 0.6 packets a second against 13 received. Afterwards ether2 received 183 and was sent 144. Rates are medians of the hourly means. RB5009UG+S+ · RouterOS 7.24.4 · Linux 5.6.3 · 2026-09-19 11:00 to 2026-09-24 11:00 UTC hourly, from the reference InfluxDB store sfp-sfpplus1 cut off ether2 cut off no loop Own-address records in the kernel log, per hour at most 12 an hour before 09-20 21:00, then a median of 1 798 0 1 000 2 000 ether2: packets a second, hourly mean, log scale received sent by the bridge ≤0.1 1 10 100 1 000 sfp-sfpplus1: packets a second, hourly mean, log scale received sent by the bridge ≤0.1 1 10 100 1 000 09-20 09-21 09-22 09-23 09-24
The loop's return, hour by hour Hourly, 2026-09-19 11:00 to 2026-09-24 11:00 UTC, from the reference InfluxDB store (RB5009UG+S+, RouterOS 7.24.4, Linux 5.6.3). Until 2026-09-20 21:00 the kernel logged at most 12 own-address records an hour, while the bridge sent sfp-sfpplus1 nothing in any hour and received 6.5 a second from it. From then until the last record, at 2026-09-23 12:06:46, it logged a median of 1 798 an hour, one every 2.0 s, and the bridge sent ether2 0.6 packets a second against 13 received. Afterwards ether2 received 183 and was sent 144. Rates are medians of the hourly means. RB5009UG+S+ · RouterOS 7.24.4 · Linux 5.6.3 2026-09-19 11:00 to 2026-09-24 11:00 UTC hourly, from the reference InfluxDB store sfp-sfpplus1 cut off ether2 cut off no loop Own-address records in the kernel log, per hour at most 12 an hour before 09-20 21:00, then a median of 1 798 0 1 000 2 000 ether2: packets a second, hourly mean, log scale received sent by the bridge ≤0.1 1 10 100 1 000 sfp-sfpplus1: packets a second, hourly mean, log scale received sent by the bridge ≤0.1 1 10 100 1 000 09-20 09-21 09-22 09-23 09-24

A real fault, not provoked ·

  • A steady, periodic stream of kernel events at idle, where the healthy state is zero.
  • received packet on <port> with own address as source address, with the router’s own MAC in it.
  • The port in that message is where the reflected frame came in, not necessarily the port the loop has cut off. On 2026-09-19..20 it named ether2 while sfp-sfpplus1 was the port receiving and sending nothing. Check each bridge port’s rx against its tx before you pull a cable.
  • Inter-arrival times of 2.00–2.01 s: the STP hello interval.
  • A busy port on which the bridge has learned no hosts, with edge=false while its neighbours are true.
  • /log/print and /interface/bridge/port/monitor both clean.
  • Alarming text is a lead; a rate is evidence; a rate before and after is a conclusion.
  • Read the inter-arrival times. They often name the protocol for you.
  • Prefer a change that discriminates between hypotheses over a change that merely might fix things.
  • Distrust a fix that only your primary instrument can confirm.
  • Keep watching after the fix. A confirmed fix says the fault was gone when you measured, and this one came back a week later.