Port losing frames
In short: an rx-overflow on one port whose overflowing intervals carry a
small fraction of the link’s capacity is a burst, not a load. Here it was a
2.5 Gbit/s sender bursting into a 1 Gbit/s port, with the overflowing intervals
at 0.36 % of the link. A smaller MTU and Ethernet
flow control both failed to stop it; pacing the sender below what the slowest
destination can drain did, from 5 399 overflows
an hour to none in the 39 minutes measured.
When Port errors in the window is red (it is green at 0, red above it), the sections below go from that one number to the port, the error, and whether the port is losing frames because it is busy or because the sender is bursting, which decides the fix. The readings are one real fault on the reference RB5009, found on 2026-09-19 and fixed the same afternoon.
Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.4 · · ether1, the 2.5 GbE port to the NAS, over 10 s counter intervals; the before figures are the three hours preceding the fix and the after figures the 39 minutes following it, at the same load
Find the port
Section titled “Find the port”The tile sums every typed MAC error on every port. It deliberately says nothing
about which: open Interface traffic and read Port errors per bin, which
draws one row per (port, error type) pair that had an error in the window and
nothing for the pairs that did not. On a healthy router it says
no port errors in this window.
On the reference device it drew exactly one row: ether1 rx overflow.
That name is the diagnosis’s first half. rx-overflow is the receive FIFO
filling faster than the switch chip can drain it — frames that arrived
correctly and were dropped for want of somewhere to put them. It is not a
cabling fault. rx-fcs-error, rx-fragment and the collision counters are,
and they send you somewhere else entirely: the cable, the duplex, the port.
Load or burst
Section titled “Load or burst”The fix depends on this, and the counters answer it. Take the receive volume in the 10-second intervals that overflowed and compare it with what the link could have carried. On the reference device the median interval with an overflow carried 8.96 Mbit/s on a 2.5 Gbit/s link: 0.36 % of its capacity.
A port cannot be overwhelmed at 0.36 % occupancy by sustained load. It can only be overwhelmed by bursts too short for a one-second average to show: the mean was 3.25 Mbit/s and the highest single second in three hours was 70 Mbit/s, while the overflow ran at 5 399 an hour, 0.502 % of every packet the sender sent.
If instead the intervals that overflow are the ones near line rate, stop here: that is a capacity problem and the answer is a faster link or less traffic.
Traffic path
Section titled “Traffic path”A burst overflows on ingress because something downstream cannot take it. The per-port counters find it by correlation: rank the overflow deltas against each other port’s transmit deltas over the same intervals.
On the reference device the strongest was ether4 at ρ 0.53, a 1 Gbit/s port,
and sfp-sfpplus1 at ρ 0.42, which is 10 Gbit/s itself but feeds a switch
whose ports are not. The CPU-bound share of the traffic correlated at ρ 0.00,
which rules out the router’s own forwarding: these frames never reached the
CPU.
ether4 also carried 6 644 tx-queue-drop. That is the same event counted
from the other end, the egress queue that could not drain fast enough, and
finding both is what turns a correlation into a mechanism: a 2.5 Gbit/s sender
bursting at line rate into a 1 Gbit/s destination.
On the dashboard it is Egress queue drops — the router’s own transmit queue, in
Interface traffic. mikroscope-egress-queue-drops fires only after ten
consecutive minutes of dropping, so bursts like these show on the panel and
deliberately do not page you.
Fifteen hours of counters
Section titled “Fifteen hours of counters”On 2026-09-15, four days before the fix, the API tier’s counters on the same
port said what kind of event the overflow is. The port counters are the only
place these overflows appear. Measured between 07:13 and 22:20 UTC, over
4 471 consecutive 10 s counter polls of ether1 (2.5 Gbps to a NAS, MTU 9000):
- 126 443
rx-overflowevents in all, present in 40 % of the intervals; per interval the median is 29, the p99 about 1 036 and the largest 3 747. Over the run that is 0.53 % of the packets the NAS sent. - Rank correlation over the 10 s deltas: 0.85 against the part of the NAS’s
receive the switch forwarded in hardware (
rx-bytesminusdriver-rx-byte), 0.00 against the part it sent to the CPU (driver-rx-byte). - Where it was going:
ether8(NGINX, 1 Gbps) carries most of the volume,ether4(Mastodon, 1 Gbps) is the most frequent destination; the SFP+ cage,ether2,ether3and the CPU path show nothing. - The frames were large: the 1024-and-up frame-size bucket on
ether1has a median of 9 331 per interval with overflow against 1 336 per interval without. - The load was not high: the median NAS receive in an interval with overflow is about 9 Mbit/s as a 10 s mean. Bursts, not sustained load.
- Nothing on the CPU side: softnet dropped 0,
time_squeezecorrelates 0.04 and theswitch0interrupts 0.05 against the overflows, with the agent’s 10 Hz data binned to 10 s. No pause frames onether1in either direction.
Read together, that is consistent with 2.5 Gbps line-rate bursts switched inside the chip toward 1 Gbps ports with no flow control in effect. Not verified: the switch chip’s exact counter semantics, and the NAS’s own retransmit count; the counters are the port’s, not the conversation’s.
The kernel tier cannot see any of it by construction. A frame the switch chip
forwards in hardware never reaches the CPU, so no /proc file on the router has
a number for it; it takes the per-port counters, and only the API has those.
Failed fixes
Section titled “Failed fixes”Two fixes that look right were tried on the reference device first, and neither worked.
A smaller MTU. Dropping the sender and the port from 9000 to 1500 cut the worst bursts by 91 % and the data lost per dropped frame by six, and did not change the frequency at all: 0.530 % of packets before, 0.502 % after. It makes each event cheaper without making events rarer.
Ethernet flow control. Negotiating pause in both directions is the
mechanism designed for exactly this, and on this hardware it never fired: 41
minutes with pause negotiated, 5 578 overflows, and rx-pause and tx-pause
both still 0. Check those two counters before believing pause is helping
you. The chip is dropping the frame rather than asking the sender to wait.
Pace the sender with a shaper whose rate is below what the slowest destination can drain. Not an AQM: the sender’s queue is empty at 0.36 % occupancy, so an AQM has nothing to manage and hands the NIC the burst unchanged. On the sender:
tc qdisc replace dev <iface> root cake bandwidth 900MbitTwo things matter in that line. The rate is under 1 Gbit/s, so the destination
drains faster than the source sends and the chip’s buffer never grows. And
cake splits GSO super-segments, which is what the sender was handing its NIC
to put on the wire back to back at line rate.
The result on the reference device, at the same traffic volume: the overflow
went from 1 054–3 681 per 20 minutes to 0, and ether4’s tx-queue-drop
to 0 with it. The cost was nothing measurable: the highest second of egress
observed was 70 Mbit/s against a 900 Mbit cap.
The reference InfluxDB store kept counting past the 39 minutes those figures
cover. Over the rest of that day, ether1’s overflow stopped in the ten minutes
to 15:30 UTC, stayed at zero for two hours and then came back a few at a time, at
much the same received rate. Whether the shaper stayed in place for the rest of
the day was not recorded, so those hours say nothing either way about whether it
holds.
Signature
Section titled “Signature”A real fault, not provoked ·
- Port errors in the window red, and Port errors per bin drawing one row:
a
rx overflowon one port and nothing else. - The intervals that overflow carry a small fraction of the link’s capacity: a burst, not a load.
- A slower port’s transmit correlates with the overflow, and carries
tx-queue-dropof its own. rx-pauseandtx-pausestay at 0 whatever the negotiated flow control says.