Skip to content

Port losing frames

In short: an rx-overflow on one port whose overflowing intervals carry a small fraction of the link’s capacity is a burst, not a load. Here it was a 2.5 Gbit/s sender bursting into a 1 Gbit/s port, with the overflowing intervals at 0.36 % of the link. A smaller MTU and Ethernet flow control both failed to stop it; pacing the sender below what the slowest destination can drain did, from 5 399 overflows an hour to none in the 39 minutes measured.

When Port errors in the window is red (it is green at 0, red above it), the sections below go from that one number to the port, the error, and whether the port is losing frames because it is busy or because the sender is bursting, which decides the fix. The readings are one real fault on the reference RB5009, found on 2026-09-19 and fixed the same afternoon.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.4 ·  · ether1, the 2.5 GbE port to the NAS, over 10 s counter intervals; the before figures are the three hours preceding the fix and the after figures the 39 minutes following it, at the same load

The tile sums every typed MAC error on every port. It deliberately says nothing about which: open Interface traffic and read Port errors per bin, which draws one row per (port, error type) pair that had an error in the window and nothing for the pairs that did not. On a healthy router it says no port errors in this window.

On the reference device it drew exactly one row: ether1 rx overflow.

That name is the diagnosis’s first half. rx-overflow is the receive FIFO filling faster than the switch chip can drain it — frames that arrived correctly and were dropped for want of somewhere to put them. It is not a cabling fault. rx-fcs-error, rx-fragment and the collision counters are, and they send you somewhere else entirely: the cable, the duplex, the port.

The fix depends on this, and the counters answer it. Take the receive volume in the 10-second intervals that overflowed and compare it with what the link could have carried. On the reference device the median interval with an overflow carried 8.96 Mbit/s on a 2.5 Gbit/s link: 0.36 % of its capacity.

A port cannot be overwhelmed at 0.36 % occupancy by sustained load. It can only be overwhelmed by bursts too short for a one-second average to show: the mean was 3.25 Mbit/s and the highest single second in three hours was 70 Mbit/s, while the overflow ran at 5 399 an hour, 0.502 % of every packet the sender sent.

If instead the intervals that overflow are the ones near line rate, stop here: that is a capacity problem and the answer is a faster link or less traffic.

A burst overflows on ingress because something downstream cannot take it. The per-port counters find it by correlation: rank the overflow deltas against each other port’s transmit deltas over the same intervals.

On the reference device the strongest was ether4 at ρ 0.53, a 1 Gbit/s port, and sfp-sfpplus1 at ρ 0.42, which is 10 Gbit/s itself but feeds a switch whose ports are not. The CPU-bound share of the traffic correlated at ρ 0.00, which rules out the router’s own forwarding: these frames never reached the CPU.

ether4 also carried 6 644 tx-queue-drop. That is the same event counted from the other end, the egress queue that could not drain fast enough, and finding both is what turns a correlation into a mechanism: a 2.5 Gbit/s sender bursting at line rate into a 1 Gbit/s destination.

On the dashboard it is Egress queue drops — the router’s own transmit queue, in Interface traffic. mikroscope-egress-queue-drops fires only after ten consecutive minutes of dropping, so bursts like these show on the panel and deliberately do not page you.

On 2026-09-15, four days before the fix, the API tier’s counters on the same port said what kind of event the overflow is. The port counters are the only place these overflows appear. Measured between 07:13 and 22:20 UTC, over 4 471 consecutive 10 s counter polls of ether1 (2.5 Gbps to a NAS, MTU 9000):

  • 126 443 rx-overflow events in all, present in 40 % of the intervals; per interval the median is 29, the p99 about 1 036 and the largest 3 747. Over the run that is 0.53 % of the packets the NAS sent.
  • Rank correlation over the 10 s deltas: 0.85 against the part of the NAS’s receive the switch forwarded in hardware (rx-bytes minus driver-rx-byte), 0.00 against the part it sent to the CPU (driver-rx-byte).
  • Where it was going: ether8 (NGINX, 1 Gbps) carries most of the volume, ether4 (Mastodon, 1 Gbps) is the most frequent destination; the SFP+ cage, ether2, ether3 and the CPU path show nothing.
  • The frames were large: the 1024-and-up frame-size bucket on ether1 has a median of 9 331 per interval with overflow against 1 336 per interval without.
  • The load was not high: the median NAS receive in an interval with overflow is about 9 Mbit/s as a 10 s mean. Bursts, not sustained load.
  • Nothing on the CPU side: softnet dropped 0, time_squeeze correlates 0.04 and the switch0 interrupts 0.05 against the overflows, with the agent’s 10 Hz data binned to 10 s. No pause frames on ether1 in either direction.

Read together, that is consistent with 2.5 Gbps line-rate bursts switched inside the chip toward 1 Gbps ports with no flow control in effect. Not verified: the switch chip’s exact counter semantics, and the NAS’s own retransmit count; the counters are the port’s, not the conversation’s.

The kernel tier cannot see any of it by construction. A frame the switch chip forwards in hardware never reaches the CPU, so no /proc file on the router has a number for it; it takes the per-port counters, and only the API has those.

Two fixes that look right were tried on the reference device first, and neither worked.

A smaller MTU. Dropping the sender and the port from 9000 to 1500 cut the worst bursts by 91 % and the data lost per dropped frame by six, and did not change the frequency at all: 0.530 % of packets before, 0.502 % after. It makes each event cheaper without making events rarer.

Ethernet flow control. Negotiating pause in both directions is the mechanism designed for exactly this, and on this hardware it never fired: 41 minutes with pause negotiated, 5 578 overflows, and rx-pause and tx-pause both still 0. Check those two counters before believing pause is helping you. The chip is dropping the frame rather than asking the sender to wait.

Pace the sender with a shaper whose rate is below what the slowest destination can drain. Not an AQM: the sender’s queue is empty at 0.36 % occupancy, so an AQM has nothing to manage and hands the NIC the burst unchanged. On the sender:

Terminal window
tc qdisc replace dev <iface> root cake bandwidth 900Mbit

Two things matter in that line. The rate is under 1 Gbit/s, so the destination drains faster than the source sends and the chip’s buffer never grows. And cake splits GSO super-segments, which is what the sender was handing its NIC to put on the wire back to back at line rate.

The result on the reference device, at the same traffic volume: the overflow went from 1 054–3 681 per 20 minutes to 0, and ether4’s tx-queue-drop to 0 with it. The cost was nothing measurable: the highest second of egress observed was 70 Mbit/s against a 900 Mbit cap.

The reference InfluxDB store kept counting past the 39 minutes those figures cover. Over the rest of that day, ether1’s overflow stopped in the ten minutes to 15:30 UTC, stayed at zero for two hours and then came back a few at a time, at much the same received rate. Whether the shaper stayed in place for the rest of the day was not recorded, so those hours say nothing either way about whether it holds.

ether1's receive overflow, before and after it stopped Per ten minutes, 2026-09-19 11:10 to 2026-09-20 00:00 UTC, from the reference InfluxDB store (RB5009UG+S+, RouterOS 7.24.4, Linux 5.6.3): the increments of ether1's rx-overflow and received-packet counters, read by the API tier a median of 56 times per ten minutes. Every ten minutes from 11:10 to 15:30 overflowed, 24 550 in the 4 h 20 min before 15:30; the two hours after it had none, and the whole 8 h 30 min after it 131. The port received a median of 313 packets a second before and 286 after. RB5009UG+S+ · RouterOS 7.24.4 · Linux 5.6.3 · 2026-09-19 11:10 to 2026-09-20 00:00 UTC per 10 minutes, from the reference InfluxDB store ether1 rx-overflow, per 10 minutes 24 550 in the 4 h 20 min before 15:30, 131 in the 8 h 30 min after 0 1 000 2 000 3 000 ether1 received packets a second, 10-minute mean 0 500 1 000 12:00 13:00 14:00 15:00 16:00 17:00 18:00 19:00 20:00 21:00 22:00 23:00
ether1's receive overflow, before and after it stopped Per ten minutes, 2026-09-19 11:10 to 2026-09-20 00:00 UTC, from the reference InfluxDB store (RB5009UG+S+, RouterOS 7.24.4, Linux 5.6.3): the increments of ether1's rx-overflow and received-packet counters, read by the API tier a median of 56 times per ten minutes. Every ten minutes from 11:10 to 15:30 overflowed, 24 550 in the 4 h 20 min before 15:30; the two hours after it had none, and the whole 8 h 30 min after it 131. The port received a median of 313 packets a second before and 286 after. RB5009UG+S+ · RouterOS 7.24.4 · Linux 5.6.3 2026-09-19 11:10 to 2026-09-20 00:00 UTC per 10 minutes, from the reference InfluxDB store ether1 rx-overflow, per 10 minutes 24 550 in the 4 h 20 min before 15:30, 131 in the 8 h 30 min after 0 1 000 2 000 3 000 ether1 received packets a second, 10-minute mean 0 500 1 000 12:00 14:00 16:00 18:00 20:00 22:00

A real fault, not provoked ·

  • Port errors in the window red, and Port errors per bin drawing one row: a rx overflow on one port and nothing else.
  • The intervals that overflow carry a small fraction of the link’s capacity: a burst, not a load.
  • A slower port’s transmit correlates with the overflow, and carries tx-queue-drop of its own.
  • rx-pause and tx-pause stay at 0 whatever the negotiated flow control says.