# Port losing frames

A MikroTik port counting rx-overflow: telling a load from a microburst with the port counters, what did not fix it, and the sender pacing that did.

Source: https://jmrplens.github.io/mikroscope/playbooks/port-errors/

**In short:** an `rx-overflow` on one port whose overflowing intervals carry a
small fraction of the link's capacity is a burst, not a load. Here it was a
2.5 Gbit/s sender bursting into a 1 Gbit/s port, with the overflowing intervals
at 0.36 % of the link. A smaller MTU and Ethernet
flow control both failed to stop it; pacing the sender below what the slowest
destination can drain did, from 5 399 overflows
an hour to none in the 39 minutes measured.

When **Port errors in the window** is red (it is green at 0, red above it), the
sections below go from that one number to the port, the error, and whether the
port is losing frames because it is busy or because the sender is bursting,
which decides the fix. The readings are one real fault on the reference RB5009,
found on 2026-09-19 and fixed the same afternoon.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.4 · 2026-09-19 · ether1, the 2.5 GbE port to the NAS, over 10 s counter intervals; the before figures are the three hours preceding the fix and the after figures the 39 minutes following it, at the same load

## Find the port

The tile sums every typed MAC error on every port. It deliberately says nothing
about which: open **Interface traffic** and read _Port errors per bin_, which
draws one row per `(port, error type)` pair that had an error in the window and
nothing for the pairs that did not. On a healthy router it says
`no port errors in this window`.

On the reference device it drew exactly one row: `ether1 rx overflow`.

That name is the diagnosis's first half. `rx-overflow` is the receive FIFO
filling faster than the switch chip can drain it — frames that arrived
correctly and were dropped for want of somewhere to put them. It is not a
cabling fault. `rx-fcs-error`, `rx-fragment` and the collision counters are,
and they send you somewhere else entirely: the cable, the duplex, the port.

## Load or burst

The fix depends on this, and the counters answer it. Take the receive volume in the 10-second intervals that overflowed and compare
it with what the link could have carried. On the reference device the median
interval with an overflow carried 8.96 Mbit/s on a
2.5 Gbit/s link: 0.36 % of its capacity.

A port cannot be overwhelmed at 0.36 % occupancy by sustained load. It can only
be overwhelmed by bursts too short for a one-second average to show: the mean
was 3.25 Mbit/s and the highest single second in three hours was 70 Mbit/s,
while the overflow ran at 5 399 an
hour, 0.502 % of every packet the sender sent.

If instead the intervals that overflow are the ones near line rate, stop here:
that is a capacity problem and the answer is a faster link or less traffic.

## Traffic path

A burst overflows on ingress because something downstream cannot take it. The
per-port counters find it by correlation: rank the overflow deltas against each
other port's transmit deltas over the same intervals.

On the reference device the strongest was `ether4` at ρ 0.53, a 1 Gbit/s port,
and `sfp-sfpplus1` at ρ 0.42, which is 10 Gbit/s itself but feeds a switch
whose ports are not. The CPU-bound share of the traffic correlated at ρ 0.00,
which rules out the router's own forwarding: these frames never reached the
CPU.

`ether4` also carried 6 644 `tx-queue-drop`. That is the same event counted
from the other end, the egress queue that could not drain fast enough, and
finding both is what turns a correlation into a mechanism: a 2.5 Gbit/s sender
bursting at line rate into a 1 Gbit/s destination.

On the dashboard it is _Egress queue drops — the router's own transmit queue_, in
**Interface traffic**. `mikroscope-egress-queue-drops` fires only after ten
consecutive minutes of dropping, so bursts like these show on the panel and
deliberately do not page you.

## Fifteen hours of counters

On 2026-09-15, four days before the fix, the API tier's counters on the same
port said what kind of event the overflow is. The port counters are the only
place these overflows appear. Measured between 07:13 and 22:20 UTC, over
4 471 consecutive 10 s counter polls of `ether1` (2.5 Gbps to a NAS, MTU 9000):

- 126 443 `rx-overflow` events in all, present in 40 % of the intervals; per
  interval the median is 29, the p99 about 1 036 and the largest 3 747. Over
  the run that is 0.53 % of the packets the NAS sent.
- Rank correlation over the 10 s deltas: 0.85 against the part of the NAS's
  receive the switch forwarded in hardware (`rx-bytes` minus `driver-rx-byte`),
  0.00 against the part it sent to the CPU (`driver-rx-byte`).
- Where it was going: `ether8` (NGINX, 1 Gbps) carries most of the volume,
  `ether4` (Mastodon, 1 Gbps) is the most frequent destination; the SFP+ cage,
  `ether2`, `ether3` and the CPU path show nothing.
- The frames were large: the 1024-and-up frame-size bucket on `ether1` has a
  median of 9 331 per interval with overflow against 1 336 per interval without.
- The load was not high: the median NAS receive in an interval with overflow is
  about 9 Mbit/s as a 10 s mean. Bursts, not sustained load.
- Nothing on the CPU side: softnet dropped 0, `time_squeeze` correlates 0.04 and
  the `switch0` interrupts 0.05 against the overflows, with the agent's 10 Hz
  data binned to 10 s. No pause frames on `ether1` in either direction.

Read together, that is consistent with 2.5 Gbps line-rate bursts switched inside
the chip toward 1 Gbps ports with no flow control in effect. Not verified: the
switch chip's exact counter semantics, and the NAS's own retransmit count; the
counters are the port's, not the conversation's.

The kernel tier cannot see any of it by construction. A frame the switch chip
forwards in hardware never reaches the CPU, so no `/proc` file on the router has
a number for it; it takes the per-port counters, and only the API has those.

## Failed fixes

Two fixes that look right were tried on the reference device first, and neither
worked.

**A smaller MTU.** Dropping the sender and the port from 9000 to 1500 cut the
worst bursts by 91 % and the data lost per dropped frame by six, and **did not
change the frequency at all**: 0.530 % of packets before, 0.502 % after. It
makes each event cheaper without making events rarer.

**Ethernet flow control.** Negotiating pause in both directions is the
mechanism designed for exactly this, and on this hardware it never fired: 41
minutes with pause negotiated, 5 578 overflows, and `rx-pause` and `tx-pause`
both still **0**. Check those two counters before believing pause is helping
you. The chip is dropping the frame rather than asking the sender to wait.

## Fix

Pace the sender with a shaper whose rate is **below what the slowest destination
can drain**. Not an AQM: the sender's queue is empty at 0.36 % occupancy, so an
AQM has nothing to manage and hands the NIC the burst unchanged. On the sender:

```sh
tc qdisc replace dev <iface> root cake bandwidth 900Mbit
```

Two things matter in that line. The rate is under 1 Gbit/s, so the destination
drains faster than the source sends and the chip's buffer never grows. And
`cake` splits GSO super-segments, which is what the sender was handing its NIC
to put on the wire back to back at line rate.

The result on the reference device, at the same traffic volume: the overflow
went from 1 054–3 681 per 20 minutes to **0**, and `ether4`'s `tx-queue-drop`
to 0 with it. The cost was nothing measurable: the highest second of egress
observed was 70 Mbit/s against a 900 Mbit cap.

The reference InfluxDB store kept counting past the 39 minutes those figures
cover. Over the rest of that day, `ether1`'s overflow stopped in the ten minutes
to 15:30 UTC, stayed at zero for two hours and then came back a few at a time, at
much the same received rate. Whether the shaper stayed in place for the rest of
the day was not recorded, so those hours say nothing either way about whether it
holds.

_ether1's receive overflow, before and after it stopped_ — Per ten minutes, 2026-09-19 11:10 to 2026-09-20 00:00 UTC, from the reference InfluxDB store (RB5009UG+S+, RouterOS 7.24.4, Linux 5.6.3): the increments of ether1's rx-overflow and received-packet counters, read by the API tier a median of 56 times per ten minutes. Every ten minutes from 11:10 to 15:30 overflowed, 24 550 in the 4 h 20 min before 15:30; the two hours after it had none, and the whole 8 h 30 min after it 131. The port received a median of 313 packets a second before and 286 after.

## Signature

**A real fault, not provoked** · 2026-09-19

- **Port errors in the window** red, and _Port errors per bin_ drawing one row:
  a `rx overflow` on one port and nothing else.
- The intervals that overflow carry a small fraction of the link's capacity: a
  burst, not a load.
- A slower port's transmit correlates with the overflow, and carries
  `tx-queue-drop` of its own.
- `rx-pause` and `tx-pause` stay at 0 whatever the negotiated flow control
  says.

> **Not measured, so not claimed**
>
> That the shaper holds. The zero above is 39 minutes at one afternoon's load, not a day's, and the
> rate was chosen against a 1 Gbit/s destination rather than derived. Whether the same cap still
> fits when something behind the 10 Gbit/s port wants more than 900 Mbit was not tested. The
> chip's exact counter semantics are MikroTik's and were not verified against its documentation.

## See also

- [Interface traffic](https://jmrplens.github.io/mikroscope/dashboards/#interface-traffic): the panels this page reads.
- [RouterOS API tier](https://jmrplens.github.io/mikroscope/sinks/api-tier/): where the per-port counters come from, and why
  the agent cannot see them.
- [Alert rules](https://jmrplens.github.io/mikroscope/dashboards/alerts/): the rule that fires on this, and what it cannot
  tell apart.
