Skip to content

Conntrack without the API

In short: inside the agent’s container nf_conntrack_count reads 0 however busy the router is, because it belongs to the container’s network namespace. The router’s real connection count is the active objects of the nf_conntrack slab cache in /proc/slabinfo, readable with privileged=yes, and its ceiling, nf_conntrack_max, reads the same from the container as from RouterOS. Over five days of hourly means in the reference store, that slab count tracked the API’s count-only and stayed a few percent above it. The ceiling and the timeouts were both read on 2026-09-14.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 · Linux 5.6.3 ·  · agent at 10 Hz in an ephemeral privileged container

The container’s own network namespace reports nf_conntrack_count = 0 no matter how busy the router is.

Measured on RB5009UG+S+ · 4 × 1.4 GHz Cortex-A72 · RouterOS 7.24.2 ·  · privileged=yes does not change the network namespace

The slab allocator is global. Under privileged=yes the agent reads /proc/slabinfo and reports the nf_conntrack cache’s active object count, which is the router’s real conntrack population:

slab: {'nf_conntrack': 6240, 'skbuff_head_cache': 1008, 'skbuff_fclone_cache': 400,
'TCP': 327, 'UDP': 180, 'TCPv6': 117, 'UDPv6': 125,
'sock_inode_cache': 645, 'kmalloc-1k': 1152, 'kmalloc-2k': 905}

In a sample it is the slab map; on /metrics it is mikroscope_slab_active_objects{cache="nf_conntrack"}. The family is absent without privileged=yes, because /proc/slabinfo is root-only. A cache the kernel does not have is omitted, not reported as a permanent 0: the agent also asks for dst_cache and ip_dst_cache, which do not appear in the output above.

/proc/slabinfo is the most expensive file the agent parses, so it is read on a 6 Hz budget floor, rounded to whole ticks: every 2nd tick (5 Hz) at the 10 Hz default, every 17th at 100 Hz. 6 Hz is also about the rate nf_conntrack was measured changing at. It is stored on change, plus a heartbeat once every 60 s.

Cross-check the slab count against the API once, so you can trust it afterwards:

/ip/firewall/connection/print count-only

On the RB5009 that count read 6 212 on 2026-09-11, the day before the slab readings. Those, on 2026-09-12, were 6 582 in the privileged discovery container and 6 287 from the agent, and the snippet above caught a third, undated moment, 6 240. A day apart, these show the same order of magnitude, not that the two track: on your own device, run the count and read the slab in the same minute.

The reference store has since done that for five days. With forward --conntrack-every set, the collector asked the API for the count while the agent read the slab, and the chart has both as hourly means, with the ratio between them underneath:

The slab count against the API's, five days Hourly means, 2026-09-19 11:00 to 2026-09-24 11:00 UTC, from the reference InfluxDB store (RB5009UG+S+, RouterOS 7.24.4, Linux 5.6.3): the RouterOS API's /ip/firewall/connection/print count-only, a median of 338 readings an hour, against the active objects of the nf_conntrack slab cache the agent read. The API count ran from 3 444 to 9 761; in every hour the slab count was 1.02 to 1.08 times it, and the two hourly series correlate at r = 0.9995. nf_conntrack_max read 966 656 throughout. RB5009UG+S+ · RouterOS 7.24.4 · Linux 5.6.3 · 2026-09-19 11:00 to 2026-09-24 11:00 UTC hourly means, from the reference InfluxDB store Tracked connections, hourly mean API count-only nf_conntrack slab objects 0 2 500 5 000 7 500 10 000 12 500 Slab objects ÷ API count 1.02–1.08 in every hour; r = 0.9995 1.00 1.05 1.10 09-20 09-21 09-22 09-23 09-24
The slab count against the API's, five days Hourly means, 2026-09-19 11:00 to 2026-09-24 11:00 UTC, from the reference InfluxDB store (RB5009UG+S+, RouterOS 7.24.4, Linux 5.6.3): the RouterOS API's /ip/firewall/connection/print count-only, a median of 338 readings an hour, against the active objects of the nf_conntrack slab cache the agent read. The API count ran from 3 444 to 9 761; in every hour the slab count was 1.02 to 1.08 times it, and the two hourly series correlate at r = 0.9995. nf_conntrack_max read 966 656 throughout. RB5009UG+S+ · RouterOS 7.24.4 · Linux 5.6.3 2026-09-19 11:00 to 2026-09-24 11:00 UTC hourly means, from the reference InfluxDB store Tracked connections, hourly mean API count-only nf_conntrack slab objects 0 2 500 5 000 7 500 10 000 12 500 Slab objects ÷ API count 1.02–1.08 in every hour; r = 0.9995 1.00 1.05 1.10 09-20 09-21 09-22 09-23 09-24

The two will not match exactly even then: they are sampled at different instants, and the slab counts objects the allocator still holds. The difference in cost is the point: a file read several times a second versus a table scan over an API session.

Next to nf_conntrack_count in /proc/sys/net/netfilter, nf_conntrack_max is not per-namespace. It reads 966 656 from inside the container, the same number a read-only /ip/firewall/connection/tracking/print reports as max-entries on the reference router (2026-09-14). So “how full is the connection table” is answerable without the API: the slab population over that ceiling.

The agent reads the ceiling once at start, since it is a sysctl a human edits and not a counter, and ships it beside the population it bounds: mikroscope_slab_limit_objects on /metrics, limit on the InfluxDB slab row.

The collector runs two detections on the population and its ceiling: conntrack-cliff, when nf_conntrack falls below half its previous stored value, and conntrack-high, when occupancy is above 80 % of nf_conntrack_max and rising over the last 60 s, with no time-to-full attached. The mikroscope-conntrack-near-limit alert rule watches the same ratio in the store and fires after five minutes above 80 %.

A trap in the same subtree: the conntrack timeouts there read Linux defaults (tcp_timeout_established 432 000 s, against RouterOS’s 1d; read 2026-09-14), so they must never be presented as the router’s configuration.

The other caches in the slab map are worth watching for their own sake:

  • skbuff_head_cache / skbuff_fclone_cache: packet buffers in flight. A spike here during a traffic event is memory pressure from the network path, not from anything you installed.
  • TCP / UDP / sock_inode_cache: sockets the router itself holds.
  • kmalloc-1k / kmalloc-2k: where large-allocation storms show up.

Needed no provoking ·

  • nf_conntrack climbing toward nf_conntrack_max while the namespace’s own count stays 0.
  • nf_conntrack falling below half its previous stored value, which is what conntrack-cliff flags; the detection says where to look, not why it fell.
  • skbuff_* rising with a traffic event, which places the memory pressure on the network path.