Skip to content

Prometheus

--prom :9124 makes the collector serve the Prometheus text format (text/plain; version=0.0.4) on GET /metrics at that address. Scrape the collector: the agent serves no exposition of its own.

The kernel-tier families on the collector are rendered by the same code the agent runs — its cumulative counters, its busy-tick histogram and its trailing windows — fed by the samples the collector received. A deployment whose collector reaches the agent only through the relay still gets metrics that do not depend on who scrapes or when.

On top of them the collector adds what only it has: the RouterOS API tier’s gauges, the derive stage’s values, its detection and gap counters, and the device-info families from the agent’s /capabilities.

The collector’s, and only the collector’s:

- job_name: "mikroscope"
scrape_interval: 5s
static_configs: [{ targets: ["<collector host>:9124"] }]

This one job carries every family the dashboard asks for. What only a sampler can know reaches the collector as data: the wake latency and read duration of every tick ride in the sample itself, and GET /sampler answers the counters that are not per-tick. There is nothing to double-count and no keep list to maintain. The scrape intervals and Prometheus versions that have been run are under Feature status.

Then the dashboard. The collector is scraped, so it cannot know the address Grafana queries: give it in --grafana-datasource-url, or name a Prometheus datasource you already have in --grafana-datasource-uid. Once Prometheus has scraped a few times:

Terminal window
export GRAFANA_TOKEN=…
mikroscope dashboards publish --prom :9124 --grafana http://grafana:3000 \
--grafana-datasource-url http://prometheus:9090 # Prometheus as Grafana reaches it

The same --grafana flags on the collector publish at every start, as long as --prom is its only sink with a dashboard: beside --influx or another, the datasource flags are refused, so publish Prometheus this way. Or import it with dashboards import --store prometheus against a datasource you already have (Datasource per sink).

Family Where it comes from
every kernel-tier family folded from the samples, by the same code the agent used to run
mikroscope_tick_interval_seconds, _wake_latency_, _read_ folded from dt_ns, wake_ns and read_ns in each sample
mikroscope_slipped_total, mikroscope_sampler_ticks_total the agent’s /sampler, read every minute
mikroscope_trigger_fired_total, _suppressed_total the same, one series per configured condition, present at 0
mikroscope_captures_held, mikroscope_capture_bytes, mikroscope_capture_budget_bytes, mikroscope_capture_bytes_served_total, mikroscope_capture_refused_total{reason} the same: what the agent holds, has served and has refused
mikroscope_api_*, mikroscope_derived_*, _collector_* the collector’s own: the API tier, the derive stage, its counters

The families read from /sampler are as fresh as the collector’s health cadence, a minute, not as fresh as the scrape. For counters of fired triggers and held captures that is the right resolution; anything per-tick travels in the samples at full rate.

Family Type Carries
mikroscope_collector_gaps_total counter ring gaps the collector saw: samples lost between pulls
mikroscope_collector_triggers_total{cause} counter capture triggers the agent fired, per cause; present once one has been seen
mikroscope_collector_detections_total{rule} counter detection events per rule, every one of the eleven rules at 0 from the first scrape
mikroscope_collector_bursts_total counter samples the derive stage flagged as a sub-sample burst

Detections and bursts are counters so a Prometheus-only user learns of an event despite a missed scrape, and every rule is rendered at 0 from the start because a family that appears only after its first event cannot be read as “none so far”.

Family Type Carries
mikroscope_derived_memory_pressure gauge the allocator’s escalation ladder at the newest sample, 0 to 4
mikroscope_derived_cycles_per_packet gauge PMU cycles per packet processed, summed over cores; absent without a PMU, in a sample with no packets, or after a counter reset
mikroscope_derived_instructions_per_packet gauge PMU instructions per packet, same conditions
mikroscope_derived_cache_misses_per_packet gauge PMU cache misses per packet, same conditions
mikroscope_derived_packets_per_interrupt gauge packets per device interrupt; absent when the timer row was not in the sample’s top-K
mikroscope_derived_fastpath_share{interface,direction} gauge fast-path share of the traffic the interface hands the CPU, between the last two counter polls; not a share of the wire; rx only while fp-tx-byte has never counted

These are the newest sample’s values, a level at scrape time; the full series is in the stores that keep every sample. What each one means, and when it is withheld, is on Derived values.

Present only once the API tier has delivered a sample; with --api-mode off or without API credentials none of these families exists.

Family Type Carries
mikroscope_api_up gauge 1 while the API tier delivers samples
mikroscope_api_cpu_load gauge RouterOS cpu-load from /system/resource
mikroscope_api_memory_bytes{kind} gauge free and total from /system/resource
mikroscope_api_uptime_seconds gauge RouterOS uptime
mikroscope_api_core_percent{cpu,kind} gauge per-core load, irq and disk percent from /system/resource/cpu
mikroscope_api_health{name} gauge each /system/health reading
mikroscope_api_interface{interface,kind} gauge rx_bps, tx_bps, rx_pps, tx_pps from monitor-traffic, plus each loss rate the router returned
mikroscope_api_interface_info{interface,label,type,role,bridge,default_name} gauge always 1; one series per interface from the configuration inventory: its comment, RouterOS type, interface lists, bridge and factory name
mikroscope_api_interface_counter_total{interface,counter} counter every per-port cumulative counter the router returned, under RouterOS’s own counter name
mikroscope_api_conntrack_entries gauge the connection count, when --conntrack-every asks for it
  • mikroscope_api_up is never rendered as 0: before the first API sample, and when the tier is off, the family is absent.
  • The conntrack count, the port counters and the fast-path shares arrive on slower cadences than the scrape. The collector holds the last value of each between polls, so a scrape between two polls still sees the family instead of a series that blinks in and out.

The API tier reads what every interface is — its comment, RouterOS type, interface lists, the bridge it is a port of, its factory name and its MTU — from three configuration-only reads at collector start and again every --labels-every (5 min by default). On /metrics that inventory is one info series per interface.

None of it is a label on the rate or counter series: a comment is edited by a human, and a changed label would start a fresh series for every rate and every one of the sixty-odd counters of that port on every edit. Join it in a query instead:

mikroscope_api_interface_counter_total * on(interface) group_left(label, type, role) mikroscope_api_interface_info
  • Every interface in the inventory gets a series, with or without a comment. An empty label value is how Prometheus spells “none” — its data model treats a label with an empty value as one that does not exist — so every series in the exposition carries the same label names.
  • type says what the counters of that interface mean: an ether port in a bridge counts its wire, including the frames the switch chip forwarded in hardware, while the bridge counts its CPU side. Neither is a subset of the other — most of what a port receives can be switched in hardware and never reach the CPU (measured) — so do not sum a port and its bridge.
  • mikroscope_api_interface_counter_total has one series per port and counter the router reports: an Ethernet port reports several times more counters than a bridge, VLAN, PPPoE, WireGuard, veth or loopback interface (counted). A counter a port does not report has no series.
  • A loss rate the router did not return has no kind, and monitor-traffic can return the drop rates with no error keys at all (verified).
  • Keys that parse as integers but count nothing — mtu, actual-mtu, l2mtu, max-l2mtu, sfp-shutdown-temperature — are sizes and configuration and get no counter series; the MTU is part of the inventory.

The collector’s exposition carries the device-info families from what it fetched from /capabilities: mikroscope_device_info, the ceilings the board publishes (mikroscope_thermal_critical_celsius, mikroscope_thermal_polling_seconds, mikroscope_cpu_frequency_limit_hertz, mikroscope_cpu_frequency_step_hertz, mikroscope_cpu_frequency_governor_info, mikroscope_cpu_frequency_cluster, mikroscope_self_cgroup_memory_max_bytes) and mikroscope_source_cadence_hz{source,reason}. See Device info.