Skip to content

Run the collector

Where the data comes from and where it goes The router runs the agent in a container that reads the shared kernel and serves it over a veth. The collector on your machine pulls that, merges the RouterOS API tier into it, derives, and writes to every sink you named. THE ROUTER Shared kernel /proc · /sys · /dev/kmsg · PMU mikroscope-agent 1–100 Hz · 300 s ring · HTTP RouterOS API 1 Hz, what the kernel cannot see YOUR MACHINE mikroscope forward pull · merge · derive · fan out veth /30 SINKS InfluxDB 3 Prometheus file · SQL Loki · OTLP Graphite · Elastic Telegraf · stdout pull every 500 ms
Where the data comes from and where it goes The router runs the agent in a container that reads the shared kernel and serves it over a veth. The collector on your machine pulls that, merges the RouterOS API tier into it, derives, and writes to every sink you named. THE ROUTER Shared kernel /proc · /sys · /dev/kmsg · PMU mikroscope-agent 1–100 Hz · 300 s ring · HTTP RouterOS API veth /30 mikroscope forward pull · merge · derive · fan out SINKS InfluxDB 3 Prometheus file · SQL Loki · OTLP Graphite · Elastic Telegraf · stdout

mikroscope forward is the collector. It pulls the kernel tier from the agent on the router, samples the RouterOS API tier, stamps both in the agent’s clock, runs the derive stage over them and writes the merged timeline to every sink you name.

Terminal window
mikroscope forward --prom :9124 --influx "$MIKROSCOPE_INFLUX_URL" --influx-db "$MIKROSCOPE_INFLUX_DB" --interfaces bridge,ether1

Name at least one sink: --file, --prom, --influx, --loki, --otlp, --graphite, --elastic, --sql, --postgres, --telegraf or --stdout. forward with none is an error, because it would read the router and throw the data away. Several at once is the normal arrangement.

  1. Ask the agent for its health. The reply carries the agent’s wall clock, its rate, its newest sequence number and its capability hash. The difference between the agent’s clock and the collector’s is the skew; the rate sizes the pull batch and the derive stage’s trailing baselines.

  2. Hand every sink the device-info stream. The agent’s /capabilities — board, kernel, ceilings, cadences — goes out at start as its own record, again whenever the capability hash changes, and otherwise every five minutes, so the device panels have a row inside any window. See Device info.

  3. Pull the ring every --poll. The first pull starts after the agent’s newest sample: forward does not replay what the ring held before it started, and refuses --from-start as an unknown flag (only record backfills). Each pull asks for the samples after the last sequence number seen. A trigger marker rides among the samples in sequence order and is forwarded as an annotation, never decoded as a sample.

  4. Derive, then fan out. Every kernel sample goes through the derive stage, and the sample, its derived values and any detection it raised go to every sink in the same order.

  5. Sample the API tier every --api-every (1 s by default) when API credentials are configured. See RouterOS API tier.

  6. Re-measure the skew every minute. A jump of more than 50 ms — a router clock step, an NTP correction — is logged and counted as a skew jump. The same health read re-checks the capability hash.

  7. On Ctrl-C or at the end of --for, pull once more, close every sink with a final flush, and print what each one did.

Nothing is interpolated anywhere in that loop: every consumer sees the cadence each source really has.

Flag Default What it sets
--poll 500ms How often the ring is pulled.
--batch 0 Samples one pull asks for. 0 is twice what one poll interval produces at the agent’s rate, never under 20, so a 50 Hz agent is not capped at 40 Hz.
  • A pull is repeated while it comes back full, up to 100 times, so the cursor catches up within one poll instead of advancing one batch per poll. A short reply is the ring’s edge.
  • The relay transport caps a pull at 13 lines. The cap is computed from the 64 512 B that /tool fetch returns at most, a ring line charged at the allocator size class of 3 456 B (the measured line of 3 230 B rounded up) and 134 % headroom for above-average lines: about 45 kB. A reply that reaches the fetch limit anyway is refused with relay reply hit the 64512-byte fetch limit; lower the batch rather than parsed truncated.
  • forward computes what the effective batch allows per --poll and warns at start when that is below the agent’s rate. For the relay against an agent at 100 Hz, 13 samples per 500 ms poll gives:
warning: at most 13 samples per pull every 500ms is 26/s, below the agent's 100 Hz; the collector will fall behind and report gaps. Raise --batch, lower --poll, or use the direct transport

A collector that falls further behind than the agent’s ring (60 s by default) receives a gap line instead of the samples, and every sink records the gap.

Record Timestamp
Kernel sample, derived values the agent’s own wall clock, as the sample carries it
Detection the wall clock of the sample that raised it
Trigger marker the agent’s wall clock of the fire
API-tier sample the collector’s clock plus the measured skew
Gap the collector’s clock when the pull that found it returned
Device-info record the collector’s clock: board facts have no timestamp of their own

forward writes to eleven sinks. Most setups need one of the first two rows; the rest exist so mikroscope fits what you already run.

If you… Use It carries Dashboard
want the whole thing, with the dashboards, and have nothing yet --influx every measurement, as line protocol yes, generated
already run Prometheus --prom every family, recomputed from the samples yes, generated
want to capture a window and look at it later --file the merged timeline as JSONL, nothing to install no
keep long-term data in PostgreSQL or TimescaleDB --sql DDL and INSERTs for psql, no driver yes, generated, on a datasource you create
write straight into a running PostgreSQL --postgres the --sql statements, down a connection yes, generated
want the kernel log and the detections where your logs are --loki events only — kmsg, detections, gaps, triggers, API errors, device records no
already run an OpenTelemetry pipeline --otlp metrics as OTLP/HTTP no
already run Graphite or Elasticsearch --graphite, --elastic every measurement, in that product’s shape yes, a smaller one
already run Telegraf --telegraf every measurement, as line protocol no
want to pipe it into something of your own --stdout line protocol or NDJSON on standard output no
  • Name several at once: --file beside a store gives you a capture to go back to, and --loki beside --influx puts the kernel log where a log query can reach it while the numbers go to the store the dashboards read.
  • Loki takes events, not metrics — the kernel-log records, the detections, the gaps, the triggers, the API tier’s errors and the device records — so a Loki-only run has no CPU or memory numbers at all.
  • --prom is scraped, not pushed: forward serves /metrics and Prometheus comes to it, so the collector has to be reachable from the Prometheus host.

Six sinks feed the five dashboards: --influx, --prom, --postgres and --sql (which share the PostgreSQL one), --graphite and --elastic. Add --grafana <url> and put a Grafana service-account token, Admin role, in GRAFANA_TOKEN: at every start, before it reaches the router, the collector creates or corrects a datasource and publishes the dashboard for each of them it writes to:

Terminal window
export GRAFANA_TOKEN=…
mikroscope forward --influx http://influx:8181 --influx-db mikroscope --grafana http://grafana:3000
  • Three describe their own datasource. --influx, --elastic and --postgres write to the server Grafana queries, so the datasource is built from their flags, at the address the collector uses. When Grafana reaches that store by another address, pass it in --grafana-datasource-url.
  • The other three need to be told. --prom is scraped and --graphite writes to carbon’s ingest port, so each needs the address Grafana queries in --grafana-datasource-url, or a datasource you already have in --grafana-datasource-uid. --sql never connects, so it takes --grafana-datasource-uid only.
  • One store per run for those two flags. Each names one datasource, so a run that writes to two stores with a dashboard and sets either is refused before anything is written. Leave them off the collector, which then publishes the stores that describe themselves and warns about the rest, and publish each of the rest once with mikroscope dashboards publish, given that store’s sink flag and its own value.
  • A failure does not stop the collector. It prints grafana: could not publish, carrying on without it: <reason>, one line per failure, and collects.
  • The Grafana comes from --grafana or MIKROSCOPE_GRAFANA_URL, never from GRAFANA_URL, which other Grafana tools may set (Grafana token).

The other five sinks have no dashboard. mikroscope dashboards publish, given the same sink flags and --grafana, publishes once without collecting. Set up in Grafana has what it creates, the other routes and the check.

Flag Destination URL or credential from the environment Shape Page
--file path.jsonl JSONL file — synchronous the file
--prom :9124 Prometheus /metrics on the collector host — in memory Prometheus
--influx URL, --influx-db DB InfluxDB 3 line protocol MIKROSCOPE_INFLUX_URL, MIKROSCOPE_INFLUX_DB, MIKROSCOPE_INFLUX_TOKEN queued InfluxDB 3
--sql path or --sql - PostgreSQL / TimescaleDB statements, for psql — synchronous SQL
--postgres DSN the same statements, into a running PostgreSQL MIKROSCOPE_POSTGRES_DSN queued --postgres
--stdout lp or --stdout json standard output — queued stdout
--loki URL Loki push API: events, not metrics MIKROSCOPE_LOKI_URL, MIKROSCOPE_LOKI_TOKEN, MIKROSCOPE_LOKI_TENANT queued Loki
--otlp URL OTLP/HTTP metrics, JSON encoding MIKROSCOPE_OTLP_URL, MIKROSCOPE_OTLP_TOKEN queued OTLP
--graphite host:port carbon plaintext over TCP MIKROSCOPE_GRAPHITE_ADDR queued Graphite
--elastic URL Elasticsearch or OpenSearch _bulk MIKROSCOPE_ELASTIC_URL, MIKROSCOPE_ELASTIC_AUTH queued Elasticsearch
--telegraf URL a Telegraf listener over HTTP, TCP or UDP MIKROSCOPE_TELEGRAF_URL, MIKROSCOPE_TELEGRAF_TOKEN queued Telegraf
  • Sink credentials never come from a flag: a flag is visible in ps and in a shell history. Each sink token is read from its MIKROSCOPE_* variable only. The agent’s own bearer token is the exception: forward takes it as --token, default MIKROSCOPE_TOKEN.
  • --host-tag (MIKROSCOPE_HOST_TAG, default router) puts the same host tag or label on every point in every sink.
  • A sink that was asked for and cannot be constructed — a port already bound, a file that cannot be opened — fails the run rather than going silently absent.

The runs the collector has made against a router and the sinks it has fed are on the evidence page, and the stores each sink has been tested against under Test suites.

The pull loop never waits on a destination. Every sink that talks to a remote renders into memory and hands the bytes to a bounded queue that a goroutine of its own drains once a second:

  • The queue holds --queue-seconds (default 60) seconds’ worth of a byte budget: 64 KiB per second for InfluxDB, PostgreSQL (--postgres), Loki, OTLP, Elasticsearch, Telegraf and stdout, 256 KiB per second for Graphite, whose one-line-per-value format is bulkier.
  • Past the budget the oldest batch is evicted and counted; the newest is always kept, because fresh telemetry beats stale.
  • A failed delivery backs off 2 s, doubling to 60 s, and logs at most one line per minute. Everything else is in the counters.
  • Each HTTP post carries a 10 s timeout, and a dead pooled connection is an error that is retried and counted, not a silent resend. --postgres sends each batch as one transaction with a 30 s timeout.

Three sinks are not queued. The file and SQL sinks write synchronously through a 64 KiB buffer, because a local file does not stall the way a remote does; a write error counts one error and one drop. The Prometheus sink updates in-memory state under a lock and serves it on scrape.

forward prints written, dropped and errors per sink, and the unit differs by shape:

  • Queued sinks count batches. written is one batch the destination accepted, dropped one batch evicted by the byte budget, errors one failed attempt — a batch that fails three times and then lands is 3 errors and 1 written. Elasticsearch adds one dropped per document the cluster refused inside a 200 reply.
  • File, SQL and Prometheus count events: one per sample, trigger, API read, gap, detection or device record accepted.

At start, on standard error: with --grafana, the grafana: lines of the publish first (Dashboard in Grafana); then one sink: <name> line per sink, the API tier’s settings, and the agent’s version, rate, sequence number, skew, transport and effective batch. Every minute, on standard error, a running report:

forwarded <n> kernel, <n> api, <n> gap(s), <n> trigger(s), <n> detection(s), <n> agent restart(s), last seq <n>; <sink>: <n> written, <n> dropped, <n> errors

On exit, on standard output, the totals and one line per sink:

forwarded <n> kernel samples, <n> api samples, <n> gap(s), <n> skew jump(s)
<sink>: <n> written, <n> dropped, <n> errors

A processor is cpu everywhere — InfluxDB tag, SQL column, Prometheus label, OTLP attribute — never core. The wire carries the kernel’s own unit and names it in the field (_khz, _kb, _ticks, _pages, _sectors); each sink converts once, to that store’s convention, and converts a value and its ceiling identically. So temperature is celsius beside critical_celsius, and block-device busy time is io_s.