Skip to content

Troubleshooting

Find the line you are looking at. Each row gives its cause and the fix; a link leads to the longer explanation below. doctor prints a check as MISSING (install stops) or WARN (install goes on), followed by a fix: line.

doctor’s checks, in the order it prints them:

You see Cause Fix
MISSING RouterOS 7.24 or later the router runs an older RouterOS upgrade RouterOS to 7.24 or later, and the container package with it (details)
MISSING architecture has a container package MikroTik publishes the container package for arm, arm64 and x86_64 only none: no agent can run on this router
MISSING architecture matches --arch <arch> --arch names another architecture than the router’s re-run with the --arch the fix names, or leave --arch out and install reads it from the router
MISSING architecture matches the --agent-tar image the tar is for another board download mikroscope-agent-<arch>.tar, the asset the fix names
WARN the router picks the image's architecture, on arm which of the two 32-bit ARM images RouterOS pulls from the index on an EN7562CT board is not known if the container stops with Exec format error, install from mikroscope-agent-armv5.tar with --agent-tar (details)
MISSING container package installed and enabled, found=0 the package is not on the router download it for the router’s architecture and RouterOS version, upload it and reboot (details)
MISSING container package installed and enabled, found=1 enabled=0 the package is there and disabled /system/package/enable container, then reboot
MISSING device-mode container=yes containers are off, and only a confirmation at the device turns them on /system/device-mode/update container=yes, then confirm it as the console asks (details)
MISSING free memory ≥ <size> less free memory than --memory-max free memory on the router; the agent needs its --memory-max plus headroom
WARN free memory leaves room for the pull less than --memory-max plus 16 MiB free, and the router extracts a pulled image before the agent starts if the pull fails, install from a tar with --agent-tar
MISSING free flash ≥ <size> the flash cannot hold the tar and the root, or the root of a pulled image, with 4 MiB to spare free space on the flash, or --disk tmpfs or --ephemeral where a tmpfs disk exists
MISSING disk <disk> exists no disk in that slot; with --ephemeral, no disk in slot tmpfs /disk/add type=tmpfs tmpfs-max-size=64M slot=tmpfs for a RAM disk, or name an existing disk with --disk
MISSING disk <disk> has ≥ <size> free not enough room on that disk free space on it, or a larger tmpfs-max-size
MISSING disk tmpfs is RAM --ephemeral, and slot tmpfs holds a disk that is not a RAM disk free the slot for a tmpfs disk, or install with --disk <slot> without --ephemeral
WARN start-on-boot suits a root in RAM the root is on a tmpfs disk, which a reboot empties, and start-on-boot=yes --start-on-boot no, or --ephemeral
MISSING veth name <veth> is free or ours a veth of that name exists without this install’s tag another --veth and --subnet, or remove the veth by hand if it is yours (details)
MISSING envlist <name>-env is free or ours an envlist of that name exists without this install’s marker another --name
MISSING install manifest <path> is free or ours a file at the manifest’s path is not this install’s manifest move it away, or pick another --name
MISSING container name <name> is free or ours a container of that name exists without this install’s tag another --container-name
MISSING subnet <subnet> does not overlap a route a route or a network on the router collides with the /30 another /30 with --subnet (details)
MISSING interface list <list> exists the list the veth would join does not exist --iface-list none when the fix offers it first; otherwise /interface/list/add name=<list>, or the --iface-list your drop rule uses
MISSING no firewall rule drops the agent's replies a firewall rule drops the agent’s replies with the plan’s list memberships the --iface-list and --addr-list the fix names, or an accept rule before that rule (details)
WARN no firewall rule drops the agent's replies, … may drop them a rule matches on something doctor does not judge: a destination, a mark, a rate if the agent does not answer after install, look at that rule first
WARN no firewall rule doctor reads is invalid or names a deleted list a rule RouterOS passes over, or one that names an interface list that was removed and matches as an empty list fix or remove what the rule names, or set its list again by name (details)
MISSING --lan-address <address> is the router's with --expose, no interface of the router holds that address the address the router has on its LAN, as /ip/address/print lists it
WARN --lan-address is not on the uplink that address’s interface carries the default route or is in the WAN list the router’s LAN address; the dst-nat would otherwise publish the agent on the Internet side
WARN no registry credential meant for another registry a /container/config username is set for another registry than the image’s Registry credentials
WARN the installed agent published on the LAN asks for a token an install of this --name has a tagged dst-nat and no TOKEN mikroscope upgrade --name <name> --token <secret>, or mikroscope uninstall --name <name> --yes, which removes the agent and its LAN rules
WARN the router does not answer DNS from its uplink allow-remote-requests=yes, and no rule of raw prerouting or filter input drops a query on the uplink a drop rule for port 53 on the uplink, or allow-remote-requests=no (details)
WARN nothing tagged for <name> that these flags do not select an earlier install with other flags left tagged objects this plan does not name mikroscope status and mikroscope uninstall with no shape flag: they read how the install was made
a fix that reads doctor could not read … the router printed something other than the answer, as an older RouterOS does for a menu it lacks run that read by hand on the router to see why

Other messages of install, upgrade and uninstall:

You see Cause Fix
iface-list "…" is a RouterOS built-in list, which takes no member --iface-list all, dynamic or static the list your drop rule uses, or none
the install named … on the router does not match the flags a flag or MIKROSCOPE_* variable contradicts the install on the router drop it: the verb uses what the router holds
the install named … asks for a token and upgrade would write its envlist without one the install has a token, and upgrade got none --token, or MIKROSCOPE_TOKEN
nothing to upgrade: run install first no install of this --name on the router install, or the --name the install was made with
no Go toolchain on PATH building the agent needs Go and a checkout --remote-image or --agent-tar
--agent-tar …: this is not a mikroscope agent image the wrong file the release’s mikroscope-agent-<arch>.tar (Offline install)
mikroscope: the image was not extracted within 120 s; mikroscope.tar stays RouterOS took longer than --extract-timeout to extract the tar a longer --extract-timeout, up to 600s; uninstall removes the tar with the container
RouterOS: unknown parameter privileged RouterOS older than 7.24, reached with --no-doctor upgrade RouterOS (details)
Container log: exec format error the image is for another architecture than the board the image for the board (details)
registry-url is https://…, from an older CLI an older CLI sent the reference without its registry host upgrade the CLI, or install from the tar (details)

doctor’s health section, which reads the running agent’s ring:

You see Cause Fix
WARN layer2-loop the bridge receives its own frames back on a port find the second path and break it (details)
WARN stp-churn a second path to the router, or a topology that keeps changing behind that port Requirements
WARN link-flap the cable, the connector, the device at the other end rebooting, or auto-negotiation failing check the link at both ends
WARN softnet-drops packets lost inside the router, arriving faster than the receive path could take them Packet flood
health … skipped: no agent answered the agent is not running, or this host cannot reach its /30 pass the --subnet and --port it was installed with; see Network access
skipped: the agent answered /healthz but its ring could not be read most often the agent has a TOKEN pass --token

MikroTik gates containers behind a switch that cannot be flipped over the network. /system/device-mode/update container=yes starts it, and the console then asks for a confirmation at the device within five minutes. doctor’s fix quotes both prompts:

  • A router prints update: please activate by turning power off or pressing reset or mode button: press the reset or mode button, or cut the power.
  • CHR prints update: turn off power in 5m to activate changes: power the VM off and on again within 5 minutes.

No flag, no script and no version of this tool can do that step for you. Arrange it first, because everything else waits on it: Requirements.

The container package is a separate download from mikrotik.com, per architecture and per RouterOS version. Upload it and reboot. doctor counts it as present only when it is installed and not disabled; a disabled one takes /system/package/enable container and a reboot.

RouterOS 7.24 added privileged=, and the container step writes it, so an earlier 7.x fails there. doctor stops an install on such a router before it writes anything (RouterOS below 7.24); with --no-doctor it fails at the container step, after the tar has been uploaded, which is why the install then takes it back with it. Upgrade RouterOS to 7.24 or later: --privileged=false does not help, because the step then writes privileged=no, the same unknown parameter (internal/router/stepspec.go). Privileged mode says what the setting changes: without it the agent cannot read /dev/kmsg, and the kernel log is where several of the fault signatures start.

The container starts and dies immediately, and the log says exec format error. The image is for another architecture than the board, and on 32-bit ARM, “arm” is not one architecture. MikroTik’s container documentation says that devices with an EN7562CT CPU support only arm32v5 container images, while its other 32-bit ARM boards run an ARMv7 userland:

Board Image that runs Release asset
hEX Refresh, hEX S (2025), any EN7562CT board linux/arm/v5 only mikroscope-agent-armv5.tar
other 32-bit ARM linux/arm/v5 or linux/arm/v7 mikroscope-agent-armv5.tar or mikroscope-agent-armv7.tar
  • With --remote-image, the router picks from the image’s index; on an EN7562CT board, which of the two it pulls is not known, and doctor warns.
  • With --agent-tar, take mikroscope-agent-armv5.tar when the board is 32-bit ARM and you are not certain which kind it is.
  • Building from a checkout, --goarm 5 is the default for the same reason.

The agent needs RouterOS 7.24 or later, and doctor checks the version first: MISSING RouterOS 7.24 or later (<version>), with the fix to upgrade RouterOS (/system/package/update) and the container package with it. The rest of the report still reads, so fix everything it names in one visit.

MISSING veth name veth-mikroscope is free or ours (found=1 ours=0): a veth with that name exists and does not carry the install’s tag. mikroscope does not build on an object it did not create, and will not remove it either. Pick another --veth, and another --subnet with it, for a second install beside something else; or remove the veth by hand if it is a leftover of your own. The envlist <name>-env and --container-name are checked the same way.

MISSING subnet 172.30.10.0/30 does not overlap a route (routes=172.30.10.0/30 via ether2): a route of the main table lies inside the /30 (an inactive or disabled one too, which can become active), or a network on another interface holds its router end, so the agent’s address would be ambiguous. Pick another /30 with --subnet, one no interface and no route of the router uses. A route that only contains the /30, such as a default route or a wider prefix to a VPN, is no overlap, and neither is a blackhole route.

MISSING no firewall rule drops the agent's replies, followed by the rule, as /ip/firewall/raw rule 0 (chain=prerouting action=drop, …) "…" drops them. doctor reads the raw prerouting and filter forward and input rules and follows the agent’s replies through them with the list memberships the install would write.

  • The fix names the --iface-list and --addr-list that let them through: install with --iface-list LAN --addr-list LANs: ….
  • When no list does, as for a rule written with src-address=!192.168.88.0/24 instead of an address list, the fix says add an accept rule for in-interface=<veth> before it, or pick a --subnet inside 192.168.88.0/24.
  • A WARN on the same line means a rule might drop them, because it matches on something doctor does not judge: if the agent does not answer after install, that rule is the first to look at.

Firewall lists explains the drop rules and the trap check.

WARN no firewall rule doctor reads is invalid or names a deleted list, followed by each such rule.

  • … is invalid (vprobe not ready): the rule names an interface that was removed or is not ready, and RouterOS passes over it, so an invalid drop drops nothing. /ip/firewall/filter/print (or raw) marks it with an I. Fix what it names, or remove the rule.
  • … is invalid, and RouterOS gives no reason: a rule changed a moment before reads so until RouterOS has applied it. Run doctor again.
  • … (in-interface-list=!*2000010) … names a deleted interface list: the list was removed after the rule was written. RouterOS keeps the rule, valid, with the list’s id in place of its name, and matches it as if the list were empty. A ! of it matches every packet: the default configuration’s in-interface-list=!LAN drop left that way drops all input, the router’s own SSH and Winbox from the LAN included. Creating a list of the same name does not repair it; set the rule’s list again by name, from the console, /ip/firewall/filter/print and then set <number> in-interface-list=!LAN, or from WebFig or Winbox.

WARN the router does not answer DNS from its uplink (allow-remote-requests=yes, uplink ether1: no rule drops a query that comes in on it). The router answers DNS for anyone its firewall lets a query in from, and with nothing dropping queries on the uplink that is the Internet: reflection attacks find such resolvers and send them a flood that fills the connection table and the CPU, and every panel then measures the flood. Either:

  • drop the queries on the uplink, before any rule that accepts them. doctor’s fix names the two rules for your uplink, at the head of the input chain:

    /ip/firewall/filter/add chain=input in-interface=ether1 protocol=udp dst-port=53 action=drop place-before=[:pick [/ip/firewall/filter/find where chain=input] 0]
    /ip/firewall/filter/add chain=input in-interface=ether1 protocol=tcp dst-port=53 action=drop place-before=[:pick [/ip/firewall/filter/find where chain=input] 0]

    On a router whose input chain has no rule, leave place-before out: it names the chain’s first rule, and there is none.

  • or stop answering remote queries, if no host on the LAN uses the router as its resolver: /ip/dns/set allow-remote-requests=no.

… may drop the queries means a rule matches on something doctor does not judge, a source address list say, and the queries pass unless it takes them. doctor reads IPv4 only: an IPv6 uplink needs the same rule in /ipv6/firewall/filter.

mikroscope sends the whole reference in remote-image=, registry host included, so /container/config registry-url, one setting for the whole device, decides nothing about the pull, and Docker Hub and GHCR serve the agent with no registry login (Registry settings). An older CLI sent the reference without its host, and RouterOS took the host from registry-url (Upgrade notes). With such a CLI, upgrade it; or set the setting as its fix line says, which changes it for every other container on the router too; or install from the tar with --agent-tar.

/container/config holds one username and password for the device. A Docker Hub account presented to GHCR fails, and the container stays in error with auth error in its log, for an image anyone can pull anonymously, when registry-url names the registry the pull goes to. Whether RouterOS also presents the username to a host named only in remote-image= is not known (Tested on), so doctor warns whenever a username is set and the host registry-url names is not the host the image comes from, and upgrade prints the same check before it removes anything.

  • Install from a tar with --agent-tar, which pulls nothing.
  • Or pull from the registry the username belongs to: the Docker Hub reference for a Docker Hub login.
  • Or clear the username if nothing else on the router needs it. The password cannot be read back once cleared, so do that only knowing where it lives.

An empty registry-url with a username set warns too, because doctor cannot tell which registry the username is for; if you know it is a Docker Hub login and you pull the Docker Hub reference, the warning is advice you can ignore.

The bridge is receiving its own frames back on that port: something behind it reaches the router by a second path. A mesh node with both a cable and a wireless backhaul is the usual cause, and so is a switch cabled twice. STP does its job and blocks the port, so nothing melts down, and RouterOS reports the port running and error-free, but whatever is behind it reaches the router some other way or not at all (Layer-2 loop is a case read end to end). Find the second path and break it; the finding goes away within one ring.

doctor reads only the ring, about the last minute, so a loop that has already cleared or comes and goes can be gone by the time it runs. The alert mikroscope-bridge-port-dark reads the same fault from history: a bridge port that received packets for ten minutes while the bridge sent it neither a unicast nor a broadcast frame.

Terminal window
mikroscope status

That prints the ownership counts and, if it can reach the agent, its health.

You see Cause Fix
after install: the container is not running on the router the container never started, or stopped /log/print where topics~"container" on the router
after install: the container runs; this host cannot reach … a firewall rule, or no route from this host to the /30 doctor’s trap check, then Network access: a static route, the relay or --expose
status: the counts, then agent: not reachable from this host (…) the container is not running, a firewall rule, or no route to the /30 /container/print and /log/print where topics~"container" on the router; then as the row above
You see Cause Fix
No events at all, ever the container runs unprivileged, and /dev/kmsg needs privileged=yes install or upgrade without -privileged=false
No PMU panels, no cycles or instructions perf_event_open is unavailable on that kernel or board none; the dashboard moves those panels to “Not available on this device”
A panel says No data and the others are fine that measurement is not produced on this device dashboards import or dashboards publish, or a collector run with --grafana at its next start, moves such panels into their row (InfluxDB, Prometheus); dashboards check lists them
A panel shows a red error badge the query failed: the datasource, not the data Stores and dashboards
forward prints … dropped for a sink the sink could not keep up Sink drops
Loki accepted everything and a query returns nothing a push is not queryable until the chunk flushes query again after the flush
Numbers stop at a round moment and resume a gap: the ring wrapped before the collector pulled it keep the collector running, or a longer --buffer
The kernel tier stops dead and the API tier carries on the agent restarted Kernel tier stops after restart
The API panels are blank and the kernel panels are fine the API connection died API panels blank

The dashboards carry a row named “Not available on this device” for exactly this: panels whose measurement the kernel or the board does not produce are moved into it rather than left to draw an empty graph among the others. mikroscope dashboards import asks the datasource which measurements it really holds and does that sorting for your store, on InfluxDB and Prometheus, and dashboards publish and a collector run with --grafana do the same. dashboards check reports what is empty but changes nothing in Grafana: Store probe.

The agent numbers its samples from 1 at every start, so an agent that restarts (an upgrade, a container restart, a reboot) has a newest sequence number far below the collector’s cursor. The collector notices on its next health read, which is once a minute, logs

agent restarted: its newest sample is 571 and the cursor was 1737212; resuming from 1

and resumes from the new ring’s oldest sample, so what the agent took while nobody was collecting is picked up rather than skipped. The minute’s worth of samples between the restart and the health read is lost with the container, not by the collector. When the router rebooted, and not only the agent’s container, the next line says so, from the kernel’s boot id the agent reports:

router rebooted: the kernel's boot id went from 6f1c3d2a-… to 0c9e6b1f-…

and, with an API tier, what RouterOS logged about the boot, read once (Boot log):

router's boot log: "router rebooted by ssh-cmd:admin@192.168.88.10/reboot"

A collector that has not noticed keeps its cursor where it was, the agent’s ring answers an empty batch to every pull, and the kernel tier stops while the API tier keeps counting, so the run looks healthy. Restarting the collector resets its cursor from the health read at start.

The mirror image of the entry above, and it has the same cause: a reboot.

The API tier holds one persistent RouterOS API connection. When the router goes away (a reboot, a RouterOS upgrade, an operator restarting the API service) that socket dies. While it is dead, everything the API feeds is blank: interface throughput and packet rate, the per-port counters, RouterOS cpu-load, and “Reboots in the window”, which reads uptime_s and so cannot count the very reboot that broke its own source.

The tier reconnects on its own. A transport failure (EOF, a broken pipe, a connection reset, a command timeout) reopens the connection, at most once every five seconds, and the round is retried on the new one. A !trap does not reconnect: that is a live router refusing a command, and asking it twice only spends the router’s CPU. A !fatal does, because that is the word RouterOS sends as it closes the session. The inventory is re-read afterwards, since an upgrade is exactly when an interface can change its name, type or bridge. You will see

api tier: reconnected (1 since start)
api tier: recovered after 137 failed round(s)

and the minute report grows an api: N failed round(s), N reconnect(s) clause for as long as either is nonzero.

Every network sink is queued, and the queue is bounded: --queue-seconds, 60 by default. A destination that cannot keep up loses the oldest batch rather than stalling the pull loop, and the count is printed in the minute report and again at the end of the run. There is no metric family for it, on purpose: Run the collector explains that the agent’s ring is what protects the data, and a collector waiting on a slow store would lose more than the store does.

You see Cause Fix
No mikroscope dashboard in Grafana after install install writes only to the router run the collector with --grafana, or dashboards publish (Set up in Grafana)
forward prints no grafana: line no Grafana named: forward reads --grafana or MIKROSCOPE_GRAFANA_URL, never GRAFANA_URL pass --grafana (Grafana variables)
--grafana needs GRAFANA_TOKEN no token in the environment of forward or dashboards publish a service-account token in GRAFANA_TOKEN (Grafana token)
dashboards publish needs --grafana … no Grafana named, in the flag or in MIKROSCOPE_GRAFANA_URL or GRAFANA_URL pass --grafana
dashboards publish needs the sink flags the collector runs with no sink flag given: each datasource is described from its sink the collector’s sink flags
no sink this builds a dashboard for is configured only sinks with no dashboard add --influx, --prom, --postgres, --graphite or --elastic (Dashboard in Grafana)
--grafana-datasource-url names one datasource for every store (or …-uid names, or both flags name) a run of several stores that sets --grafana-datasource-url or -uid, or its variable publish that store on its own (details)
--grafana-datasource-sslmode "…" is not a mode Grafana's PostgreSQL datasource has a mode the datasource does not have, such as prefer disable, require, verify-ca or verify-full, in lower case
--grafana-dry-run needs --grafana a dry run of forward with no Grafana named: there is no publish to preview add --grafana, or leave the dry run out
import/check need --grafana, GRAFANA_TOKEN and --datasource-uid one of the three is missing pass --grafana (or set GRAFANA_URL) and --datasource-uid (Import with the CLI)
could not ask … which measurements it holds the reason in brackets: a store with no data yet (the datasource reported no mikroscope measurements), a datasource Grafana cannot reach or query, or a store that cannot be probed (cannot probe a "…" datasource: PostgreSQL, Graphite, Elasticsearch) no data yet: publish again once it holds data; unreachable: the address Grafana reaches the store at, in --grafana-datasource-url for a datasource the collector made; cannot probe: nothing (Store probe)
red table … not found badges on an InfluxDB dashboard uploaded by hand the file carries the compiled defaults dashboards import or dashboards publish (Import manually)
flightsql: Unauthenticated on every InfluxDB panel the datasource has only one of its two credential fields set set both, or let forward --grafana build it (details)
tls: first record does not look like a TLS handshake the FlightSQL side tries TLS against a plain-HTTP InfluxDB insecureGrpc (details)
Empty panels, and the SQL runs fine outside Grafana a result with no time-typed column a time column (details)
the <name> sink does not know the address Grafana would query --prom and --graphite cannot describe their own datasource --grafana-datasource-url or --grafana-datasource-uid (details)
--influx is a write URL this cannot take apart a write URL that is not InfluxDB 3’s the server and --influx-db, or adopt a datasource (details)
standard_conforming_strings = off the server would misread the statements turn it on for the connection, or use --sql (details)
database "…" does not exist --postgres does not create databases createdb mikroscope (details)
grafana: could not publish, carrying on without it the publish failed, and the collector keeps collecting read what the server said; --grafana-dry-run (details)
uninstall lists the same tables every time InfluxDB 3 renames a deleted table nothing: they are filtered (details)
uninstall --targets data says a sink stores nothing it can remove --prom, --graphite and --sql store nothing this can delete remove it where it lives (details)
--targets dashboard left the datasource behind it was adopted nothing (details)
uninstall takes no --grafana-dry-run uninstall’s dry run is leaving out --yes drop the flag: without --yes it lists what would go and removes nothing
uninstall: asking Grafana whether dashboard … is there Grafana could not be read: unreachable, the token refused, or a server error fix the URL or the token and run it again; nothing was removed

flightsql: Unauthenticated, from an InfluxDB datasource

Section titled “flightsql: Unauthenticated, from an InfluxDB datasource”

The InfluxDB plugin has two transports and reads the credential from a different place on each: the HTTP calls take the Authorization header, the FlightSQL ones take token. A datasource with only one of them set answers this on every panel while the store itself is healthy. Set both secure fields, token = the token and httpHeaderValue1 = Bearer <token>, or let forward --grafana build the datasource, which sets both.

tls: first record does not look like a TLS handshake

Section titled “tls: first record does not look like a TLS handshake”

The same plugin, the same two transports, and the other half of the same trap: the FlightSQL side attempts TLS unless insecureGrpc is set, so a datasource pointed at a plain-HTTP InfluxDB answers this on every panel. Set insecureGrpc on the datasource, or use forward --grafana, which follows the scheme of the URL the sink writes to.

Grafana’s InfluxDB SQL plugin rejects a result whose columns are all numeric and none is time-typed, so a panel whose query returns one row of numbers renders its no-value text, which looks exactly like a quiet device. If you are writing a panel, give it a time column even when the panel does not plot one; mikroscope dashboards check runs every panel’s query through Grafana’s own API and is the fastest way to tell a broken query from a quiet one.

the <name> sink does not know the address Grafana would query

Section titled “the <name> sink does not know the address Grafana would query”

forward --grafana derives a datasource from the sink’s own address for --influx, --elastic and --postgres. --prom is scraped and --graphite writes to the carbon ingest port, so neither knows where Grafana would query: pass that address in --grafana-datasource-url, or create the datasource in Grafana and name it in --grafana-datasource-uid, which also tells forward to leave it alone. Either flag is for that store alone: a run of several stores that sets one is refused (One store per run). --sql never connects and fails with its own message, which points at --postgres or a datasource named in --grafana-datasource-uid.

--grafana-datasource-url and --grafana-datasource-uid each name one datasource, and two stores are two servers read by two plugins, so a run that publishes more than one store and sets either, or its MIKROSCOPE_GRAFANA_DATASOURCE_* variable, is refused before any request, dry run included. Publish the store that needs one on its own, given only its sink flag:

Terminal window
mikroscope dashboards publish --prom :9124 --grafana http://grafana:3000 \
--grafana-datasource-url http://prometheus:9090

and run the collector without the datasource flag. With --grafana it still publishes the stores of --influx, --elastic and --postgres, which describe their own datasource, and warns at every start about the one it cannot describe.

--influx is a write URL this cannot take apart

Section titled “--influx is a write URL this cannot take apart”

--influx may be a server (http://host:8181, with --influx-db) or a full write URL, and a full one is used verbatim. A URL whose shape is not InfluxDB 3’s /api/v3/write_lp?db=…, a v2 /api/v2/write say, writes as well as ever and cannot describe a datasource, because nothing can read the database name back out of it. Pass the fields, or adopt a datasource by uid.

--postgres refuses a server that answers this, rather than writing to it. The statements quote by doubling the single quote and nothing else, so with the setting off a kernel-log message ending in a backslash escapes its own closing quote and every statement after it is parsed as string content. Turn it on for the connection, ALTER ROLE … SET standard_conforming_strings = on, or use --sql and apply the file, whose header sets it itself.

--postgres connects to a database; it does not create one. The sink declares its tables on the first batch and nothing else, because creating databases is not a collector’s job. Run createdb mikroscope first. The collector goes on collecting meanwhile: the failure is counted against the sink and logged once a minute, and the other sinks are unaffected.

The collector says grafana: could not publish, carrying on without it

Section titled “The collector says grafana: could not publish, carrying on without it”

That is the designed behaviour, not a half-failure to chase. Refusing to start would trade the samples of the hour spent not running, which cannot be recovered, for a dashboard published on the next restart, which can. The line carries what the server said. --grafana-dry-run prints what it would write without writing anything, and runs before the router is touched.

Each store that failed has a line of its own that names it, the datasource for <store>: … or the dashboard for <store>: …, and the stores after it were still published. The reasons it prints most: --grafana needs GRAFANA_TOKEN, no sink this builds a dashboard for is configured, the <name> sink does not know the address Grafana would query, --grafana-datasource-url names one datasource for every store, and what the server said when it refused a write (Grafana token). dashboards publish prints the same reasons after mikroscope:, one line per store that failed, and exits 1.

InfluxDB 3 deletes a table by renaming it and leaving the entry in information_schema under a name carrying the deletion instant. uninstall filters those out; without the filter it would list them on every run, delete them successfully (the store accepts a delete of a name it has already retired) and never empty. If you are looking at mikroscope_cpu-<timestamp> in a query result, that table is already gone.

uninstall --targets data says a sink stores nothing it can remove

Section titled “uninstall --targets data says a sink stores nothing it can remove”

Three of them do not. --prom is scraped rather than written to, so the series live in a Prometheus this has never heard of; --graphite offers no delete at all, and its whisper files have to go by hand; --sql writes a file, and rows already loaded from it into a real database have to be removed there, or with --postgres pointed at that database. Each prints its own reason, because silence would read as nothing to remove.

It was adopted. A datasource named in --grafana-datasource-uid was somebody else’s before the collector ran and is somebody else’s after: publishing leaves it alone and so does removing.

Once the data is arriving, the question changes from “why is this broken” to “what is this telling me”. Diagnose faults has the fault signatures and the checks to make before you trust a reading.