Set up in Grafana
mikroscope install puts nothing in Grafana: it writes only to the router. The dashboards reach
Grafana by one of three routes, and the collector’s is the shortest because it creates the
datasource as well. Whichever route, mikroscope dashboards check then runs every panel’s query
through Grafana’s own API. That proves the queries return rows, not that a reader can read the
result.
Choose a route
Section titled “Choose a route”| Route | Command | Datasource | Use it when |
|---|---|---|---|
| From the collector | forward … --grafana <url>, or once without collecting: dashboards publish with the same flags |
created from the sink flags, or adopted with --grafana-datasource-uid |
you want the datasource made for you; the collector publishes again at every start |
| With the CLI | dashboards import --store <store> --datasource-uid <uid> |
one that already exists | you have the datasource, or the collector cannot reach Grafana |
| By hand | Grafana → Dashboards → New → Import | one that already exists | no token for mikroscope |
Scroll sideways to see every column
Every route needs a store the collector writes to (Run the collector). The
dashboard’s uid is mikroscope-<store> on every route, so importing again replaces the dashboard
instead of adding a second one.
No store or Grafana yet: deploy/ has two compose stacks, each with Grafana and the
collector. The InfluxDB one publishes the dashboard when the collector starts, once GRAFANA_TOKEN
is set.
Grafana token
Section titled “Grafana token”The collector and the CLI read a Grafana service-account token from GRAFANA_TOKEN only; no flag
takes it (Grafana’s service accounts).
What it must be allowed to do depends on the verb:
forward --grafanaanddashboards publishcreate the folder and create or correct the datasource, anduninstall --targets dashboarddeletes it: give that service account the Admin role.dashboards importwrites the dashboard and nothing else, anddashboards checkwrites nothing: it runs queries.--grafana-dry-run, given with--grafana, contacts no Grafana and needs no token.
forward takes the Grafana from --grafana or MIKROSCOPE_GRAFANA_URL only, never GRAFANA_URL;
which variable the other verbs read is in
Environment variables.
Publish from the collector
Section titled “Publish from the collector”Point forward at a Grafana and, when it starts, before the first sample, the collector creates
the datasource and publishes the dashboard for each store it writes to:
export MIKROSCOPE_INFLUX_TOKEN=… # the sink's token; leave it out for a store without authexport GRAFANA_TOKEN=…mikroscope forward \ --influx http://influx:8181 --influx-db mikroscope \ --grafana http://grafana:3000On standard error:
grafana: folder "mikroscope" (bfyr3khdp41z4b) createdgrafana: influxdb: datasource mikroscope-influxdb (influxdb) createdgrafana: influxdb: could not ask the datasource which measurements it holds (…); using the compiled defaultsgrafana: influxdb: dashboard http://grafana:3000/d/mikroscope-influxdb/…The third line appears while the store holds nothing yet (below). --grafana-dry-run prints what
it would write and stops. Check what it published with
mikroscope dashboards check --grafana http://
(Check every panel).
mikroscope dashboards publish, given the same sink and --grafana flags, does the same once
without collecting and exits; a refusal from Grafana is its exit status, where the collector
carries on. It prints the same lines to standard output, without the grafana: prefix, and
contacts no router. The sink flags only tell it which stores to publish and how to describe their
datasources: --prom serves no /metrics there, and nothing is written to a store.
mikroscope dashboards publish \ --influx http://influx:8181 --influx-db mikroscope \ --grafana http://grafana:3000What it makes has fixed names, so a second run writes to the same places: the folder
--grafana-folder names, found by its title or created, and per store a datasource and a
dashboard, both with the uid mikroscope-<store> (mikroscope-influxdb, mikroscope-elasticsearch,
mikroscope-prometheus, mikroscope-postgres, mikroscope-graphite). The datasource is also
named mikroscope-<store>; one you adopt keeps its own uid.
| Flag | Default | What it does |
|---|---|---|
--grafana |
empty | the Grafana to publish to (its variables). Empty, the default, publishes nothing |
--grafana-folder |
mikroscope |
the folder to publish into; --grafana-folder "" is Grafana’s General folder |
--grafana-datasource-uid |
empty | adopt an existing datasource by uid instead of creating one, and leave it untouched; one store per run |
--grafana-datasource-url |
empty | the address Grafana queries (host:port for PostgreSQL, a URL for the others), for the two sinks that cannot know it; it also overrides one a sink does know; one store per run |
--grafana-datasource-sslmode |
empty | disable, require, verify-ca or verify-full for the PostgreSQL datasource; any other value is refused |
--grafana-dry-run |
off | with --grafana, print what it would write, contact no Grafana and need no token; forward then stops before collecting, and refuses it without a Grafana |
Scroll sideways to see every column
Each flag but --grafana-dry-run also has a MIKROSCOPE_GRAFANA_* variable
(Environment variables). An empty variable counts as
unset, so the General folder takes the flag, --grafana-folder "".
- It is off unless you ask. A collector that wrote to somebody’s Grafana because it could is not
a collector anyone should run.
GRAFANA_TOKENis required for the same reason: some Grafanas accept an anonymous request, and one that did would write as whoever the server thinks is asking. - A failure here does not stop the collector. An unreachable Grafana, a rejected token and a
datasource the server will not take are each a warning,
grafana: could not publish, carrying on without it: <reason>, and a collector that goes on collecting. Every store is tried, and each one that fails gets a warning of its own that names it. Refusing to start would trade the samples of the hour spent not running, which cannot be recovered, for a dashboard published on the next restart, which can.dashboards publishprints the same reasons, one per store, and exits 1. - It publishes again at every start. The dashboard is imported over itself, so an edit made in
Grafana is replaced at the next start: save an edited copy as a new dashboard. The datasource it
created is corrected back to what the sink flags describe, and one that carries a token or a
password is rewritten at every start and reported
updated. - An empty store gets the compiled defaults. It publishes before the first sample, so on a
store that holds nothing yet the probe finds nothing and the dashboard carries the
compiled defaults. Restart the collector once the store holds data, or run
dashboards publishwith the collector’s flags, and it publishes the dashboard your store supports. PostgreSQL, Graphite and Elasticsearch are never probed and print the same line at every start. - It copies the sink’s credential. The InfluxDB datasource carries the sink’s own token, which
can write; the Elasticsearch one sends
MIKROSCOPE_ELASTIC_AUTHthe way the sink does, and the PostgreSQL one carries a password written in--postgres. For a read-only credential in Grafana, create the datasource yourself and adopt it with--grafana-datasource-uid. - It deletes nothing. A dashboard or datasource of a sink the collector no longer writes to
stays where it is: removing somebody’s dashboard unasked is not a thing a collector should do.
uninstall --targets dashboardremoves them (Update and remove). It publishes no alert rules either (Alert rules).
Datasource per sink
Section titled “Datasource per sink”Three sinks know an address Grafana can query, because it is the address
they write to or dial. Two cannot know it, and no amount of reading their
flags would find it — but told the address in --grafana-datasource-url,
there is nothing else to derive: a Prometheus datasource is a URL, and so is a
Graphite one.
| Sink | Datasource | Why |
|---|---|---|
--influx |
derived | the server it writes to, at the address the collector uses; --influx-db names the database |
--elastic |
derived | the base URL it _bulks to, at the address the collector uses; --elastic-index gives the index pattern |
--postgres |
derived | it dials the server, so the connection string has the host, port, database, user and sslmode |
--prom |
told the URL | it serves /metrics and is scraped: the Prometheus Grafana asks is one it has never heard of |
--graphite |
told the URL | it speaks the carbon ingest port, which is not the web API Grafana queries — usually a different port on the same host |
--sql |
never | it writes statements to a file and never connects, so no host, port, user or password exists anywhere in the flags |
Scroll sideways to see every column
mikroscope forward --prom :9124 --grafana http://grafana:3000 \ --grafana-datasource-url http://prometheus:9090 # Prometheus as Grafana reaches itOr name a Prometheus datasource you already have with --grafana-datasource-uid. The scrape job is
still yours to add (Prometheus).
Each of those two flags names one datasource, so a run that would publish two stores or more and
sets either is refused before anything is sent, dry run included. Publish such stores one at a time
with dashboards publish, each with its own sink flag and its own --grafana-datasource-url or
--grafana-datasource-uid, and run the collector without those two flags. With --grafana it
still publishes the stores that describe themselves (--influx, --elastic, --postgres) at every
start, and warns about each of the others, which it leaves as dashboards publish made them.
--sql is the one that can never be described. Its error names --postgres, the sink that can,
or a datasource you create in Grafana and name in --grafana-datasource-uid. An adopted datasource
is never corrected: reconciling one the collector did not create would overwrite settings someone
else chose.
The InfluxDB datasource has settings of its own:
InfluxDB URL fields
Section titled “InfluxDB URL fields”--influx takes the server and --influx-db the database, and the write URL is assembled from
them. A whole write URL, query string and all, is still taken as it is — the form
MIKROSCOPE_INFLUX_URL may already hold. But a write URL is the sink’s shape and the wrong shape
for everything else: Grafana wants the server and the database apart, and it will not take a write
path at all.
--influx http://influx:8181 --influx-db mikroscope # the URL is assembled--influx "http://influx:8181/api/v3/write_lp?db=mikroscope&precision=nanosecond" # taken as it isWhen --influx carries a path, the datasource’s database is read back out of that URL, never
from --influx-db: a datasource pointed at a different database than the one the sink fills is
worse than no datasource. A write URL this cannot take apart — a v2 /api/v2/write, say — still
writes fine and simply cannot describe a datasource, and forward --grafana says so instead of
building one that answers nothing.
Import with the CLI
Section titled “Import with the CLI”export GRAFANA_URL=http://grafana:3000 GRAFANA_TOKEN=…mikroscope dashboards import --store influxdb --datasource-uid <uid>| Flag | Default | Used by | What it does |
|---|---|---|---|
--store |
influxdb |
import, check | influxdb, prometheus, postgres, graphite or elasticsearch |
--grafana |
$GRAFANA_URL, then $MIKROSCOPE_GRAFANA_URL |
import, check | Grafana’s base URL |
--datasource-uid |
none, required | import, check | the datasource DS_MIKROSCOPE is bound to |
--no-probe |
off | import, check | skip asking the datasource what it holds; use the compiled defaults |
--window |
15m |
check | length of the query window |
--end |
now | check | the window’s right edge, RFC 3339 |
--var |
none | check | set a dashboard variable, name=value, repeatable; Graphite needs prefix and host, Elasticsearch host |
--out |
dashboards |
gen | the directory gen writes the eight files into, made if it is missing |
Scroll sideways to see every column
The token is GRAFANA_TOKEN (Grafana token). import creates no datasource:
bind it to one from Datasources, or to mikroscope-<store> once the collector or
dashboards publish has created it. Without a Grafana URL, a token and a datasource UID, both
commands stop with import/check need --grafana, GRAFANA_TOKEN and --datasource-uid.
import posts the dashboard to Grafana’s /api/dashboards/import with the datasource input
resolved to your UID, overwrite on, into the General folder (folderId 0). The dashboard’s uid
is fixed, so importing again replaces the same dashboard at the same URL. It prints that URL.
import has no folder flag.
Import manually
Section titled “Import manually”Take the file from the CLI you run, so it matches your version: mikroscope dashboards gen writes
the five dashboards and the three alert files into ./dashboards, making the directory if it is
missing, and the files are readable only by you. The release archives do not carry them, and
the dashboards/ directory in the repository follows main, which can be ahead of
your release.
Grafana → Dashboards → New → Import, as Grafana’s own import
guide describes it: upload the
mikroscope-<store>.json that matches your datasource (influxdb, prometheus, postgres, graphite or
elasticsearch), and pick the datasource when Grafana asks for DS_MIKROSCOPE.
A file uploaded this way carries the compiled defaults: the PSI and block-device panels (five;
three on Graphite and Elasticsearch) sit in the not-available row. On InfluxDB their queries are
switched off; on the other four stores they stay on, because a missing measurement is not an error
there (why). Every other panel ships with
its query on, whether your store holds its measurement or not. On InfluxDB a panel whose table or
column is missing then shows InfluxDB 3’s planning error, as a red badge, when its section is
opened. dashboards import and the collector’s publish avoid that. The detections annotation stays
on in a file uploaded this way; before the first detection it fails without showing anything, as
the annotations section describes.
Datasources
Section titled “Datasources”The CLI and manual routes bind the dashboard to a datasource that already exists, and the collector
adopts one with --grafana-datasource-uid. A datasource’s UID is the last path segment of its
settings URL in Grafana, /connections/.
InfluxDB 3
Section titled “InfluxDB 3”The database is the one the sink writes to, and the sink’s first write creates it (InfluxDB 3); the datasource’s token must be able to read it.
The datasource is type influxdb, version: SQL, dbName: mikroscope — the fields of Grafana’s
own provisioning example for InfluxDB 3 with SQL — with
both secure fields set:
httpHeaderValue1=Bearer <token>(the HTTP path)token=<token>(the FlightSQL path)
Without the second, panels fail with flightsql: Unauthenticated.
Prometheus
Section titled “Prometheus”Scrape the collector with one job, and point a Prometheus datasource at that Prometheus:
- job_name: "mikroscope" scrape_interval: 5s static_configs: [{ targets: ["<collector host>:9124"] }]mikroscope forward --prom :9124 carries every family: the kernel tier recomputed from the samples
it received, the collector’s own derived and detection families, and what only the sampler can
produce. The agent serves no /metrics; its tick timing rides in the samples and its own counters
come over GET /sampler, which the collector reads every minute. A second job would have nothing to
add.
PostgreSQL
Section titled “PostgreSQL”Two sinks fill it. --postgres writes the rows into a running PostgreSQL, and the collector can
build this datasource from its connection string. --sql writes a script instead:
forward --sql out.sql, then psql -f out.sql, or --sql - | psql. Until that script has been
applied, a datasource on that database answers every panel with “relation does not exist”.
The datasource is Grafana’s grafana-postgresql-datasource, with the database and the user the
sink wrote as. sslmode is yours to choose; postgresVersion only decides which syntax the plugin
may emit, and every query in this dashboard is plain SQL.
The panels are the InfluxDB ones, rewritten: the bucket macro, the percentiles, the casts and the
column names the SQL sink had to change because user, from and to are reserved words. Some
panels are not rewritten and are dropped silently rather than moved into the “not available”
row: the kernel-log panels, because the SQL schema has no count table for it, and the panels over
mikroscope_buddy and the RouterOS interface counters, which are wide in InfluxDB and long in SQL,
where a pivot is a different question, and “Headroom under the container memory cap”, because the
cap rides on InfluxDB’s mikroscope_self row and is a device fact in the SQL schema. The committed
PostgreSQL dashboard therefore
carries 161 panels against InfluxDB’s 177.
Graphite
Section titled “Graphite”Graphite has no labels: every dimension is a path node, so a query IS a path — and the first two
nodes are yours. --graphite-prefix (default mikroscope) and --host-tag are therefore
dashboard variables, read from Graphite’s own metric tree, and the dashboard asks for them at
the top rather than being hard-coded to whoever generated it.
The datasource is type graphite; nothing else is needed. From the CLI, check cannot read a
browser’s variable picker, so it takes them:
mikroscope dashboards check --store graphite --datasource-uid <uid> \ --var prefix=mikroscope --var host=routerElasticsearch
Section titled “Elasticsearch”Type elasticsearch, with the index the --elastic-index you forwarded with produces and
@timestamp as the time field.
This dashboard is the smallest of the five, and the reason is in the documents rather than in the
queries: the sink writes a sample’s per-core and per-device readings as arrays — cpu is an
array of one object per core — and a dynamically mapped array is a multi-valued field with no
correspondence between its members. avg(cpu.busy_ratio) is the mean over the cores, which is a
real number; “core 2’s busy ratio” is not expressible at all without a nested mapping the sink does
not declare. So the Elasticsearch panels are the scalar aggregates, and the per-core ones are
absent rather than wrong.
The dashboard has one variable, Host, filled by a terms lookup on host.keyword, because one
index can hold several routers. The lookup is the JSON string the datasource parses, so Host fills
from the index (rendered). check never runs that lookup:
it takes the value from --var host=.
Store probe
Section titled “Store probe”Before generating, dashboards import, dashboards publish and forward --grafana ask the
datasource which of mikroscope’s measurements it holds, through Grafana’s /api/ds/query.
dashboards check asks the same question, so it checks the dashboard import would store:
SELECT table_name, column_name FROM information_schema.columns WHERE table_schema = 'iox'Columns and not only tables: InfluxDB 3 refuses a query naming a missing column at planning time exactly as it refuses a missing table, and a store written before a field existed has the table and not the field. A panel that reads a field added later declares it, and the probe checks it.
group by(__name__) ({__name__=~"mikroscope_.+"})An instant query that returns one series per metric name that exists, with no samples to transfer.
A histogram counts as present when its _bucket, _count or _sum series is.
import and check print datasource holds N measurements; on InfluxDB, N counts tables plus
table.column pairs, so it is larger than the number of measurements. The dashboard is then
generated against the answer:
- A panel whose measurements and required fields are all present ships in its own section with its query — including a panel the compiled defaults put in the not-available row.
- A panel with anything missing moves into the collapsed “Not available on this device” row with
its queries hidden. It runs nothing, so it cannot paint a red
table … not foundbadge; its no-value text names what this store does not hold. The queries stay in the panel, so they can be switched back on in its editor. - On InfluxDB, a store with no
mikroscope_detection(or nomikroscope_trigger) gets that annotation layer switched off, its query kept. Prometheus answers an absent counter with an empty result, so the probe leaves its layers as they are. - A probe that fails — an error from Grafana, or an answer with no
mikroscope_names in it — is a warning, not an error.importandcheckprintwarning: could not ask <store> which measurements it holds, and the collector printsgrafana: <store>: could not ask the datasource which measurements it holds(dashboards publishthe same line withoutgrafana:), each with the reason. All of them carry on with the compiled defaults, so you are told which dashboard you got.
--no-probe skips the question and uses the compiled defaults, which is also what plain gen does,
since it has no datasource to ask.
Only InfluxDB and Prometheus can be probed. The query above is written per datasource type,
and the file internal/ has one for those two and
cannot probe a "…" datasource for the rest. With --store postgres, --store graphite or
--store elasticsearch, and on a collector publishing those stores, the probe therefore always
fails, always warns, and always ships the compiled defaults — the same result as passing
--no-probe. That is a gap, not a design: those three stores get the dashboard the panel list
declares rather than the one their data supports.
Check every panel
Section titled “Check every panel”export GRAFANA_TOKEN=…mikroscope dashboards check --grafana http://grafana:3000 --store influxdb --datasource-uid mikroscope-influxdb --window 15mmikroscope-<store> is the datasource the collector or dashboards publish created; after an import
of your own, pass your datasource’s UID. check writes nothing to Grafana: it regenerates the dashboard, probe
included, and runs its queries.
What check verifies
Section titled “What check verifies”check generates the dashboard exactly as import would — probe included — and then, for every
panel, including every panel nested inside a collapsed row, sends each of its queries that is not
hidden through Grafana’s /api/ds/query against your datasource over the window, and counts the
rows that come back. A hidden query is skipped because Grafana does not run it either, so a panel
in the not-available row reads none … rows=0 frames=0 with no error. The request carries the step Grafana would compute for that panel: the window divided by 900
data points, raised to the panel’s own minimum interval where it has one. Without that step, an
increase(x[$__interval]) target returns an empty frame, because a step below the scrape interval
leaves fewer than two points in the range. Every InfluxDB panel that bins with $__dateBin has a
minimum interval of at least 1 s: Grafana writes that macro’s bin in whole seconds, so a step under
1 s draws a bin 0 seconds wide and the panel comes back empty, in check and in a browser zoomed in
to a few minutes (measured).
It prints one line per panel and a verdict:
ok <panel title> rows=<n> frames=<n> none <panel title> rows=<n> frames=<n> <error, if any> FAIL <panel title> rows=<n> frames=<n> <error, if any>every panel returns data (<k> known-empty tolerated)- ok: the panel returned at least one row and no query reported an error.
- none: the panel is marked known-empty and did not qualify as ok. Two kinds of panel carry the mark: those whose emptiness is the healthy state (for example the detection and trigger panels, the two port-event panels, the gaps table, the opt-in conntrack poll, the sub-sample burst panel and the worst kernel-log severity timeline), and those in the not-available row. A known-empty panel is tolerated whether it returned no rows or an error.
- FAIL: anything else — no rows, or an error on any of the panel’s queries even if another returned rows.
With one or more failures check exits non-zero with N panel(s) return no data (K known-empty tolerated). A dashboard is not done until every panel that is not known-empty returns rows.
Check limitations
Section titled “Check limitations”check counts rows. It does not see a legend, an axis, a unit, a colour, a threshold or the grid,
and a passing check says nothing about whether the dashboard is readable
(an example). Readability is established by rendering the
dashboard in a browser.
What else is outside its reach, from the code:
- The dashboard stored in Grafana.
checkregenerates the dashboard locally and runs those queries. It does not read back whatimportor the collector stored, so a dashboard edited in Grafana’s UI is not what it checks. - Whether a number is right. One row is a pass. Wrong arithmetic that returns rows passes.
- Annotations. Only panels are walked; the detections and triggers annotation queries are not run.
- Alert rules. The provisioning files
genwrites are not loaded or evaluated. - A variable’s own query.
checktakes each variable’s value from--varand never runs the lookup that fills the picker, so a lookup that fails in the browser still passes. - What the browser does to a query. Some variables are substituted in the browser, not by the
server
checktalks to. The InfluxDB datasource can escape a macro such as$__interval_msin the browser into SQL InfluxDB 3 cannot parse (seen); only the rendered dashboard shows it. - The real panel width.
checkpretends every panel is 900 data points wide, the value a browser sent for this dashboard’s graphs. A narrower panel gets a wider bin.
Check a past window
Section titled “Check a past window”--window is the length of the query window and --end moves its right edge, so the panels can be
checked against a capture that has already finished rather than against an idle now:
mikroscope dashboards check --store influxdb --datasource-uid <uid> \ --window 1h --end <RFC 3339 time>A panel answers differently over a window with data than over one without, and a check is only as good as the window it is pointed at.
The Grafana versions, stores and windows that check and the browser renders have run against
are on Tested on.
Update and remove
Section titled “Update and remove”Update. The dashboards come from the CLI, not from the agent, so upgrade leaves them as they
are. After a new CLI or collector image, publish them again the way you first did: restart the
collector run with --grafana, run dashboards publish or dashboards import again, or upload
the new gen output. Each keeps the uid, so the dashboard is replaced in place and an edit made in
Grafana is lost. Regenerate and provision the alert rules again too
(Alert rules).
Remove. uninstall --targets dashboard, given the collector’s sink flags and --grafana, lists
per store the dashboard mikroscope-<store>, however it was imported, and the datasource
mikroscope-<store>, each only if it is there; --yes removes them:
export GRAFANA_TOKEN=…mikroscope uninstall --targets dashboard \ --influx http://influx:8181 --influx-db mikroscope --grafana http://grafana:3000 # lists what would gomikroscope uninstall --targets dashboard \ --influx http://influx:8181 --influx-db mikroscope --grafana http://grafana:3000 --yes # and removes itLeaving out --yes is its dry run; it refuses --grafana-dry-run. With --grafana-datasource-uid
it lists no datasource: an adopted one stays. The folder and the
provisioned alert rules stay too, and so does a datasource under another uid
(Remove dashboards and data). A Grafana
it cannot read, because it is unreachable, refuses the token or fails, stops it before anything is
removed, and the error names what it was asking about:
asking Grafana whether dashboard mikroscope-influxdb is there: ….