Skip to content

InfluxDB

The InfluxDB sink writes every ghchronicle point, dated when it happened, to InfluxDB 2 or 3, and it is the store the dashboards are built against.

sinks:
influxdb:
url: http://localhost:8181
token: ${INFLUX_TOKEN}
org: default
bucket: github
batch: 5000
# exclude: [gh_job_log]

Line protocol posted to the v2 write endpoint, which InfluxDB 2 and 3 both serve, so one sink covers both. The token goes in an Authorization header, org and bucket in the query string.

Precision is nanoseconds, because the traffic points are days and the workflow ones are seconds and one precision has to cover both.

Batches default to 5000 lines per request. batch lowers that for a server with a smaller body limit.

Everything dated, which is most of the project:

  • The traffic of a particular Tuesday, months later.
  • The star curve since the first star, drawn from one row per repository and day.
  • The merge time of a pull request closed in July.
  • Lines added and removed per commit, per author, over years.

InfluxDB keys a point by measurement, tag set and timestamp, so rewriting a point that already exists is not a duplicate. Replaying the same fourteen-day traffic window every six hours converges instead of accumulating, which is what makes the whole backfill design work.

The dashboards are generated against this store first; the other four query sets are translations of it.

Rewriting is free in rows and not in files

Section titled “Rewriting is free in rows and not in files”

InfluxDB 3 Core writes one Parquet file per partition per write request and never compacts them, and it refuses any query that would open more than its file limit. So the row that overwrites harmlessly still costs a file, and a sweep that offers the same history every six hours buys “Query would scan 10000 Parquet files, exceeding the file limit” a few weeks later.

That is why the write ledger exists, why dedupe is on by default here, and why turning it off is a decision rather than a tidy-up: only what changed is written.

Names measurements this sink should not receive. It defaults to gh_job_log, which is text meant for a log store: writing thousands of lines of build output into a metrics database is a lot of storage for something nobody will query as a number.

An excluded measurement is still collected and still offered to this sink, so the sweep log says what the database took rather than what it was handed:

level=INFO msg=written sink=influxdb family=joblogs points=0 unchanged=0 filtered=440

filtered here is a measurement named above, or a point carrying no field the line protocol can render. It is the same key every sink uses. Until this was counted the line read points=440 for a measurement the database has never held a row of.

A write that comes back 400 is bisected: the sink halves the batch, retries, and narrows down to the individual lines the server refuses to parse. Those are logged one by one with the server’s own reason, everything else is written, and the sweep reports a warning rather than a failure.

level=WARN msg="sink rejected some lines" sink=influxdb family=actions rejected=2

That behaviour matters because a batch is five thousand lines. Failing the whole batch on one malformed value would lose four thousand nine hundred and ninety-nine good points.

Use the InfluxDB 3 datasource in SQL mode for the shipped dashboard, and point it at the database the sink writes to. The dashboard file declares DS_INFLUXDB as an input, so importing asks you to choose your own datasource rather than carrying somebody else’s uid.

SELECT time, "count" FROM gh_traffic WHERE kind = 'views' AND repo = 'ghchronicle'

When a release changes what a measurement’s rows are keyed by, -migrate asks InfluxDB 3’s catalog whether the old tag is a tag column of the live table, and applying the change deletes that one table:

DELETE /api/v3/configure/table?db=<bucket>&table=<measurement>

That delete is InfluxDB’s own way of setting a table aside. Measured against InfluxDB 3 Core 3.0.0 to 3.11.5: the table is renamed <measurement>-<instant>, the instant in UTC, for example gh_discussion_comment-20261001T091004, stays listed and answers queries under that name, and the next write under the old name creates the table afresh, in the new shape, even where a column changes from a tag to a field. The binary reads the name back from the catalog, prints it, and keeps it in the state file, with when the server has scheduled its hard deletion, read from the system table of its _internal database. That is 72 hours after the delete (measured on 3.2.1 to 3.11.5), and the server keeps the name in its catalog for its --delete-grace-period after that, 24 hours by default; where the time cannot be read, 3.2.0 among them, 72 hours is assumed, which is what 3.2.0 schedules too. A server before 3.2 never purges it: see below. ghchronicle leaves the copy to the server, and forgets it once the server’s time for it has come. It never asks for it to go sooner: hard_delete_at is never sent, since measured on 3.11.5 now had not removed the rows eleven minutes later, and it takes the days to undo the change away; and from 3.10.0 on, a delete of the renamed table is answered with a 409, with or without it (measured on 3.10.0 to 3.11.5). Until then the copy’s rows can be read with SQL and written back through the sink.

InfluxDB 3 Core refuses a query that would open more Parquet files than its --query-file-limit, 432 by default, which a table written every ten minutes passes in days. The check reads the catalog, which opens none, and counts the rows only of a table that holds the old tag; a count Core refuses leaves the refill with no bound, and a refused read of whose rows the table holds leaves the change needing your word, the refusal quoted in the plan. See the query file limit.

The token has to be allowed to delete a table. A refused delete leaves the change pending, and the error names the same request to send by hand.

InfluxDB 2 has no rename. Applying the change there deletes every row of the measurement in the bucket, over every instant it can hold, through POST /api/v2/delete with the predicate _measurement="<measurement>", and that is final. Measured against InfluxDB 2.7.12: it took that measurement and left gh_discussion_comment_x, the other measurements of the bucket and the same measurement in another bucket as they were. So a start never applies it on its own; -migrate -yes does.

-uninstall data leaves out the tables InfluxDB 3 has already deleted, and says of each whether the server purges it or it stays: on a server before 3.2 every one stays, and on a later one each that a release before 3.2 deleted, asked of the server’s system table, with the request that removes it where the release takes one (see below). The rest are the server’s to purge, and only a release from 3.10 on is said to refuse to be asked for that sooner. 3.2.0’s system table cannot say which, and there the note says so.

Measured against InfluxDB 3 Core 3.0.0, 3.0.3 and 3.1.0, none of which has a hard deletion: no deleter runs, hard_delete_at is taken and ignored, _internal has no system.tables, and a delete of the renamed table answers 200 and renames it once more, <measurement>-<instant>-<instant>. The copy stays, and answers queries, for as long as the server runs one of those releases. 3.2.0 is the first release with a deleter: 3.2.0 and 3.3.0 were measured scheduling a hard deletion and carrying it out as 3.4.0 and later do.

The binary reads the release from /ping, and on a server before 3.2 says so where it matters: the plan says set aside for good before anything is applied, applying says it with the copy’s name, and -migrate lists each copy such a server keeps as kept aside ... for good, asked of the server itself, so the list holds a copy the state file has forgotten, or a new state file never knew. -uninstall data says the same of every table it deletes there. Each ends in the request that removes the copy once the server runs a release that takes it:

Terminal window
curl -X DELETE '<url>/api/v3/configure/table?db=<bucket>&table=<copy>&hard_delete_at=now' \
-H 'Authorization: Bearer <token>'

A copy made before 3.2 keeps no hard deletion time through an upgrade: measured with copies 3.0.3 and 3.1.0 made, opened by 3.2.0, 3.4.0, 3.9.13 and 3.11.5. So it stays after the upgrade too, until it is told to go. 3.2.0, 3.4.0 and 3.9.13 took the request for such a copy and dropped it from their catalog once their delete grace period had passed, 3.2.0 after renaming it once more; 3.10.0 to 3.11.5 answer it with a 409. So the request is sent while the server runs a release from 3.2 to 3.9, on the way up. From 3.2.1 on, -migrate lists the copies the system table shows with no hard deletion time, which is how such a copy looks there, and -uninstall data says each stays, with that request where the release takes it (measured on 3.4.0 with a copy 3.1.0 made). 3.2.0’s system table has no such column, and answers the question with a 500, so on 3.2.0 -migrate lists none and -uninstall data says it cannot tell. Measured as well: 3.11.5, started on the data directory of a 3.0.3, began with an empty catalog, while 3.4.0 and 3.9.13 read the catalogs 3.0.3 and 3.1.0 had left.

  • Choosing a store compares InfluxDB with the others, and holds the write ledger every one of them shares.
  • The dashboards says which of the five is drawn against which store, and what a panel a store cannot answer becomes.
Written and maintained by
MIT licenceRelease history