Skip to content

Loki

The Loki sink sends the GitHub activity that is an event rather than a number (a star, a release, a failed run) to Grafana Loki as log lines.

sinks:
loki:
url: http://loki:3100/loki/api/v1/push
tenant_id: ""
labels:
job: ghchronicle
max_age: 1h
batch: 1000

Some of what GitHub reports is a measurement and some of it is an event. “The repository has 148 stars” is a measurement. “Someone starred it at 03:03, this release was published, that workflow failed on main, this alert was raised” are events: each happened once, at a known moment, and what you want later is to read them in order and search them, not to average them.

Twenty-two measurements have an event rendering: stars in both directions, forks, releases, published package versions, pull requests, reviews, review threads, issues, commits, workflow runs, job logs, repository activity, Dependabot alerts, code scanning analyses, the event feed, notifications, discussions, webhook deliveries, deployments, ruleset versions and external contributions. Everything else is a gauge in disguise and is not sent.

A measurement with no rendering is dropped in silence, which is right for a gauge and wrong for an event nobody has got round to: deployments and review threads were dated events with no log line for months, and nothing said so. So every dated measurement now has to appear in one of two tables in internal/sink/loki.go, the renderings or the refusals, and each refusal carries the reason it is not a log line. A test fails on a dated measurement that appears in neither.

It is not dropped in silence in the sweep log any more. This sink writes a fraction of what it is offered, so most of its lines carry filtered and points=0, which is what it stored rather than what it was given. An entry older than max_age, or too far behind the newest entry of its own stream for Loki to accept, is counted there too and reported as a dropped-entries line.

Each line reads as a sentence first and carries every tag and field after it in logfmt, so the same line is greppable in a terminal and queryable in Grafana without keeping two copies of the data.

someone starred acme/telemetry full_name="acme/telemetry" user="someone" starred=1

Two renderings changed in 2.6.0, and one of them again in 2.6.1. A line already in Loki keeps the text it was sent with:

  • An external contribution is a line when the item closes, at the moment it closed, and says what happened to the item and not who did it, since the row does not say who merged or closed it: USER's pull request OWNER/REPO#N was merged or USER's pull request OWNER/REPO#N was closed without merging, and USER's issue OWNER/REPO#N was closed. A closed row whose state and merged flag disagree reads USER's contribution OWNER/REPO#N. An item still open is no line at all from 2.6.1 on: its row is stamped at the start of every day it is seen open, which is a reading and not something that happened, and opening it is already a line of the event feed, with action="opened". 2.6.0 sent that row as ... is open, and only on a day an outbound pass wrote in the first hour of the UTC day. Every line used to say USER merged OWNER/REPO#N, open issues included.
  • A release reads published release TAG of OWNER/REPO, or published prerelease TAG of OWNER/REPO, at the moment it was published. It used to read release TAG of OWNER/REPO, N downloads, so a query filtering on downloads matches only the old lines.

batch is how many entries go in one push, 1000 unless it is lowered.

The stream label is kind, which is what you filter on first.

{job="ghchronicle", kind="workflow_run"} |= "failure"
{job="ghchronicle", kind="job_log"}

max_age, and why the reason is not the obvious one

Section titled “max_age, and why the reason is not the obvious one”

Loki refuses an entire push when one entry predates reject_old_samples_max_age, a week by default, and half of what this collector produces is older than that on purpose: a star from 2020, a pull request from 2024.

But the limit that actually bites is the other one. Loki also refuses an entry more than its out-of-order window behind the newest entry already in that stream, which is half of the ingester’s max_chunk_age, one hour by default.

Measured against a real Loki 3: once the stream held an entry from 19:14, one from 00:35 the same day came back as “entry too far behind”. Against Loki 3.7.7 with its default limits, a stream holding an entry five minutes old took one 55 minutes old and refused one 75 minutes old.

So the horizon is applied two ways:

  1. against the wall clock,
  2. against the newest entry that stream has been sent before.

Inside one push nothing is behind, because the sink sends each stream oldest first and Loki judges each entry against the newest one before it: one push of entries from 23 hours, 12 hours and a minute ago into an empty stream was taken whole.

What falls outside is left out and counted, at debug level, rather than costing the whole push.

max_age defaults to one hour, which is Loki’s default window. Raise it only if you have raised Loki’s max_chunk_age to match, since the window is half of it.

A release is the event that shows what the horizon costs, and one of the two streams that look further back. Its line is rendered from gh_release_published, at the moment the release was published. It used to come from gh_release, which is stamped at the sweep because its downloads move, so every repository pass pushed every release again. Measured on 2.5.1 in production, over 30.9 hours and 27 repo passes, that was 4,313 of the 10,467 lines the sink sent, 41 per cent, for the 2 releases published in those hours. Dated at the publication, a release is first seen by the repo pass after it, so one published just after a pass read its repository is a whole cadence old when the next pass writes, plus however late that pass runs: a tick it lost, the slower families that ran before it, a restart. With repo and max_age both at their default hour, max_age alone left such a release out, and every later pass only saw it older.

So the release stream looks back the repo cadence plus max_age: two hours at the defaults, seven with repo: 6h. An external contribution is the other stream that does, for the same reason: its line is dated when the item closed, and the first pass to see it is the outbound pass after that, so the stream looks back the outbound cadence plus max_age, two hours at the defaults. Before 2.6.1 it had max_age alone, and a closing was sent only when the next pass wrote within the hour of it, which an hourly pass does not do for an item closed in the seconds after the last one read the searches, nor, when it runs late, for one closed in the minutes it is late by.

Each lookback is capped at six days, a day short of the week reject_old_samples_max_age allows, and it never shortens max_age: one set past six days is these streams’ horizon as it is every other’s. Loki refuses an old line only for being behind a newer one in its stream, which the second check above still makes. A release published in the hour before a pass is sent by that pass and again by the next, the same line at the same instant, which Loki keeps once. A contribution is sent twice the same way, and its line carries the item’s comment count and title, which can move in between: measured against Loki 3.7.7, a line sent again with comments moved from 1 to 2 was kept as a second line at the same instant. Under -once the cadence that matters is the schedule that runs the binary, so set every.families.repo and every.families.outbound to it. Both are in the metrics store either way.

The dated history. That is what a metrics store is for, and it is why the two run together rather than one replacing the other. A log answers “what happened recently, in order”; a time series answers “how much, over which period”.

The joblogs family, off until every.families.joblogs gives it a cadence, collects the last forty lines of every failed GitHub Actions job. It is text rather than a measurement, so the InfluxDB sink excludes it by default and the Prometheus exporter skips it. Loki is where it belongs, and the query is:

{job="ghchronicle", kind="job_log"}

The exported dashboards do not show it, because a dashboard bound to one datasource cannot query two and an importer may have no Loki: they carry a text panel, “Where failure output went”, with that query. Publishing the dashboard from the collector swaps that panel for the lines, newest first, filtered by the dashboard’s repository variable where the store’s variable can be read as a regular expression. A Loki sink whose address ends in /loki/api/v1/push is enough: a Loki datasource Grafana already has at that address, with the path dropped, is adopted, and otherwise the datasource is made from it. One writing anywhere else says so and takes a grafana.datasource.loki_uid instead, and so does a Grafana that reaches Loki by another address. A token that may not make the datasource costs this panel and nothing else: the run warns once and publishes every dashboard with the note.

A stream is its job, its kind and the configured labels, and a point’s tags go into the line, so a release that moves a tag changes no stream, and -migrate has nothing to do here. A line is what was said when it was said, and Loki refuses entries older than its window anyway.

  • Choosing a store compares Loki with the others, and holds the write ledger every one of them shares.
  • The dashboards says which of the five is drawn against which store, and what a panel a store cannot answer becomes.
Written and maintained by
MIT licenceRelease history