Prometheus
The Prometheus sink is an exporter that serves the current value of each GitHub metric for Prometheus to scrape.
sinks: prometheus: listen: 127.0.0.1:9605 path: /metricsAn exporter, not a pusher: point a scrape at it. It is the one sink in the project that is not outbound, and it exists because Prometheus insists on pulling.
What it serves
Section titled “What it serves”Metric names are github_<measurement>_<field>, with the tags as labels.
github_repo_stars{repo="ghchronicle",language="Go",visibility="public"} 283github_workflow_runs_count{repo="ghchronicle",conclusion="success"} 412Can Prometheus keep the history?
Section titled “Can Prometheus keep the history?”No, and not by choice: it cannot serve the dated history. Prometheus stamps a sample at scrape time
and rejects anything meaningfully older: measured against Prometheus 3.14 with
--web.enable-otlp-receiver and a thirty-minute out-of-order window, a sample
dated two days back comes back as HTTP 400.
Through this sink, GitHub’s fourteen-day traffic window collapses to its most recent day, and the star history to the current total. That is worth stating plainly rather than hiding. Run it alongside a history store rather than instead of one; both can run at once and the collection happens only once.
The reduction
Section titled “The reduction”Before serving, Summarize reduces each measurement according to a rule.
| Rule | What survives |
|---|---|
keep | The most recent value per label set. Snapshots |
sum | The rows added up. Windows, such as views over the fourteen days |
count | A count plus the mean of each numeric field. Dated items |
skip | Nothing |
A measurement with no rule is skipped, so a new collector cannot quietly flood the exporter with one series per star.
The batch is read the way the stores hold it. A point with the same
measurement, tags and time as an earlier one is the same row, so it is added,
counted and averaged once, as the history stores keep it once;
dating a point
says which of its fields each store keeps. A price is the one number a sum
does not add: the bill repeats a SKU’s price on each row, one per repository
and day, so github_billing_usage_price_per_unit is the highest of them, the
MAX(price_per_unit) of the SQL dashboards, and not the price times the days
billed.
Each mean is over the items that carried the field, not over the count. A
collector leaves a field out when it has no honest value for it: a job with no
start time has no queued_seconds, and a job GitHub no longer lists steps for
has no steps. Counted as zeros, a sweep in which half the jobs had lost their
steps would halve github_workflow_jobs_steps_mean. Three fields are the
exception, because leaving them out is itself the answer: merged on an
outside contribution is written only on a merge, advanced on a fork only when
it has a push date, and pull_requests on a run only when it ran for any. Each
of those is averaged over every item, so the mean of merged is the share of
contributions merged rather than 1.
The reducer also publishes total, a running distinct-item count per series.
That is what lets a Prometheus dashboard say “per day” through increase(),
since it has no rows to count.
Two measurements are skipped for size
Section titled “Two measurements are skipped for size”The commit punch card is one series per repository, weekday and hour, and the release assets one per file ever published. Measured, together they were four fifths of the exporter’s entire output. Both are drawn properly by the InfluxDB dashboard.
Ten more are skipped for a reason other than size, nine of them as history and one as text, so twelve measurements in all never reach the exporter. The list, with the reason for each, is on dating a point.
The exporter holds only what the last sweep collected, and for workflow runs that is the newest thirty per repository between builds: an ordinary sweep reads the run list in pages of thirty, and pages on only while a page is full of runs newer than its two hour window. The stores keep every run the walk ever saw; this is the one place the smaller page is visible.
Workflow jobs are the other: a sweep carries the jobs of the runs it listed
for the first time, so gh_workflow_jobs counts the jobs new to this process
rather than those of the newest twenty runs, and a sweep in which no run
finished carries none, leaving the last count standing until it goes stale a
day later. The total beside it is unaffected, since it counts every distinct
job the process has ever seen; the stores are unaffected too, since a job is
written once and dated when it finished.
Scraping it
Section titled “Scraping it”scrape_configs: - job_name: ghchronicle static_configs: - targets: ["127.0.0.1:9605"]The exporter holds its samples in memory, so a restart empties it. That is why the first sweep after start-up runs every enabled family whatever the state file says: without it, a twelve-hour family would leave its panels reading zero for half a day.
sinks: prometheus: listen: 0.0.0.0:9605 path: /metrics no_prime: true # leave the first sweep on its ordinary scheduleno_prime: true switches that priming sweep off, for an account whose quota is
tight enough that a full sweep on every restart is not affordable. The price is
exactly the behaviour the priming exists to avoid: until each family’s cadence
comes round, the panels that read it have nothing, and for a twelve-hour family
that is half a day of zeros. The stores are unaffected either way, since they
keep what was collected before the restart.
The priming sweep records as run only the families that were due, and the others run again when they would have without the restart. Recorded all at once, the families of six hours or more would be due together a day later and take turns, and the last of them would pass the day after which the exporter drops a series.
A series that has not been rewritten for 24 hours is dropped, so a repository that leaves the sweep stops being reported as if it were still there. The horizon is not configurable.
Errors at start-up, not in a goroutine
Section titled “Errors at start-up, not in a goroutine”prometheus exporter: listen tcp :9605: bind: address already in useReported when the process starts rather than swallowed in a background goroutine, so a port clash cannot leave you with a running collector and a silently missing exporter.
In a container
Section titled “In a container”The example listens on 127.0.0.1, which inside a container is the container’s
own loopback and unreachable from the host. Use 0.0.0.0:9605 there and let
the port publication decide who can reach it.
Nothing to migrate
Section titled “Nothing to migrate”The exporter holds today’s values in memory, and a restart of the new release
serves them in its shape, so -migrate
has nothing to do here. The Prometheus server keeps the old series until its
retention, and they end where the new ones begin, so no query counts an item
twice.
Where to go next
Section titled “Where to go next”- Choosing a store compares Prometheus with the others, and holds the write ledger every one of them shares.
- The dashboards says which of the five is drawn against which store, and what a panel a store cannot answer becomes.