Backfill
A backfill walks every GitHub surface back as far as the API answers, once, so the history starts before the first sweep.
ghchronicle -config config.yaml -backfillA sweep is an increment. A backfill is a walk. They are opposite intentions and the tool treats them as such.
The difference in one table
Section titled “The difference in one table”| Sweep | Backfill | |
|---|---|---|
| Pages per collector | the collector’s own small default | until the API runs out, or the bound is reached |
| When the reserve is reached | skip the family and warn | wait for the window to reset, then carry on |
| Which families run | those whose interval has elapsed, the slow ones taking turns in the service | every enabled family, whatever the state file says |
| When it is stopped | nothing to keep; the next tick sweeps | a checkpoint keeps its place, and the same command carries on |
| How often | on a schedule, forever | deliberately, usually once |
A normal sweep must never block. It protects the reserve so whatever else uses the same token keeps working, and it skips a family rather than sleeping. A backfill is run on purpose and the only thing that matters is that it finishes, so it parks until the budget is whole again. And a backfill that is stopped half way keeps its place, so the same command started again carries on from there rather than from the first repository.
Stopping one, and picking it up again
Section titled “Stopping one, and picking it up again”A walk of a real account is hours long, and a reason to stop it turns up: the author’s own had been going six and a half hours when the binary underneath it had to be replaced. Stopping it used to cost the whole walk.
The place is kept in a file beside the state file, named after it with
-progress.json instead of .json. Nothing configures it. It is written after
every repository rather than after every family, and that is measured rather
than preferred: on an account of fifty nine repositories the account wide
families take four minutes between them, and then actions takes two hours and
three quarters and commits longer still, so a checkpoint at the end of each
family would keep the four minutes and throw away the rest.
What it records is what has been delivered. A repository goes into it only once its whole walk has reached every configured sink, which is what makes skipping it on the way back safe: there is no tail of its own pagination left anywhere. A repository whose collector failed, or whose rows a store refused, is not recorded, and the walk comes back to it. Coming back to it costs requests and nothing else, because a point carries the date the thing happened and re-collection rewrites the same rows.
level=INFO msg="backfill stopped, and what it had written is kept" file=/var/lib/ghchronicle/backfill-progress.json running_for=6h40m7s families_complete=9 family=actions repositories_written=23 resume="run the same command again"The file itself is meant to be read, and says the same thing at more length: the families that finished and when, the family that was in flight, and which of its repositories are done. It is removed when the walk has covered every family it was asked for, so a checkpoint that exists is a walk with work left in it.
Reaching the end without an error is not the same thing. A family truncated by a secondary rate limit is handed back as a pass rather than as a failure, and a family that failed on every repository is deliberately left unmarked; a walk can finish tidily with either of those behind it. The checkpoint stays, and the line at the end names the families still to do.
An ordinary sweep keeps none of this. It writes no checkpoint and reads none: the state file it keeps is about cadences, and a half walked backfill has no business in it.
The two things it refuses
Section titled “The two things it refuses”A checkpoint it cannot read. A machine that loses power can leave an empty or half written file behind. The run stops with a sentence naming it rather than starting the walk over in silence: what the file lists may be hours of somebody’s history, and starting over is the answer that looks like success.
A checkpoint from a different walk. The file records what the walk was asked for: the API it reads, the targets, the families that were enabled, the date bound exactly as it was spelled, and the stores it writes to. If one of those changed between the stop and the resume, the list of repositories not to walk again means something else, and the run stops and names the setting that changed.
The stores are there because a record in this file means the rows reached every sink, and every sink means the ones that process had. Add a store while the walk is stopped and the resume skips everything the first half recorded, so the new store ends up holding the tail of the walk and nothing before it. Point an existing one somewhere else and the same hole opens in a different place. Only the settings you can read are compared: a store’s password is never written into this file.
The bound is the one that would go wrong quietly. A walk bounded at a year records its repositories as written; an unbounded resume would skip every one of them and leave a store that looks complete and stops a year back, with nothing in the rows to say so.
Both refusals leave the file where it is. Put the setting back to carry on, or delete the file to walk again from the first repository.
Replacing the binary is not one of them. That is the thing that stopped the walk this was written for, so a checkpoint written by another build is resumed, with a line saying which build wrote what it names.
Seeing how far it has got
Section titled “Seeing how far it has got”ghchronicle -config config.yaml -backfill-statusIt reads the checkpoint and prints it: when the walk began, how long ago it
last recorded anything, how many families are complete out of how many, the
family it was inside and how far into it, and the command that carries it on,
with the -backfill-since and the -families it was started with. A
migration’s refill in progress is printed after it, with what it writes where.
It asks GitHub nothing and writes nothing, so it is safe to run while a walk is
going, and it needs no token: the credential is there because a sweep asks
GitHub, and this asks nobody. When there is no walk in progress it says so and
names the file it looked for.
Going back on its own
Section titled “Going back on its own”A pass can end with families left and no error at all. The usual reason is a
few minutes of bad weather at the other end: a 502 on two repositories out of
sixty leaves their family unmarked, and the walk ends tidily with one family
short.
ghchronicle -config config.yaml -backfill -backfill-retry 1hwaits an hour and goes back for whatever is left, and keeps doing that until there is nothing left. It costs almost nothing, because a resume walks only what the checkpoint does not already hold: two repositories out of sixty, not sixty.
It stops on its own when a pass records nothing new. That is the whole rule, and it is deliberately not a list of which errors are worth retrying: a repository that was deleted fails the same way every hour, and a list of retryable statuses is wrong the moment GitHub answers something it did not answer before. A pass that gains nothing is a pass whose obstacle waiting does not clear. There is also a cap of ten passes, as a backstop rather than a knob.
Without the flag nothing waits and nothing goes back, which is what it did before this existed.
The cooldown
Section titled “The cooldown”When a bucket is at or below its reserve, the backfill waits. The wait is not guessed: every GitHub response says exactly when its window resets, so the collector sleeps until that instant plus a second of slack, and logs what it is doing.
level=INFO msg="rate limit reserve reached, waiting for the window to reset" family=commits wait=23m11sTwo guards on that wait. It is capped at one hour, so a clock skew or a stale
header cannot turn into an unbounded sleep, and it has a floor of one second,
so it cannot become a busy loop. A cancelled context ends the wait immediately,
which is what makes Ctrl+C work during an overnight backfill.
How far back
Section titled “How far back”-backfill-since on the command line, or backfill.since in the configuration
file. Four spellings, because people reach for different ones:
| Value | Means |
|---|---|
2024- | that date |
90d | ninety days ago |
2y | two years ago |
720h | a Go duration before now |
empty, all, unlimited | no bound at all |
No bound means the walk stops only where the API does, however many hours or days that takes, pausing at every rate limit reset along the way.
ghchronicle -config config.yaml -backfill -backfill-since 2ySome families only
Section titled “Some families only”ghchronicle -config config.yaml -backfill -families discussions,outboundwalks the families named and no other, into every configured store. It is what
a store that lost one family’s history wants: dropped by hand, or sent through
a Telegraf whose store was cleared, without walking every other family it
still holds. The names are the families -groups prints. One that is not a
family, and the flag without -backfill, are refused with 2, as a command line
that does not parse; a family the configuration switches off is refused with 1
before anything is asked, of GitHub or of the stores, the start’s migration
check included, since the walk would read none of it and end complete.
Its checkpoint is the backfill’s, and records the families it was asked for, so
a backfill of other families, or of every family, refuses it and names the
family that differs, and -backfill-status prints the resume line with the
same -families in it.
The refill of a migration
Section titled “The refill of a migration”A migration that cleared a store reads its history back with a backfill of its
own, the refill: the families that write what was cleared, writing that alone
into the stores it was cleared from; see reading the history
back. It keeps its
checkpoint beside the backfill’s, named with -refill.json, so a backfill in
progress and a refill neither refuse nor overwrite each other’s, and
-backfill-status prints both.
Run it once, first
Section titled “Run it once, first”A new install should run a backfill before, or right after, starting the service. A sweep’s first pass is a wider increment, not a history: a month of workflow runs, the star history in full, and the newest page of everything else. The points are dated, so every store is fine with that; a dashboard at ninety days or two years is not, because the history it draws begins on the day the collector was installed.
Measured after a day of sweeps and no backfill, over repositories holding hundreds of pull requests each: pull requests, about a fifth of what the repositories report, because a sweep reads one page of fifty per repository however many it holds; issues, about half; commits, well under a tenth and none older than thirty days; jobs and steps for a tenth of the workflow runs, so the queue wait, the slowest jobs and the steps that fail were computed over that tenth. At two years, Pull requests merged read a fifth of what Pull requests merged, ever said. Stars and forks were complete, because the first sweep walks those to the end anyway, and so were releases, deployments and the alerts, because none of those repositories had more than the hundred of each that a sweep reads.
What it reaches that a sweep does not
Section titled “What it reaches that a sweep does not”-
The whole commit history, rather than the last page. This is the one family where a backfill is qualitatively different rather than merely wider: without it the lines-changed series begins on the day you installed the collector.
-
Every workflow run GitHub still holds, each expanded into its jobs, and into its steps while GitHub still serves them. Measured on 24 September 2026, GitHub listed every job of a run 278 days old, but no steps for any run created before 12 April, about five and a half months back. Such a job, if it finished as success, failure or timed out, is written with no
stepsfield rather than a 0, since it ran at least one. -
The archived repositories, in full, whatever
include_archivedsays. Their history is the account’s history and it never moves again, which is exactly why a sweep leaves them out and why one walk of them is enough. Forks stay as configured. What does still move on an archived repository, its stars, forks and watchers, a sweep reads without a backfill: the listing it already pays for says which repositories are archived, and everytotalssweep asks about all of them, one query per twenty five, for two rows each. One isgh_repo_archived, dated when the repository was archived. The other is itsgh_repo_total, stamped at the sweep like a collected repository’s, and that is the row the account’s star and fork totals are read from. -
Every pull request and issue, in pages of fifty, where a sweep reads what changed since the sweep before, and once a day every open item and what moved since the day before.
-
Every pull request and issue the account opened in other people’s repositories that has since been merged or closed, in pages of a hundred ordered by when each one last moved, where a sweep reads each of the three closed states back to a cadence before the sweep before. That is enough for a sweep because a merge or a close moves the item to the top, however long ago it was opened, and it is one page unless more than a hundred items moved in that time; a sweep that read one page and no more lost the item that a hundred later updates, a bot locking old threads or a relabel, had pushed onto the second. The ones still open are bounded by neither a page nor a date: every sweep reads them all, since each one gets a row for every day it stays open. In 2.5.1 and earlier each of the five searches read its newest hundred and stopped, on a sweep and on a backfill alike.
GitHub serves a thousand results of any search and no more, so an account past a thousand in one of those states keeps the thousand that moved most recently. That is said in the log, once per count, rather than left to show as two panels that disagree:
level=WARN msg="outbound search read fewer items than it counts, GitHub serves a thousand at most" kind=pull_request state=merged count=2860 read=1000 -
Every page of artifacts, of repository activity and of code scanning analyses, where a sweep reads five, two and one.
-
Every release, deployment and discussion, and every Dependabot and code scanning alert, where a sweep reads the newest page of each list.
-
Every issue and discussion comment the account left anywhere, and every answer of its own that was accepted, where a sweep reads the newest hundred comments of each kind and the newest five hundred accepted answers.
-
Every star the account gave, where a sweep reads the newest five hundred.
-
The co-authored pull requests behind Pair Extraordinaire’s progress, walked over the account’s whole life, where a pass adds the days since the last one and walks the whole history once a week.
-
Every page of every stargazer list, where a sweep reads the newest hundred stars of a repository whose list it has walked whole once. Up to 2.5.0 a first walk that failed was recorded as done, and a backfill read a recorded list by its first and last pages, so the stars in between were never read; a backfill now reaches them.
-
The whole inbox, read threads included, where a sweep reads the unread threads that moved and, once a day, twenty pages of fifty with the read ones.
-
Every month of billing GitHub still answers, until three empty months in a row, where a sweep reads two.
-
The newest hundred webhook deliveries per hook, where a sweep reads thirty.
-
At most 500 failed job logs per repository, none older than ninety days, once
joblogshas a cadence; a sweep reads at most ten.
Two things a sweep remembers, a backfill leaves alone. A 403 or 404 that a sweep remembers for a day is asked again, since a backfill consults no such memory. And it reads the cache file beside the state file but writes nothing to it: the pages it walks are ones no sweep asks for, and kept there they would push out the sweeps’ own, so a service started after it would start colder than it stopped.
What it does not switch on
Section titled “What it does not switch on”A backfill runs every enabled family whatever the state file says. It enables
none of them. The three families that ship with a cadence of 0, deps,
history and joblogs, stay off unless they have been given one by name under
every.families, and -backfill does not change that.
history is the one that surprises people, because walking every past year of
the contribution calendar is exactly what somebody asking for “the whole
history” has in mind. Give it a cadence first:
every: families: history: 24hSee cadences for why default and
groups cannot switch these three on either.
Two endpoints that needed their own handling
Section titled “Two endpoints that needed their own handling”Both were found by running it, not by reading documentation.
Dependabot refuses page numbers. It answers an error to page= outright
and pages by cursor instead, so the alert walk is written against cursors.
The GraphQL gateway gives up on a hundred pull requests, and on fifty
commits. Asking for a hundred pull requests with their reviews in one query
answers an HTML 502 after about ten seconds, and so does asking for fifty
commits of a busy repository with the checks each one carries: twelve pages of
the author’s own commit walks met it, in six repositories. Both walks halve
their page size and retry on the same cursor, silently: neither collector
carries a logger, so a backfill of a busy repository shows this only as a
slower family, never as a line. The page is halved while it is larger than
ten, so fifty is asked again at twenty-five, twelve and six. A page the gateway
still gives up on at that smallest size, and a timeout in any walk with no
smaller page to ask, is a failure: the rows read before it are written, the
collector failed line says the query is too large for one request, and the
repository is not recorded in the checkpoint, so a resume walks it again. Up
to 2.6.3 the commit walk took the timeout for the end of the history and
recorded the repository as walked. Only that answer is halved for: a 503, or
a 502 that comes back sooner than ten seconds, is not a query the gateway ran
out of time on, and the client asks it once
more
at the same size.
Running one safely
Section titled “Running one safely”It is idempotent. Points are keyed by measurement, tags and timestamp, so a backfill run twice rewrites the same rows rather than doubling them, in every store that keeps the history. What it costs is API quota and time.
Two things worth doing first: run -list to confirm the repository set, and
check that the store you are writing to is the one that keeps dates. Backfilling
into Prometheus collects a great deal of history and then reduces all
of it to a single current value.