Skip to content

Backfill

A backfill walks every GitHub surface back as far as the API answers, once, so the history starts before the first sweep.

Terminal window
ghchronicle -config config.yaml -backfill

A sweep is an increment. A backfill is a walk. They are opposite intentions and the tool treats them as such.

SweepBackfill
Pages per collectorthe collector’s own small defaultuntil the API runs out, or the bound is reached
When the reserve is reachedskip the family and warnwait for the window to reset, then carry on
Which families runthose whose interval has elapsed, the slow ones taking turns in the serviceevery enabled family, whatever the state file says
When it is stoppednothing to keep; the next tick sweepsa checkpoint keeps its place, and the same command carries on
How oftenon a schedule, foreverdeliberately, usually once

A normal sweep must never block. It protects the reserve so whatever else uses the same token keeps working, and it skips a family rather than sleeping. A backfill is run on purpose and the only thing that matters is that it finishes, so it parks until the budget is whole again. And a backfill that is stopped half way keeps its place, so the same command started again carries on from there rather than from the first repository.

A walk of a real account is hours long, and a reason to stop it turns up: the author’s own had been going six and a half hours when the binary underneath it had to be replaced. Stopping it used to cost the whole walk.

The place is kept in a file beside the state file, named after it with -progress.json instead of .json. Nothing configures it. It is written after every repository rather than after every family, and that is measured rather than preferred: on an account of fifty nine repositories the account wide families take four minutes between them, and then actions takes two hours and three quarters and commits longer still, so a checkpoint at the end of each family would keep the four minutes and throw away the rest.

What it records is what has been delivered. A repository goes into it only once its whole walk has reached every configured sink, which is what makes skipping it on the way back safe: there is no tail of its own pagination left anywhere. A repository whose collector failed, or whose rows a store refused, is not recorded, and the walk comes back to it. Coming back to it costs requests and nothing else, because a point carries the date the thing happened and re-collection rewrites the same rows.

level=INFO msg="backfill stopped, and what it had written is kept" file=/var/lib/ghchronicle/backfill-progress.json running_for=6h40m7s families_complete=9 family=actions repositories_written=23 resume="run the same command again"

The file itself is meant to be read, and says the same thing at more length: the families that finished and when, the family that was in flight, and which of its repositories are done. It is removed when the walk has covered every family it was asked for, so a checkpoint that exists is a walk with work left in it.

Reaching the end without an error is not the same thing. A family truncated by a secondary rate limit is handed back as a pass rather than as a failure, and a family that failed on every repository is deliberately left unmarked; a walk can finish tidily with either of those behind it. The checkpoint stays, and the line at the end names the families still to do.

An ordinary sweep keeps none of this. It writes no checkpoint and reads none: the state file it keeps is about cadences, and a half walked backfill has no business in it.

A checkpoint it cannot read. A machine that loses power can leave an empty or half written file behind. The run stops with a sentence naming it rather than starting the walk over in silence: what the file lists may be hours of somebody’s history, and starting over is the answer that looks like success.

A checkpoint from a different walk. The file records what the walk was asked for: the API it reads, the targets, the families that were enabled, the date bound exactly as it was spelled, and the stores it writes to. If one of those changed between the stop and the resume, the list of repositories not to walk again means something else, and the run stops and names the setting that changed.

The stores are there because a record in this file means the rows reached every sink, and every sink means the ones that process had. Add a store while the walk is stopped and the resume skips everything the first half recorded, so the new store ends up holding the tail of the walk and nothing before it. Point an existing one somewhere else and the same hole opens in a different place. Only the settings you can read are compared: a store’s password is never written into this file.

The bound is the one that would go wrong quietly. A walk bounded at a year records its repositories as written; an unbounded resume would skip every one of them and leave a store that looks complete and stops a year back, with nothing in the rows to say so.

Both refusals leave the file where it is. Put the setting back to carry on, or delete the file to walk again from the first repository.

Replacing the binary is not one of them. That is the thing that stopped the walk this was written for, so a checkpoint written by another build is resumed, with a line saying which build wrote what it names.

Terminal window
ghchronicle -config config.yaml -backfill-status

It reads the checkpoint and prints it: when the walk began, how long ago it last recorded anything, how many families are complete out of how many, the family it was inside and how far into it, and the command that carries it on, with the -backfill-since and the -families it was started with. A migration’s refill in progress is printed after it, with what it writes where. It asks GitHub nothing and writes nothing, so it is safe to run while a walk is going, and it needs no token: the credential is there because a sweep asks GitHub, and this asks nobody. When there is no walk in progress it says so and names the file it looked for.

A pass can end with families left and no error at all. The usual reason is a few minutes of bad weather at the other end: a 502 on two repositories out of sixty leaves their family unmarked, and the walk ends tidily with one family short.

Terminal window
ghchronicle -config config.yaml -backfill -backfill-retry 1h

waits an hour and goes back for whatever is left, and keeps doing that until there is nothing left. It costs almost nothing, because a resume walks only what the checkpoint does not already hold: two repositories out of sixty, not sixty.

It stops on its own when a pass records nothing new. That is the whole rule, and it is deliberately not a list of which errors are worth retrying: a repository that was deleted fails the same way every hour, and a list of retryable statuses is wrong the moment GitHub answers something it did not answer before. A pass that gains nothing is a pass whose obstacle waiting does not clear. There is also a cap of ten passes, as a backstop rather than a knob.

Without the flag nothing waits and nothing goes back, which is what it did before this existed.

When a bucket is at or below its reserve, the backfill waits. The wait is not guessed: every GitHub response says exactly when its window resets, so the collector sleeps until that instant plus a second of slack, and logs what it is doing.

level=INFO msg="rate limit reserve reached, waiting for the window to reset" family=commits wait=23m11s

Two guards on that wait. It is capped at one hour, so a clock skew or a stale header cannot turn into an unbounded sleep, and it has a floor of one second, so it cannot become a busy loop. A cancelled context ends the wait immediately, which is what makes Ctrl+C work during an overnight backfill.

-backfill-since on the command line, or backfill.since in the configuration file. Four spellings, because people reach for different ones:

ValueMeans
2024-01-01that date
90dninety days ago
2ytwo years ago
720ha Go duration before now
empty, all, unlimitedno bound at all

No bound means the walk stops only where the API does, however many hours or days that takes, pausing at every rate limit reset along the way.

Terminal window
ghchronicle -config config.yaml -backfill -backfill-since 2y
Terminal window
ghchronicle -config config.yaml -backfill -families discussions,outbound

walks the families named and no other, into every configured store. It is what a store that lost one family’s history wants: dropped by hand, or sent through a Telegraf whose store was cleared, without walking every other family it still holds. The names are the families -groups prints. One that is not a family, and the flag without -backfill, are refused with 2, as a command line that does not parse; a family the configuration switches off is refused with 1 before anything is asked, of GitHub or of the stores, the start’s migration check included, since the walk would read none of it and end complete.

Its checkpoint is the backfill’s, and records the families it was asked for, so a backfill of other families, or of every family, refuses it and names the family that differs, and -backfill-status prints the resume line with the same -families in it.

A migration that cleared a store reads its history back with a backfill of its own, the refill: the families that write what was cleared, writing that alone into the stores it was cleared from; see reading the history back. It keeps its checkpoint beside the backfill’s, named with -refill.json, so a backfill in progress and a refill neither refuse nor overwrite each other’s, and -backfill-status prints both.

A new install should run a backfill before, or right after, starting the service. A sweep’s first pass is a wider increment, not a history: a month of workflow runs, the star history in full, and the newest page of everything else. The points are dated, so every store is fine with that; a dashboard at ninety days or two years is not, because the history it draws begins on the day the collector was installed.

Measured after a day of sweeps and no backfill, over repositories holding hundreds of pull requests each: pull requests, about a fifth of what the repositories report, because a sweep reads one page of fifty per repository however many it holds; issues, about half; commits, well under a tenth and none older than thirty days; jobs and steps for a tenth of the workflow runs, so the queue wait, the slowest jobs and the steps that fail were computed over that tenth. At two years, Pull requests merged read a fifth of what Pull requests merged, ever said. Stars and forks were complete, because the first sweep walks those to the end anyway, and so were releases, deployments and the alerts, because none of those repositories had more than the hundred of each that a sweep reads.

  • The whole commit history, rather than the last page. This is the one family where a backfill is qualitatively different rather than merely wider: without it the lines-changed series begins on the day you installed the collector.

  • Every workflow run GitHub still holds, each expanded into its jobs, and into its steps while GitHub still serves them. Measured on 24 September 2026, GitHub listed every job of a run 278 days old, but no steps for any run created before 12 April, about five and a half months back. Such a job, if it finished as success, failure or timed out, is written with no steps field rather than a 0, since it ran at least one.

  • The archived repositories, in full, whatever include_archived says. Their history is the account’s history and it never moves again, which is exactly why a sweep leaves them out and why one walk of them is enough. Forks stay as configured. What does still move on an archived repository, its stars, forks and watchers, a sweep reads without a backfill: the listing it already pays for says which repositories are archived, and every totals sweep asks about all of them, one query per twenty five, for two rows each. One is gh_repo_archived, dated when the repository was archived. The other is its gh_repo_total, stamped at the sweep like a collected repository’s, and that is the row the account’s star and fork totals are read from.

  • Every pull request and issue, in pages of fifty, where a sweep reads what changed since the sweep before, and once a day every open item and what moved since the day before.

  • Every pull request and issue the account opened in other people’s repositories that has since been merged or closed, in pages of a hundred ordered by when each one last moved, where a sweep reads each of the three closed states back to a cadence before the sweep before. That is enough for a sweep because a merge or a close moves the item to the top, however long ago it was opened, and it is one page unless more than a hundred items moved in that time; a sweep that read one page and no more lost the item that a hundred later updates, a bot locking old threads or a relabel, had pushed onto the second. The ones still open are bounded by neither a page nor a date: every sweep reads them all, since each one gets a row for every day it stays open. In 2.5.1 and earlier each of the five searches read its newest hundred and stopped, on a sweep and on a backfill alike.

    GitHub serves a thousand results of any search and no more, so an account past a thousand in one of those states keeps the thousand that moved most recently. That is said in the log, once per count, rather than left to show as two panels that disagree:

    level=WARN msg="outbound search read fewer items than it counts, GitHub serves a thousand at most" kind=pull_request state=merged count=2860 read=1000
  • Every page of artifacts, of repository activity and of code scanning analyses, where a sweep reads five, two and one.

  • Every release, deployment and discussion, and every Dependabot and code scanning alert, where a sweep reads the newest page of each list.

  • Every issue and discussion comment the account left anywhere, and every answer of its own that was accepted, where a sweep reads the newest hundred comments of each kind and the newest five hundred accepted answers.

  • Every star the account gave, where a sweep reads the newest five hundred.

  • The co-authored pull requests behind Pair Extraordinaire’s progress, walked over the account’s whole life, where a pass adds the days since the last one and walks the whole history once a week.

  • Every page of every stargazer list, where a sweep reads the newest hundred stars of a repository whose list it has walked whole once. Up to 2.5.0 a first walk that failed was recorded as done, and a backfill read a recorded list by its first and last pages, so the stars in between were never read; a backfill now reaches them.

  • The whole inbox, read threads included, where a sweep reads the unread threads that moved and, once a day, twenty pages of fifty with the read ones.

  • Every month of billing GitHub still answers, until three empty months in a row, where a sweep reads two.

  • The newest hundred webhook deliveries per hook, where a sweep reads thirty.

  • At most 500 failed job logs per repository, none older than ninety days, once joblogs has a cadence; a sweep reads at most ten.

Two things a sweep remembers, a backfill leaves alone. A 403 or 404 that a sweep remembers for a day is asked again, since a backfill consults no such memory. And it reads the cache file beside the state file but writes nothing to it: the pages it walks are ones no sweep asks for, and kept there they would push out the sweeps’ own, so a service started after it would start colder than it stopped.

A backfill runs every enabled family whatever the state file says. It enables none of them. The three families that ship with a cadence of 0, deps, history and joblogs, stay off unless they have been given one by name under every.families, and -backfill does not change that.

history is the one that surprises people, because walking every past year of the contribution calendar is exactly what somebody asking for “the whole history” has in mind. Give it a cadence first:

every:
families:
history: 24h

See cadences for why default and groups cannot switch these three on either.

Two endpoints that needed their own handling

Section titled “Two endpoints that needed their own handling”

Both were found by running it, not by reading documentation.

Dependabot refuses page numbers. It answers an error to page= outright and pages by cursor instead, so the alert walk is written against cursors.

The GraphQL gateway gives up on a hundred pull requests, and on fifty commits. Asking for a hundred pull requests with their reviews in one query answers an HTML 502 after about ten seconds, and so does asking for fifty commits of a busy repository with the checks each one carries: twelve pages of the author’s own commit walks met it, in six repositories. Both walks halve their page size and retry on the same cursor, silently: neither collector carries a logger, so a backfill of a busy repository shows this only as a slower family, never as a line. The page is halved while it is larger than ten, so fifty is asked again at twenty-five, twelve and six. A page the gateway still gives up on at that smallest size, and a timeout in any walk with no smaller page to ask, is a failure: the rows read before it are written, the collector failed line says the query is too large for one request, and the repository is not recorded in the checkpoint, so a resume walks it again. Up to 2.6.3 the commit walk took the timeout for the end of the history and recorded the repository as walked. Only that answer is halved for: a 503, or a 502 that comes back sooner than ten seconds, is not a query the gateway ran out of time on, and the client asks it once more at the same size.

It is idempotent. Points are keyed by measurement, tags and timestamp, so a backfill run twice rewrites the same rows rather than doubling them, in every store that keeps the history. What it costs is API quota and time.

Two things worth doing first: run -list to confirm the repository set, and check that the store you are writing to is the one that keeps dates. Backfilling into Prometheus collects a great deal of history and then reduces all of it to a single current value.

Written and maintained by
MIT licenceRelease history