The sweep
A sweep is one pass over every family whose interval has elapsed, except that the running service starts one family of six hours or more in each sweep, more only where one could not keep every cadence, and leaves the others that are due for the next ticks: the slow families take turns. It is an increment, not a rebuild: it asks for the little that can have changed since last time, writes what it got to every configured store, and records when each family ran.
The shape of one pass
Section titled “The shape of one pass”Two things in that diagram are the whole design. Every collector produces dated points, and the stores that cannot hold a date get them reduced to current values first, by the same reducer, before they ever see the data. That is the subject of dating a point.
Why families and not one interval
Section titled “Why families and not one interval”The surfaces move at wildly different speeds. Workflow runs finish every few minutes on a busy account. Labels and milestones are edited a few times a week. The list of forks changes a few times a year. One interval for all of them would either waste the rate limit on the slow ones or lose the fast ones, so each family carries its own.
| Family | Default | Collects |
|---|---|---|
actions | 15m | Workflow runs, jobs, steps, the Actions cache |
activity | 15m | The repository activity log, where a force push is recorded |
events | 15m | The account event feed, which keeps only the last 300 |
notifs | 15m | The notification inbox |
ratelimit | 15m | What the collector has left to spend, in each budget |
deployments | 30m | Deployments and their environments, batched over every repository |
account | 1h | Profile, contribution calendar, contribution totals |
achievements | 1h | The profile badges, and the distance to each next tier |
analyses | 1h | Code scanning analyses, which GitHub prunes |
artifacts | 1h | Artifacts and their expiry |
billing | 1h | Usage per day, product, SKU and repository |
commits | 1h | Lines changed and signature state, per commit |
discussions | 1h | The forum half of a repository |
issueevents | 1h | The timeline of what moved: labels, assignments, transitions |
issues | 1h | Pull requests, issues and reviews, per item |
outbound | 1h | Stars given, and work in other people’s repositories |
repo | 1h | Stars, forks, languages, topics, releases, rulesets |
security | 1h | Dependabot and code scanning alerts |
stars | 1h | Stars per day for every repository, and the stargazer walk once, then the newest hundred |
totals | 1h | The lifetime numbers, asked of GitHub rather than added up here |
planning | 6h | Labels and milestones |
settings | 6h | Webhooks and their deliveries, environments, deploy keys |
traffic | 6h | The whole 14-day window, rewritten |
forks | 12h | Who forked, and when |
profile | 12h | Packages, gists, social accounts |
stats | 12h | Commits per week, the punch card, the workflow definitions |
branches | 24h | Which branches are live and how stale each tip is |
inventory | 24h | What a workflow’s own token may do, both secret stores, default code scanning |
keys | 24h | The account’s SSH and GPG keys, and when each expires |
policyfiles | 24h | SECURITY.md, CODEOWNERS, dependabot.yml and FUNDING.yml |
rulesets | 24h | Every version of every ruleset’s changelog |
deps | off | The dependency SBOM of each repository, and what changed |
history | off | Each year’s contribution calendar, the current one included |
joblogs | off | The tail of every failed job’s log |
Setting any of them to 0 switches it off entirely. See
cadences.
Discovery
Section titled “Discovery”The repository list is rebuilt once an hour, by the first sweep that finds it an hour old, give or take half a tick: a sweep reads its clock a few milliseconds either side of the hour, and without that margin a repository created in between waited a tick more. Repositories are created rarely and listing them costs a page per hundred, so anything shorter spends quota to learn nothing. Forks and archived repositories are excluded by default, for the reason set out in targets.
Failure is per repository, not per sweep
Section titled “Failure is per repository, not per sweep”A family that fails on one repository is logged and skipped; the sweep continues. This matters more than it sounds, because “failure” here is usually a feature being switched off: of fifty repositories, most have Dependabot off, and each of them answers 403. Treating that as an error would lose the other forty-nine.
There is one exception, and it is deliberate. A family where every repository failed is not marked as run. Marking it would hide the outage until the next cadence, which for the twelve-hour families is half a day.
level=WARN msg="family failed everywhere, not marking it as run" family=securityThe state file
Section titled “The state file”state_file holds nine things, and
the two a sweep is judged by are when each family last ran and when each
repository was first seen.
The second is what makes the one-off full walk of the stargazer list happen
once instead of on every sweep, as history_read does for the daily star
history. It is written through a temporary file and renamed,
so a crash mid-write cannot leave a truncated state that would trigger a full
re-collection.
Beside it the collector keeps what makes a sweep cheap rather than what makes it
complete: the ETag cache, the workflow runs whose jobs were written, the
refusals and the pull request page sizes, in
<name>-cache.bin. A restart
reads it back and starts where the last process stopped, not from nothing, and
deleting it costs one full-price pass of each family and loses nothing.
The brake
Section titled “The brake”Before each family the collector checks the three rate buckets it actually spends from. If any of them is at or below its reserve and the window has not reset yet, the family is skipped and a warning is logged rather than the budget being spent to the last call. See rate limits.
A backfill is the opposite intention: it waits for the window to turn over instead of skipping.