Cost of a sweep
Measured on 2026-09-11 with the real binary against the live API, every family
switched on and every request logged with the x-ratelimit-resource and
x-ratelimit-used headers of its own response. Cold is the first sweep of
a fresh install: an empty ETag cache, a month of workflow runs. A restarted
service no longer starts there: it reads its cache back from the file beside
the state file, and only one
that has lost that file pays the empty cache again. Steady is the third
sweep of the same process a few minutes later, when about four fifths of its
requests were answered 304 and cost nothing; the second sweep is the one that
pays the fill-in the actions row describes. A 304 is free, and it is only ever
available to REST: GraphQL carries no ETag, so its column is the same in both
sweeps.
The figures are per repository where the family asks per repository, and per
sweep where it asks about the account. core is the REST bucket, pt a
GraphQL point. Families with a request that is not core say which bucket it
went to: search (30 a minute), webhook_deliveries (500 a minute) and
dependency_sbom (100 a minute) each have their own. A row the round of
2026-09-11 changed describes the requests the code makes now and says what the
two sweeps measured where that differs. There is no total row on purpose: a
total belongs to one account and ages with it, so multiply the per-repository
rows by the number of repositories -list prints and you have your own.
Per family
Section titled “Per family”| Family | Cadence | Cold | Steady | Notes |
|---|---|---|---|---|
| account | 1h | 1 pt, 7 core | 1 pt, 1 core | one query: the calendar, the totals, the pins, the sponsors block and the star lists; 66 KB, never compressed, never conditional; beside it the profile, for the organizations GraphQL’s following leaves out, which GitHub all but never answered with a 304, once in 48 conditional reads from 2026-09-12 to 2026-09-27, and the six package listings GraphQL cannot see, which answer one until a package changes |
| achievements | 1h | 1 page off budget, 1 pt, plus the co-authored walk | 1 page off budget, 2 pt | the badges come from the public profile page, read without the token and charged to no bucket. The progress rows beside them do come from the API: one counts query, and a walk over the merged pull requests whose co-authored commits no count exposes. The first pass walks the whole life of the account, and since search stops at a thousand results it splits any range that overflows into two, one point per page and one per split, and one per query that reads on past the first hundred commits of a pull request with more, ten pull requests to a query, which is tens of points on a long-lived account and one on a new one: 35 queries and 23.7 MB over 2,315 public merged pull requests, measured on 2026-09-27. The count it finds is kept in the state file with the day it covers, so every later pass walks only the pull requests merged since, one page, 468 to 513 KB for each hourly pass the production proxy logged on 2026-09-27 on the same account, a size that grows through the day with what is merged, where the whole walk had been 18 to 24 MB every day; the whole history is walked again once a week, since a repository made private or deleted takes its pull requests out of the count. The page and the two points every hour, the whole walk once a week |
| totals | 1h | 1 search, 1 pt, plus 1 pt per 10 repos and 1 per 25 archived ones set aside | the same | ten of the eleven counts are one GraphQL query of ten aliases, 406 bytes at cost 1, measured identical to the REST answers of the same minute; only the commit count is still a search, charged every time; a refused counts query is asked again one count at a time, ten points on that sweep only, so one failed search loses one number as it did in REST |
| ratelimit | 15m | 1 pt | 1 pt | GET /rate_limit is free; the point is the GraphQL half of the same question |
| events | 15m | 3 core | 1 core | before the round: three pages of a feed that changes every time, so If-None-Match never matches. Now: the walk stops at the page that carries the newest event of the previous sweep, remembered as last_event in the state file, which is usually the first page: 1 core a sweep, 192 a day fewer at the quarter hour; a first sweep and a backfill still read the three |
| notifs | 15m | 20 core | 1 core | before the round: twenty pages of fifty; one new notification shifts every page, so the ETag never matches, for six or seven rows that were new. Now: since= the newest updated_at seen minus two cadences (last_notified in the state file), which at the quarter hour is one page for a window that opens half an hour before the newest thread seen: 1 core a sweep, and the twenty once a day (last_full), on a backfill, and on the first sweep after an upgrade, read threads included (all=true): that bounds to a day whatever GitHub’s filter left out of the windowed reads, and it is the read that closes a thread, since a sweep lists unread threads only and a thread read without a reply leaves that listing rather than coming back with its unread tag changed. about 1,800 core a day fewer at the quarter hour |
| billing | 1h | 2 core | 1 core | one per month walked; the month in progress changes, the previous one is a 304. Measured from 2026-09-13 to 2026-09-26, six hours apart, the month in progress answered 200 to all 46 conditional reads and the one before 304 to all 46 |
| profile | 12h | 4 core plus 1 per package | 0 core | the profile, the social accounts, the gists, the packages and one page of versions per package; all 304 after the first sweep |
| outbound | 1h | 9 pt, plus 1 per further hundred stars given, items still open or moved since the sweep before, or accepted answers | the same | all GraphQL: the starred list is a query of a hundred a page (13.5 KB against 617 KB decompressed from REST), newest first, up to five pages on every sweep and every page on a backfill, the five searches one point a page (7 KB against 100 KB), the two comment walks one each, and the account’s accepted answers one more, read on their own so an answer accepted after its comment left the newest hundred is still written (15 of them on 2026-09-26, one page, and five pages at most on a sweep). The two searches for what is still open are read to the end on every sweep, one page until an account has more than a hundred open of a kind, and the three for what was merged or closed back to a cadence before the sweep before, ordered by what moved last, which is one page each unless more than a hundred moved in that time; a backfill reads those to the end too, ten pages at most, since search serves a thousand results and no more. Nothing here has an ETag on either path, so the two columns are the same |
| history | off | 1 pt per past year | the same | the account’s whole calendar, once, at a point per year it has existed |
| keys | 24h | 2 core | 0 core | the SSH and the GPG keys |
| traffic | 6h | 4 core per repo | 0 core | views, clones, referrers, paths; the fourteen-day window is a 304 until it moves |
| repo | 1h | 3 core per repo, 2 pt per 10 repos | 0 core for a repo that did not change, 2 pt per 10 repos | the repository, its community profile and one page of releases each; languages, topics, rulesets and branch protection ride in one GraphQL query per ten repositories. A fork has no community profile, and the 404 it answers, charged in full because it carries no ETag, is asked once a day rather than on every pass: fifteen forks were charged 404 of them in 30.9 hours, and with the memory are charged 15 a day |
| branches | 24h | 1 pt per 14 repos | 1 pt per 14 repos | one query per fourteen repositories |
| stars | 1h | 1 core per repo, plus 1 core per 30 weeks of its life | 1 pt per 10 repos, plus 1 core per repo, usually a free 304 | the full walk through REST the first time a repository is seen, then the newest hundred of every repository in one GraphQL query per ten, about a kilobyte per repository; a restart with a state file starts at that query. Before the round the last page was asked of every repository on every sweep, all but one of them 304. The daily star history came after the round and is not in its figures: 1 core per repository, a 304 unless a day of the last thirty weeks gained or lost a star or a week began; its first read is a page per thirty weeks of the repository’s life, 180 pages for the nineteen starred repositories of the account measured, once |
| issues | 1h | up to 4 pt a page of a month of moves, 1 pt or more for open items | 1 to 2 pt per repo that moved | before the round: a page of fifty with ten review threads each, 8 points, which the gateway refused with a 502 once a sweep on the busiest repository. Now: 2 points of GraphQL per repo for what changed in two cadences, ten at a time (1 point where the repository holds five or fewer). The first sweep of each UTC day also reads every open item of every repository, however long ago it moved, in pages sized from the open counts of the last totals, and what moved since the day’s read before it, a month back on a new state file, in pages of up to twenty-five (5, 10, 20, 25 or 50 cost 1, 2, 3, 4 or 9, asked with dryRun on 2026-09-30). Until 2.6.5 the day’s read was one page of the fifty items of any state updated last, which missed an open item once fifty others had moved past it, and met the gateway’s ten seconds on the busiest repository every day. Since 2.6.0 the read of what changed asks only the repositories where an issue or a pull request was updated in its window, which one query for all of them says first (see Asking first what moved); the read of every open item is made in every repository |
| issueevents | 1h | several pt per repo updated in the month, plus 1 core per stacked pull request in it | 1 pt per repo that moved, 0 to 1 core | before the round: one page of a hundred events each, up to a megabyte per repository because every event embeds its whole issue; a 304 on all but the repository that moved. Now: one GraphQL query per repository, the timeline of the ten most recently updated issues and ten pull requests for the events of the last two cadences (1 pt, 2 to 6 KB, a second page only when more than ten items moved), plus one core per pull request in a stack, whose added_to_stack event the timeline cannot name; thirty days on a first sweep, once, which is minutes rather than seconds where the pull requests are nearly all stacked, because almost every one updated in the month takes the per-issue road for its added_to_stack; a steady sweep is a point per repository that moved and one such read at most. And from a cadence before the last run after a gap, so a stopped process does not leave its hours out of the series. Measured against the list over a week of two repositories: every event of every type agrees, field for field, except a commit referencing an issue nobody has touched, 3 of 2,217, which does not move the issue and so is not asked for. A backfill walks /issues/{n}/events per item instead of the list, 1.4 KB compressed per twelve events against 45 KB per event. Since 2.6.0 a sweep asks only the repositories where an issue or a pull request was updated in its window (see Asking first what moved) |
| actions | 15m | a month of runs per repo, plus 1 core per run | 1 core per repo that had a run, plus 1 per run completed since | before the round: pages of a hundred runs reaching a month back, a job list per run, and the caches and workflows of every repository, which was three fifths of the cold sweep’s bytes, and those job lists again on every sweep, nearly all of them 304. Now: pages of 30, one on a quiet repository and up to 7 while they come full of runs newer than the window, plus one per run not yet expanded, jobs listed once per attempt, at most 20 new runs a sweep; pages of 100 on the first sweep and in a backfill. Measured over three sweeps of one process, the second is the one that pays the fill-in, because the page of thirty is a new URL for every repository and holds runs the first sweep’s twenty did not cover; from there a sweep costs one page per repository that had a run and one job list per run completed since, so what it costs is how many runs the account completes. The cache listing is read in pages of a hundred entries, up to ten, so only a repository past a hundred pays for a second, and a page is a 304 while the listing has not changed: from 2026-09-25 12:58Z to 2026-09-27 01:39Z the account’s listings answered 304 to 3,059 of 3,256 requests |
| artifacts | 1h | up to 5 core per repo | 0 to 5 core per repo | a busy repository fills its pages, and all five of them are charged again whenever an artifact was added or expired: 5 on one sweep, 0 on the next |
| security | 1h | 2 core per repo | 0 core | Dependabot and code scanning alerts; most are the 403 of a repository with Dependabot switched off or the 404 of one without code scanning, and a refusal carries no ETag, so each was charged again on every sweep before the round remembered them. Now: a refusal is answered from memory for a day, per family, repository and endpoint, so the steady figure is 0 core on the second sweep and one request per refusal once a day. The rest are conditional requests, all 304. Only a list whose first page comes back full, in a repository past a hundred alerts of that kind, asks for more: its open alerts, a page per hundred still open, and for code scanning a page of one alert whose last page is the total. Measured on 2026-09-26 that is two lists of this account and three requests a sweep, conditional like the rest and a 304 while nothing moves |
| stats | 12h | 3 core per repo | 0 core | participation and the punch card; GitHub recomputes them slowly and answers 304 |
| discussions | 1h | 2 pt per repo with a forum | the same | before the round: fifty threads with twenty comments and twenty replies each, 11 points, asked of every repository, including every one with no forum. Now: 2 points per repo with a forum, the ten most recently updated threads with the same twenty comments and twenty replies each (comments come oldest first, so a shorter page there would stop recording the eleventh comment of a thread); a repository whose forum is off is never asked |
| commits | 1h | 1 pt per repo with a commit in the month | 1 pt per repo that moved | the last two cadences of the default branch, since 2.6.0 only in the repositories whose head was committed in them (see Asking first what moved); a month back on the first sweep, which is a few hundred kilobytes for a busy repository; where the gateway gives up on that page of fifty, as it did on the busiest repository measured, it is read in pages of twenty-five |
| activity | 15m | 2 core per repo | 0 to 2 core per repo | the repository log, a hundred entries per page; charged only for a repository whose log moved, 2 on one sweep and 0 on the next |
| analyses | 1h | 1 core per repo | 0 core | most are the 403 and 404 of repositories without code scanning, charged again each sweep before the round; now remembered for a day, so the steady sweep is conditional requests and 0 core |
| forks | 12h | 1 core per repo | 1 pt per 10 repos | one page each through REST on the first sweep of a fresh install and in a backfill, then the newest hundred of every repository in one GraphQL query per ten, under a kilobyte per repository; a repository the batch reports holding more than a hundred forks is walked through REST as well, since a fork row’s stars and last push move and the batch cannot refresh the rows past its page. Before the round: one request per repository a sweep, all 304 |
| planning | 6h | 1 pt per repo | 1 pt per repo | labels and milestones |
| joblogs | off | 1 core per repo, 1 blob per failed job | 0 core | the failed runs of the last hour per repository, asked of a list filtered to the month a re-run can reach back to, and one blob per failed job from object storage, outside the API’s quota; the filter’s URL changes once a day, so a day costs one charged page per repository and the rest are 304 |
| settings | 6h | 2 core per repo, 1 webhook_deliveries per hook | 0 core, 1 webhook_deliveries per hook that moved | webhooks, environments and deploy keys; the deliveries of every hook are charged to their own bucket |
| rulesets | 24h | 1 core per repo plus 1 per ruleset | 0 core | one list per repository and one history per ruleset, both with an ETag; none of it charged on a day nobody edited a ruleset, when the steady sweep asks the same questions and every one of them is conditional |
| inventory | 24h | 4 core per repo | 0 core | the workflow token policy, the secrets, the code scanning setup; the refusals were the 403 of repositories without it, now remembered for a day, which at this cadence is the next sweep anyway; the rest are conditional requests, all 304 |
| deployments | 30m | 1 pt per 5 repos | 1 pt per 5 repos | one query per five repositories, the newest hundred deployments each |
| policyfiles | 24h | 1 pt per 5 repos | 1 pt per 5 repos | one query per five repositories, the history of four paths each |
| deps | off | 1 core and 1 dependency_sbom per repo | 0 core, 1 dependency_sbom per repo that received a commit | one commit read and one SBOM per repository; the SBOM has its own bucket and most were 404, now remembered for a day. GitHub regenerates a SBOM on every read and its ETag never matches, so the photograph is taken only when the head moved since the last sweep (the head itself is a free 304 when it did not): 0 dependency_sbom on a repository without a commit, one per repository that received one, and one of the SBOM reads timed out on GitHub’s side on each of the three sweeps measured |
What weighed before the round was not REST but GraphQL: issues and
discussions were 88 % of the points, because both queries were sized for a
backfill and asked on every sweep. The only search is the commit count of
totals, one a sweep against a budget of thirty a minute.
The round of 2026-09-11 lowered several of these rows, and the first
measurement is the baseline it is measured against: the job lists a run has
already had expanded are not asked for again, the runs page is thirty on an
ordinary sweep, discussions is asked only of repositories with a forum and
for ten threads, pulls is sized to the repository and windowed to two
cadences, notifications and the event feed stop at what the previous sweep saw,
a 403 or 404 is remembered for a day instead of being charged again every hour,
and the newest stars and forks, the starred list, the outbound searches and ten
of the eleven totals counts moved from REST to GraphQL, the same rows for a
point apiece instead of a request apiece. Measured again after the round, three
sweeps of one process on the evening of the same day, the steady sweep charges
about a third of the core and a third of the GraphQL points it charged
before, moves half the bytes on the wire, and answers about four fifths of its
requests with a 304. Projected to a day at the cadences of the time that is
roughly an eighth of the REST calls and a third of the points, plus one job
list per workflow run completed, which is the one term that grows with how busy
the repositories are rather than with how many there are. Two things did not
fall. The cold sweep charges more core than before, because the first sweep
of issue events now reads the whole month through the timeline and, where pull
requests are stacked, one per-issue list for nearly every pull request updated
in it: minutes, once per process. And a sweep that runs every family at once
still takes minutes rather than seconds, because its conditional requests cost
a third of a second each and its GraphQL queries nine tenths, one after the
other; the sweeps production runs are the 15 minute one and the hourly one,
each a fraction of the whole.
The cadences of 2.6.0 run ten families more often, events, notifs and
activity every quarter of an hour, deployments every half hour, and stars,
billing, analyses, account, totals and discussions every hour, because
their extra passes cost next to nothing. Costed from the request log of the
production process from 2026-09-25 12:58Z to 2026-09-26 19:50Z, 37 repositories,
each family’s first pass left out so that every pass counted had a warm ETag
cache, they add about 250 core requests and 500 GraphQL points a day, 0.2 and
0.4 per cent of the two daily budgets, and 22 search requests. The points are
mostly deployments, eight a pass there, and totals, six; the core requests
mostly activity, events and analyses, which are charged only for what
moved, since a 304 is free: in the same log, a core request answered 304 left
x-ratelimit-used where the request before it in the same rate window had put
it 11,692 times out of 11,771. Two more families run every hour since the same
release, each priced in its own row above: outbound, which ran every twelve,
and achievements, which ran once a day while every pass walked the whole
merged history.
Asking first what moved
Section titled “Asking first what moved”commits, issueevents and the incremental pass of issues read only what
moved since a window of their own, one GraphQL query per repository, and GitHub
charges a query for the page it asks for, not for what comes back. So a sweep
that runs any of the three first asks every repository, twenty-five to a query
at one point each, when the head of its default branch was committed and when
its newest issue and its newest pull request were updated, and each family
leaves unread a repository where nothing moved since the start of its own
window. That is the condition under which the read comes back empty, not a
guess at it: measured on 2026-09-27, history(since:) holds a head committed at
the very second it names and nothing one second later. A repository the query
does not answer for is read as it always was, and neither the day’s read of
every open item of issues nor a backfill asks it at all.
No row of the table above carries the query, since it belongs to none of the
three: it is one point per twenty-five repositories on every sweep that runs
any of them. A family that leaves repositories unread says how many, once a
pass, and its gh_collector_family
row counts them in repos beside the ones it read, with no rows written for
them:
level=INFO msg="nothing moved since the window, repositories left unread" family=commits repos=34A query that failed, or answered for some repositories and not the others, is said at warning, and every repository it gave no answer for is read as if it had not been asked:
level=WARN msg="which repositories moved could not be asked, reading those it did not answer for" answered=36 repos=37 err="graphql: the gateway gave up (502); the query is too large for one request (36 repositories were asked successfully)"From the request log of the production process between 2026-09-25 12:58Z and
2026-09-26 19:50Z, 27 passes of each family over 37 repositories, before the
query existed: 961 of the 999 commits answers held no commit, and 468 of the
968 incremental issues answers and 486 of the 1,005 issueevents answers held
no item at all, 2,077 points in all, 43 per cent of the 4,786 the process spent
on GraphQL. For those 37 repositories the query is two points a sweep, 54 over
the 27 sweeps, so it saves 2,023 points, about 66 an hour: 42 per cent of all
the GraphQL points and 52 per cent of the three families’. That is a floor. An
answer is counted empty when it is under 300 bytes, the budget block and an
empty list; a repository whose issues and pull requests all predate the window
answers with a page of them anyway, is not counted here, and is skipped all the
same.
GraphQL is the cheap one, by a wide margin
Section titled “GraphQL is the cheap one, by a wide margin”One query returns the full 366-day contribution calendar, every contribution total, the per-repository commit breakdown and the social counts, for one point of a five thousand point budget. The same data over REST would be dozens of calls and would not include the calendar at all, because the calendar exists nowhere else.
That is why the account family can run every hour and still cost almost nothing, and why the expensive families are the REST ones that scale with how busy the repositories are.
What scales with activity, not with size
Section titled “What scales with activity, not with size”Everything else costs a roughly fixed number of calls per repository. These do not:
actionscosts a page of thirty runs, up to seven while they come full of runs newer than the window, plus one request per run whose jobs this process has not written yet, at most twenty a sweep. A repository with continuous integration on every push generates runs continuously; a quiet one costs the one page, answered 304, and nothing else.artifactswalks up to five pages per repository, and a busy repository fills them and has them charged again whenever an artifact was added or expired.activityis charged only for a repository whose log moved: the page of one that did not is a free 304.commits,issueeventsand the incremental pass ofissuesread only the repositories where something moved in their window, a point or two each, after asking which for one point per twenty-five repositories. The day’s read of every open item ofissuesis made in every repository, busy or not.outboundcosts a point more for each further hundred of what it reads: stars given, items still open or moved since the sweep before, and accepted answers.
Switching a family off
Section titled “Switching a family off”Set its interval to 0.
every: families: artifacts: 0 joblogs: 0A family that is off writes nothing and costs nothing. Its dashboard panels go empty, which is the honest reading. See cadences.
Reading the arithmetic for your own account
Section titled “Reading the arithmetic for your own account”A first sweep is the expensive one: the full stargazer walk, the whole daily
star history of every repository, a month of workflow runs, the co-authored
pull requests of the account’s whole life, which achievements walks again
once a week (35 queries and 23.7 MB over 2,315 public merged pull requests,
measured on 2026-09-27), and (with every.families.history set) every past
year’s contribution calendar. After that, multiply the per-repository rows
above by the number of repositories -list prints, and divide the hourly
budget by the cadence.
A card runs every family whatever the cadences say,
because every number it draws comes from that one sweep, so N cards are N
sweeps, and a workflow drawing three of them pays three. Its GraphQL is a whole
sweep’s every time, since GraphQL has no ETag. Its REST requests are the cold
column’s only the first time: every run reads the cache file beside its
state_file and asks with the validators stored there, so from the second card
on a page that did not change is a free 304. A run whose state file is not kept
from one run to the next pays them cold every time: a container with nothing
mounted where the file goes, or the Action given no config:, which starts
each run with a state directory of its own. A -card-only run reads the cache
file and never writes it, so it is warm only when another run keeps that file.
The signal that the sum came out wrong is a warning, not a guess:
level=WARN msg="rate limit reserve reached, family skipped" family=actionsOnce in a while is fine. Every sweep means the cadences are too fast for the number of repositories.