Skip to content

Rate limits

GitHub does not have a rate limit. It has fifteen of them, and every response says which one it just charged in the x-ratelimit-resource header.

BucketLimitSpent by
core5000 per hourEvery REST call
graphql5000 points per hourEvery GraphQL query: the pull requests and issues, the account, commits, discussions, deployments and the rest, which the cost table prices family by family, and the query that asks which repositories moved
search30 per minuteThe commit count of totals, and nothing else: GraphQL search has no COMMIT type

Two more are charged by one family each: webhook_deliveries (500 a minute) by the delivery list of every hook in settings, and dependency_sbom (100 a minute) by the SBOM in deps. Neither is one of the fifteen /rate_limit reports; they exist only in the headers of the endpoints that charge them, and both are named in the cost table where they apply. The brake looks at the three above, not at whichever bucket was charged last, which matters for the reason below.

github.reserve_rate is how many calls are never spent. It defaults to 500. The collector stops a family rather than crossing that line, so whatever else uses the same token keeps working.

github:
token: ${GITHUB_TOKEN}
reserve_rate: 500

The reserve is scaled to each bucket: a fifth of its limit, or the configured value, whichever is smaller.

When a bucket is below its reserve, the collector waits for the reset the response already told it about, rather than sleeping a guessed interval and retrying. In a normal sweep the family is skipped with a warning:

level=WARN msg="rate limit reserve reached, family skipped" family=artifacts

Once is fine. Every sweep means the cadences are too fast for the number of repositories; see cost of a sweep for what to lengthen first.

Every response is cached by its ETag. A repeat request sends If-None-Match, and GitHub answers 304 Not Modified when nothing has changed.

A 304 costs no quota at all. That is not an optimisation detail, it is the thing that makes short cadences affordable. Most of what this collects barely moves: a repository’s language breakdown, its topics, its community profile, the list of workflows. Asking for them every hour would be unaffordable if each question cost a call. Asking whether they changed costs nothing.

The consequence is that the measured cost of a sweep in the next page is an upper bound reached on the first sweep and after a change, not the steady-state figure.

What the cache keeps beside each ETag is the value the collector decoded, encoded again, not the body GitHub sent. A page of a hundred workflow runs is 1.4 MB of which the collector keeps a few hundred bytes per run, so a sweep’s cache holds about a ninth of what the raw bodies would, and a 304 decodes a ninth of the bytes. A test runs every REST collector twice against a fake GitHub that answers the repeat 304 and fails on the first point that differs, which is what keeps the replay the same answer as the original. An entry is keyed by the URL and the type that decoded it, so the one URL two collectors read differently, GET /repos/{owner}/{repo} (four flags for discovery of a repository named in targets.repos that no listing returned, forty fields for the repo family), has an entry per reader and neither is answered from the other’s.

The cache is bounded, and the bound is 256 MB. It is an LRU, and an entry is charged the bytes it holds plus two hundred for the map slot, the list element and the struct around them, so a cache full of small answers is accounted at what it really costs rather than at half of it. There is no key for it in the configuration file: a deployment whose live set does not fit raises it from Go with SetCacheLimit, and CacheStats().Evicted is the number that says whether it had to, because it stays at zero for as long as one sweep’s live set fits.

Two things follow. A limit below one sweep’s live set is not a smaller cache but no cache, since every sweep would evict what the next one is about to ask for. And the bound is resident memory once the cache is full, which is the number to know before giving the process a container with a memory limit: it can hold that much of decoded bodies on top of its own footprint.

The cache outlives the process. What was asked for within twice the longest cadence is kept in a file beside the state file, written at most every five minutes while the collector runs and once more when it stops, and a restart reads it back, so its first sweep asks with the validators the last one stored. Before that file, the first sweep after a restart of the author’s service made 130 requests and not one of them was answered 304. An entry is kept under a digest of everything its type decodes rather than under the type’s name, so an upgrade that changes what a collector reads asks that collector’s URLs again in full instead of answering them with a body stored without the new field.

A 403 or a 404 is how GitHub says a feature is switched off: Dependabot on a repository that does not use it, code scanning where it was never enabled, a forum on a repository with discussions off, the community profile of a fork, which GitHub does not serve at all (all 28 forks of the account it was measured on answered 404, and all 39 other repositories 200). That is not an error, and the collector turns it into “there is nothing here”.

What it is not is free. A refusal carries no ETag, so where a page that did not change costs nothing, a feature that is off is charged in full on every sweep, for ever. So a refusal is remembered for a day, keyed by family, repository and endpoint, and the sweeps inside that day ask nothing.

Two consequences worth knowing:

  • Switch a feature on and it is noticed a day later at the latest, not on the next sweep.
  • Restarting the process does not ask again. The refusals are kept with the ETag cache in the file beside the state file, each until the end of its own day, and deleting that file is how to ask again straight away.

A spent budget is a 403 as well and is never remembered as one: the client types it apart precisely so that a rate limit is not read as a feature that is off.

An answer GitHub could not finish is asked once more

Section titled “An answer GitHub could not finish is asked once more”

Four statuses say GitHub did not finish an answer: a 502 or a 504 is the gateway in front of it giving up on an answer that did not come back in time, a 500 is the application giving up on it itself, and a 503 is a server that could not take the request just then. All four are intermittent. Measured on the author’s own sweeps between 2026-09-11 and 2026-09-29, of 457,098 REST requests GitHub answered 50 with a 502, each after ten to eleven seconds, 2 with a 504 and 9 with a 500. Seven of those 500s were the artifact listing of the repository with the longest artifact history, after eight and a half seconds: the same slow listing, giving up at the application’s own limit rather than at the gateway’s. The size of the page is not the cause: asking for one artifact took as long as asking for a hundred.

So a REST request that answers any of the four is asked once more, two seconds later, before the collector hears about it. Of the 21 502s and 504s asked again that way since 2.6.0, 20 were answered. A 500 has been asked again that way only by hand, once, and answered 500 again; it is asked again all the same, because it is the same listing giving up one step further in, and a retry that fails costs what a 502’s does. Three things hold for that second attempt:

  • It is charged. A 500 and a 502 each cost a request from core, even when asked conditionally, where a 304 costs none. So the retry goes through the brake like any other request, and a budget already at its reserve is not spent on it.
  • It is conditional. It carries the same If-None-Match as the first attempt, so a page that has not changed still comes back as a free 304.
  • It is the only one. Each attempt can hold the family for ten seconds, and a request that fails twice is logged and left to the next sweep as before. What the collector read before it is kept either way.

A failed job’s log is two requests: the API answers with a redirect, which is charged, and object storage sends the text. When storage answers one of the four, storage alone is asked again, on the same signed URL and without the brake, because it is not GitHub’s API and spends no budget. The API is not asked for a new redirect.

A GraphQL query is asked once more the same way, two seconds later and through the brake, with one answer left out. GitHub documents its GraphQL timeout as a 502 or a 504 once a query has run for more than ten seconds, and that one is a query too large for one request. The rest of the four are asked again as they were, on the same page: a 500, a 503, and a 502 or a 504 that came back sooner than ten seconds, which cannot be that timeout. Measured on the same sweeps, of 66,821 queries GitHub answered 42 with a 502 after 10.45 to 11.20 seconds, 9 with a 504 after 11.03 to 11.23, and 3 with a 503 after 0.70 to 1.08 seconds. Each of those three cost a point and was a page of a commit history, which is why a quick failure is not read as a query too large: the commit walk took one, at the time, as the end of what it could read, and each of the three stopped its walk there and reported success. A query that fails twice is a failure the collector reports, and the next sweep asks again; so is an answer that is not JSON and is not that timeout, which is no GraphQL answer at all.

What is not asked again:

  • A refusal. Any 4xx, and a 501 or a 505, is the answer the same request gets every time. A spent budget is the brake’s to wait out.
  • A request that got no answer at all. The only kind measured was a name that could not be looked up, six times, three of them on the way to a job log’s storage. Each took ten seconds to fail, which is the resolver asking every server twice before it gives up, so the lookup had already been asked again. No connection refused, reset or TLS handshake timeout appeared, and a request whose kept-alive connection closed before it was answered is sent again by Go’s HTTP client itself.
  • GitHub’s GraphQL timeout. A 502 or a 504 to a query that ran for ten seconds means a query too large to answer in them, which the same query would meet again. The pull request, commit and co-authored pull request walks ask again on the same cursor at half the page while it is larger than ten, so a page of fifty is asked again at twenty-five, twelve and six, and a query about several repositories at once is asked again with half of them, down to one. What still times out at the smallest, and every other walk or single query, is a failure the collector reports with the rows it read before it, never the end of what there is to read: a backfill does not record that repository as walked, and the next sweep reads it again. Up to 2.6.3 a walk with no smaller page to ask took the timeout for the end of its data and reported success, which on the author’s account stopped twelve commit walks, all at a page of fifty; asked by hand, twenty-five of the same commits answered in under five seconds.
level=INFO msg="rate budget" bucket=core remaining=4354 limit=5000

Worth an alert on the warning above rather than on this line. A budget that dips is normal; a family skipped every sweep is a configuration problem.

A backfill is the opposite intention and says so. When a bucket runs out it waits for the window to reset rather than giving up, because what it was asked for is the whole walk, and a family skipped for want of budget would be a family left to walk again. Stopped half way, it loses little: a checkpoint beside the state file keeps its place, repository by repository, and the same command carries on from there.

Written and maintained by
MIT licenceRelease history