Measurements
Ninety-five measurements. Each row says how a point is dated, because that is the thing that decides which questions it can answer.
How to read the tables
Section titled “How to read the tables”| Dating | Means |
|---|---|
| dated | The point carries the moment the thing happened, so the history is real and re-collecting rewrites the same rows |
| daily | A snapshot with no date of its own, stamped at the start of the UTC day so a day’s sweeps converge on one row rather than piling up |
| now | A current state, which only makes sense as “what is true at this moment” |
Every measurement carries owner, repo and full_name as tags unless it is
account-wide, in which case it carries user. The three travel together or not
at all, and repo is always the short name, on somebody else’s repository as
much as on your own: owner=golang, repo=go, full_name=golang/go. That is
what lets a filter written for one measurement answer in every other, and it is
what the dashboards’ repository variable is built from. The short name is not an
identity, since two owners can use the same one, so a query that needs identity
groups by full_name and one that needs a person groups by owner.
That has not always been true of all of them. Twelve
carried no owner and put the full name inside repo, and gh_billing_usage
had a short repo, no owner and an org. The three shapes did not overlap at
all, which is worse than overlapping badly: a filter built from one of them
matched nothing whatsoever in the others, and a union of two counted every
repository twice. They are one shape now. org is gone with them, and it loses
nothing: organizationName is not a property of the billing endpoint this tool
calls, so the tag could only ever hold (none), and on the organisation form of
that report it names the organisation whose report was asked for, which is the
owner. owner carries it.
Upgrading from a release before this shape is a breaking change to what is already stored, and it is larger than a seam. Read the four paragraphs below before upgrading a database you want to keep.
Your existing rows are not converted, and most of them are written again.
Rows already stored keep the old tags and nothing rewrites them, because in
InfluxDB the tag set is part of a point’s identity. What is easy to miss is
that most of the twelve are dated at the item’s own time and re-offered on
every sweep, and the write ledger keys on the tag set. Change the tag set and
every one of those points is a miss, so the first sweep after the upgrade
writes the whole retained history again, at the same timestamps, under the new
tags, beside the old rows. That is not a seam you can wait out. A sum over a
range that covers the rewritten history doubles immediately and stays
doubled, for as long as GitHub keeps serving those items. On the account this
was developed against that was a full year of gh_contribution_day_repo, which
a shipped panel sums, and several hundred rows each of gh_issue_comment,
gh_external_contribution and gh_discussion_comment. Only gh_event and
gh_notification behave like a seam, because they are windows GitHub forgets.
So there are two honest options, and waiting is not one of them. Recreate the database, which is what the author did; or delete the old-shape rows yourself, which in InfluxDB 3 means dropping the thirteen tables, since a tag cannot be renamed in place and a delete predicate cannot name a tag the new rows do not carry. There is no migration and there will not be one: renaming a tag means rewriting every affected row under a new identity, which is a restore rather than an update.
The PostgreSQL sink stops loading until its tables are changed. It emits
CREATE TABLE IF NOT EXISTS, which does nothing against a table an earlier
release created, so the INSERT that follows names owner and full_name
columns that do not exist and an ON CONFLICT key that does not exist either,
and psql rejects it for all thirteen measurements. Drop those tables and let
the next file recreate them, or add the two columns and rebuild each primary
key:
ALTER TABLE gh_event ADD COLUMN IF NOT EXISTS "owner" TEXT NOT NULL DEFAULT '', ADD COLUMN IF NOT EXISTS "full_name" TEXT NOT NULL DEFAULT '';ALTER TABLE gh_event DROP CONSTRAINT IF EXISTS gh_event_pkey;ALTER TABLE gh_event ADD PRIMARY KEY ("time", "action", "full_name", "owner", "ref_type", "repo", "type");The rebuilt key must be exactly the key this tool would declare for that
measurement: time followed by the tag columns of the table above, in name
order. Do not assemble it from the columns your table happens to have. An
upgraded gh_billing_usage still carries an org column from the old shape
that nothing writes any more and that is not part of the key; include it and
the ON CONFLICT the sink emits matches no unique constraint, so the load
fails again with a message that points nowhere. The check on your own work is
the CREATE TABLE the sink writes for that measurement into a fresh file: its
PRIMARY KEY list is the answer.
Adding the columns keeps the old rows but does not convert them: they hold the empty string where the new rows hold an owner, so they double-count exactly as they do in InfluxDB. Dropping is the clean option in both stores.
Queries you wrote yourself need updating, and only some of them fail loudly.
gh_billing_usage.org no longer exists, and in InfluxDB 3 naming a column no
row has written fails at planning, so that one tells you. A filter such as
gh_event WHERE repo = 'owner/name' keeps parsing and quietly returns nothing;
it becomes full_name = 'owner/name'.
Most measurements also carry a url field: the page on GitHub for the thing the
row is about, so a dashboard row that names an item can also open it. A url is
absolute or absent, since the dashboards link to the value itself. Every
measurement that has one is linked item by item from at least one table, with
two kinds of exception. The rows of some are only ever counted or added up, into
a curve, a bar, a stat or a table line per repository, reviewer, job or check,
where a per-item url has no row to sit on: gh_pull_request_review,
gh_workflow_job, gh_event, gh_issue_event, gh_artifact,
gh_contribution_day, gh_contribution_day_repo, gh_commit_check,
gh_issue_comment, gh_billing_usage, gh_dependabot_alert,
gh_code_scanning_alert, and the account-wide gh_account, gh_account_total,
gh_contributions_total and gh_sponsors_listing. And two carry a url no table
links: gh_release_published, which no panel reads, since the releases table
links each release from gh_release, and gh_upstream_repo, which Work
elsewhere joins for its Stars while each of its rows links the item rather than
the repository.
A tag GitHub leaves empty is written as (none), one spelling on every
measurement, and the same (none) goes into a few string fields that say
nothing on most rows, a pull request’s decision or an issue’s assigned_to:
InfluxDB 3 creates a column the first time a row carries it, and a query naming
a column no row has written fails outright, so those are written on every row.
The same convention writes (ghost) where a login belonged to an account that
has since been deleted.
The parentheses are the point of the spelling. GitHub answers unknown itself
in a Dependabot alert’s relationship, and none is a value in more than one
of its enums, so a fallback spelled either way could not be told from an
answer; (none) is never a value GitHub returns. One value does read none
without them and means it: gate on gh_commit is the state of the commit’s
gate, and none is the state of a commit no gate ever ran on, beside
SUCCESS and FAILURE.
A value that moves after the row’s own date is a field, never a tag. A tag
is part of a row’s identity, so a tag that changes after the fact opens a
second series at the same instant and the stale row stays beside the new one
for ever: measured after eleven hours of sweeps, one artifact in fifty had a
row with expired=false and another with expired=true at the same
timestamp, and a commit seen PENDING by one sweep and FAILURE by the next
was counted twice. Every such tag is a field now, under a new name because
InfluxDB 3 fixes a column as tag or field at first write: gh_artifact.expired
is live, gh_commit.checks is gate, state on both alert items is
alert_state and the code scanning reason is resolution, gh_issue’s
state_reason, assignee, milestone and parent are resolution,
assigned_to, milestone_title and parent_issue, gh_pull_request.draft
and review_decision are is_draft and decision,
gh_discussion.answered is has_answer, gh_pull_request_review.state is
review_state, gh_deployment.state is outcome and gh_notification.unread
is is_unread. gh_discussion_comment.is_answer, which 2.6.1 stopped writing,
needed no new name: it is the answers field every row of that measurement
already carried, 1 for the accepted answer and 0 for any other comment.
state on a pull request, an issue and an external
contribution stays a tag because its date moves with it: the open row is
stamped at the start of the day and the closed row when it closed.
The Prometheus exporter reads the demoted values back as labels; Graphite,
which keeps no strings, cannot group by them and its panels say so.
A store written before 2.6.1 and since holds gh_discussion_comment in two
shapes. A comment an earlier release read before it was accepted and again
after is two rows at one instant, one per value of is_answer. A comment an
earlier release read gains one more row without the tag when a 2.6.1 sweep reads
it again: a second for most, a third for one already held twice. The two
dashboard tables over the measurement read one row per comment in every store,
accepted when any of its rows says so, so an answer accepted before 2.6.1 and
taken back since still reads accepted. Dropping the measurement and running a
-backfill leaves one row per comment: the table in InfluxDB 3 and in
PostgreSQL, the github.discussion_comment paths in Graphite, the
ghchronicle-gh_discussion_comment index in Elasticsearch. The PostgreSQL sink
that connects keeps writing into a table an earlier release made, conflicting on
that table’s own key, and its new rows hold the empty string for is_answer.
The SQL file cannot know that key, so replayed into such a table its
gh_discussion_comment statements are refused while the rest of the file loads,
until the table is dropped. The Prometheus exporter keeps nothing across a
restart and counts each comment once. -migrate says which of the configured
stores hold the two shapes, under 2.6.1/gh_discussion_comment/is_answer: see
Migrations.
From a row to a query
Section titled “From a row to a query”Every row here is a table in the store, its tags are columns you filter and group by, and its fields are the numbers. Read against InfluxDB 3 in SQL mode, the three datings turn into three shapes of query.
A dated measurement is history, so it is read over a range:
SELECT time, "count" FROM gh_trafficWHERE kind = 'views' AND repo = 'telemetry' AND time > now() - INTERVAL '90 days'A daily snapshot is one row per day, so the newest row is the answer and the difference between two days is the movement:
SELECT time, downloads FROM gh_release_assetWHERE asset = 'ghchronicle_linux_amd64.tar.gz' ORDER BY time DESC LIMIT 30A dated item carries one row per thing that happened, which is what lets a question be asked of the items rather than of a count:
SELECT date_trunc('week', time) AS week, count(*) AS merged, avg(seconds_to_merge) / 3600 AS hoursFROM gh_pull_request WHERE state = 'MERGED' GROUP BY week ORDER BY weekThe same three shapes work in the other history stores; the dashboards carry one query set per store for every panel, which is the place to copy from.
Every measurement, alphabetically
Section titled “Every measurement, alphabetically”Ninety-five, each link landing on the table it is in.
gh_account · gh_account_total ·
gh_achievement · gh_achievement_progress ·
gh_actions_cache ·
gh_actions_cache_entry ·
gh_actions_policy · gh_artifact ·
gh_artifact_total · gh_billing_usage ·
gh_branch ·
gh_branch_protection ·
gh_code_scanning_alert ·
gh_code_scanning_alert_item ·
gh_code_scanning_analysis ·
gh_code_scanning_setup ·
gh_collector_family ·
gh_commit ·
gh_commit_check · gh_commit_punchcard ·
gh_commits_week · gh_contribution_day ·
gh_contribution_day_repo · gh_contribution_repo ·
gh_contribution_year · gh_contributions_total ·
gh_dependabot_alert · gh_dependabot_alert_item ·
gh_dependabot_ecosystem ·
gh_dependency ·
gh_dependency_change ·
gh_dependency_license ·
gh_deploy_key ·
gh_deployment ·
gh_discussion · gh_discussion_comment ·
gh_environment · gh_event ·
gh_external_contribution · gh_fork ·
gh_gist · gh_issue ·
gh_issue_comment · gh_issue_event ·
gh_job_log · gh_key · gh_label ·
gh_milestone · gh_notification ·
gh_package · gh_package_version ·
gh_pinned_item · gh_policy_file ·
gh_profile_flag · gh_pull_request ·
gh_pull_request_review ·
gh_rate_limit · gh_release ·
gh_release_asset ·
gh_release_published · gh_repo ·
gh_repo_activity ·
gh_repo_archived · gh_repo_community ·
gh_repo_created · gh_repo_language ·
gh_repo_policy ·
gh_repo_topic ·
gh_repo_total ·
gh_review_thread ·
gh_ruleset ·
gh_ruleset_rule ·
gh_ruleset_version · gh_secret ·
gh_security_feature · gh_security_setting ·
gh_social_account · gh_sponsors_listing ·
gh_sponsors_tier · gh_sponsorship ·
gh_star · gh_star_day ·
gh_star_given · gh_star_list ·
gh_traffic ·
gh_traffic_path · gh_traffic_referrer ·
gh_upstream_repo ·
gh_webhook ·
gh_webhook_delivery ·
gh_workflow ·
gh_workflow_job ·
gh_workflow_run ·
gh_workflow_run_total ·
gh_workflow_step
Audience
Section titled “Audience”| Measurement | Dated | Tags | Fields |
|---|---|---|---|
gh_ | dated, one point per day | kind (views, clones) | count, uniques, url |
gh_ | daily | referrer | count, uniques, url, referrer_url |
gh_ | daily | path | count, uniques, title, url |
GitHub serves fourteen days and the whole window is rewritten on every sweep, so a collector that was down for a day repairs itself on the next run. The referrers and paths are the top ten of that same window with no dates attached, which is why they are a snapshot rather than a series.
Stars and forks
Section titled “Stars and forks”| Measurement | Dated | Tags | Fields |
|---|---|---|---|
gh_ | dated, when the star was given | user | starred, url, user_url |
gh_ | dated, one point per day: the Pacific day, at 00:00 UTC | stars | |
gh_ | dated | user, language | stars, repo_stars, url |
gh_ | dated, when the fork was created | by | forks, stars, seconds_to_push, advanced, url |
gh_star_day is the stars a repository gained on each day, read from GitHub’s
star history
for every repository a sweep collects, whatever the token may see of its
stargazers. It names nobody, so it carries no tag of its own and no url:
gh_star is the row per star, with who gave it and the second, and it exists
only where the token may read the stargazer list, which since July 2026 means a
repository’s admins and collaborators. No panel reads both, so a repository
with both is counted once, and the panels that count stars from rows read
gh_star_day, so one whose list is hidden is counted all the same.
A day is GitHub’s own calendar day in America/Los_Angeles, and its row is
stamped at 00:00 UTC of that date, as a gh_contribution_day row is stamped at
00:00 UTC of its own. Measured against the stargazer lists of nineteen
repositories, 440 stars, the Pacific day matched every one, and the UTC day put
38 of one repository’s 127 days wrong. So a star given on a
European morning can sit a day before the date gh_star gives it. Each day is
anchored to the week the API returns, never to the clock, no day after the sweep
is written, and today may be written as 0 before it has begun in Pacific time,
to be corrected by the next sweep.
The count is the repository’s current stargazers, each on the day they starred,
so an unstar takes the star off the day it was given rather than the day it was
taken back. A sweep reads the newest thirty weeks, one page, and writes every
day of them, zeros included, so a day that drops from one to none is rewritten
the next time. Older pages are read on a repository’s first sight, once after
upgrading to 2.5.0, and in a backfill; the state file records a whole read as
history_read, and only after a walk that reached the end of the history, so a
walk cut short, by an error or by a later page answering 403 or 404, is read
whole again on the next sweep. Those pages write only the days that have stars:
zero rows back to each repository’s creation would multiply the files InfluxDB 3
opens, and a day older than thirty weeks that lost its only star therefore keeps
it, the way gh_star keeps every star it ever saw. The sum usually sits at
gh_repo.stars or a little below it, since that count also includes accounts
GitHub no longer lists: one short on three of those nineteen repositories. A
star given more than thirty weeks ago and taken back after the history was read
can make it sit above: only a backfill that reaches back to its day lowers that
day, and only while the day holds another star, so one that was its day’s only
star stays for good. Prometheus and an OTLP backend with raw: false never see
it: an unstar lowers a past day, which no counter can do, and
github_repo_stars already carries every repository’s current count.
GitHub Enterprise Server does not serve the star history: its REST API, checked
against 3.21 and 3.22, has the stargazer list and no stargazers/history.
There every repository answers it with a 404, which is read as nothing to
collect, so gh_star_day stays empty, and so do the panels that count stars
from it: Stars gained over time and Stars over time in InfluxDB and the SQL
stores, Stars gained over time in Graphite and Elasticsearch, and Graphite’s
Recent stars. Recent stars still lists names from gh_star in the other
stores, and Prometheus still counts from it. A history that answered 404 is not
recorded as read, so a server that starts serving it has each repository’s
history read whole on the next sweep.
gh_star_given is the outbound direction: what this account starred in other
people’s repositories. advanced on a fork separates a real derivative from a
bookmark, which most forks are, and seconds_to_push says how long after the
fork its last push came. It is negative when GitHub reports a push older than
the fork itself.
How long a fork has been idle is not stored, because the row is dated when the
fork was created: it is that date subtracted from now, less seconds_to_push,
and a query computes it. Stored, it was one day larger every day written on to
a row dated years earlier, which meant “when the sweep ran” rather than
anything about the fork, and every rewrite cost a file in the fork’s own
partition.
Repositories
Section titled “Repositories”| Measurement | Dated | Tags | Fields |
|---|---|---|---|
gh_ | now | language, visibility, license, archived, fork, default_branch | stars, forks, watchers, open_issues, size_kb, age_days, days_since_push, days_since_config_change, network, repo_id, is_template, has_pages, web_commit_signoff_required, allow_update_branch, pull_request_creation_policy, url |
gh_ | now | language | bytes |
gh_ | now | topic | present, url |
gh_ | now | health_percentage, url, has_readme, has_license, has_contributing, has_code_of_conduct, has_issue_template, has_pull_request_template | |
gh_ | dated, when the repository was archived | archived, age_days_at_archive, url | |
gh_ | now | tag, draft, prerelease | downloads, assets, age_days, url |
gh_ | daily | tag, asset | downloads, size_bytes, digest, content_type, uploader, age_days, url |
gh_ | dated, when published | tag | published, prerelease, url |
open_issues is GitHub’s field and GitHub counts pull requests in it. Use
gh_issue to count issues.
gh_release_published is when each release was published, to the second, and
it is the one to group by for a calendar of releases or to read “the latest
stable release” from, the newest row with prerelease false. gh_release
stays stamped at the sweep, because its downloads move, and its age_days is
floored to whole days counted back from the sweep: the date it reconstructs is
a day late for any release published later in the day than the sweep ran.
published is the integer 1 on every row, so counting releases is a sum.
prerelease is a boolean field here, where it is a tag on gh_release: a
pre-release is promoted by unticking the box on the published release, and as a tag a promotion that
keeps the publication’s date would write a second row at the same instant,
which the sum would count twice. A release taken back to a draft and published
again writes a second row if GitHub gives it a new published_at, so the exact
count is the distinct tag values of each repository. A draft has no
publication, so it writes no row here, and its gh_release row carries no
age_days: GitHub sends a draft with published_at null, and the age measured
from nothing held 106751 days on every draft row. The first sweep after an
upgrade dates the page of releases it reads; -backfill dates the rest.
The url on gh_release_asset is the asset’s download address, not a page:
following it fetches the binary. The assets are inventory, anchored to the
start of the UTC day like the cache entries: stamped at the sweep, every asset
was a fresh row every hour, which was 15 per cent of the whole database after
eleven hours. One row per asset per day still answers “downloads per day”, and
the newest row is still the value.
gh_repo_archived is the one row about a repository that carries a date rather
than a state: dated at archivedAt, a clear-out is visible as the batch it
was, and a live repository produces no row at all, so counting the rows is
counting the archive. It does not need include_archived. The listing a sweep
already pays for says which repositories are archived, and the totals family
asks the date of all of them, beside their lifetime row, in one GraphQL query
per twenty five repositories, on every totals sweep: a point per twenty five
at that cadence, and the rows it rewrites are the same rows, which is what an
exporter that keeps only what is rewritten needs; the listing cannot supply the
date itself, since REST
carries no archived_at and its updated_at was measured two seconds to
eight minutes after the archive. An archived fork under the default fork rule
is the one kind with no row.
Development
Section titled “Development”| Measurement | Dated | Tags | Fields |
|---|---|---|---|
gh_ | dated when closed, daily while open | number, state, author | is_draft, decision, title, labels, label_names, author_association, additions, deletions, churn, changed_files, commits, comments, total_comments, reviews, review_requests, review_threads, base_ref, head_ref, merged_by, merge_commit, mergeable, merge_state, stack, stack_size, stack_position, seconds_to_first_review, seconds_to_first_human_review, seconds_to_merge, seconds_open, url |
gh_ | dated, when submitted | number, author, reviewer, bot, self | review_state, reviews, seconds_to_review, url |
gh_ | dated when closed, daily while open | number, state, author | resolution, assigned_to, milestone_title, parent_issue, comments, reactions, labels, label_names, sub_issues_total, sub_issues_completed, pull_request, seconds_to_close, seconds_open, url |
gh_ | dated, when committed | sha, author, branch, signature | gate, additions, deletions, churn, changed_files, commits, signed, oid, headline, url, pull_request, checks_total, checks_failed |
gh_ | dated, when the check finished | sha, app, check, conclusion | checks, failed, url |
gh_ | dated, when it happened | event, actor, kind, bot, label, milestone, requested_reviewer, review_requester, mentioned | events, number, title, url, commit_id, rename_from, rename_to |
gh_ | dated, when the thread’s first comment was written | thread, number, author, bot | path, comments, resolved, outdated, subject_type, resolved_by |
gh_ | dated, when created | category, answerable, author, number | has_answer, comments, replies, reactions, upvotes, closed, state_reason, seconds_to_answer, seconds_to_close, title, url |
gh_ | daily | label | issues, pull_requests, used, url |
gh_ | daily | milestone, state | progress, issues, pull_requests, days_to_due, seconds_to_close, url |
gh_ | dated | user, number, kind, state | contributions, merged, private, additions, deletions, changed_files, title, comments, seconds_to_merge, seconds_open, url |
gh_ | now | stars, forks, private, language, url |
gh_commit is what replaces stats/code_frequency, which returns 202 with an
empty body forever on a personal account. signature is unsigned when there
is no signature at all, which is a different fact from one that failed to
verify.
gate is the state of the whole gate on that commit, which is not the same
claim as a workflow run having failed: a run says one job failed, the rollup
says the commit came out red. It is a field because the verdict lands after
the commit’s own date. gh_commit_check holds only the checks that are not
GitHub Actions, since everything Actions runs is already collected in far more
detail.
seconds_to_first_review counts any review, and on an account with review
bots that is the bot: measured over 140 pull requests, its median was five
seconds, because 132 were first reviewed by sourcery-ai or coderabbitai within
a minute of opening. seconds_to_first_human_review is the wait for somebody
else, over the first twenty reviews the query fetches: not a bot, and not the
author. The author’s reply in a review thread arrives as a review of state
COMMENTED under their own name, and on this account it was the earliest
non-bot review on every one of the 91 pull requests that had one, so a wait
that counted it measured how fast the owner answers sourcery-ai. A pull
request whose fetched reviews are all bots and the author’s own replies
carries no such field rather than a wrong one. A bot is a GitHub App
(__typename Bot) or a login ending in [bot]; a deleted account is not one.
gh_pull_request_review.bot draws the same line per review and self marks
the author’s own, so a reviewers table can leave both out or show them apart;
bot is what gh_review_thread.bot already does for threads.
title, label_names and author_association are fields because a title is
unbounded and nine labels on one pull request are one row, not nine series.
labels is the count and label_names the names joined by commas, absent when
there are none, on pull requests and issues alike. author_association is
OWNER, MEMBER, COLLABORATOR, CONTRIBUTOR, FIRST_TIME_CONTRIBUTOR or
NONE: what separates an outside contribution from the owner’s own work.
mergeable and merge_state are written only while a pull request is open.
A merged one keeps answering CONFLICTING long after it was merged, which is
stale rather than false but reads as a repository full of conflicts. A sweep
reads only what was updated in the last two cadences, so an open pull request
nobody touches has its seconds_open, mergeable and merge_state rewritten
once a day by the day’s read of every open item rather than every hour, however
long ago it last moved; anything that moves
updatedAt, a review, a comment, a push, a close, is rewritten by the sweep
that follows it.
stack, stack_size and stack_position describe a stack of dependent pull
requests and are absent on a pull request that is in none. stack is the
stack’s own number, not a member’s, so the honest way to count deliveries is
distinct stack values plus the rows carrying no stack fields at all.
review_requests and review_threads are the two counts a reviewing flow is
measured with, and total_comments counts every comment on the pull request
rather than the ones in comments, which are the ones on the conversation.
sub_issues_total and sub_issues_completed are how far an epic has got, from
the checklist GitHub keeps on the parent. parent_issue is the other end of
the same relation, on the child, and is 0 on an issue with no parent.
pull_request is the pull request that closed the issue, 0 when none did.
gh_issue_event is the transition rather than the state. gh_issue and
gh_pull_request say what something ended up as; this says when it was
labelled, closed, reopened, renamed or had a review requested. A reopening
exists nowhere else. mentioned is the person a mentioned or subscribed
event happened to, which GitHub files as the actor without saying who wrote
the comment; here that person has a tag of their own and actor reads
(none) on those two types, so the account named in “@coderabbitai” no longer
shares a column with the app that reviews. An app is spelled the way REST
spells it, with the [bot] suffix, on every measurement.
gh_discussion counts comments and replies apart: comments answer the
discussion, replies answer those, and GitHub’s own number on the page is the
two added together.
gh_external_contribution is the work the account did in repositories it does
not own, from five searches, one per kind and state: pull requests
merged, open and closed unmerged, and issues open and closed. An item
still open is rewritten at the start of every day it stays open, so a sweep
reads both open states whole. A closed item is written once, stamped when it
closed, so the three closed states are read by when each item last moved: a
sweep reads each back to a cadence before the sweep before, which holds
whatever merged or closed since then however long ago it was opened and however
many other items moved in between, and is one page of a hundred unless more
than a hundred did. A backfill reads on until the pages run out or reach
backfill.since. GitHub serves a thousand
results of any search and no more, so an account past a thousand in one state
has the thousand that moved most recently, and the log says so at warning.
gh_account_total.pulls_merged_elsewhere and issues_elsewhere are GitHub’s
own counts and are right whatever the cap. An item that was open and later
closed keeps its open rows beside the closed one, since state is a tag.
private says whether the repository the item went to is private. The
searches run with the account’s token, so an organisation’s private
repositories come back beside the public ones, and this field is what lets a
page shown to others leave them out; a field and not a tag, so it added no
second identity to rows already stored. A pull request carries additions,
deletions and changed_files, the names gh_pull_request uses, and an issue
carries none of the three. All of it comes on the nodes the searches already
return, and measured on 2026-09-26 the query costs one point a page with it as
without it.
gh_upstream_repo is the repository side of the same searches: one row per
repository they reached in the pass, stamped at the sweep, with its stars and
forks as integers, private as a boolean, its primary language, absent when
GitHub detects none, and its page. The star count is not on
gh_external_contribution because that row is dated when the item closed, and a
count that moves nearly every day would rewrite a row of the past each time. In
InfluxDB and PostgreSQL the “Work elsewhere” table joins the newest of these
rows inside the range onto each item as Stars, and in Prometheus its Stars is
the count the exporter holds; Graphite and Elasticsearch cannot join one
measurement onto another, and their tables have no Stars column. A repository
gets a row whenever a search reads an item in it: each open state is read whole
on every sweep, and each closed state from its most recent page, so a repository
whose only items are closed and further back than that page keeps the row of the
last sweep that read one of them, or of the last backfill.
Continuous integration
Section titled “Continuous integration”| Measurement | Dated | Tags | Fields |
|---|---|---|---|
gh_ | dated, when it finished | workflow (the file path), event, conclusion, actor | duration_seconds, queued_seconds, attempt, success, run_id, run_number, pull_request, pull_requests, headline, head_repo, head_sha, head_branch, name, title, workflow_id, initial_actor, url |
gh_ | now | runs | |
gh_ | dated, when it finished | workflow, job_name, attempt, conclusion, runner_group, labels | duration_seconds, queued_seconds, steps, success, runner, run_id, head_sha, head_branch, url |
gh_ | dated, when it finished | workflow, job_name, attempt, step, conclusion | duration_seconds, step_number |
gh_ | now | workflow, path, state | active, age_days, days_since_change, url |
gh_ | dated, when created | artifact | live, size_bytes, retention_days, digest, run_id, head_sha, head_branch, url |
gh_ | now | live_bytes, live_count, count, walked | |
gh_ | now | size_bytes, count | |
gh_ | daily | cache, ref | size_bytes, caches, key, days_since_use, age_days |
gh_ | dated | activity, actor | events, id, ref_name |
branch was a tag on gh_workflow_run and on gh_artifact, and is now the
field head_branch on both, and on gh_workflow_job as well. A branch is an
identity, but not a reusable one: every pull request and every Dependabot bump
mints a name that never comes back, so the tag grows without bound, and the
bounded question a reader actually asks, whether this was a push or a pull
request, is already the event tag. InfluxDB 3 also fixes a column as a tag or
a field the first time it sees it and refuses every later write that disagrees,
so keeping the name would have meant dropping both tables to publish a value
the API itself calls head_branch.
attempt is a tag on the job and on the step because the job listing is now
asked for every attempt rather than the last one. Without it the two tries of a
re-run are one series, told apart only by the second they finished in, and a
flaky test cannot be distinguished from a broken one.
queued_seconds on a run is written on first attempts only. GitHub keeps the
run’s created_at across re-runs, so on a second attempt the gap to
run_started_at is the time a person took to press the button, not a runner
queue: 0.4 s on average on first attempts against 1,747 s on second ones,
measured. The queue of a retry exists only per job.
run_number is the “#1483” GitHub shows and people quote; run_id is what the
API keys by. pull_request is the number of the first pull request GitHub
linked to the run and pull_requests how many it linked, both absent when it
linked none. headline is the first line of the commit that ran, and
head_repo is written only when the run came from another repository, which
is what a fork’s pull request looks like.
gh_workflow_run_total is the run listing’s own total_count, which is the
whole history rather than the few hundred runs the walk sees. It is current
state, so it is stamped now, and it is the only place “how many runs ever” can
be answered without scanning the table.
An ordinary sweep asks for the run list in pages of thirty rather than a hundred. The page is thirteen kilobytes a run, of which the collector keeps six hundred bytes, and at a hundred runs it was a megabyte and a half per active repository every quarter of an hour, forty six percent of everything a day downloads. The stores lose nothing: the walk still pages on while a page is full of runs newer than the window, up to seven pages, which is the two hundred and ten runs two pages of a hundred reached, and the first sweep after start and a backfill still ask for a hundred. What changes is the Prometheus exporter, which holds only what the last sweep collected and shows the newest thirty runs between builds rather than the newest hundred.
The jobs of a run are listed once. The jobs of a completed attempt never change, and listing them again every sweep was a request per run in the window, nearly all of them 304s that cost no quota but a third of a second of waiting each, ninety six times a day. The collector remembers each attempt whose jobs it wrote, and keeps that memory with the ETag cache in the file beside the state file, so a restart asks only for new runs and new attempts, as the sweep before it would have; a backfill lists every run regardless. A re-run keeps the run’s id and is a new attempt, so it is listed again. The cap of twenty bounds what a sweep pays, not which runs get jobs: a window with more runs than that fills in twenty a sweep. A run is remembered only once every store has taken the pass that listed its jobs: a pass a store refused forgets its runs, and the next one lists and writes them again. The file keeps a run while the listing keeps returning it, and forgets it on the same horizon as an answer nobody asks for any more. A start that finds no write ledger to say what the stores hold, or a store added since, recalls none of them and lists their jobs again.
A run’s jobs outlive their steps. Measured on 24 September 2026, GitHub listed
every job of a run 278 days old, with its times and its runner, and gave every
run created before 12 April, about five and a half months back, an empty step
list. A job that finished as success, failure or timed out ran at least one
step, so when its list comes back empty it is written with no steps field
rather than with a 0 that says it had none, and with no gh_workflow_step rows:
a zero there would pull every mean of steps toward nothing as far back as a
backfill reached. Every other job keeps steps as the length of its list,
because its zero can be true: a skipped job runs no step (measured), and a
cancelled one can stop before its first.
retention_days is the retention an artifact actually got, which is rarely the
configured default: eighty-eight of a hundred artifacts measured lived one day
against a setting of ninety.
gh_repo_activity is one row per activity type, actor and second. The branch
is the field ref_name, one name per pull request and per Dependabot bump, and
without it in the key the branches one push moved in the same second would be
one row in every store, the last one written standing for all of them: eight
force pushes in one second, measured. The entries that share a key are folded
into one point, events counting them, ref_name naming every branch joined
by commas, id the newest entry’s, so a sum of events is the number of
activities everywhere.
Queue time only exists at the job level. The run-level figure folds the wait into the duration, and the job-level one includes waiting for a dependency, so a job that waits nineteen minutes for another job to finish is not evidence of a runner shortage.
runner is a field, not a tag: a hosted runner is named uniquely per run, so
as a tag it would create a series for every job ever executed. The workflow job
tag is job_name rather than job, because job collides with the labels
Prometheus adds at scrape time.
gh_artifact_total carries three counts because its size is on none of the
obvious ones. count is GitHub’s own total for the repository and it counts
the artifacts GitHub has already expired: measured on jmrplens/jmrp.io on
2026-09-17, page 40 of the listing was expired to the last row against a
declared 29,405. walked is how far the five-page cap let the walk go.
live_bytes is the size of the artifacts GitHub still holds among the ones
walked, and live_count is how many those are, which is the count the size is
over. When walked is below count the live figures are a floor rather than a
total, which on that repository was short by a factor of fifty six. A walk that
a failed page cut short still writes the row, from the pages before it, and its
walked stops short of count in the same way; only a first page that failed
writes none, because then there is no count to be short of.
gh_actions_cache says a repository holds twelve gigabytes;
gh_actions_cache_entry says which key holds them and which has not been
touched for a week, which is what decides what GitHub evicts at the ten
gigabyte ceiling. The cache tag is the key without its content hash, because
the whole key is a series per build.
A row of gh_actions_cache_entry is one cache on one ref for the day, with
its entries summed into it: caches is how many there are, size_bytes their
total, days_since_use and key those of the one used most recently, and
age_days that of the oldest. Up to 2.5.2 it was a row per entry, and every
entry of one cache on one ref had the same tags and the same day, so the store
kept whichever was written last: on 2026-09-26 the fifteen CodeQL caches on
main of jmrplens/jmrplens, 57.9 MB between them, were stored as one of 3.8 MB.
No API lists a past day’s caches, so those days stay as they are, and
-migrate only notes them, under 2.6.0/gh_actions_cache_entry/sum.
The listing is read a hundred entries a page, up to ten pages, where it used to
stop at the first: on 2026-09-27 two repositories of the account held 118 and
232 entries. It is read newest created first, an order a cache hit does not
change. A pass that reads fewer entries than the listing said it held, because
one was deleted between two of its pages, writes no gh_actions_cache_entry
row, and nor does one whose later page failed: summed, a part of a cache would
be written over the whole of it that the day’s earlier passes stored, and the
day’s next pass writes the rows. The gh_actions_cache total, read before the
listing, is written either way.
Security
Section titled “Security”| Measurement | Dated | Tags | Fields |
|---|---|---|---|
gh_ | now | severity, ecosystem | open, url |
gh_ | dated, when raised | number, severity, ecosystem, package, ghsa, scope, relationship, manifest | alert_state, alerts, cvss, cvss_v4, epss, epss_percentile, cve, cwe, summary, vulnerable_range, first_patched, dismissed_reason, dismissed_by, dismissed_comment, seconds_to_detect, seconds_to_resolve, url |
gh_ | now | severity, tool | open, url |
gh_ | dated, when raised | number, severity, tool, rule, path, category, ref | alert_state, resolution, alerts, commit, line, cwe, seconds_to_resolve, url |
gh_ | dated, when the scan ran | tool, version, ref, category | analyses, results, rules, commit |
gh_ | now | feature | enabled, open_alerts, alerts, url |
gh_ | now | setting, status | enabled |
gh_ | daily | state, query_suite, schedule | setups, languages, days_since_change |
gh_ | daily | kind (actions, dependabot), secret | secrets, age_days, days_since_rotation |
gh_ | daily | permissions | policies, can_approve_pr |
The tags on the two item measurements cost nothing: both listings were already
carrying them, and the series was always keyed by number, so they group rows
that exist one per alert rather than multiplying them. All of them are always
written, falling back to (none) when GitHub omits one. A tag written only
sometimes gives the measurement two Graphite path depths, and the panels index
their nodes from one fixed table. alert_state on both items and resolution
on the code scanning one are fields, since an alert is dated when it was
raised and both move when it closes; resolution reads open until then, so
the column exists before any alert has closed. A build before 1.0.0 tagged both
items with state, and the code scanning one with reason, so a store it
wrote holds those rows beside today’s: -migrate finds them, under
1.0.0/gh_dependabot_alert_item/state and
1.0.0/gh_code_scanning_alert_item/state, in the stores that can be asked.
The fields are the opposite: cve, cwe, first_patched, epss,
epss_percentile and seconds_to_detect are written only when the advisory
carries them, because a missing EPSS score is not a score of zero. cvss and
cvss_v4 need the same guard for a different reason: GitHub always sends both
keys and fills the one it lacks with 0.0, and an advisory published with a v4
vector only was 78 of the 225 alerts of one repository, enough for “worst CVSS”
over a severity group of them to read zero. Neither score is written unless it
is above zero; a panel that wants one number per alert reads
COALESCE(cvss_v4, cvss).
summary is the advisory’s title, vulnerable_range the range it covers,
which next to first_patched is the action to take, and dismissed_reason,
dismissed_by and dismissed_comment say why a person closed an alert without
fixing it; they exist only on an alert in state dismissed.
seconds_to_detect is the gap between the advisory being published and the
alert being raised here. It is negative when the alert came first, which happens
when an advisory is written up after the fact.
The two cwe fields cannot be joined. Dependabot writes CWE-400 and code
scanning writes cwe-079, both GitHub’s own spelling, and neither is normalised
here. line is the start line of the alert’s most recent instance, and its zero
is GitHub’s own value for an alert about a whole file, not a missing reading.
A Dependabot alert closes three ways, not two: auto_dismissed_at is how GitHub
closes a development-dependency alert on its own, leaving the other two null.
An alert closed that way used to be counted as still open forever.
An alert still open carries no field for how long it has been open, and that is
deliberate. The row is dated when the alert was raised, so the answer is now()
less the row’s own timestamp and a panel computes it when it is asked. Written
by the collector instead, as seconds_open, it was true only at the instant of
the sweep that wrote it and it moved on every sweep: a reader querying last
week got whatever the last sweep decided, and every rewrite filed another
parquet file in the partition of the alert’s original date. Measured on
2026-09-17, the two alert families were writing about 234 files a day between
them for 1,700 rows, whether or not GitHub had anything new to say.
gh_security_feature exists so that no data and no alerts are distinguishable.
Without it, a repository with Dependabot switched off looks exactly like one
with nothing to fix. enabled is read from the first full page of the listing,
which answers 403 when the feature is off and, for code scanning, 404 when
nothing has been analysed yet.
open_alerts, and open on gh_dependabot_alert and gh_code_scanning_alert,
count every alert of the repository that is still open, however old. A sweep
reads the newest page of each list, a hundred alerts in every state, and when
that page comes back full the list is read again with state=open, to its end,
and the counts are taken from that. An alert still open behind a hundred newer
ones that were fixed is counted; up to 2.5.2 it was not.
alerts is not the same number for both features. For code scanning it is the
repository’s total: when the page comes back full it is the last page GitHub
declares for a page of one alert, which on one repository read 1,393 where the
page had said 100. Dependabot’s list pages by cursor and declares no last page,
so there alerts is what the sweep read: 100 on a sweep means a hundred or
more, and a backfill that walks the whole list writes the total. The two item
measurements are the alerts the walk read either way, and complete after a
backfill.
Account
Section titled “Account”| Measurement | Dated | Tags | Fields |
|---|---|---|---|
gh_ | now | followers, following, following_users, public_repos, gists, packages, projects, starred, watching, sponsors, sponsoring, account_age_days, pronouns, url | |
gh_ | now | calendar_total, commits, pull_requests, reviews, issues, repositories, restricted, repos_with_commits, repos_with_issues, repos_with_pulls, repos_with_reviews, url | |
gh_ | dated, one point per calendar day | contributions, level, url | |
gh_ | dated, end of the year; the year in progress daily | year | contributions, commits, issues, pull_requests, reviews, repositories, restricted, repos_with_commits, repos_with_issues, repos_with_pulls, repos_with_reviews, partial |
gh_ | now | kind (commits, issues, pulls, reviews) | contributions, commits, days, commits_dated, url |
gh_ | dated, the day the commits belong to | private, own | commits, url |
gh_ | dated, the Sunday of its week | commits, owner_commits | |
gh_ | now | weekday, hour | commits |
gh_ | now | package, type, visibility | versions, tagged_versions, age_days, days_since_update, url |
gh_ | dated, when published | package, type, visibility, tag | digest, published, url |
gh_ | now | gist, public | files, comments, size_bytes, description, url, age_days, days_since_update |
gh_ | daily | achievement | name, tier_number, tier_name, present, image, url |
gh_ | daily | achievement | name, count, tier_number, next_threshold, percent, page_tier, agrees, image, url |
gh_ | now; the orcid row daily | provider | url, present |
gh_ | now | pinned, position, kind, stars, days_since_push, url | |
gh_ | now | flag | enabled, message, age_days, url |
gh_ | dated, when the sponsorship was made | direction (sponsor, maintainer), sponsorable | sponsorship, active, one_time, privacy, tier, amount_cents, url |
gh_ | now | has_listing, listing_name, listing_public, listing_age_days, tiers, monthly_income_cents, next_payout_cents, next_payout_date, sponsor_spend_cents, lifetime_received_cents, sponsorships_received, goal_kind, goal_title, goal_target, goal_percent, url | |
gh_ | daily | tier | tiers, price_cents, one_time, retired, age_days, url |
gh_ | daily | list | lists, items, private, name, age_days, days_since_add, url |
gh_ | now | pulls_opened, pulls_merged, pulls_open_now, pulls_merged_elsewhere, pulls_reviewed, issues_opened, issues_closed, issues_elsewhere, commented_elsewhere, commits, repositories, url | |
gh_ | dated, when created | fork | created, private, url |
gh_ | daily | kind (ssh, gpg), key | keys, age_days, days_since_use, never_used, days_to_expiry, verified, revoked, can_sign, emails, url |
gh_ | dated | own, is_reply, author, comment, number | comments, answers, upvotes, title, reply_to, private, discussion_answered, discussion_answerable, discussion_closed, answered_by, answer_chosen_by, state_reason, category, seconds_to_answer, seconds_to_close, url |
gh_ | dated | own, number | comments, private, url |
following is the profile’s own number, and it counts organisations as well as
people. GraphQL’s following connection counts only users, which on this
account read four where the profile read nine, so more than half of it was
invisible. Both are kept: following is what the profile page shows,
following_users is the connection’s people-only count. When the profile
request fails the connection’s count fills both, which is the tell that the
profile was not reached.
versions on a package is the count the package element declares, with the
walked count as the fallback. tagged_versions counts named releases only. It
used to count every container tag, and about half of those are the OCI referrers
fallback tag GitHub publishes for each attestation and signature manifest:
sha256- followed by the digest the row already carries. Nobody pulls one,
there is a fresh one on every build, and excluding them halved the count on two
packages, from a hundred and twenty six to fifty seven on one.
gh_package_version no longer writes a row for one either. A package GitHub
attaches to no repository writes (none) in all three of the tags that name
one, rather than leaving them out and landing in a series with no repository
column for a query to name.
gh_pinned_item and gh_profile_flag are the profile page itself as data. A
pin has no date of its own, so both are stamped now. A pinned gist is not in a
repository at all: GitHub names it by its hash, so repo is that hash and
owner is the account, which is the only owner a pin can have. The kind field
says which of the two a row is. position is a field and
not a tag: a repository that moves from slot two to slot three is the same pin,
and as a tag every rearrangement would fork the series. flag is a closed list
of eight: hireable, developer_program, campus_expert, github_star,
bounty_hunter, employee, sponsors_listing, which is whether the account
has a Sponsors profile at all, and limited_availability, which is the
availability status carried as an eighth flag rather than a measurement of its
own, with the message it displays and the age_days since that status was
set. age_days is written on that row alone and is the age of the status
message, not of the flag; GitHub says when none of the other flags was
granted, so no row carries a date for them.
gh_sponsorship is the only dated record of the money. gh_account.sponsors
and gh_account.sponsoring are counts as of now that say neither when nor to
whom, and gh_sponsors_listing.lifetime_received_cents is a total with no dates
in it. Both connections are read with activeOnly off, which is what recovers a
lapsed one. sponsorable is the other party, and it is the literal word
private when the sponsorship hides it, in which case no URL is written rather
than one being guessed at.
gh_sponsors_tier is standing inventory, the way an SSH key is. Dating a tier
at its creation would put all eight of them in 2021, outside every dashboard
range, where they would read as “no tiers”; anchored to the start of the UTC day
they converge on one row per tier per day, and age_days keeps the creation
date recoverable.
gh_star_list is the same shape for the same reason: the lists the account
files its stars into, one row per list with how many it holds, anchored to the
start of the UTC day. A list carries two dates, when it was made and when a
star last went into it, and both survive as age_days and days_since_add
rather than dating the row, which would put a list made in 2024 outside every
dashboard range. The tag is the slug, which the list’s page is addressed by;
the display name is a field. Whether the slug outlives a rename is not
verified, since checking it means renaming a list. It rides in the
account query that was already being paid for: measured on 2026-09-11, eleven
lists with their item counts added nothing to a cost of one.
gh_contribution_day is the only place the green squares exist as data. With
every.families.history set, it reaches back to the year the account was
created, at one GraphQL point per year. Its level is the square’s shade, GitHub’s own
quartile of the year as the 0 to 4 the profile draws, which is not a function
of the count: on one account 83 contributions on one day and 52 on another
were both the second quartile. The quartile is of the window asked for, the
trailing twelve months for the sweep and the calendar year for history, so
a day both write can change shade between the two, as it does on the profile
when a year is picked. Which is why the dashboard’s grid shades a day by its
own count instead: measured on 2026-09-14, the profile page shades by the
fifths of the busiest day of the window, a rule that reproduced all 366 of its
squares from GitHub’s own counts, where level disagreed with the page on 33
of those days.
gh_achievement is the one measurement that does not come from the API.
GitHub lists achievements nowhere in REST or GraphQL, so the family reads the
public profile page, https://github.com/<login>?tab=achievements, every hour
as an anonymous visitor: no token travels to it and it is charged to no
budget. One row per badge, stamped at the start of the UTC day: name is the
badge, tier_number is the number on its label (1 with no label, 2 to 4 for
x2 to x4) and tier_name the colour that goes with it (default, bronze,
silver, gold); the number is not called tier because Elasticsearch maps a
field name once across every measurement’s index and tier is already a
string on sponsorships. The parser is strict about the markup it accepts and holds each part
of a card to the others, so when GitHub changes the page the family logs one
warning and writes nothing until the parser is updated; the rows it wrote
before stay, and a panel reading the newest row per badge goes stale rather
than wrong.
image is the badge image the page shows at that tier, for a panel to draw.
The site the page is read from is derived from github.base_url. A base_url
that is a proxy in front of the API has to name the site with
github.web_url, because the API host
answers the page’s url with a JSON 404, and the family refuses that rather than
reading it as no badges.
An account with no badge at all has no achievements tab: its url answers 404 while the profile answers 200, which is no rows and no warning, the same reading every family gives a 404. A 200 with no card in it is refused as a changed page rather than read as none.
gh_achievement_progress is written by the same family beside the badges: one
row per badge that has tiers (Pull Shark, Galaxy Brain, Starstruck, Pair
Extraordinaire), whether or not the page shows it yet, saying how far the
account is from the next tier. GitHub publishes neither the rule a badge is
earned by nor the count it has reached, so the count is recomputed from the API
and the thresholds are the community’s, the Tiers table of
Schweinepriester/github-profile-achievements
as read on 2026-09-12: Pull Shark counts merged pull requests anywhere and its
tiers begin at 2, 16, 128 and 1024; Galaxy Brain counts the discussions whose
accepted answer the account wrote, at 2, 8, 16 and 32; Starstruck takes the
stars on the most starred repository of the account’s own, forks left out, at
16, 128, 512 and 4096; Pair Extraordinaire counts merged pull requests in
public repositories with a co-authored commit, one per pull request, at 1, 10,
24 and 48, cross-checked the same day against a hand count (the two counts
differed by two, the difference falling in a range where the hand count ran
past the thousand results a search pages, and the page showed the same tier
either way; the co-authored pull requests of a private repository moved
nothing). The single-tier badges and the two GitHub is still testing have no
row: there is no next tier to measure against. tier_number is the tier the
count implies (0 below the first threshold), page_tier the tier the profile
page shows (0 when the badge is not on it) and agrees whether the two are the
same; when they are, next_threshold is where the next tier begins (0 at the
top) and percent the count against it (100 at the top). A row that disagrees
is a rule the page contradicts, or a page GitHub has not recomputed yet, said
once per process in the log, and it carries neither field, so no bar is drawn
from a rule the page contradicts. Three of the counts are one GraphQL query;
the fourth is a walk over the account’s merged pull requests in public
repositories with their commit messages, split by merge date where a range
holds more than the thousand results a search will page, one point a page. Over
a whole account’s life that is a few dozen points and 24 MB (35 queries and
94 seconds over 2,315 pull requests, measured on 2026-09-27), so it is done
once and then kept: the state file holds the count with the last UTC day it
covers, and each pass walks only the pull requests merged since, one page a
pass: 468 to 513 KB on that account for each hourly pass the production proxy
logged on 2026-09-27, a size that grows through the day with what is merged.
The day a pass runs on is still being merged into, so its pull requests are in
that day’s row and walked again by the next pass. The whole history is walked
again once a week, because the count can
go down (a repository made private or deleted takes its pull requests out of
is:public), and whenever the rule the count was kept by has changed.
A pass the API will not answer writes the badges and no progress rows, and says
which read failed at warning, with the error: achievement counts unavailable, no progress rows this pass or co-authored pull requests unavailable, no progress rows this pass. It costs that pass and nothing more. The rows are the
day’s, so the pass an hour later writes them, and a walk cut short leaves the
kept count as it was, for the next pass to walk its days again. A count the walk
could not settle whole is written as a floor, and said once per process while
its numbers stay the same:
level=WARN msg="co-authored pull request count is a floor" capped=false truncated=3capped is a day that alone held more than the thousand results a search
pages, and truncated how many pull requests had more commits than a page of a
hundred, no trailer in the ones read, and the rest of their commits unread
because the query that asked for them failed, which is the most the count can
be short by. Both are kept in the state file with the count, so a pass that
adds to a floor still says it is one.
A pull request with more than a hundred commits and no trailer in the first
hundred is read on, a hundred commits at a time and ten pull requests to a
query, one point each, until a trailer turns up or the commits run out. Up to
2.6.0 it was left there and counted as a floor: on the account measured, three
pull requests of 143, 144 and 248 commits made every pass warn truncated=3,
and reading the rest of their commits on 2026-09-28, two queries and 121 KB,
found no trailer in any, so the count of 33 had been exact all along. The weekly
whole walk pays those queries, and so does each pass that walks the day such a
pull request merged.
gh_social_account carries one more row than the social accounts listing:
the homepage, under the provider website, from the blog of the profile.
The ORCID iD the profile page shows is in no endpoint, so the achievements
family, which already reads that page every hour, writes it from the page’s
vcard under the provider orcid, stamped at the start of the UTC day like
the badges. The links the API does list (Mastodon, LinkedIn, Bluesky) are
written by the profile family from the API and skipped on the page, so no
account is written twice. The read is as strict as the badges’: a page
without the vcard is a changed page and one warning, never “no accounts”.
pronouns is the profile’s pronouns line, he/him, as a field on the
headline row, absent when the profile shows none. A field and not a tag: it
is free text the owner can edit, and as a tag every edit would fork the
account’s one series.
gh_issue_comment and gh_discussion_comment are read from the newest end
of their connections, which list oldest first: a sweep’s one page is the
hundred newest comments. The account’s accepted answers are then read on
their own, the newest five hundred on a sweep and all of them in a backfill,
so a comment accepted as the answer after it left that hundred is still
written with answers 1 on the next sweep. A comment both reads return
is written once. Both carry private, whether the repository the comment was
left in is private, for the reason gh_external_contribution does: own does
not say it, since the account’s own repositories can be private or public and
an organisation’s are not the account’s own at all.
gh_contribution_year has one row per past year, dated the thirty-first of
December, and one for the year in progress, asked for on every run from the
first of January to now. That row is a snapshot, stamped at the start of the
UTC day and marked partial, so a panel comparing years can tell a bar that is
still growing from one that is finished; read it as the newest row per year.
The first run of the next year replaces it with the final row dated December.
gh_account.packages is counted from the REST listings the profile family
walks, not from GraphQL’s packages connection, which does not see the
container registry and answered 0 for an account whose four packages are all
containers. If a listing fails the GraphQL count stands.
gh_account_total is the answer to “how many ever”. Every other measurement
here is a row per fact, which is the right shape for “how many in July” and the
wrong one for a lifetime count. GitHub counts them itself: ten of them are one
GraphQL query with a search alias for each, and commits, which GraphQL search
cannot count, is a REST search of its own. So each number is one row and is
right on the first sweep of a fresh install.
The three fields that end in _elsewhere share one meaning of elsewhere: in a
repository the account does not own, since the search qualifier -user:LOGIN
leaves out the ones it does, so an organisation’s repositories count as
elsewhere even when the account belongs to it. pulls_merged_elsewhere is its pull
requests merged there and issues_elsewhere the issues it opened there.
commented_elsewhere is the issues and pull requests there that carry at least
one comment of its own (commenter:LOGIN -user:LOGIN): threads, not comments,
each counted once however many it left. One it opened counts only if it also
commented, since the opening post is not a comment: measured on 2026-09-26, 33
of the 102 threads the account had opened elsewhere. In 2.5.1 and earlier it
was commenter:LOGIN -author:LOGIN, which left out the threads the account
opened and kept every one in its own repositories: 125 where the definition
above gives 55 on the account this was read against, 103 of them at home and 54
of those Dependabot pull requests. A store holding rows from before the upgrade
sees the series drop by that difference on the first sweep after it.
gh_dependency_change writes a row on every sweep of the deps family, not
only when a dependency moved: a range with no change, a head that did not move
and the first sweep, which has no base yet, each write one row with change
and ecosystem at (none) and both counters at zero. InfluxDB 3 creates a
table at its first point and answers a query naming a table it has not seen
with an error, so a measurement written only on a change did not exist until
the first bump, and the panel over it was an error until then. The zero row
costs no request and the panels leave the (none) series out.
The three measurements about other people’s repositories,
gh_discussion_comment, gh_issue_comment and gh_repo_created, exist because
a sweep over one’s own repositories cannot see any of it. Each carries own so
the two can be told apart.
Activity
Section titled “Activity”| Measurement | Dated | Tags | Fields |
|---|---|---|---|
gh_ | dated | type, action, ref_type | events, public, commits, url |
gh_ | dated, last update | reason, private, subject_type | is_unread, notifications, title, url |
Both are windows, not histories. GitHub keeps the last three hundred events
of the past thirty days, and drops notifications after three months unless they
are saved. What is captured is what was there when the sweep ran. gh_event carries no actor, since the
feed is the account’s own and the actor was the login on every row; action
and ref_type are written on every row, (none) on a push. is_unread is a
field: reading a thread does not move its updated_at, so the daily read with
all=true, the one that lists a thread read without a reply, rewrites the same
row rather than opening a second one beside it.
A notification’s url is derived from the API address of its subject, and a
subject shape the mapping does not recognise is left without one rather than
guessed at, so a good part of the rows carry no link.
Configuration and delivery
Section titled “Configuration and delivery”| Measurement | Dated | Tags | Fields |
|---|---|---|---|
gh_ | daily | hook (the id), host, active | events, hooks |
gh_ | dated, when delivered | hook, host, event, status, code, ok | deliveries, duration_seconds, redelivery |
gh_ | daily | ruleset, target, enforcement | rulesets, active, days_since_change, url |
gh_ | daily | ruleset, rule | rules, bypass_actors, bypass_always, bypass_sampled, ref_include, ref_exclude |
gh_ | dated, when saved | ruleset, target, actor_type | versions, version_id, ruleset_id, actor_id, url |
gh_ | daily | pattern | rules, admin_enforced, allows_deletions, allows_force_pushes, blocks_creations, dismisses_stale_reviews, requires_approving_reviews, required_reviews, requires_code_owner_reviews, requires_commit_signatures, requires_conversation_resolution, requires_linear_history, requires_status_checks, requires_strict_status_checks, required_checks, requires_deployments, restricts_pushes, restricts_review_dismissals, url |
gh_ | daily | branch, is_default | branches, oid, days_since_commit |
gh_ | dated, when the deployment was created | deployment, environment, task | outcome, deployments, deployment_state, success, superseded, creator, commit, ref, log_url, environment_url, run_id, seconds_to_status, seconds_live, url |
gh_ | dated, when the path last changed | file (dependabot, codeowners, security, funding) | present, bytes, changes, path, blocks, ecosystems, url |
gh_ | dated, when dependabot.yml last changed | ecosystem, interval | blocks |
gh_ | daily | environment | environments, days_since_change, age_days, protection_rules, has_branch_policy, protected_branches, custom_branch_policies, can_admins_bypass, url |
gh_ | daily | key, read_only | keys, days_since_use |
gh_ | now | security_policy, forking_allowed, discussions, issues, wiki, sponsorships, blank_issues, auto_merge, delete_branch_on_merge, merge_commit, rebase_merge, squash_merge, funding_links, issue_templates, pull_request_templates, branch_protection_rules, codeowners, codeowners_errors, vulnerability_alerts, url | |
gh_ | now | visibility, archived, fork | commits, stars, forks, watchers, issues, issues_open, issues_closed, pulls, pulls_open, pulls_merged, pulls_closed, releases, discussions, labels, milestones, branches, tags, size_kb, repo_id, age_days, days_since_push, url |
gh_ | daily | ecosystem | packages |
gh_ | daily | license | packages |
gh_ | now | change, ecosystem | packages, vulnerable, base, head |
gh_ | now | resource | limit, used, remaining, used_ratio, seconds_to_reset, own_cost, own_queries |
gh_ | now | family, scope (family, repo), reason | repos, failed, points, error |
Webhooks fail silently. Measured, one hook had been answering 403 for seventy-eight of its last hundred deliveries and nothing anywhere said so.
A 404 from branch protection does not mean unprotected: a repository can be governed entirely by rulesets, which that endpoint knows nothing about.
gh_ruleset_version is the changelog behind gh_ruleset: one row per saved
version of a ruleset, dated the moment GitHub saved it, with the actor that
saved it. days_since_change only summarizes that history: a ruleset switched
off on a Tuesday and back on the Friday after reads as “changed three days
ago”, and nothing else collected says a protection was ever absent. GitHub
names the actor by id and type and not by login, so the row carries
actor_type as a tag and actor_id as a field. The family is rulesets,
daily: one list request per repository and one history request per ruleset,
both with an ETag, so a day on which nobody edited a protection costs nothing
from the budget. Measured on 2026-09-11 against the ruleset guarding the
busiest repository measured: twenty versions across five months, 3 KB, one
core request.
Only the host of a webhook URL is stored. The path usually carries a secret.
gh_repo_policy and gh_repo_total arrive in one batched GraphQL query that
costs a single point for ten repositories, which is why settings that a REST
sweep would price at a hundred and ninety eight calls are collected at all.
codeowners_errors is the one that fails silently: a broken CODEOWNERS file
stops requesting reviews and says nothing. vulnerability_alerts rides in that
same query at no extra cost and is a second, independent reading of the switch
gh_security_feature{feature="dependabot"}.enabled reports: one is the
repository’s own setting, the other is whether the listing actually answered.
Two sources that disagree is the case worth seeing.
gh_repo_total is where an archived repository’s stars and forks are read
from, on the Overview and in Every repository, ever, and gh_repo is not.
One the default filter sets aside for being archived gets no gh_repo row from
a sweep, since no family walks it, and has only the one a backfill wrote, but
it is still starred, unstarred and forked, so the totals family writes its
gh_repo_total on every sweep from the query that dates its archive: the same
tags and fields a collected repository’s row has, archived true, stamped at
the sweep. Up to 2.5.1 only a backfill wrote it, once, and on 2026-09-26 one
such row said 4 stars where GitHub said 3. A live repository’s stars and forks
on the Overview still come from gh_repo, the row the Inventory table and the
card read too. It gets no gh_repo_policy. That
query asks about twenty five repositories at a time: measured the same day,
the gateway answered the lifetime row of fifty archived repositories once in
9.2 seconds and refused it twice after about eleven, and answered twenty five
in under seven.
The three dependency measurements are off by default. The SBOM is one call and a megabyte or two per repository, and only the aggregate is kept: a single dependency bump is three hundred and seventy changes, and what is stored is six rows.
The commit a diff ends at, and the next one starts from, is read as the bare
SHA of HEAD under the application/vnd.github.sha media type: forty bytes,
where the one-commit listing it used to read was five and a half kilobytes.
The answer carries an ETag and is asked for conditionally, so on a repository
nobody pushed to the day’s read is a free 304, as the listing’s was. The SBOM
is read only when that head moved: GitHub regenerates it on every request, so
its ETag never matches and each read is charged from its own bucket, and a
repository without a commit has the packages it had.
gh_rate_limit and gh_collector_family are what the collector measures of
itself. The first is what it has left to spend: GitHub runs fifteen independent
budgets, and without it a family skipped for want of budget looks exactly like a
family with nothing to report.
The second is what each sweep managed to do. One row per family it ran, always,
with how many repositories it was asked about (repos), how many of them it
could not collect (failed) and how many rows it produced (points). For
commits, issueevents and issues, repos also counts the repositories the
movement query found nothing
new in, which the family left unread and wrote no row for. And one
row more per repository it lost, naming that repository the way every other
measurement names one and carrying reason, a bounded word for what stopped it
(the HTTP status, rate limited, query too large, canceled), with the whole
message in the error field. scope is what tells the two apart: family for
the first kind, whose repository tags hold (none), and repo for the second.
error holds (none) too where there is no message, which is not decoration:
the line protocol drops an empty string field, so a column written only on a
failure would not exist at all until one happened, and a query naming it would
be refused rather than answered with no rows.
The rows that always arrive are the point of it. A family with no row at all in a sweep did not run in that sweep, which an empty panel could never say, and they are also what makes the measurement exist on an account where nothing has ever failed: a table InfluxDB has never been written to is not drawn empty, it is refused.
One value of family is not a family. discover is the repository listing,
which is not configurable and cannot be switched off, and it is here because
every family depends on it: a sweep that cannot list the repositories runs none
of them, and without this the page would show sixteen families that never ran
and no reason for any of it. It writes a row only when it failed, because a
listing that worked is already stated by every other row of the same sweep.
This exists because of one measured failure. On 2026-09-16 gh_workflow_run and
gh_workflow_job held nothing at all for the five busiest repositories of this
account, each because one /repos/<repo>/actions/runs/<id>/jobs call had
answered 502 once and the runner had thrown away everything that family had
already collected for that repository. The Continuous integration row of the
dashboard was computed over an account missing its five busiest repositories,
the Cost row on the same page reported one of them burning 27.6 K macOS minutes,
and the only record of the cause was one line in a journal. The collector keeps
what it gathered before a failure now, and it writes down what failed.
GET /rate_limit reports the budgets and charges for none of them, which is
what makes almost all of this free. Not every one it reports is true: measured with the token this runs under, the endpoint answered
graphql as used=0, remaining=5000 in the same minute GraphQL itself answered
used=162 and moved by one on every query, and the two do not even share a
clock. So the graphql row is built from the rateLimit block GraphQL answers
with, and the endpoint’s version of it is dropped rather than published beside
it. That reading is a request every fifteen minutes and costs nothing in points,
measured.
graphql is not the only bucket the endpoint invents. Measured on 2026-09-12,
an SBOM request’s headers said dependency_sbom used 1, remaining 99, reset in
59 s, and GET /rate_limit two seconds later said used 0, remaining 100, its
reset sliding forward a second per call. Every REST answer names the bucket it
charged in its headers, so a row is built from the newest headers the client
saw whenever they are still inside their own window and say more was spent
than the endpoint admits. The windows of dependency_sbom and search are one
minute, so a dependency_sbom row that reads zero between runs of the deps
family is a refilled bucket, not the defect.
The graphql row can therefore be absent, where the others are written
whenever the endpoint answers at all. It is not written when nothing has been
read from GraphQL yet and no earlier reading is still inside its window: a
missing row says “not measured” where a zero says “nothing spent”, and the zero
was the defect.
own_cost and own_queries are on that row alone. limit, used and
remaining describe the whole token’s window, shared with whatever else holds
it; these two are the part this process is answerable for. They count from
process start, so a restart returns them to zero and a panel has to read them as
a counter rather than a value.
Job logs
Section titled “Job logs”| Measurement | Dated | Tags | Fields |
|---|---|---|---|
gh_ | dated, when the line was printed | workflow, job_name, run | line, head_branch |
Off by default: set every.families.joblogs. It is text rather than a
measurement, so it is excluded from the InfluxDB sink by default and skipped by
the Prometheus exporter; Loki is where it belongs. The exported dashboards carry
a text panel, “Where failure output went”, in the place the lines would take,
since an importer may have no Loki; cmd/publish_dashboard -loki <datasource-uid> publishes the dashboard with the lines drawn from Loki in that
panel’s place (see the
dashboards).
Only failed jobs, and only the last forty lines of each. A successful job’s output is thousands of lines nobody will read, each log costs a request, and the tail is where a failure explains itself. GitHub deletes logs after the repository’s retention period, ninety days by default, and answers 410 once they are gone; a backfill asks only for the last ninety days, whatever the setting, and at most 500 failed jobs per repository.
workflow here is the same file path, through the same helper, so a log line
joins to the run that printed it; and branch became the field head_branch
for the same two reasons as on the run. run is deliberately a series per run,
which is what a log line genuinely belongs to, and it is affordable only because
this family is off by default.
Colour codes are stripped and the byte order mark GitHub writes before the first timestamp is removed, so a search for a word does not fail because the word happened to be coloured.
The failure list is asked only for the runs created in the thirty one days before the window opened, rounded down to the day. Thirty one days because a re-run keeps the created_at of its first attempt and GitHub allows one for thirty days: measured on 2026-09-11, the newest failure of the busiest repository measured was the third attempt of a run created two hours before it finished, and a margin the length of a job would have missed every re-run of a failure older than a morning. Unfiltered the list was the newest hundred failures the repository ever had, six hundred kilobytes per repository per sweep for a window of an hour that is nearly always empty; filtered, a month of failures, sixty eight rows and a megabyte decompressed on that busiest repository, a few rows or none on most. Rounding keeps the URL, and with it the ETag, the same across the sweeps of a day, which is what makes the repeat a free 304 rather than a charged 200 on a fresh URL; within the day the page changes only when a failure is created or re-run. The cut at the window itself is still made here, by when the run finished.
| Measurement | Dated | Tags | Fields |
|---|---|---|---|
gh_ | dated, per day | product, sku, unit | quantity, price_per_unit, gross, discount, net, url |
unit is GitHub’s own unitType, capitalised as GitHub sends it: Minutes,
GigabyteHours, AICredits, Requests. It is passed through rather than
normalised, and the panels that read minutes filter on the capital.
There is no org tag. The only billing endpoint a personal account can read is
its own, and that report has no organizationName: the field belongs to the
organization report, which needs an organization to ask about. Checked against
the published OpenAPI description and against the live endpoint, where none of
487 usage items carried the key. Written anyway it was (none) on every row
ever collected, which is a column and a legend entry that only ever says there
is nothing here.
repo is (none) on a charge that belongs to no repository, which is what a
Copilot seat is. That is a real row of the bill and not a repository, so the
cost table by repository leaves it out; the spend totals above it include it.
net is not always zero. On the account this was developed against it carries
the monthly credit, which is why gross, discount and net are all stored rather
than one being derived from the others. The row is stamped at the start of its
day: GitHub’s date arrives as the first billed minute of the day on half the
rows, which would double the row had GitHub reported a different minute on the
next read.
Columns that exist only once written
Section titled “Columns that exist only once written”InfluxDB 3 creates a column the first time a row carries it, and a query that
names a column no row has written fails at planning rather than answering
null: the whole panel goes red. So a field written only when GitHub has a
value for it does not exist on a database where that has never happened. The
ones a fresh database is most likely to lack: gh_discussion.state_reason,
seconds_to_answer and seconds_to_close; gh_milestone.days_to_due and
seconds_to_close; gh_ruleset_rule.ref_exclude; and
gh_workflow_run.initial_actor. The same rule covers every seconds_to_*
that needs a closing, every url on an item GitHub sends no address for, the
optional advisory fields on an alert, checks_total and checks_failed on a
commit a gate ran on, label_names on an item with a label, resolved_by on
a resolved thread, queued_seconds on a run’s first attempt,
pull_request, pull_requests, headline and head_repo on a run GitHub
linked, described or took from a fork, merged on work elsewhere that was
merged and additions, deletions and changed_files on work elsewhere that
is a pull request, which an account that has only opened issues in other
people’s repositories has never written, and language on an upstream
repository GitHub detected one in. Two more are worth naming, because the
tables above list them beside fields that are always there:
gh_dependabot_alert_item.dismissed_comment, written only when whoever
dismissed an alert typed a reason, and gh_event.commits, written only on a
push event. Neither column exists on the production database this
documentation was checked against. gh_label writes only the labels somebody
has used; gh_repo_total.labels is the declared count.
One of these is named by a shipped panel, and it is the one most likely to be
missing: gh_pull_request.seconds_to_first_human_review, which exists only
once somebody other than the author and other than a bot has reviewed a pull
request. On a database where that has never happened, the stat that reads it
reports a schema error rather than No data, and it takes the five values
beside it in the same panel with it. Nothing in a query can ask whether a
column exists, so this is a property of the store rather than a defect to
repair: the repair, if it bites, is one row of any kind carrying the field.
How rare it is, measured on 2026-09-17 against the account this was developed on: fourteen rows in the whole store, over fourteen pull requests of four repositories, the newest raised on 2026-07-05, and none of them inside the last fortnight, against 703 pull requests of 825 in that fortnight that had a first review from a bot or from their own author. The field is not broken; it is the answer to a narrower question than a reader expects, which is why the panel is named for that question.
The same reading applies to the two on gh_workflow_run.
initial_actor is written only when GitHub’s actor and triggering_actor
differ, which is a re-run somebody else asked for: 0 of 10,201 runs in a
fortnight here, and 0 of the 300 newest runs of three repositories checked
against the API at the same time. head_repo is written only for a run that
came from another repository: 18 of those 10,201, all from one fork’s pull
request. Both are correct and both are rare, which is what a field written
only when GitHub has something to say looks like.