Skip to content

Short answers

Each answer below is the short form of a page that has the detail, and links to it. Where an answer quotes GitHub rather than this project, it links to GitHub’s own documentation.

How do I keep GitHub traffic data longer than 14 days?

Section titled “How do I keep GitHub traffic data longer than 14 days?”

Collect it more often than every fourteen days and store each day under its own date. GitHub’s traffic API serves views and clones per day for the last fourteen days only. ghchronicle’s traffic family rereads the whole window every six hours by default and writes each day at that day’s date, so the rows converge rather than pile up, and a collector that was down for a week loses nothing. Older days cannot be recovered: GitHub never kept them.

The detail is in what is collected and dating a point.

Read GitHub’s daily star history, which gives the stars a repository gained on each day, back to the week it was created, to anyone who can see the repository. ghchronicle’s stars family reads it whole on its first sweep, for every repository it collects, and writes each day as a gh_star_day point, so the curve starts at the repository’s first star, not on the day it was installed, and a repository whose stargazer list GitHub hides from the token still gets its daily curve. Where the token may read that list, the family also walks it with the star media type and writes each star as a gh_star point, dated to the second and naming who gave it. Since July 2026 GitHub serves the list only to a repository’s admins and collaborators.

The token always has that access on the account’s own repositories and on those of an organisation its user administers. After the first sweep, ghchronicle asks for the newest thirty weeks of each history, one request per repository and usually a free 304, and the newest hundred stars of each list it may read, ten repositories to a GraphQL query. The history is GitHub’s stargazers/history endpoint: stars per day grouped by week, thirty weeks to a page, paging back to the repository’s first week, thirteen pages for cli/cli, back to 2019. Its days are Pacific calendar days, and it counts today’s stargazers, so an unstar takes a star off the day it was given. See what is collected and what GitHub will not give.

Can I collect GitHub history from before I installed it?

Section titled “Can I collect GitHub history from before I installed it?”

Mostly, with one deliberate run of ghchronicle -config config.yaml -backfill. It walks every enabled family until the API runs out, or back to the bound -backfill-since sets, waits for the rate limit to reset instead of skipping, and, stopped half way, carries on from its checkpoint when run again. Three things no backfill reaches: traffic older than fourteen days, the event feed past its three hundred events or thirty days, and job logs older than ninety days, joined from 1 October 2026 by workflow runs past the repository’s retention period.

The walk, its checkpoint and its bound are on backfill.

Can I archive GitHub notifications and the event feed?

Section titled “Can I archive GitHub notifications and the event feed?”

Yes, as long as the collector runs while they are still there. The events family keeps the account’s activity feed, which GitHub serves up to three hundred events and none older than thirty days, and notifs keeps the inbox, which GitHub says holds notifications for three months unless they are saved. Both run every fifteen minutes by default, each item a point dated when it happened, and a Loki sink turns them into log lines.

GitHub states both windows in its documentation on the event feed and the inbox. The families are in what is collected, and the log lines in Loki.

How long does GitHub keep workflow run history?

Section titled “How long does GitHub keep workflow run history?”

From 1 October 2026, ninety days by default, according to GitHub’s documentation. Logs and artifacts already had that default; from that date the repository’s retention setting also deletes workflow runs, checks and commit statuses, which until then were kept for 400 days or more. A public repository can keep them between one and ninety days, a private one up to 400. ghchronicle’s actions family writes every run with its jobs and steps, dated when it finished.

GitHub’s statement is under retention for checks, workflow runs and artifacts. A backfill walks every run GitHub still holds, back to -backfill-since when one is set, and expands each into its jobs, and into its steps while GitHub still serves them; from 1 October 2026 it can reach no further back than the repository’s retention period, ninety days by default.

Can Prometheus store ghchronicle’s history?

Section titled “Can Prometheus store ghchronicle’s history?”

No. Prometheus stamps a sample when it is scraped and refuses one much older: measured against Prometheus 3.14 with a thirty-minute out-of-order window, a sample dated two days back came back as HTTP 400. So the Prometheus exporter, and the OpenTelemetry sink with raw: false, serve current values reduced from the dated rows, which suits alerting. The history belongs in InfluxDB, PostgreSQL, Graphite or Elasticsearch, alongside Prometheus if you run both.

The measurement and the reduction are on dating a point.

Which database should I store GitHub metrics in?

Section titled “Which database should I store GitHub metrics in?”

One that keys a row by its timestamp, because that is what lets the same fourteen days be rewritten on every sweep without duplicates. InfluxDB, PostgreSQL, Graphite and Elasticsearch all keep the dated history, and ghchronicle generates a Grafana dashboard for each of them and for Prometheus, which keeps current values only. Running more than one sink is the normal arrangement: a history store for the charts, Prometheus for alerts.

Each store is weighed on choosing a store, and the dashboards are on importing.

Why does stats/code_frequency always return 202?

Section titled “Why does stats/code_frequency always return 202?”

On a personal account, GitHub answers stats/code_frequency and stats/contributors with 202 and an empty body indefinitely. A 202 normally means the numbers are still being computed and a later request will get them; for these two, the next answer is another 202. ghchronicle does not call them: lines added and removed come from the commits collector instead, per commit, attributed to an author and dated to the commit.

The ten-second check and the other endpoints that do not answer are on what GitHub will not give.

Why does the GitHub traffic API return 403 with my token?

Section titled “Why does the GitHub traffic API return 403 with my token?”

Because traffic is gated on push access to the repository, not on being able to read it. GitHub serves views and clones only to someone who can push to the repository, and a fine-grained token also needs the repository permission Administration (read); without both, the answer is a 403. A workflow’s automatic GITHUB_TOKEN cannot be granted Administration, so it gets a 403 even in the repository it runs in. ghchronicle records the 403 as unavailable and carries on, so the symptom is an empty traffic panel, not a failed sweep.

Every scope and what its absence looks like is on the token.

For its repositories, yes: name the organisation under targets.orgs and its repositories are discovered and swept like the account’s own, which needs the read:org scope. The account-wide families, the contribution calendar, the event feed, notifications, billing, packages and the outbound ones, need targets.user and describe a user. Organisation-only surfaces such as the audit log are not collected: the only organisation endpoint ghchronicle calls is the repository list.

The keys are on targets, and what an organisation has that a personal account cannot see is on what GitHub will not give.

Can it run in GitHub Actions without a server?

Section titled “Can it run in GitHub Actions without a server?”

Yes, apart from the store it writes to. The repository ships a composite Action, jmrplens/ghchronicle@v2, which downloads a release binary and runs one sweep, a backfill or a card render. It needs a personal access token, because the automatic GITHUB_TOKEN lacks push access to other repositories, cannot read traffic even in its own, and is not a user. A hosted runner keeps no state file between runs, so unless the state is cached, which needs a configuration file, every run is a first run: it collects every family whatever its cadence, walks the stargazers, the whole star history and the account’s co-authored pull requests again, and pays for every answer in full, with no ETag to be answered a free 304.

The modes, the inputs and the cache step are on GitHub Actions.

Not by default. The collector never spends the last 500 calls of a budget, or a fifth of the bucket’s limit when that is smaller, and a sweep skips a family with a warning rather than crossing that line, so whatever else shares the token keeps working. Every response is cached by its ETag, and a 304 Not Modified costs no quota at all, which is what makes short cadences affordable.

The brake is on rate limits, and every family is priced on cost of a sweep.

  • Compared with the alternatives: KipHub Traffic, repohistory, github-repo-stats, star-history and the Prometheus exporters, cell by cell.
  • Troubleshooting: the messages that look like errors and are not, and the ones that are.
Written and maintained by
MIT licenceRelease history