The file
ghchronicle reads one YAML file, and the documented example in the repository is the place to start it from.
cp config.example.yaml config.yamlconfig.example.yaml documents every option in comments. It is the file to
copy, and it is kept in step with the code: adding a setting means adding it
there too.
The blocks
Section titled “The blocks”github: # the token, the reserve, the timeouttargets: # which repositoriesgroups: # which categories of metric are collected at allsinks: # where the points goevery: # how often things run: a default, per group, per familyheartbeat: # how often the sweep loop wakes, for a test runlog: # level, format, and an optional rotating filestate_file: # what a restart remembersbackfill: # the bound, when run with -backfillmigrate: # what a start does about what an upgrade left in the storesgrafana: # where the dashboard is published, when askedOnly github.token and one of targets.user, targets.orgs or
targets.repos are genuinely required, plus at least one sink. Everything else
has a default.
${VAR} expansion
Section titled “${VAR} expansion”A credential, an address or a file path is expanded from the environment at
start-up. ${GITHUB_TOKEN} becomes the value of that variable, or an empty
string if it is not set, except in a file path, below.
github: token: ${GITHUB_TOKEN}sinks: influxdb: url: http://localhost:8181 token: ${INFLUX_TOKEN}Expansion is per key, not over the whole file, and these are the keys it reaches:
- credentials:
github.token,sinks.influxdb.token, every value ofsinks.otlp.headers,sinks.telegraf.usernameandpassword,sinks.postgres.dsn,sinks.elasticsearch.api_key,usernameandpassword, andgrafana.token,userandpassword; - addresses:
github.base_urlandweb_url, theurlofinfluxdb,loki,telegrafandelasticsearch,sinks.otlp.endpoint,sinks.graphite.addr,grafana.urlandgrafana.datasource.url; - the identifiers of what Grafana already has:
grafana.dashboard_uid,grafana.datasource.uidandgrafana.datasource.loki_uid; - file paths:
state_file,sinks.dedupe_file,log.file,sinks.file.pathandsinks.sql.path, where a leading~is also the home directory, sostate_file: ~/.ghchronicle/state.jsonis a file under it. Only~on its own or before a separator is read that way,~nameis left as written, and a~where the process has no home directory is refused at start-up naming the key. So is a${VAR}in a path that is unset or empty, since the path left without it is not the one written:state_file: ${STATE_DIRECTORY}/state.jsonwould be/state.json, with the ledger and the cache file beside it at the root of the filesystem. Before 2.6.1 no path was expanded, and that~in astate_filewas a directory called~under the working directory.
Every other value is read as written, so a ${VAR} in sinks.loki.tenant_id,
sinks.influxdb.bucket, sinks.prometheus.listen, a label, a prefix or a
cadence reaches the collector as those characters.
This is the whole reason the file can be committed. The configuration is the shape of the deployment and belongs in version control; the tokens are credentials and belong in an environment file with mode 600, or in a secret store.
Directory/etc/ghchronicle/
- config.yaml world readable, version controlled
- ghchronicle.env mode 600, never committed
github
Section titled “github”github: token: ${GITHUB_TOKEN} reserve_rate: 500 timeout: 30s # base_url: https://github.example.com/api/v3 # web_url: https://github.example.com| Key | Default | Meaning |
|---|---|---|
token | required | A classic or fine-grained personal access token |
reserve_ | 500 | Calls never spent, so whatever else uses the token keeps working. A non-positive value falls back to the default |
timeout | 30s | Per request. GraphQL over a large account can be slow. A Go duration; anything unparseable or not positive falls back to the default |
base_ | api.github.com | A GitHub Enterprise instance uses https://<host>/api/v3. GitHub Enterprise Server does not serve the daily star history, so there gh_star_day and the star-count panels drawn from it stay empty |
web_ | derived from the API | The site the profile page is on, read by achievements without the token. Needed only when base_url is a proxy in front of the API |
The reserve is scaled per bucket; see rate limits.
state_file
Section titled “state_file”state_file: /var/lib/ghchronicle/state.jsonNine things, and deleting the file costs a different one for each:
last_run, when each family last ran. Without it every family is due at once, so the next sweep is a full one, except that the service starts the families of six hours or more one a sweep.first_saw, when each repository’s stargazer list was first walked whole. It is written only after a walk that came back without an error. Without it the one-off full walk of the stargazer list is done again.history_read, when each repository’s daily star history was last read whole. It is written only after a walk that reached the end of the history, not after one an error or a later page’s 403 or 404 cut short. Without it that history is read whole again, a page per thirty weeks of the repository’s life.last_head, the commit each repository was on when the dependency diff last ran. Without it the dependency changes in the gap are gone: the next sweep has the photograph and no diff.last_full, when each family that normally reads what changed last took the day’s read: forissues, every open item and what moved since the day’s read before it, recorded for each repository, so one that failed takes it again and no other does. Without it a repository reads as due, so the next sweep takes it, reaching back a month.last_notified, where the inbox window was cut. Without it zero asks for the whole inbox.last_event, the newest event the feed had. Without it empty reads the whole feed.coauthored, the Pair Extraordinaire count theachievementsfamily has settled, the last UTC day it covers, the version of the rule it was counted by and the day the whole history was last walked. Each pass walks only the pull requests merged since that day, and the whole history again once a week or when the rule has changed. Without it the next pass walks the account’s merged pull requests whole, which on an account with 2,315 of them was 35 queries and 23.7 MB, where a pass that has it walks only the day in progress: the counts query and one page of the walk, two points, and 468 to 513 KB for each hourly pass the production proxy logged against the same account on 2026-09-27, a size that grows through the day with what is merged.stores, for each store a run writes to, where it points, the release that first wrote it, the release that last did, the migrations applied to it, the copies of old rows a migration set aside there, and the history a migration cleared there and has not read back yet. Without it a SQL file, a Graphite or a Telegraf is taken to be the running release’s own, so what an earlier release left in it is no longer shown, and a store cleared and not read back looks like one that never held the old shape, so nothing reads its history back until a-backfill -familiesof the families the refill named does.
Seven of the nine cost only quota, because what is collected again is keyed by
measurement, tags and timestamp and overwrites what is already stored.
last_head loses something: the dependency changes between the head it held
and the next one are read from a range that nothing can name once the head is
gone. So does stores: which release first wrote a store that cannot be asked
what its rows are keyed by, and a refill a migration still owes, are questions
nothing can answer afterwards.
A file that is there and cannot be read or parsed is not taken for a new one:
a run that writes to the stores stops and names it, since read as new it would
forget a refill still owed, and its first save would put a new file in its
place. A run as another user, root most often, is what leaves it unreadable;
give it back to the user the collector runs as. A run that saves the file
keeps what another process recorded of the stores since it read it, so two
runs that overlap, a -once from cron across a -migrate -yes, do not undo
each other’s records of the stores.
A run with -card-only writes none of the nine. Its points reach the
card and no store, so a mark it left behind would make the
next collection skip a family, or narrow a read, whose data went into a picture
and nowhere else. It reads the file as any other run does.
The other file a sweep remembers itself in
Section titled “The other file a sweep remembers itself in”sinks: dedupe_file: /var/lib/ghchronicle/state-written.bin dedupe_horizon: 720h| Key | Default | Meaning |
|---|---|---|
sinks. | beside state_file, as <name>-written.bin | The ledger of what has already been written. off disables it for every sink |
sinks. | 720h | How long the ledger remembers a point nothing offers any more. A Go duration; anything unparseable or not positive falls back to the default |
Losing it costs one sweep of rewriting and nothing else, which is exactly what a store that has been wiped and needs filling again wants. Both files want a persistent path: only what changed is written.
The cache beside it
Section titled “The cache beside it”/var/lib/ghchronicle/state-cache.binBeside state_file, as <name>-cache.bin, is what a sweep learned about GitHub
that makes the next one cheap. There is nothing to configure: its place is the
state file’s. It holds four things:
- the conditional cache: every
answer asked for within twice the longest cadence the configuration runs, and
never less than a day, 48 hours at the default cadences, with the ETag it came
with, so a restart asks with
If-None-Matchand is answered a free 304; - the workflow runs whose jobs were written, so a restart does not list them again;
- the refusals, each until the end of its own day;
- the page sizes the
totalsfamily gave the pull request query.
All four used to live in the process, and every restart paid for them again. Measured on the author’s service on 2026-09-26, the first 38 minutes after a restart spent 1,092 charged core requests on passes that cost about 66 with the cache warm.
The runs are a claim that every store holds their jobs, so a start does not
recall them where that is not known: where no write ledger remembers what the
stores hold, which is the case once the ledger is deleted to fill a wiped store
again, with dedupe_file: off or a store’s own dedupe: false, and in a run
that ends with its sweep; and where a destination has been added since the file
was written. Their jobs are then listed and offered again with every other
point. Two of those cases are news, a ledger the configuration keeps that reads
empty and a destination added, and start-up says them:
level=INFO msg="the write ledger remembers nothing, listing the jobs of the runs the cache file remembers again" runs=666level=INFO msg="a store was added since the cache file was written, listing the jobs of the runs it remembers again" stores=influxdb,loki,postgres written_to=influxdb,loki runs=666The others are how the configuration always runs, a store kept without a
ledger or a run that ends with its sweep, and are said only at debug, as
not every store keeps a write ledger.
Every start that finds the file says what it read, at its first sweep, and
one that finds none says so at debug:
level=INFO msg="cache file read" file=/var/lib/ghchronicle/state-cache.bin written=2026-09-27T15:27:41+02:00 answers=1170 runs=672 refusals=100 page_sizes=54 age=1sIt is written at most every five minutes while the collector runs, and once more
when it stops: at the end of a -once, and when the service is asked to stop
with SIGTERM or SIGINT, which is what systemctl stop, docker stop and
Ctrl-C send. A process killed outright
cannot, so SIGKILL, the kernel’s out-of-memory killer, docker kill,
taskkill /F and Stop-ScheduledTask lose what it learned since the last save,
five minutes at most, and the next start asks those answers again. A save that
fails says so, and so does a start that finds a file it cannot read:
level=WARN msg="cache file not saved" file=/var/lib/ghchronicle/state-cache.bin err="open /var/lib/ghchronicle/state-cache.bin.tmp: permission denied"level=WARN msg="cache file not read, the first pass of each family pays in full" file=/var/lib/ghchronicle/state-cache.bin err="..."A run with -card-only reads it and does not write it, for the reason it
leaves the state file alone. A -backfill reads it
and does not write it either: the pages it walks are pages no sweep asks for,
and kept they would crowd the sweeps’ own answers out of the file. An upgrade
keeps it: an answer is kept under the shape of what its collector reads, so the
only answers asked for again in full are those of a collector that now reads
something else.
It is bounded. An answer larger than a megabyte is left out, and so is whatever comes past 64 MB, the least recently asked for first. Measured against the author’s account of 37 repositories, one sweep of every family keeps 1,065 answers, 10.3 MB of them and 1.6 MB of file, none of them larger than 404 KB. Over days the cache in memory grows by the answers whose query carries a moving window: the author’s service held 6,695 after five and a half days, about 99 MB, and the 2,517 of them asked for in the last 48 hours, which are what the file keeps, come to about 31 MB. It holds what GitHub answered about private repositories as well as public ones, and is written with mode 600, as the state file is.
Deleting it costs one pass of each family at full price, the price every restart
paid before it existed, and loses nothing. Every answer is asked for again,
every refused feature is asked about again, which is also how to have a feature
switched on today noticed before its day is out, the jobs of the newest runs are
listed once more, and the first sweep runs totals before the pull requests,
whatever its cadence says, so their pages are sized again. When totals was
not due anyway, it says so:
level=INFO msg="no page sizes remembered, running totals before the pull requests it sizes"A file that does not load in full, cut short, damaged or written in another format, is not read at all, with the warning above, and costs the same; the next save writes over it. It wants the same persistent path as the other two.
backfill
Section titled “backfill”backfill: since: 2yOnly applies to a run started with -backfill, and is overridden by
-backfill-since. See backfill.
A backfill keeps one more file beside state_file, <name>-progress.json,
which is where it records what it has already written so a stop costs one
repository rather than the walk. There is nothing to configure: it appears when
a backfill starts, it is removed when the walk reaches the end, and a sweep
neither writes nor reads it. See stopping one, and picking it up
again.
migrate
Section titled “migrate”migrate: autoWhat a run that writes to the stores does, before its first sweep, about a
store an earlier release left in a shape this one no longer writes. -migrate
prints what each configured store holds of every such change; see
Migrations.
| Value | What a start does |
|---|---|
auto | The default. Applies on its own a pending change that loses nothing, says so at WARN, reads back what it cleared, and what an earlier run cleared and did not finish reading, and warns about every other one at every start |
warn | Applies nothing and reads nothing back, and warns about every pending change, and every store still owed its history, at every start |
A change loses nothing when all three hold: GitHub still serves the whole
history, so reading it again brings back every row the store holds; the old
rows are set aside for at least 24 hours rather than deleted, which InfluxDB 3,
PostgreSQL and Elasticsearch do; and every row in the store is this
configuration’s, which the store is asked. There is no value that applies the
rest on its own. A one-shot run, -once or -backfill, on a new state file
applies nothing on its own either: that is every run of the Action without a
state file restored, whose record of a refill still owed goes with its runner.
Each warning names the two commands, -migrate for the plan and
-migrate -yes to apply it, with the service stopped. Any other value is
refused at start-up.
The service also holds a lock beside the state file, <name>-lock, for as long
as it runs, which is what keeps -migrate -yes, and -uninstall -yes of the
data or the state, from changing the stores under it. There is nothing to configure: it follows state_file, the operating
system lets go of it however the process ends, and the file left behind only
names the last process that held it.
When it is wrong, it says so at start-up
Section titled “When it is wrong, it says so at start-up”Configuration is validated before the first call is made, and the messages name the key and what it needs.
| Message | Means |
|---|---|
github. | Exactly what it says |
state_ | A path names a variable with no value: set it, or write the path out |
targets: | Nothing to collect |
sinks: | A run that collects and discards is almost never what anyone meant |
every. | The name is not a family. The message lists the ones that exist |
sinks. | Each sink validates its own required keys and says which |
sinks. | The value is not one of the accepted ones, and the message lists them |
migrate: | The same, for migrate |
The rest
Section titled “The rest”- Targets: which repositories, and the fork and archived defaults.
- Cadences: the
everyblock, and what0means. - Logging: level, format and the rotating file.
- Choosing a store: the
sinksblock, one page per store. - Letting the binary publish the dashboard:
the
grafanablock.