Skip to content

The file

ghchronicle reads one YAML file, and the documented example in the repository is the place to start it from.

Terminal window
cp config.example.yaml config.yaml

config.example.yaml documents every option in comments. It is the file to copy, and it is kept in step with the code: adding a setting means adding it there too.

github: # the token, the reserve, the timeout
targets: # which repositories
groups: # which categories of metric are collected at all
sinks: # where the points go
every: # how often things run: a default, per group, per family
heartbeat: # how often the sweep loop wakes, for a test run
log: # level, format, and an optional rotating file
state_file: # what a restart remembers
backfill: # the bound, when run with -backfill
migrate: # what a start does about what an upgrade left in the stores
grafana: # where the dashboard is published, when asked

Only github.token and one of targets.user, targets.orgs or targets.repos are genuinely required, plus at least one sink. Everything else has a default.

A credential, an address or a file path is expanded from the environment at start-up. ${GITHUB_TOKEN} becomes the value of that variable, or an empty string if it is not set, except in a file path, below.

github:
token: ${GITHUB_TOKEN}
sinks:
influxdb:
url: http://localhost:8181
token: ${INFLUX_TOKEN}

Expansion is per key, not over the whole file, and these are the keys it reaches:

  • credentials: github.token, sinks.influxdb.token, every value of sinks.otlp.headers, sinks.telegraf.username and password, sinks.postgres.dsn, sinks.elasticsearch.api_key, username and password, and grafana.token, user and password;
  • addresses: github.base_url and web_url, the url of influxdb, loki, telegraf and elasticsearch, sinks.otlp.endpoint, sinks.graphite.addr, grafana.url and grafana.datasource.url;
  • the identifiers of what Grafana already has: grafana.dashboard_uid, grafana.datasource.uid and grafana.datasource.loki_uid;
  • file paths: state_file, sinks.dedupe_file, log.file, sinks.file.path and sinks.sql.path, where a leading ~ is also the home directory, so state_file: ~/.ghchronicle/state.json is a file under it. Only ~ on its own or before a separator is read that way, ~name is left as written, and a ~ where the process has no home directory is refused at start-up naming the key. So is a ${VAR} in a path that is unset or empty, since the path left without it is not the one written: state_file: ${STATE_DIRECTORY}/state.json would be /state.json, with the ledger and the cache file beside it at the root of the filesystem. Before 2.6.1 no path was expanded, and that ~ in a state_file was a directory called ~ under the working directory.

Every other value is read as written, so a ${VAR} in sinks.loki.tenant_id, sinks.influxdb.bucket, sinks.prometheus.listen, a label, a prefix or a cadence reaches the collector as those characters.

This is the whole reason the file can be committed. The configuration is the shape of the deployment and belongs in version control; the tokens are credentials and belong in an environment file with mode 600, or in a secret store.

  • Directory/etc/ghchronicle/
    • config.yaml world readable, version controlled
    • ghchronicle.env mode 600, never committed
github:
token: ${GITHUB_TOKEN}
reserve_rate: 500
timeout: 30s
# base_url: https://github.example.com/api/v3
# web_url: https://github.example.com
KeyDefaultMeaning
tokenrequiredA classic or fine-grained personal access token
reserve_rate500Calls never spent, so whatever else uses the token keeps working. A non-positive value falls back to the default
timeout30sPer request. GraphQL over a large account can be slow. A Go duration; anything unparseable or not positive falls back to the default
base_urlapi.github.comA GitHub Enterprise instance uses https://<host>/api/v3. GitHub Enterprise Server does not serve the daily star history, so there gh_star_day and the star-count panels drawn from it stay empty
web_urlderived from the APIThe site the profile page is on, read by achievements without the token. Needed only when base_url is a proxy in front of the API

The reserve is scaled per bucket; see rate limits.

state_file: /var/lib/ghchronicle/state.json

Nine things, and deleting the file costs a different one for each:

  • last_run, when each family last ran. Without it every family is due at once, so the next sweep is a full one, except that the service starts the families of six hours or more one a sweep.
  • first_saw, when each repository’s stargazer list was first walked whole. It is written only after a walk that came back without an error. Without it the one-off full walk of the stargazer list is done again.
  • history_read, when each repository’s daily star history was last read whole. It is written only after a walk that reached the end of the history, not after one an error or a later page’s 403 or 404 cut short. Without it that history is read whole again, a page per thirty weeks of the repository’s life.
  • last_head, the commit each repository was on when the dependency diff last ran. Without it the dependency changes in the gap are gone: the next sweep has the photograph and no diff.
  • last_full, when each family that normally reads what changed last took the day’s read: for issues, every open item and what moved since the day’s read before it, recorded for each repository, so one that failed takes it again and no other does. Without it a repository reads as due, so the next sweep takes it, reaching back a month.
  • last_notified, where the inbox window was cut. Without it zero asks for the whole inbox.
  • last_event, the newest event the feed had. Without it empty reads the whole feed.
  • coauthored, the Pair Extraordinaire count the achievements family has settled, the last UTC day it covers, the version of the rule it was counted by and the day the whole history was last walked. Each pass walks only the pull requests merged since that day, and the whole history again once a week or when the rule has changed. Without it the next pass walks the account’s merged pull requests whole, which on an account with 2,315 of them was 35 queries and 23.7 MB, where a pass that has it walks only the day in progress: the counts query and one page of the walk, two points, and 468 to 513 KB for each hourly pass the production proxy logged against the same account on 2026-09-27, a size that grows through the day with what is merged.
  • stores, for each store a run writes to, where it points, the release that first wrote it, the release that last did, the migrations applied to it, the copies of old rows a migration set aside there, and the history a migration cleared there and has not read back yet. Without it a SQL file, a Graphite or a Telegraf is taken to be the running release’s own, so what an earlier release left in it is no longer shown, and a store cleared and not read back looks like one that never held the old shape, so nothing reads its history back until a -backfill -families of the families the refill named does.

Seven of the nine cost only quota, because what is collected again is keyed by measurement, tags and timestamp and overwrites what is already stored. last_head loses something: the dependency changes between the head it held and the next one are read from a range that nothing can name once the head is gone. So does stores: which release first wrote a store that cannot be asked what its rows are keyed by, and a refill a migration still owes, are questions nothing can answer afterwards.

A file that is there and cannot be read or parsed is not taken for a new one: a run that writes to the stores stops and names it, since read as new it would forget a refill still owed, and its first save would put a new file in its place. A run as another user, root most often, is what leaves it unreadable; give it back to the user the collector runs as. A run that saves the file keeps what another process recorded of the stores since it read it, so two runs that overlap, a -once from cron across a -migrate -yes, do not undo each other’s records of the stores.

A run with -card-only writes none of the nine. Its points reach the card and no store, so a mark it left behind would make the next collection skip a family, or narrow a read, whose data went into a picture and nowhere else. It reads the file as any other run does.

The other file a sweep remembers itself in

Section titled “The other file a sweep remembers itself in”
sinks:
dedupe_file: /var/lib/ghchronicle/state-written.bin
dedupe_horizon: 720h
KeyDefaultMeaning
sinks.dedupe_filebeside state_file, as <name>-written.binThe ledger of what has already been written. off disables it for every sink
sinks.dedupe_horizon720hHow long the ledger remembers a point nothing offers any more. A Go duration; anything unparseable or not positive falls back to the default

Losing it costs one sweep of rewriting and nothing else, which is exactly what a store that has been wiped and needs filling again wants. Both files want a persistent path: only what changed is written.

/var/lib/ghchronicle/state-cache.bin

Beside state_file, as <name>-cache.bin, is what a sweep learned about GitHub that makes the next one cheap. There is nothing to configure: its place is the state file’s. It holds four things:

  • the conditional cache: every answer asked for within twice the longest cadence the configuration runs, and never less than a day, 48 hours at the default cadences, with the ETag it came with, so a restart asks with If-None-Match and is answered a free 304;
  • the workflow runs whose jobs were written, so a restart does not list them again;
  • the refusals, each until the end of its own day;
  • the page sizes the totals family gave the pull request query.

All four used to live in the process, and every restart paid for them again. Measured on the author’s service on 2026-09-26, the first 38 minutes after a restart spent 1,092 charged core requests on passes that cost about 66 with the cache warm.

The runs are a claim that every store holds their jobs, so a start does not recall them where that is not known: where no write ledger remembers what the stores hold, which is the case once the ledger is deleted to fill a wiped store again, with dedupe_file: off or a store’s own dedupe: false, and in a run that ends with its sweep; and where a destination has been added since the file was written. Their jobs are then listed and offered again with every other point. Two of those cases are news, a ledger the configuration keeps that reads empty and a destination added, and start-up says them:

level=INFO msg="the write ledger remembers nothing, listing the jobs of the runs the cache file remembers again" runs=666
level=INFO msg="a store was added since the cache file was written, listing the jobs of the runs it remembers again" stores=influxdb,loki,postgres written_to=influxdb,loki runs=666

The others are how the configuration always runs, a store kept without a ledger or a run that ends with its sweep, and are said only at debug, as not every store keeps a write ledger.

Every start that finds the file says what it read, at its first sweep, and one that finds none says so at debug:

level=INFO msg="cache file read" file=/var/lib/ghchronicle/state-cache.bin written=2026-09-27T15:27:41+02:00 answers=1170 runs=672 refusals=100 page_sizes=54 age=1s

It is written at most every five minutes while the collector runs, and once more when it stops: at the end of a -once, and when the service is asked to stop with SIGTERM or SIGINT, which is what systemctl stop, docker stop and Ctrl-C send. A process killed outright cannot, so SIGKILL, the kernel’s out-of-memory killer, docker kill, taskkill /F and Stop-ScheduledTask lose what it learned since the last save, five minutes at most, and the next start asks those answers again. A save that fails says so, and so does a start that finds a file it cannot read:

level=WARN msg="cache file not saved" file=/var/lib/ghchronicle/state-cache.bin err="open /var/lib/ghchronicle/state-cache.bin.tmp: permission denied"
level=WARN msg="cache file not read, the first pass of each family pays in full" file=/var/lib/ghchronicle/state-cache.bin err="..."

A run with -card-only reads it and does not write it, for the reason it leaves the state file alone. A -backfill reads it and does not write it either: the pages it walks are pages no sweep asks for, and kept they would crowd the sweeps’ own answers out of the file. An upgrade keeps it: an answer is kept under the shape of what its collector reads, so the only answers asked for again in full are those of a collector that now reads something else.

It is bounded. An answer larger than a megabyte is left out, and so is whatever comes past 64 MB, the least recently asked for first. Measured against the author’s account of 37 repositories, one sweep of every family keeps 1,065 answers, 10.3 MB of them and 1.6 MB of file, none of them larger than 404 KB. Over days the cache in memory grows by the answers whose query carries a moving window: the author’s service held 6,695 after five and a half days, about 99 MB, and the 2,517 of them asked for in the last 48 hours, which are what the file keeps, come to about 31 MB. It holds what GitHub answered about private repositories as well as public ones, and is written with mode 600, as the state file is.

Deleting it costs one pass of each family at full price, the price every restart paid before it existed, and loses nothing. Every answer is asked for again, every refused feature is asked about again, which is also how to have a feature switched on today noticed before its day is out, the jobs of the newest runs are listed once more, and the first sweep runs totals before the pull requests, whatever its cadence says, so their pages are sized again. When totals was not due anyway, it says so:

level=INFO msg="no page sizes remembered, running totals before the pull requests it sizes"

A file that does not load in full, cut short, damaged or written in another format, is not read at all, with the warning above, and costs the same; the next save writes over it. It wants the same persistent path as the other two.

backfill:
since: 2y

Only applies to a run started with -backfill, and is overridden by -backfill-since. See backfill.

A backfill keeps one more file beside state_file, <name>-progress.json, which is where it records what it has already written so a stop costs one repository rather than the walk. There is nothing to configure: it appears when a backfill starts, it is removed when the walk reaches the end, and a sweep neither writes nor reads it. See stopping one, and picking it up again.

migrate: auto

What a run that writes to the stores does, before its first sweep, about a store an earlier release left in a shape this one no longer writes. -migrate prints what each configured store holds of every such change; see Migrations.

ValueWhat a start does
autoThe default. Applies on its own a pending change that loses nothing, says so at WARN, reads back what it cleared, and what an earlier run cleared and did not finish reading, and warns about every other one at every start
warnApplies nothing and reads nothing back, and warns about every pending change, and every store still owed its history, at every start

A change loses nothing when all three hold: GitHub still serves the whole history, so reading it again brings back every row the store holds; the old rows are set aside for at least 24 hours rather than deleted, which InfluxDB 3, PostgreSQL and Elasticsearch do; and every row in the store is this configuration’s, which the store is asked. There is no value that applies the rest on its own. A one-shot run, -once or -backfill, on a new state file applies nothing on its own either: that is every run of the Action without a state file restored, whose record of a refill still owed goes with its runner. Each warning names the two commands, -migrate for the plan and -migrate -yes to apply it, with the service stopped. Any other value is refused at start-up.

The service also holds a lock beside the state file, <name>-lock, for as long as it runs, which is what keeps -migrate -yes, and -uninstall -yes of the data or the state, from changing the stores under it. There is nothing to configure: it follows state_file, the operating system lets go of it however the process ends, and the file left behind only names the last process that held it.

Configuration is validated before the first call is made, and the messages name the key and what it needs.

MessageMeans
github.token is empty and GITHUB_TOKEN is unsetExactly what it says
state_file: "${X}/state.json" names ${X}, which is unset, ...A path names a variable with no value: set it, or write the path out
targets: set at least one of user, orgs or reposNothing to collect
sinks: enable at least one of ...A run that collects and discards is almost never what anyone meant
every.families.<name>: unknown collectorThe name is not a family. The message lists the ones that exist
sinks.influxdb: url and bucket are requiredEach sink validates its own required keys and says which
sinks.sql.dialect: "mysql" is not postgres, the only dialect so farThe value is not one of the accepted ones, and the message lists them
migrate: "off" is not auto or warnThe same, for migrate
Written and maintained by
MIT licenceRelease history