Skip to content

Upgrading

Install the new release the way the old one was installed, over it, and start the collector again. The configuration, the state file and the write ledger carry over as they are, and nothing has to be migrated by hand. What follows is what the first hours after the upgrade look like, so that none of it reads as a fault, and the few things worth doing about them.

The dashboards are generated from the binary, so publish them again after an upgrade, with -publish-dashboard or grafana.publish_on_start: see letting the binary do it.

A release sometimes changes what a stored row is keyed by. A tag that becomes a field is the usual case: every row written before the change is a series the rows written after it are not, so the store keeps both for ever, and a count over them grows by one for every item read on both sides of the change. The binary carries the list of every such change a store written by a 2.x release can still hold, each under an ID that names the release and the measurement, such as 2.6.1/gh_discussion_comment/is_answer, and -migrate checks every configured store against it:

Terminal window
ghchronicle -config /etc/ghchronicle/config.yaml -migrate

It prints one block per store and one line per change, and changes nothing: the stores are asked questions, GitHub is asked for the repository list alone, and the state file is read and never written. A store that can be asked decides for itself; one that cannot is decided by what the state file remembers of it.

StoreWhat decides
InfluxDB 3Whether the old tag is a tag column of the live table, asked of the catalog. A table InfluxDB has set aside is listed under another name and is not asked
InfluxDB 2Whether any row carries the old tag. The tag keys stay listed after a delete, so they are not asked
PostgreSQLWhether any row holds a value in the old tag’s column, in the schema the sink writes to
ElasticsearchA count of the documents that carry the old tag. The mapping keeps a field after its documents are gone
SQL file, Graphite, TelegrafThe release the state file records as the first to write the store
Loki, Prometheus, OTLP, file, stdoutNothing: none of them keeps a row whose identity a release could change

Each line starts with what the check found:

  • not needed: the store holds nothing of the old shape.
  • pending: it holds the old shape, or, for a store that cannot be asked, may hold it. The lines under it say why, what bringing it along would do there, which families would read the measurement again and from when, and what would not come back: the rows of repositories the configuration no longer covers, and of families it has switched off. For the comments, those are other people’s comments on such a repository: the account’s own come back through outbound wherever they are. The last line says whether applying it needs nobody’s word: GitHub still serves the whole history, so every row the store holds would come back, the old rows would be set aside for at least 24 hours rather than deleted, and every row in the store is this configuration’s. A store shared with another collector, one holding rows the refill would not bring back, one whose rows could not be compared with the configuration, an InfluxDB 2, whose only way is a delete, and every store that cannot be asked need somebody’s word, and the line says which reason applies. A row that names no account belongs to a configuration of organisations alone, which writes no user: it is this configuration’s only when this one has no targets.user either.
  • note: the rows are there and nothing can put them right, so the plan says what they mean and changes nothing.
  • frozen: nothing this configuration runs writes the measurement any more, so its rows are history and are left as they are.
  • applied: the state file records the change as applied to this store, and the store, where it can be asked, agrees.
  • unreachable: the store did not answer, so nothing about it is known.

Every run that writes to the stores, the service, -once, -backfill and a card drawn beside the stores, checks them the same way before its first sweep, and what it does next is the migrate setting’s:

  • auto, the default, applies on its own every pending change marked safe to apply unattended, and nothing else. It says so first, at WARN, with what it found, why the change exists and what it does in that store, including where the old rows are set aside; then it reads the measurement again from GitHub and only then sweeps.
  • warn applies nothing.

A one-shot run, -once or -backfill, on a new state file applies nothing on its own under auto either, and warns instead. That is every run of the Action that does not restore its state file with actions/cache, whose state file goes with its runner, and with it the record that a refill is still owed: a refill that failed, or a job cancelled half way, would leave the store cleared and nothing anywhere saying so. The next run on a host, which has a state file by then, or -migrate -yes, applies it.

A pending change a start does not apply is a WARN at every start, named by store, with the reason it was left and the two commands, copied from the configuration the run was given. One line in the log, wrapped here:

level=WARN msg="migration pending" sink=graphite measurement=gh_discussion_comment
migration=2.6.1/gh_discussion_comment/is_answer why="is_answer was a tag, ..."
not_applied="only whoever runs the Graphite host can remove its files"
plan="ghchronicle -config /etc/ghchronicle/config.yaml -migrate"
apply="ghchronicle -config /etc/ghchronicle/config.yaml -migrate -yes"
first="stop this service: -migrate -yes refuses to run beside it"

There is no setting that applies the rest on its own: a change that could lose rows waits for somebody to read the plan and give the word. A note is said once at INFO and then at DEBUG, since nothing will ever change it, and so is a change whose measurement nothing the configuration runs writes any more. A store that does not answer within 30 seconds is a WARN too, and the sweep goes on without waiting longer.

What a start finds not needed, and what it has applied, it records, and the next start takes the record’s word for it: once every change is settled for a store, a start does not ask that store anything. -migrate does not take that word and asks again every time.

Terminal window
systemctl stop ghchronicle
ghchronicle -config /etc/ghchronicle/config.yaml -migrate -yes
systemctl start ghchronicle

Run it as the user the service runs as, with the environment the service has: it saves the state file and the files beside it, and a file root saves is one the service can no longer read. A run that finds the state file there and cannot read it, or cannot parse it, stops and names it rather than starting from a new one, which would forget a refill still owed: give the file back to the service’s user. systemd and Docker show how.

Pause a cron job or a timer that runs -once as well, and let a -backfill that is running finish first. Neither holds the lock while it sweeps, and a SQL file two processes write at once is not one psql can replay. Their state file is safe either way: a run that saves it keeps what another process recorded of the stores since it read it, so a -once that ran across -migrate -yes does not put back the record it read before.

-migrate -yes prints the same plan and then applies every pending change, the unsafe ones too, since -yes is the word the plan asked for. Each one is recorded in the state file as it is applied, so a stop half way costs nothing the next run cannot pick up: what was applied is not done twice. A line under the plan says what happened to each:

  • applied: done. Where the old rows are kept, it names them; where somebody else has to act, on a Graphite host or behind a Telegraf, it prints the commands to run there.
  • failed: the store refused, with its reason. The others still go ahead.
  • held back: the store holds rows of accounts this configuration does not collect, or whose rows it holds could not be compared with this configuration, and the line says which and why. Set aside, another’s rows come back only when whoever collects them reads them again, so this takes -migrate-others as well as -yes.
  • unreachable: the store did not answer, so nothing was done there.
  • refill: what was read back from GitHub, from when and into which stores, or why reading it back did not finish: see reading the history back.
  • reconciled: each copy of the old rows compared with what came back, and the items GitHub no longer serves.

It exits 0 when everything pending was applied and read again, and 1 when anything was left, with the sentence that says how to go on. It refuses before it changes anything, and exits 1, with no GitHub token or no repository list, since nothing it clears could be read again, and while another process holds the state file: the service holds it for as long as it runs, beside the state file as <name>-lock, and -migrate -yes names that process rather than changing the stores it writes. A service started while -migrate -yes runs waits for it to finish, and a second service on the same state file is refused.

In the GitHub Action, mode: migrate runs -migrate -yes: see the inputs.

StoreWhat applying does
InfluxDB 3Deletes the one table, which InfluxDB keeps as <measurement>-<instant>, queryable, and purges itself 72 hours later, or never before 3.2
InfluxDB 2Deletes every row of the measurement in the bucket. Nothing is kept, so a start never does it on its own
PostgreSQLRenames the table <measurement>-<instant> in the sink’s schema; ghchronicle drops it 24 hours later
ElasticsearchBlocks writes to the index, clones it to <index>-<instant> and deletes it; ghchronicle deletes the clone 24 hours later
SQL fileWrites DROP TABLE IF EXISTS for the measurement into the file, ahead of the rows written after it
GraphitePrints the commands that remove the old paths on the Graphite host
TelegrafSays what to do in the store behind it

Each store touches that one measurement and nothing else: the table, index or paths named exactly, in the database, bucket, schema or prefix the sink writes to. A measurement that looks like it, another collector’s tables and another prefix’s indices are never reached. A store that can be asked is only changed once it has been asked and found holding the old shape, and a dry run sends it nothing but questions.

Where a store keeps the old rows aside, the plan names each copy under the store, with a kept aside line saying when it was set aside and when it goes, and the state file keeps it until then. ghchronicle purges its own copies once they have been kept 24 hours: the service after a sweep, any run at its next start, and -migrate -yes whenever it runs. A run on a new state file, every run of the Action that does not restore one among them, asks PostgreSQL and Elasticsearch for copies named the way a migration names them, since its state file cannot name them. InfluxDB 3 purges its own on a schedule of its own, which is read back from the system table of its _internal database: 72 hours after the delete by default (measured on 3.2.1 to 3.11.5), and it keeps the name in its catalog for its delete grace period after that, 24 hours by default. A server before 3.2 never purges one: the plan says so before applying, and -migrate asks the server for every copy it keeps that way, with the request that removes it once the server runs a release from 3.2 to 3.9 (see before 3.2 the copy stays). Until a copy goes, undoing the change is on each store’s page.

A store that was cleared is written again whole. The write ledger forgets the measurement in that store alone, every other measurement and every other store keeping what it remembers, and the cache file forgets what it claims about the families that write the measurement, their refusals and, when actions is among them, the workflow runs whose jobs it had written, so the sweeps after it ask everything a first sweep asks.

A store that was cleared holds none of the measurement’s history until it is read again from GitHub. Right after applying, the same run reads it back with one backfill of its own, the refill: the families that write the measurement, discussions and outbound for the comments, security for the alert items, and no other; writing that measurement alone; into the stores that were cleared and no other.

refill read discussions and outbound again, since 2023-11-14, writing gh_discussion_comment to
elasticsearch, influxdb and postgres
  • Only what was cleared. A family asks GitHub everything it always asks, and only the measurement cleared is written; the rest of what the family collects is in the stores already. Each store gets what was cleared in that store, so a store still holding a measurement in its old shape is never handed that measurement’s history beside it.
  • Only where it was cleared. InfluxDB, PostgreSQL, Elasticsearch, the SQL file, after its DROP, and Graphite, at the new depth. Never Loki, the Prometheus exporter, OTLP, the file sink or stdout, which were not cleared and would hold every row of the refill a second time, nor a Telegraf, whose store is somebody else’s: the plan names the -backfill -families that sends the history through it once that store has been cleared.
  • As far back as the store held. The day of the oldest row the store held, read before it was cleared: a bound later than that would lose the difference for good once the copy is purged. A store that cannot say how far back its rows go, the SQL file, Graphite, or an InfluxDB 3 Core that would not count them (below), is read back with no bound: what applying takes there is every row whatever its date, and backfill.since bounds what a backfill reaches, not what a store holds. One refill for several stores reads back as far as the furthest.
  • Waiting, as a backfill waits. It has a GitHub client of its own that waits for a spent rate limit to turn over, where a sweep skips, and -backfill-retry goes back for what it leaves as it does for a backfill. The containerised suite’s whole -migrate -yes, clearing InfluxDB 3.11.2, PostgreSQL 18.6 and Elasticsearch 9.5.3 and reading the comments back from the fake GitHub, takes about 2 seconds; on the author’s account, the outbound walk it needs took 33 seconds when 2.6.1’s was done by hand.

The refill is owed before a store is touched, and the state file records it then, under the store’s refill key: a store cleared and not read back looks, to anyone who asks it, exactly like one that never held the old shape, and only that record says the history is still to come. A clear that failed and left the store as it was, which the store says, takes the debt back; one whose answer did not arrive, a proxy’s 502 or a timeout, is looked at again, and when the store cannot say whether it was carried out the refill stays owed, which at worst reads the history once for nothing. It goes when the refill reaches the end of every family. A refill cut short, by a stop, a store that refused a write or GitHub not answering, keeps a checkpoint of its own beside the state file, named with -refill.json, apart from a backfill’s so neither refuses or overwrites the other, and is resumed from there: by -migrate -yes run again, even with nothing left to apply, and by any start under migrate: auto. Under migrate: warn a start says it is owed, at every start, with the command. -backfill-status prints how far it has got, and -migrate lists it under its store as refill owed.

Under migrate: auto the service applies what is safe and reads it back after its sinks are built and before its first sweep, so its sweeps start when the refill ends. The Prometheus exporter is up in that time and holds nothing until the first sweep, as after any restart. A service stopped during the refill exits at once, keeping what it wrote, and its next start carries on. The write ledger remembers what the refill wrote, so the first sweep after it does not send those rows again. Under migrate: warn nothing is read back and the service sweeps as it always did, writing this release’s shape beside the old one, and -migrate -yes refuses to run beside it.

A refill brings back what GitHub still serves for the targets configured now. When the refill ends, each store that kept a copy of the old rows and can be asked, InfluxDB 3, PostgreSQL and Elasticsearch, compares the copy with the table, item by item: the comment of each comment, the full_name and number of each alert. What the copy holds and the table does not is what GitHub no longer served, a repository deleted or no longer covered, a comment deleted, an alert whose feature was switched off; the report and the log name a few, and those rows stay only in the copy until it is purged. From the containerised suite, whose seeded comment the fake GitHub does not serve:

reconciled gh_discussion_comment in postgres: 1 item in gh_discussion_comment-20260928T234418, 6 now; 1 GitHub
no longer serves, whose rows are only in the copy until it is purged: 1

Carrying them over is left to the reader, from the copy, while it is there: for an item whose old shape was two rows at one instant, which of the two was right is not something a program can tell offline. Measured read-only against production’s InfluxDB 3.11.5, the copy InfluxDB kept of the table dropped by hand for 2.6.1 held 114 comments, and the table read back 121, none of the 114 missing.

InfluxDB 3 Core refuses a query that would open more Parquet files than its --query-file-limit, 432 by default, and a table a sweep writes every ten minutes is past that in days. The check reads the catalog, which opens no file, so it still finds the old shape in such a table, and a table with no old tag is not counted at all. What it cannot do there is count the rows, which leaves the refill with no bound, nor read whose rows the table holds, which leaves the change needing your word and -migrate -yes holding it back for -migrate-others; the plan quotes the server’s refusal. Measured on 3.11.2 with the limit lowered to 3: before this, the same table was unreachable at every start and -migrate -yes could not apply it at all. Raising the limit on the server, for the length of the migration, lets it answer everything.

What the state file remembers of each store

Section titled “What the state file remembers of each store”

The stores key of the state file keeps, for every store a run writes to, where it points, the release that first wrote it and the release that last did. A first start, on a state file with no history, records the running release as the first writer of every store, so a fresh install has nothing to migrate. The first start after an upgrade from 2.6.1 or earlier records the first writer as unknown, which for a SQL file, a Graphite or a Telegraf reads as older than any change, so those show each change of a 2.x release as pending, or as a note where nothing can put it right, until it is applied. A sink pointed at another store starts a new record, and the one it leaves is kept while it still owes a refill or names a copy there, said at every start and by -migrate as owed there, and taken back if the sink points there again. A URL written another way, with a trailing slash, a capital or the default port, is the same store. A state file deleted after an upgrade takes the record with it, and the stores that cannot be asked are then taken to be the running release’s; a refill still owed goes with it too, and nothing reads that history back until a -backfill -families of the families it names does.

A change made before 1.0.0, the state and reason tags of the alert items, was never written by a release, so only a store that can be asked can show it.

2.6.0 keeps what a sweep learned about GitHub in a cache file beside the state file, and 2.5.x wrote none. So the first start finds nothing to read, says so at debug as no cache file yet, the first pass of each family pays in full, and asks everything once without an ETag. That is the price every 2.5.x restart paid: measured on the author’s service on 2026-09-26, the first 38 minutes after a restart spent 1,092 charged core requests on passes that cost about 66 with the cache warm. From the second start on the file is there, and the log says cache file read.

achievements walks the whole merged history, once

Section titled “achievements walks the whole merged history, once”

The co-authored pull request count behind Pair Extraordinaire is kept in the state file from 2.6.0 on, and a 2.5.x state file holds none. The first achievements pass therefore walks every pull request the account has merged: 35 queries and 23.7 MB over 2,315 of them, measured on 2026-09-27. After that a pass walks the days since the count it keeps, and the whole history again once a week.

The pull request query is sized per repository from the counts totals reads, and 2.6.0 keeps those sizes in the cache file. The first sweep has none, so it runs totals before the pull requests whatever its cadence says. At its hourly cadence totals is usually due anyway after an upgrade, and then it simply runs first; only when it is not due, and the sweep is not a primed one, which runs every family, does the log say why it runs:

level=INFO msg="no page sizes remembered, running totals before the pull requests it sizes"

A 2.5.x state file has the account’s daily families marked at one instant and its twelve-hour ones at another, because that is when they ran, together. 2.6.0 starts at most one family of six hours or more a sweep, so the first day after the upgrade brings them in one a tick: over up to two and a half hours at the built-in cadences, and longer with deps, history or joblogs switched on. The log names who starts and who waits, under slow families due together take turns, and from then on each keeps its own time. See the slow families take turns.

The search behind gh_account_total.commented_elsewhere used to count the threads the account commented on in its own repositories too. From 2.5.2, whose changes shipped in 2.6.0, it leaves them out, as its two siblings already did, and the field keeps its name. So every store’s series drops by the difference on the first totals sweep after the upgrade, from 125 to 55 on the account it was measured on. It is a correction, not a loss.

2.6.0 adds fields to measurements whose tables an earlier release made, and a table an earlier release made has no column for them. The PostgreSQL sink that connects reads each table’s columns the first time its process meets it and adds the missing ones with ALTER TABLE ... ADD COLUMN IF NOT EXISTS. The SQL file sink cannot ask, so its file carries that statement for every field, and replaying it into a database an earlier release filled makes psql print a notice for every table and every column that is already there. They are expected:

NOTICE: relation "gh_repo" already exists, skipping
NOTICE: column "stars" of relation "gh_repo" already exists, skipping

See how the declarations arrive.

The Stars column of Work elsewhere waits for outbound

Section titled “The Stars column of Work elsewhere waits for outbound”

The “Work elsewhere” table joins gh_upstream_repo, a measurement 2.6.0 adds, for the stars of each repository. Until the first outbound pass of the new version writes it, which is within the hour, InfluxDB and PostgreSQL refuse the whole table rather than leave the column empty, and a range that ends before the upgrade has no stars to show in any store. Wait for that pass, or publish the dashboards after it; see the data looks wrong.

A configuration copied from an older example keeps the old cadences

Section titled “A configuration copied from an older example keeps the old cadences”

Up to 2.6.0 the example configuration set every family under every.families to the built-in value of its release, so a config.yaml copied from it sets them all. One copied from 2.5.x keeps twelve of them slower than 2.6.0 runs them: account, outbound and totals at 12h, achievements at 24h, stars, billing and analyses at 6h, discussions at 2h, deployments at 1h, and activity, events and notifs at 30m. Nothing warns, since the warning is for a cadence four times faster than the built-in one. Deleting those lines is what picks up the current values; from 2.6.1 the example shows them commented out, so a copy of it sets none. See every family, its group and its built-in cadence.

Whether a comment is the accepted answer is a field

Section titled “Whether a comment is the accepted answer is a field”

gh_discussion_comment carried is_answer as a tag, and a maintainer accepts an answer days after the comment was written, so one comment read before and after that was two rows at the same instant. 2.6.1 writes no is_answer at all: the answers field every row already carried, 1 for the accepted answer and 0 for any other comment, says the same thing. A store written before 2.6.1 and since holds the measurement in two shapes until the measurement is dropped and filled again with a backfill. The dashboards read both shapes, one row per comment, and the measurements page says what each store keeps and how to drop it. -migrate says which of the configured stores still hold the old shape, under 2.6.1/gh_discussion_comment/is_answer: see Migrations.

What the old shape still costs a reader is an answer accepted and taken back since: its old row says accepted, so both comment panels read it accepted until the store is brought along. Measured in the containerised suite with an answer taken back on the fake GitHub, in InfluxDB 3.11.2, PostgreSQL 18.6, Elasticsearch 9.5.3, Graphite 1.1.10-5 and the SQL file replayed into PostgreSQL: Answers elsewhere and Discussion answers read it accepted, one row for the comment, before the upgrade in every store; not accepted, still one row, in the first three after a start under migrate: auto, which leaves Graphite and the SQL file to their operator; and not accepted in all five after -migrate -yes, once the Graphite commands had run and the file had been replayed from where it was.

state_file, sinks.dedupe_file, log.file, sinks.file.path and sinks.sql.path take a leading ~ as the home directory and a ${VAR} from the environment, as credentials and addresses always did. A configuration that wrote either one used to get a directory named with those characters, under the working directory, and now gets the path it meant. Where that was the state file, the state the old release kept is in the directory with the odd name, and the new one starts without it: move the files across before the first start if the walks it saves are worth keeping. A ${VAR} in a path that is unset or empty stops the start and names the key, rather than leaving the path without that part. See ${VAR} expansion.

Up to 2.6.0 the image had no /var/lib/ghchronicle, so the volume a compose stack or a docker run mounted there was created owned by root, and the collector, uid 65532, saved neither its state nor its cache in it until the volume was handed over with a chown. From 2.6.1 the image carries the directory, owned by uid 65532, and Docker gives a new volume mounted there that owner. An empty volume an earlier image left to root is handed over the same way the first time a container of 2.6.1 is created on it, which docker compose pull followed by docker compose up -d does; one already handed over is left as it is. A host directory mounted there still needs its chown. See what has to be writable.

achievements walks the whole merged history once more

Section titled “achievements walks the whole merged history once more”

The co-authored count behind Pair Extraordinaire now reads every commit of a pull request with more than a hundred, where 2.6.0 read the first hundred and counted the rest as a floor. That is a change to the rule the count is kept by, so the count a 2.6.0 state file holds is set aside and the first achievements pass after the upgrade walks every pull request the account has merged: 41 queries with the counts query, 41 points, in 1 minute 53 seconds over the account measured, on 2026-09-28. A co-authored pull request count is a floor warning that came back on every pass stops with it, unless a pull request’s remaining commits really could not be read. See gh_achievement_progress.

Written and maintained by
MIT licenceRelease history