Upgrading
Install the new release the way the old one was installed, over it, and start the collector again. The configuration, the state file and the write ledger carry over as they are, and nothing has to be migrated by hand. What follows is what the first hours after the upgrade look like, so that none of it reads as a fault, and the few things worth doing about them.
The dashboards are generated from the binary, so publish them again after an
upgrade, with -publish-dashboard or grafana.publish_on_start: see letting
the binary do it.
Migrations
Section titled “Migrations”A release sometimes changes what a stored row is keyed by. A tag that becomes a
field is the usual case: every row written before the change is a series the
rows written after it are not, so the store keeps both for ever, and a count
over them grows by one for every item read on both sides of the change. The
binary carries the list of every such change a store written by a 2.x release
can still hold, each under an ID that names the release and the measurement,
such as 2.6.1/gh_discussion_comment/is_answer, and -migrate checks every
configured store against it:
ghchronicle -config /etc/ghchronicle/config.yaml -migrateIt prints one block per store and one line per change, and changes nothing: the stores are asked questions, GitHub is asked for the repository list alone, and the state file is read and never written. A store that can be asked decides for itself; one that cannot is decided by what the state file remembers of it.
| Store | What decides |
|---|---|
| InfluxDB 3 | Whether the old tag is a tag column of the live table, asked of the catalog. A table InfluxDB has set aside is listed under another name and is not asked |
| InfluxDB 2 | Whether any row carries the old tag. The tag keys stay listed after a delete, so they are not asked |
| PostgreSQL | Whether any row holds a value in the old tag’s column, in the schema the sink writes to |
| Elasticsearch | A count of the documents that carry the old tag. The mapping keeps a field after its documents are gone |
| SQL file, Graphite, Telegraf | The release the state file records as the first to write the store |
| Loki, Prometheus, OTLP, file, stdout | Nothing: none of them keeps a row whose identity a release could change |
Each line starts with what the check found:
not needed: the store holds nothing of the old shape.pending: it holds the old shape, or, for a store that cannot be asked, may hold it. The lines under it say why, what bringing it along would do there, which families would read the measurement again and from when, and what would not come back: the rows of repositories the configuration no longer covers, and of families it has switched off. For the comments, those are other people’s comments on such a repository: the account’s own come back throughoutboundwherever they are. The last line says whether applying it needs nobody’s word: GitHub still serves the whole history, so every row the store holds would come back, the old rows would be set aside for at least 24 hours rather than deleted, and every row in the store is this configuration’s. A store shared with another collector, one holding rows the refill would not bring back, one whose rows could not be compared with the configuration, an InfluxDB 2, whose only way is a delete, and every store that cannot be asked need somebody’s word, and the line says which reason applies. A row that names no account belongs to a configuration of organisations alone, which writes no user: it is this configuration’s only when this one has notargets.usereither.note: the rows are there and nothing can put them right, so the plan says what they mean and changes nothing.frozen: nothing this configuration runs writes the measurement any more, so its rows are history and are left as they are.applied: the state file records the change as applied to this store, and the store, where it can be asked, agrees.unreachable: the store did not answer, so nothing about it is known.
What a start does about it
Section titled “What a start does about it”Every run that writes to the stores, the service, -once, -backfill and a
card drawn beside the stores, checks them the same way before its first sweep,
and what it does next is the migrate
setting’s:
auto, the default, applies on its own every pending change marked safe to apply unattended, and nothing else. It says so first, atWARN, with what it found, why the change exists and what it does in that store, including where the old rows are set aside; then it reads the measurement again from GitHub and only then sweeps.warnapplies nothing.
A one-shot run, -once or -backfill, on a new state file applies nothing on
its own under auto either, and warns instead. That is every run of the Action
that does not restore its state file with actions/cache, whose state file goes
with its runner, and with it the record that a refill is still owed: a refill
that failed, or a job cancelled half way, would leave the store cleared and
nothing anywhere saying so. The next run on a host, which has a state file by
then, or -migrate -yes, applies it.
A pending change a start does not apply is a WARN at every start, named by
store, with the reason it was left and the two commands, copied from the
configuration the run was given. One line in the log, wrapped here:
level=WARN msg="migration pending" sink=graphite measurement=gh_discussion_comment migration=2.6.1/gh_discussion_comment/is_answer why="is_answer was a tag, ..." not_applied="only whoever runs the Graphite host can remove its files" plan="ghchronicle -config /etc/ghchronicle/config.yaml -migrate" apply="ghchronicle -config /etc/ghchronicle/config.yaml -migrate -yes" first="stop this service: -migrate -yes refuses to run beside it"There is no setting that applies the rest on its own: a change that could lose
rows waits for somebody to read the plan and give the word. A note is said once
at INFO and then at DEBUG, since nothing will ever change it, and so is a
change whose measurement nothing the configuration runs writes any more. A
store that does not answer within 30 seconds is a WARN too, and the sweep
goes on without waiting longer.
What a start finds not needed, and what it has applied, it records, and the
next start takes the record’s word for it: once every change is settled for a
store, a start does not ask that store anything. -migrate does not take that
word and asks again every time.
Applying it by hand
Section titled “Applying it by hand”systemctl stop ghchronicleghchronicle -config /etc/ghchronicle/config.yaml -migrate -yessystemctl start ghchronicleRun it as the user the service runs as, with the environment the service has: it saves the state file and the files beside it, and a file root saves is one the service can no longer read. A run that finds the state file there and cannot read it, or cannot parse it, stops and names it rather than starting from a new one, which would forget a refill still owed: give the file back to the service’s user. systemd and Docker show how.
Pause a cron job or a timer that runs -once as well, and let a -backfill
that is running finish first. Neither holds the lock while it sweeps, and a SQL
file two processes write at once is not one psql can replay. Their state file
is safe either way: a run that saves it keeps what another process recorded of
the stores since it read it, so a -once that ran across -migrate -yes does
not put back the record it read before.
-migrate -yes prints the same plan and then applies every pending change, the
unsafe ones too, since -yes is the word the plan asked for. Each one is
recorded in the state file as it is applied, so a stop half way costs nothing
the next run cannot pick up: what was applied is not done twice. A line under
the plan says what happened to each:
applied: done. Where the old rows are kept, it names them; where somebody else has to act, on a Graphite host or behind a Telegraf, it prints the commands to run there.failed: the store refused, with its reason. The others still go ahead.held back: the store holds rows of accounts this configuration does not collect, or whose rows it holds could not be compared with this configuration, and the line says which and why. Set aside, another’s rows come back only when whoever collects them reads them again, so this takes-migrate-othersas well as-yes.unreachable: the store did not answer, so nothing was done there.refill: what was read back from GitHub, from when and into which stores, or why reading it back did not finish: see reading the history back.reconciled: each copy of the old rows compared with what came back, and the items GitHub no longer serves.
It exits 0 when everything pending was applied and read again, and 1 when
anything was left, with the sentence that says how to go on. It refuses before
it changes anything, and exits 1, with no GitHub token or no repository list,
since nothing it clears could be read again, and while another process holds
the state file: the service holds it for as long as it runs, beside the state
file as <name>-lock, and -migrate -yes names that process rather than
changing the stores it writes. A service started while -migrate -yes runs
waits for it to finish, and a second service on the same state file is refused.
In the GitHub Action, mode: migrate runs -migrate -yes: see the
inputs.
What applying does in each store
Section titled “What applying does in each store”| Store | What applying does |
|---|---|
| InfluxDB 3 | Deletes the one table, which InfluxDB keeps as <measurement>-<instant>, queryable, and purges itself 72 hours later, or never before 3.2 |
| InfluxDB 2 | Deletes every row of the measurement in the bucket. Nothing is kept, so a start never does it on its own |
| PostgreSQL | Renames the table <measurement>-<instant> in the sink’s schema; ghchronicle drops it 24 hours later |
| Elasticsearch | Blocks writes to the index, clones it to <index>-<instant> and deletes it; ghchronicle deletes the clone 24 hours later |
| SQL file | Writes DROP TABLE IF EXISTS for the measurement into the file, ahead of the rows written after it |
| Graphite | Prints the commands that remove the old paths on the Graphite host |
| Telegraf | Says what to do in the store behind it |
Each store touches that one measurement and nothing else: the table, index or paths named exactly, in the database, bucket, schema or prefix the sink writes to. A measurement that looks like it, another collector’s tables and another prefix’s indices are never reached. A store that can be asked is only changed once it has been asked and found holding the old shape, and a dry run sends it nothing but questions.
Where a store keeps the old rows aside, the plan names each copy under the
store, with a kept aside line saying when it was set aside and when it goes,
and the state file keeps it until then. ghchronicle purges its own copies once
they have been kept 24 hours: the service after a sweep, any run at its next
start, and -migrate -yes whenever it runs. A run on a new state file, every
run of the Action that does not restore one among them, asks PostgreSQL and
Elasticsearch for copies named the way a migration names them, since its state
file cannot name them. InfluxDB 3 purges its own on a schedule of its own, which
is read back from the system table of its _internal database: 72 hours after
the delete by default (measured on 3.2.1 to 3.11.5), and it keeps the name in
its catalog for its delete grace period after that, 24 hours by default. A
server before 3.2 never purges one: the plan says so before applying, and
-migrate asks the server for every copy it keeps that way, with the request
that removes it once the server runs a release from 3.2 to 3.9 (see before 3.2
the copy stays). Until a
copy goes, undoing the change is on each store’s page.
A store that was cleared is written again whole. The write
ledger forgets the
measurement in that store alone, every other measurement and every other store
keeping what it remembers, and the cache file
forgets what it claims about the families that write the measurement, their
refusals and, when actions is among them, the workflow runs whose jobs it had
written, so the sweeps after it ask everything a first sweep asks.
Reading the history back
Section titled “Reading the history back”A store that was cleared holds none of the measurement’s history until it is
read again from GitHub. Right after applying, the same run reads it back with
one backfill of its own, the refill: the families that write the measurement,
discussions and outbound for the comments, security for the alert items,
and no other; writing that measurement alone; into the stores that were cleared
and no other.
refill read discussions and outbound again, since 2023-11-14, writing gh_discussion_comment to elasticsearch, influxdb and postgres- Only what was cleared. A family asks GitHub everything it always asks, and only the measurement cleared is written; the rest of what the family collects is in the stores already. Each store gets what was cleared in that store, so a store still holding a measurement in its old shape is never handed that measurement’s history beside it.
- Only where it was cleared. InfluxDB, PostgreSQL, Elasticsearch, the SQL
file, after its
DROP, and Graphite, at the new depth. Never Loki, the Prometheus exporter, OTLP, the file sink or stdout, which were not cleared and would hold every row of the refill a second time, nor a Telegraf, whose store is somebody else’s: the plan names the-backfill -familiesthat sends the history through it once that store has been cleared. - As far back as the store held. The day of the oldest row the store held,
read before it was cleared: a bound later than that would lose the
difference for good once the copy is purged. A store that cannot say how far
back its rows go, the SQL file, Graphite, or an InfluxDB 3 Core that would
not count them (below), is read back with no bound: what applying takes
there is every row whatever its date, and
backfill.sincebounds what a backfill reaches, not what a store holds. One refill for several stores reads back as far as the furthest. - Waiting, as a backfill waits. It has a GitHub client of its own that
waits for a spent rate limit to turn over, where a sweep skips, and
-backfill-retrygoes back for what it leaves as it does for a backfill. The containerised suite’s whole-migrate -yes, clearing InfluxDB 3.11.2, PostgreSQL 18.6 and Elasticsearch 9.5.3 and reading the comments back from the fake GitHub, takes about 2 seconds; on the author’s account, theoutboundwalk it needs took 33 seconds when 2.6.1’s was done by hand.
The refill is owed before a store is touched, and the state file records it
then, under the store’s refill key: a store cleared and not read back looks,
to anyone who asks it, exactly like one that never held the old shape, and only
that record says the history is still to come. A clear that failed and left
the store as it was, which the store says, takes the debt back; one whose
answer did not arrive, a proxy’s 502 or a timeout, is looked at again, and when
the store cannot say whether it was carried out the refill stays owed, which at
worst reads the history once for nothing. It goes when the refill reaches the
end of every family. A refill cut
short, by a stop, a store that refused a write or GitHub not answering, keeps a
checkpoint of its own beside the state file, named with -refill.json, apart
from a backfill’s so neither refuses or overwrites the other, and is resumed
from there: by -migrate -yes run again, even with nothing left to apply, and
by any start under migrate: auto. Under migrate: warn a start says it is
owed, at every start, with the command. -backfill-status prints how far it has
got, and -migrate lists it under its store as refill owed.
What the service does meanwhile
Section titled “What the service does meanwhile”Under migrate: auto the service applies what is safe and reads it back after
its sinks are built and before its first sweep, so its sweeps start when the
refill ends. The Prometheus exporter is up in that time and holds
nothing until the first sweep, as after any restart. A service stopped during
the refill exits at once, keeping what it wrote, and its next start carries on.
The write ledger remembers what the refill wrote, so the first sweep after it
does not send those rows again. Under migrate: warn nothing is read back and
the service sweeps as it always did, writing this release’s shape beside the
old one, and -migrate -yes refuses to run beside it.
What GitHub no longer serves
Section titled “What GitHub no longer serves”A refill brings back what GitHub still serves for the targets configured now.
When the refill ends, each store that kept a copy of the old rows and can be
asked, InfluxDB 3, PostgreSQL and Elasticsearch, compares the copy with the
table, item by item: the comment of each comment, the full_name and
number of each alert. What the copy holds and the table does not is what
GitHub no longer served, a repository deleted or no longer covered, a comment
deleted, an alert whose feature was switched off; the report and the log name a
few, and those rows stay only in the copy until it is purged. From the
containerised suite, whose seeded comment the fake GitHub does not serve:
reconciled gh_discussion_comment in postgres: 1 item in gh_discussion_comment-20260928T234418, 6 now; 1 GitHub no longer serves, whose rows are only in the copy until it is purged: 1Carrying them over is left to the reader, from the copy, while it is there: for an item whose old shape was two rows at one instant, which of the two was right is not something a program can tell offline. Measured read-only against production’s InfluxDB 3.11.5, the copy InfluxDB kept of the table dropped by hand for 2.6.1 held 114 comments, and the table read back 121, none of the 114 missing.
InfluxDB 3 Core’s query file limit
Section titled “InfluxDB 3 Core’s query file limit”InfluxDB 3 Core refuses a query that would open more Parquet files than its
--query-file-limit, 432 by default, and a table a sweep writes every ten
minutes is past that in days. The check reads the catalog, which opens no file,
so it still finds the old shape in such a table, and a table with no old tag is
not counted at all. What it cannot do there is count the rows, which leaves the
refill with no bound, nor read whose rows the table holds, which leaves the
change needing your word and -migrate -yes holding it back for
-migrate-others; the plan quotes the server’s refusal. Measured on 3.11.2
with the limit lowered to 3: before this, the same table was unreachable at
every start and -migrate -yes could not apply it at all. Raising the limit on
the server, for the length of the migration, lets it answer everything.
What the state file remembers of each store
Section titled “What the state file remembers of each store”The stores key of the state file
keeps, for every store a run writes to, where it points, the release that first
wrote it and the release that last did. A first start, on a state file with no
history, records the running release as the first writer of every store, so a
fresh install has nothing to migrate. The first start after an upgrade from
2.6.1 or earlier records the first writer as unknown, which for a SQL file, a
Graphite or a Telegraf reads as older than any change, so those show each
change of a 2.x release as pending, or as a note where nothing can put it
right, until it is applied. A sink pointed at another store starts a new
record, and the one it leaves is kept while it still owes a refill or names a
copy there, said at every start and by -migrate as owed there, and taken
back if the sink points there again. A URL written another way, with a
trailing slash, a capital or the default port, is the same store. A state file
deleted after an upgrade takes the record with it, and the stores that cannot
be asked are then taken to be the running release’s; a refill still owed goes
with it too, and nothing reads that history back until a -backfill -families
of the families it names does.
A change made before 1.0.0, the state and reason tags of the alert items,
was never written by a release, so only a store that can be asked can show it.
From 2.5.x to 2.6.x
Section titled “From 2.5.x to 2.6.x”The first start pays in full, once
Section titled “The first start pays in full, once”2.6.0 keeps what a sweep learned about GitHub in a cache file beside the state
file, and 2.5.x wrote none. So
the first start finds nothing to read, says so at debug as no cache file yet, the first pass of each family pays in full, and asks everything once
without an ETag. That is the price every 2.5.x restart paid: measured on the
author’s service on 2026-09-26, the first 38 minutes after a restart spent
1,092 charged core requests on passes that cost about 66 with the cache warm.
From the second start on the file is there, and the log says cache file read.
achievements walks the whole merged history, once
Section titled “achievements walks the whole merged history, once”The co-authored pull request count behind Pair Extraordinaire is kept in the
state file from 2.6.0 on, and a 2.5.x state file holds none. The first
achievements pass therefore walks every pull request the account has merged:
35 queries and 23.7 MB over 2,315 of them, measured on 2026-09-27. After that a
pass walks the days since the count it keeps, and the whole history again once
a week.
totals runs first
Section titled “totals runs first”The pull request query is sized per repository from the counts totals reads,
and 2.6.0 keeps those sizes in the cache file. The first sweep has none, so it
runs totals before the pull requests whatever its cadence says. At its hourly
cadence totals is usually due anyway after an upgrade, and then it simply
runs first; only when it is not due, and the sweep is not a primed one, which
runs every family, does the log say why it runs:
level=INFO msg="no page sizes remembered, running totals before the pull requests it sizes"The slow families spread out
Section titled “The slow families spread out”A 2.5.x state file has the account’s daily families marked at one instant and
its twelve-hour ones at another, because that is when they ran, together. 2.6.0
starts at most one family of six hours or more a sweep, so the first day after
the upgrade brings them in one a tick: over up to two and a half hours at the
built-in cadences, and longer with deps, history or joblogs switched on.
The log names who starts and who waits, under slow families due together take turns, and from then on each keeps its own time. See the slow families take
turns.
commented_elsewhere drops
Section titled “commented_elsewhere drops”The search behind gh_account_total.commented_elsewhere used to count the
threads the account commented on in its own repositories too. From 2.5.2, whose
changes shipped in 2.6.0, it leaves them out, as its two siblings already did,
and the field keeps its name. So every store’s series drops by the difference
on the first totals sweep after the upgrade, from 125 to 55 on the account it
was measured on. It is a correction, not a loss.
PostgreSQL tables gain columns
Section titled “PostgreSQL tables gain columns”2.6.0 adds fields to measurements whose tables an earlier release made, and a
table an earlier release made has no column for them. The PostgreSQL sink that
connects reads each table’s columns the first time its process meets it and
adds the missing ones with ALTER TABLE ... ADD COLUMN IF NOT EXISTS. The SQL
file sink cannot ask, so its file carries that statement for every field, and
replaying it into a database an earlier release filled makes psql print a
notice for every table and every column that is already there. They are
expected:
NOTICE: relation "gh_repo" already exists, skippingNOTICE: column "stars" of relation "gh_repo" already exists, skippingSee how the declarations arrive.
The Stars column of Work elsewhere waits for outbound
Section titled “The Stars column of Work elsewhere waits for outbound”The “Work elsewhere” table joins gh_upstream_repo, a measurement 2.6.0 adds,
for the stars of each repository. Until the first outbound pass of the new
version writes it, which is within the hour, InfluxDB and PostgreSQL refuse the
whole table rather than leave the column empty, and a range that ends before the
upgrade has no stars to show in any store. Wait for that pass, or publish the
dashboards after it; see the data looks
wrong.
A configuration copied from an older example keeps the old cadences
Section titled “A configuration copied from an older example keeps the old cadences”Up to 2.6.0 the example configuration set every family under every.families
to the built-in value of its release, so a config.yaml copied from it sets
them all. One copied from 2.5.x keeps twelve of them slower than 2.6.0 runs
them: account, outbound and totals at 12h, achievements at 24h,
stars, billing and analyses at 6h, discussions at 2h, deployments
at 1h, and activity, events and notifs at 30m. Nothing warns, since
the warning is for a cadence four times faster than the built-in one. Deleting
those lines is what picks up the current values; from 2.6.1 the example shows
them commented out, so a copy of it sets none. See every family, its group and
its built-in
cadence.
To 2.6.1
Section titled “To 2.6.1”Whether a comment is the accepted answer is a field
Section titled “Whether a comment is the accepted answer is a field”gh_discussion_comment carried is_answer as a tag, and a maintainer accepts
an answer days after the comment was written, so one comment read before and
after that was two rows at the same instant. 2.6.1 writes no is_answer at
all: the answers field every row already carried, 1 for the accepted answer
and 0 for any other comment, says the same thing. A store written before 2.6.1
and since holds the measurement in two shapes until the measurement is dropped
and filled again with a backfill. The dashboards
read both shapes, one row per comment, and the measurements
page says what
each store keeps and how to drop it. -migrate says which of the configured
stores still hold the old shape, under
2.6.1/gh_discussion_comment/is_answer: see Migrations.
What the old shape still costs a reader is an answer accepted and taken back
since: its old row says accepted, so both comment panels read it accepted
until the store is brought along. Measured in the containerised suite with an
answer taken back on the fake GitHub, in InfluxDB 3.11.2, PostgreSQL 18.6,
Elasticsearch 9.5.3, Graphite 1.1.10-5 and the SQL file replayed into
PostgreSQL: Answers elsewhere and Discussion answers read it accepted, one row
for the comment, before the upgrade in every store; not accepted, still one
row, in the first three after a start under migrate: auto, which leaves
Graphite and the SQL file to their operator; and not accepted in all five after
-migrate -yes, once the Graphite commands had run and the file had been
replayed from where it was.
A path setting is expanded
Section titled “A path setting is expanded”state_file, sinks.dedupe_file, log.file, sinks.file.path and
sinks.sql.path take a leading ~ as the home directory and a ${VAR} from
the environment, as credentials and addresses always did. A configuration that
wrote either one used to get a directory named with those characters, under
the working directory, and now gets the path it meant. Where that was the
state file, the state the old release kept is in the directory with the odd
name, and the new one starts without it: move the files across before the
first start if the walks it saves are worth keeping. A ${VAR} in a path that
is unset or empty stops the start and names the key, rather than leaving the
path without that part. See ${VAR}
expansion.
A Docker state volume needs no chown
Section titled “A Docker state volume needs no chown”Up to 2.6.0 the image had no /var/lib/ghchronicle, so the volume a compose
stack or a docker run mounted there was created owned by root, and the
collector, uid 65532, saved neither its state nor its cache in it until the
volume was handed over with a chown. From 2.6.1 the image carries the
directory, owned by uid 65532, and Docker gives a new volume mounted there that
owner. An empty volume an earlier image left to root is handed over the same
way the first time a container of 2.6.1 is created on it, which
docker compose pull followed by docker compose up -d does; one already
handed over is left as it is. A host directory mounted there still needs its
chown. See what has to be
writable.
achievements walks the whole merged history once more
Section titled “achievements walks the whole merged history once more”The co-authored count behind Pair Extraordinaire now reads every commit of a
pull request with more than a hundred, where 2.6.0 read the first hundred and
counted the rest as a floor. That is a change to the rule the count is kept by,
so the count a 2.6.0 state file holds is set aside and the first achievements
pass after the upgrade walks every pull request the account has merged: 41
queries with the counts query, 41 points, in 1 minute 53 seconds over the
account measured, on 2026-09-28. A co-authored pull request count is a floor warning that came
back on every pass stops with it, unless a pull request’s remaining commits
really could not be read. See
gh_achievement_progress.