systemd
The unit
Section titled “The unit”[Unit]Description=ghchronicle, GitHub metrics collectorAfter=network-online.targetWants=network-online.target
[Service]Type=simpleUser=ghchronicleGroup=ghchronicleEnvironmentFile=/etc/ghchronicle/ghchronicle.envExecStart=/usr/local/bin/ghchronicle -config /etc/ghchronicle/config.yamlRestart=alwaysRestartSec=30s
# This is the only process on the host holding a GitHub token, so it gets# nothing it does not need.NoNewPrivileges=truePrivateTmp=truePrivateDevices=trueProtectSystem=strictProtectHome=trueProtectKernelTunables=trueProtectKernelModules=trueProtectControlGroups=trueProtectClock=trueProtectHostname=trueProtectProc=invisibleRestrictNamespaces=trueRestrictRealtime=trueRestrictSUIDSGID=trueLockPersonality=trueMemoryDenyWriteExecute=trueSystemCallArchitectures=nativeSystemCallFilter=@system-serviceCapabilityBoundingSet=AmbientCapabilities=RestrictAddressFamilies=AF_INET AF_INET6
StateDirectory=ghchronicleReadWritePaths=/var/lib/ghchronicle
[Install]WantedBy=multi-user.targetWhy it is hardened
Section titled “Why it is hardened”The threat model is short and it is the whole justification: this is very likely the only process on the host holding a GitHub token with read access to every repository of an account. A token is a bearer credential. Anything that can read this process’s memory or its environment file has the account.
So the unit gives the process exactly what it needs, which turns out to be almost nothing: an outbound TCP socket and one writable directory.
| Directive | What it removes |
|---|---|
CapabilityBoundingSet=, AmbientCapabilities= | Every Linux capability. It binds no privileged port and owns no device |
No | Any path to gaining privileges through exec, including setuid binaries |
Protect | Write access to the entire filesystem, except ReadWritePaths |
Protect | Every home directory, which is where the interesting credentials on a host usually live |
PrivateTmp=true, PrivateDevices=true | Shared temporary files, and the physical device nodes |
Protect | The ability to see other users’ processes in /proc, so it cannot read another process’s command line |
Restrict | Unix and netlink sockets. It talks HTTPS and nothing else |
MemoryDenyWriteExecute=true, LockPersonality=true | The usual shellcode primitives |
System | Every syscall outside the ordinary service set, including the module and kernel-tuning ones |
ProtectKernelTunables, ProtectKernelModules, ProtectControlGroups, ProtectClock, ProtectHostname, RestrictNamespaces, RestrictRealtime, RestrictSUIDSGID | Every remaining route to changing the host from inside the service |
StateDirectory=ghchronicle makes systemd create /var/lib/ghchronicle with
the right ownership on start, so the state file has somewhere to live without a
manual mkdir and a chown that someone will forget after a reinstall.
Four files live there, not one. Beside state.json the sweep keeps its write
ledger, state-written.bin by default, which is what stops an unchanged point
being written again, and its cache, state-cache.bin, which is what lets a
restart ask GitHub only for what changed; and the service holds state-lock
for as long as it runs, which is how -migrate -yes knows not to change the
stores under it. A backfill, or the reading back of a
migration, that stops half way
leaves its checkpoint there too. ReadWritePaths covers the directory, so all
of them are already allowed. Put the state file or the ledger
somewhere else and that path needs adding here, and the cache follows the state
file wherever it goes. Losing the ledger costs one sweep of rewriting:
only what changed is written;
losing the cache, one sweep at full price:
the cache beside it.
Installing it
Section titled “Installing it”-
Create the user and the directories.
Terminal window sudo useradd --system --no-create-home --shell /usr/sbin/nologin ghchroniclesudo mkdir -p /etc/ghchronicle -
Put the configuration in place.
Directory/etc/ghchronicle/
- config.yaml world readable, no secrets in it
- ghchronicle.env mode 600, the tokens
Directory/var/lib/ghchronicle/
- state.json created by the service
- state-written.bin the write ledger, beside it
- state-cache.bin the cache, beside it too
- state-lock held by the service while it runs
-
Write the environment file, and nothing else in it.
/etc/ghchronicle/ghchronicle.env GITHUB_TOKEN=github_pat_...INFLUX_TOKEN=...Terminal window sudo chmod 600 /etc/ghchronicle/ghchronicle.envEverything in
config.yamlreads these through${VAR}, which is what lets the config be world readable and version controlled while the secrets are not. -
Start it.
Terminal window sudo systemctl daemon-reloadsudo systemctl enable --now ghchroniclesudo systemctl status ghchronicleThe first sweep runs every family, since none has run yet, except that the service starts at most one family of six hours or more a sweep:
traffic,stats,forksand the rest of the eleven reach the store over the first two and a half hours, and only this once. See the slow families take turns.
What to watch
Section titled “What to watch”The log says what was written and where.
level=INFO msg=written sink=influxdb family=repo points=934 unchanged=1955level=INFO msg="rate budget" bucket=core remaining=4477 limit=5000What the log says on a good day has every routine line.
Three warnings are worth an alert:
rate limit reserve reachedmeans a family was skipped to protect the budget. Once is fine; every sweep means the cadences are too fast for the number of repositories.family failed everywhere, not marking it as runmeans every repository failed for one family, so it will be retried rather than treated as done.migration pending, at every start after an upgrade, means a store still holds rows in a shape this release no longer writes, and the start left them: see after an upgrade.
journalctl -u ghchronicle -fjournalctl -u ghchronicle -p warning --since todayAfter an upgrade
Section titled “After an upgrade”Replace the binary and restart the service. Under the default
migrate: auto, the start checks every
store against the changes the new release carries and, before its first sweep,
applies on its own each one that loses nothing, saying so at WARN. A change
it leaves is a WARN at every start, naming the store, the reason and the two
commands: what a start does about
it.
Run those as the service’s own user and with its environment file, with the
service stopped, since -migrate -yes refuses to run beside it:
sudo systemctl stop ghchroniclesudo systemd-run --uid=ghchronicle --gid=ghchronicle --pipe --wait --collect \ --property=EnvironmentFile=/etc/ghchronicle/ghchronicle.env \ /usr/local/bin/ghchronicle -config /etc/ghchronicle/config.yaml -migratesudo systemd-run --uid=ghchronicle --gid=ghchronicle --pipe --wait --collect \ --property=EnvironmentFile=/etc/ghchronicle/ghchronicle.env \ /usr/local/bin/ghchronicle -config /etc/ghchronicle/config.yaml -migrate -yessudo systemctl start ghchronicleThe first prints the plan and changes nothing; the second applies it and reads
back what it cleared. systemd-run reads the environment file as root, as the
unit does, so it stays mode 600, and runs the binary as ghchronicle: measured
on systemd 257, a command started that way saw the token of a file only root
could read and ran as the user named. Run by root instead, -migrate -yes
saves the files it rewrites beside the state file as root’s, mode 600,
state.json among them, and the service, which runs as ghchronicle, then
stops at its start and names the file rather than start from a new one, which
would forget a refill still owed: chown ghchronicle:ghchronicle it back.
With -once run from cron rather than a service, comment the line out for the
length of the two commands: a -once holds no lock while it sweeps, so nothing
stops it running beside -migrate -yes.
cron instead of a service
Section titled “cron instead of a service”-once runs a single sweep and exits, which is all a scheduler needs.
0 * * * * /usr/local/bin/ghchronicle -config /etc/ghchronicle/config.yaml -onceKeep the state file on a persistent path even in this mode. Without it every
run collects every family, whatever its cadence, and walks the stargazer list,
the whole star history and the co-authored pull requests again; the cache file
beside it is what lets a run ask GitHub only for what changed. Note that an
hourly cron gives every family an hourly cadence at best, so the five that run
every fifteen minutes, actions, events, notifs, activity and
ratelimit, run four times less often: a workflow run’s queue time is read
after it is over, and a notification thread moved twice inside the hour shows
only its second move.