The test layers
The tests come in three layers, and they are worth telling apart because only the first one is free. Two of them start nine containers, and anyone deciding whether to wait for that deserves to know what it buys.
| Layer | Command | Docker | Time | What it proves |
|---|---|---|---|---|
| L1, the contract | make test | no | about 5 s | the exact bytes each sink puts on the wire |
| L2, the stores | make test-e2e-docker | yes | 53 s warm | a real store accepts those bytes, and keeps the date of the event |
| L3, the dashboards | the same target | yes | included above | the five dashboards’ own queries answer against what the sinks wrote |
L1 runs on every push. L2 and L3 are one suite behind the dockere2e build
tag, so go test ./... never starts a container; in CI they run weekly, on
demand, and as a release gate.
L1: the bytes on the wire
Section titled “L1: the bytes on the wire”test/e2e builds the real binary, runs it against a fake GitHub whose fixtures
live in test/e2e/testdata and whose route table is test/e2e/fakegh, and
points every sink at an httptest capture server. Then it asserts the bytes:
the line protocol, the _bulk envelope, the SQL statements, the Graphite path,
the OTLP payload.
The fake also prices its answers the way api.github.com does: one ETag per
REST fixture and none on a GraphQL answer, a 304 charged nothing for a request
that presents it, one for every other answer on the API, nothing for the object
storage a job log redirects to, and the budget block on every GraphQL answer.
That is what lets TestTheSecondSweepIsPricedByTheCache run two
sweeps in a single process and hold that every URL answered 200 the first
time came back 304 the second, that the second sweep charged fewer than
half the core requests of the first, and that own_cost on the
gh_rate_limit row is the number of queries the process made rather than
zero.
It searches the way GitHub does, too. Each of the five outbound searches asks
for one kind in one state, is:pr is:merged or is:issue is:open, and the
fake serves it only the items of its fixture those qualifiers select, counting
the others out of issueCount. A kind or a state it does not know, an is:,
type: or state: it has no rule for, fails the test; the author, the owners
left out and the order are the fixture’s own already, and it reads past them.
Until 2.6.2 it answered every search with every item, so a sweep wrote the one
merged pull request five times, once as an open issue, and the dashboards drew
those rows.
A fixture never writes out a recent date. It spells one as an offset the fake
resolves when it serves the file, "@DAYS_AGO_12@T00:00:00Z", beside the
@NOW@, @TODAY@ and @SOON@ it already resolved, and @DAYS_AHEAD_n@ for
whatever has to still be in the future. A dashboard asks for the last day, week
or month, so a written-out date is correct only until it falls out of that
window, and nothing announces the day it does: the release of 2.4.0 failed with
four panels of every store answering nothing, twelve hours after the same suite
had passed, because the day one pull request merged on had crossed now-30d in
between. A test that has to name one of those days asks fakegh.DaysAgo, and a
date written out by hand fails TestNoFixtureWritesOutARecentDate on the spot.
That is the right test for a format, and it is fast enough to run while a sink is being changed. What it cannot catch is anything the receiver has an opinion about. A capture server answers 204 to everything. It has no column types, no mapping, no query planner and no schema.
L2: the stores themselves
Section titled “L2: the stores themselves”test/e2e/docker starts the real stores in containers, runs one sweep from the
same fake GitHub into all of them, then reads each store back and asserts the
value, the tags and above all the timestamp.
The dating rule is the product of this tool: a star is stamped when it was given, a workflow run when it finished, a traffic day at that day’s own date. No capture test can prove a store kept the date of the event rather than the date of the sweep, because storing it is the store’s job.
These are the defects this layer exists for, each of them real:
InfluxDB fixes a column’s type on first sight. InfluxDB 3 decides that a
column is a tag or a field the first time it sees it and refuses every later
write that disagrees: 400 invalid column type for column 'owner', expected iox::column_type::tag. A capture server answers 204 and notices nothing. This
has already cost a database wipe, and reproducing it was the first thing the
containerised stack was used for.
Elasticsearch’s dynamic mapping decides whether the dashboards can
aggregate. One panel pulls a url through a top_metrics, which normally
needs a keyword field rather than a text one. Two separate audits recorded that
as unverifiable for want of a real Elasticsearch. This layer answers it, by
indexing through _bulk and reading back the mapping the cluster built for
itself.
PostgreSQL has to accept the DDL. The SQL sink emits statements rather than
speaking the wire protocol, so until now nothing ever had them parsed. The
suite pipes them through psql, inserts, and plans the panels’ queries against
the schema the sink created rather than against one transcribed from InfluxDB.
Graphite paths have to have the depth the dashboards index. The dashboards address path nodes by position, and the agreement between those positions and what the sink writes was kept by a hand-maintained table that nothing checked. Here the sink writes to carbon and the render API is asked for the path back.
L3: the dashboards’ own queries
Section titled “L3: the dashboards’ own queries”With the stores loaded, the five generated dashboards are run through Grafana’s
/api/ds/query, which is the path cmd/check_dashboards takes against a live
Grafana. Every datasource is provisioned at boot with a fixed uid and the
harness mints a service account token, so a panel query goes through Grafana
exactly as it would for a person looking at the dashboard.
That is what turns three manual checkers into something CI runs, and it is what settles the Elasticsearch question above: a panel that cannot aggregate returns no frame.
Each panel is posted with its range resolved to the two instants a browser sends, and asked three questions, and every stat and gauge two more below: did the datasource refuse it, did a panel over something the sweep wrote inside the range answer with anything, and do the five stores draw the same thing. For the third, each answer is replayed through what Grafana does between the query and the screen, the Prometheus datasource’s own reshaping of a table, the panel’s transformations and its field overrides, and every pair of stores is compared on what a reader would see: a tile’s number, its unit and the words it shows for nothing, a table’s rows over the columns both stores draw, and a bar’s name and length. A table is held as well to one order of the columns both stores draw, which rows matched over those columns cannot show: on the 2.6.2 branch 28 Prometheus tables and 4 Graphite ones headed them in another order than the SQL stores, “Every bucket” with Most used last. The review of 2.6.1 found eleven differences by putting the dashboards side by side, among them a table that drew seven rows for one repository and stat tiles that had lost their units. Nine of them lived in the dashboards, and run against that release’s dashboards this fails on eight. The ninth was four columns Elasticsearch’s “Open the longest” went without, which no comparison of values can see. What a store can hold at all is its own description’s business, so another assertion holds every column the SQL stores draw in a table to being drawn by each other store that draws the table, or named in that store’s own words about the panel: on the 2.6.2 branch 33 tables of the other three stores lacked a column their descriptions did not name, and each now draws it or says why not.
Every stat and gauge is then asked about nothing, twice: about a repository no
sweep wrote anything for, and over thirty days no point of any sweep falls in,
four hundred days back or more. Over no rows a SQL count is 0 and anything else
is null, which a tile draws as the words its panel gives a value that is not
there, and every store is held to drawing the tiles the SQL stores draw, with
the same words. The range is asked because the account’s own snapshots do not
follow the repository picker: over a range four hundred days back Graphite drew
three stat groups as panels with nothing in them, not even the names. A tile a
store cannot draw over nothing is listed in tilesLeftOverNothing, in
test/e2e/docker/tiles_over_nothing_test.go, under words of that store’s own
description, and held to them as dashboardsDiffer is.
InfluxDB and PostgreSQL run one statement over the same rows, so they are also
held to drawing the rows of a table or a bar chart in one order, which a
comparison of the rows as a set cannot see: before the statements named their
tie-breakers, twelve of the 49 panels both draw with more than one row put
two of them the other way round in one run, and eleven in the next.
TestEverySQLListOrdersItsRowsCompletely, in internal/dashboards, holds
every list’s ORDER BY to naming what tells two of its rows apart, so that
the next tie does not wait for a run that happens to draw it.
Two of the rules held of what is drawn are held of the specification as well,
since a table the fixture leaves empty in both SQL stores is never drawn at
all: every column the SQL stores draw that another store does not is named in
that store’s own description, and a chart the SQL stores fold into other folds
in each other store or says it does not (TestEveryColumnAStoreLacksIsNamedInItsDescription
and TestEveryStoreFoldsTheRestIntoOtherOrSaysItDoesNot). And a Prometheus
query that is wrong only for series the fixture never makes, a label value
nobody wrote or a series that stood still over the range, is put to promtool
inside the stack’s own Prometheus over series written for it
(TestPrometheusAnswersSeriesTheFixtureHasNoneOfAsTheSQLStoresDo): until it
was, the Signed commits of an account that never signs read “no commits”
beside a count of 57.
A store that draws a panel differently on purpose says why in its own
description of the panel, and the difference is listed in dashboardsDiffer,
in test/e2e/docker/dashboards_agree_test.go, under those words. The test fails
when the words are no longer in the description, when the stores have come
to draw the panel alike, and when a store the entry names draws the panel as
the stores it does not name, so an entry can neither outlive its reason nor
excuse a store that needs no excuse. What the harness
itself causes is absorbed where it arises rather than listed: the exporter’s
half minute of history, values the collector computes from its own clock across
three sweeps a minute apart, values a query computes from now() across stores
asked seconds apart, and names Graphite holds as path nodes. A separate
assertion fails on any panel, time series included, that draws a field under the
name its datasource gave it, such as p50.0 seconds_to_merge.
What none of them catch
Section titled “What none of them catch”Every layer runs against the fake GitHub, so nothing here notices GitHub
changing a payload, retiring an endpoint or throttling differently. That is
what ghchronicle -once against a real token is for.
They also prove nothing about a store the suite does not start. The answer covers InfluxDB 3 Core, PostgreSQL 18, Elasticsearch 9, Graphite 1.1, Prometheus 3, Loki 3, the OpenTelemetry collector and Telegraf, at the pinned versions. OpenSearch, TimescaleDB and anything behind the Telegraf or OTLP hop are still an inference from the format.
There are five more things, and none of them is a layer. Each is switched on by
an environment variable and skipped when it is absent, so an ordinary go test ./... stays offline.
The store you actually run. test/live pushes a handful of points at a Loki
or an OpenTelemetry collector named in GHC_LIVE_LOKI or GHC_LIVE_OTLP. It
answers the one question containers cannot: whether your instance accepts them.
GHC_LIVE_LOKI=http://localhost:3100 go test ./test/live/The real API, end to end. GHC_E2E_LIVE=1 runs TestLiveAPI against GitHub
itself rather than the fake, with a real GITHUB_TOKEN, sweeping the account
named in GHC_E2E_USER.
GHC_E2E_LIVE=1 GHC_E2E_USER=octocat GITHUB_TOKEN=ghp_... go test ./test/e2e/ -run TestLiveAPIWhat a sweep costs in cache. GHC_LIVE_CONFIG points
TestLiveSweepCacheFootprint at a configuration file and sweeps the account it
names, reporting the entries and the bytes the conditional-request cache holds
after each sweep. Those are the figures the 256 MB bound and the
cost of a sweep rest on, and this is how to reproduce
them for your own account. GHC_LIVE_DUMP=1 adds the per-URL list to standard
output.
GHC_LIVE_CONFIG=config.yaml go test ./internal/ghapi/ -run TestLiveSweepCacheFootprint -vOne repository, one family. cmd/probe runs the collectors against a single
repository and prints the line they would write, writing nothing anywhere.
GHC_DUMP=<family> prints every point of that family in full, which is the
fastest way to see what a collector actually produces.
go run ./cmd/probe owner/nameGHC_DUMP=actions go run ./cmd/probe owner/nameThe pictures of the card. GHC_CARD_GALLERY names an existing directory
and TestCardGallery renders one card per layout into it, from the fake
GitHub rather than from anybody’s account. That is where the pictures on
the layouts page come from, and a layout that
changes shape is one command away from a set that agrees with it. The account
is the base fixtures with test/e2e/testdata/gallery/ laid over them: a year
of contributions, GitHub’s whole fourteen days of traffic, five repositories to
rank and one of them in six languages, which the smaller account every other
suite asserts on cannot give a picture. A fixture named <repo>~<fixture>
there answers for that one repository, and any other repository borrows
hello-world’s. Each layout comes out of one sweep under -card-theme both as
two files, card-<layout>.svg in the light palette and card-<layout>_dark.svg
in the dark one, which is what the site’s ThemeImage and the README’s
<picture> read. The two layouts that loop come out a second time under
-card-motion loop, as card-<layout>-loop.svg and its _dark twin. Only
those two: on every other layout loop draws the same card as once, so a
looping picture of one would be a second copy of the first under a name that
promises something else. Which layouts they are is the registry’s Loops, and
the gallery reads it rather than keeping its own list.
mkdir -p /tmp/cardsGHC_CARD_GALLERY=/tmp/cards go test ./test/e2e/ -run TestCardGallerymake check-gallery renders the gallery into a scratch directory and fails,
naming every difference, if the committed set no longer matches it byte for
byte; make gallery regenerates it in place. CI’s “Generated artifacts” job
runs the check on every pull request.
Running the stack
Section titled “Running the stack”Docker with the compose plugin, and room for the images. Then:
make test-e2e-dockerUp, run, down on every path including a failing assertion, and then a check
that docker ps shows nothing of the project left. A suite that leaves nine
containers behind on a failure is a suite nobody runs twice.
Boot, measured cold with the images already pulled: Elasticsearch 29 s, Loki 21 s, Grafana 13 s, the Graphite render API 10 s, InfluxDB 8 s, PostgreSQL 6 s, the rest 6 s. The stack is ready in 30 s; the target end to end, teardown included, is 53 s.
Debugging with the stack up
Section titled “Debugging with the stack up”The reason to fail an assertion is to go and look at the store, and a suite that tore the store down first cannot be looked at. So the two halves are separate targets, and the harness reuses a stack it finds already running and leaves it running.
-
Start the stores and leave them up. The command prints the port each service ended up on.
Terminal window make e2e-docker-up -
Run the suite, or one test of it, as many times as it takes.
Terminal window go test -count=1 -tags dockere2e -timeout 30m -v ./test/e2e/docker/GHCHRONICLE_E2E_KEEP=1also stops the test binary tearing down a stack it started itself, which is what you want when a single-runis failing. -
Ask the store what it thinks, then tear it down.
Terminal window make e2e-docker-logs SERVICE=influxdbmake e2e-docker-down
With the ports from step 1:
# What InfluxDB thinks each column is. This is the answer to a 400 on write.curl -s "http://127.0.0.1:<influx>/api/v3/query_sql?db=ghchronicle" \ --data-urlencode "q=SELECT * FROM information_schema.columns WHERE table_name = 'gh_repo'"
# The mapping Elasticsearch built for itself.curl -s "http://127.0.0.1:<es>/ghchronicle-*/_mapping?pretty"
# What the SQL sink actually created.psql "postgres://ghchronicle:ghchronicle@127.0.0.1:<pg>/ghchronicle" -c '\d+ gh_repo'
# The Graphite path, node by node.curl -s "http://127.0.0.1:<graphite>/metrics/find?query=github.repo.*"Grafana is at the port it published, with admin and admin, and every
datasource is already provisioned, so a panel query can be pasted into Explore
and run by hand.
Three things the stack had to be told
Section titled “Three things the stack had to be told”Each of these silently produced a wrong answer before it was found, and each is in the compose file or its configuration with the measurement beside it:
- Carbon drops a point older than its longest archive without saying so. A
star dated 2020 vanished under a six year retention and the write was
reported as accepted. The retention is
1d:12yfor that reason. - Carbon’s default
MAX_CREATES_PER_MINUTEis 50, fewer paths than one sweep creates, so most of a first sweep would be dropped. - Loki answers a push with 204 and will not serve it until the chunk is
flushed. The test polls rather than asking once, and
chunk_idle_periodis 5 s.
A machine that adds the rule runs those assertions for the first time, which is
when promNeedsHistory in dashboards_test.go starts to matter: the exporter
is alive only for the length of the test, so every timeseries panel and every
panel built on increase() is held to nothing and only the instant panels are
asserted.
.github/workflows/e2e.yml runs make test-e2e-docker and has three ways in:
manual dispatch with an optional ref, a weekly schedule on main, and
workflow_call, so a release pipeline gates a tag on it with one line rather
than a copy of the job that drifts from the original.
It is not a required check on a pull request: nine containers and around 10 GB of image is too much for every push, and L1 is what covers every push. The weekly run is the point of the schedule. Nothing else in the repository ever starts a container, so without it the suite would run only when somebody remembered it, which is how a suite ends up broken for weeks with nobody knowing.
The same suite also runs under the race detector, in
.github/workflows/race.yml: weekly, at every release beside the E2E gate, and
by hand. The harness builds the collector with -race and starts it with
GORACE=halt_on_error=1, so a race inside the collector fails the test that
started it, with the report. Locally it is make test-e2e-docker-race.