Skip to content

The test layers

The tests come in three layers, and they are worth telling apart because only the first one is free. Two of them start nine containers, and anyone deciding whether to wait for that deserves to know what it buys.

LayerCommandDockerTimeWhat it proves
L1, the contractmake testnoabout 5 sthe exact bytes each sink puts on the wire
L2, the storesmake test-e2e-dockeryes53 s warma real store accepts those bytes, and keeps the date of the event
L3, the dashboardsthe same targetyesincluded abovethe five dashboards’ own queries answer against what the sinks wrote

L1 runs on every push. L2 and L3 are one suite behind the dockere2e build tag, so go test ./... never starts a container; in CI they run weekly, on demand, and as a release gate.

test/e2e builds the real binary, runs it against a fake GitHub whose fixtures live in test/e2e/testdata and whose route table is test/e2e/fakegh, and points every sink at an httptest capture server. Then it asserts the bytes: the line protocol, the _bulk envelope, the SQL statements, the Graphite path, the OTLP payload.

The fake also prices its answers the way api.github.com does: one ETag per REST fixture and none on a GraphQL answer, a 304 charged nothing for a request that presents it, one for every other answer on the API, nothing for the object storage a job log redirects to, and the budget block on every GraphQL answer. That is what lets TestTheSecondSweepIsPricedByTheCache run two sweeps in a single process and hold that every URL answered 200 the first time came back 304 the second, that the second sweep charged fewer than half the core requests of the first, and that own_cost on the gh_rate_limit row is the number of queries the process made rather than zero.

It searches the way GitHub does, too. Each of the five outbound searches asks for one kind in one state, is:pr is:merged or is:issue is:open, and the fake serves it only the items of its fixture those qualifiers select, counting the others out of issueCount. A kind or a state it does not know, an is:, type: or state: it has no rule for, fails the test; the author, the owners left out and the order are the fixture’s own already, and it reads past them. Until 2.6.2 it answered every search with every item, so a sweep wrote the one merged pull request five times, once as an open issue, and the dashboards drew those rows.

A fixture never writes out a recent date. It spells one as an offset the fake resolves when it serves the file, "@DAYS_AGO_12@T00:00:00Z", beside the @NOW@, @TODAY@ and @SOON@ it already resolved, and @DAYS_AHEAD_n@ for whatever has to still be in the future. A dashboard asks for the last day, week or month, so a written-out date is correct only until it falls out of that window, and nothing announces the day it does: the release of 2.4.0 failed with four panels of every store answering nothing, twelve hours after the same suite had passed, because the day one pull request merged on had crossed now-30d in between. A test that has to name one of those days asks fakegh.DaysAgo, and a date written out by hand fails TestNoFixtureWritesOutARecentDate on the spot.

That is the right test for a format, and it is fast enough to run while a sink is being changed. What it cannot catch is anything the receiver has an opinion about. A capture server answers 204 to everything. It has no column types, no mapping, no query planner and no schema.

test/e2e/docker starts the real stores in containers, runs one sweep from the same fake GitHub into all of them, then reads each store back and asserts the value, the tags and above all the timestamp.

The dating rule is the product of this tool: a star is stamped when it was given, a workflow run when it finished, a traffic day at that day’s own date. No capture test can prove a store kept the date of the event rather than the date of the sweep, because storing it is the store’s job.

These are the defects this layer exists for, each of them real:

InfluxDB fixes a column’s type on first sight. InfluxDB 3 decides that a column is a tag or a field the first time it sees it and refuses every later write that disagrees: 400 invalid column type for column 'owner', expected iox::column_type::tag. A capture server answers 204 and notices nothing. This has already cost a database wipe, and reproducing it was the first thing the containerised stack was used for.

Elasticsearch’s dynamic mapping decides whether the dashboards can aggregate. One panel pulls a url through a top_metrics, which normally needs a keyword field rather than a text one. Two separate audits recorded that as unverifiable for want of a real Elasticsearch. This layer answers it, by indexing through _bulk and reading back the mapping the cluster built for itself.

PostgreSQL has to accept the DDL. The SQL sink emits statements rather than speaking the wire protocol, so until now nothing ever had them parsed. The suite pipes them through psql, inserts, and plans the panels’ queries against the schema the sink created rather than against one transcribed from InfluxDB.

Graphite paths have to have the depth the dashboards index. The dashboards address path nodes by position, and the agreement between those positions and what the sink writes was kept by a hand-maintained table that nothing checked. Here the sink writes to carbon and the render API is asked for the path back.

With the stores loaded, the five generated dashboards are run through Grafana’s /api/ds/query, which is the path cmd/check_dashboards takes against a live Grafana. Every datasource is provisioned at boot with a fixed uid and the harness mints a service account token, so a panel query goes through Grafana exactly as it would for a person looking at the dashboard.

That is what turns three manual checkers into something CI runs, and it is what settles the Elasticsearch question above: a panel that cannot aggregate returns no frame.

Each panel is posted with its range resolved to the two instants a browser sends, and asked three questions, and every stat and gauge two more below: did the datasource refuse it, did a panel over something the sweep wrote inside the range answer with anything, and do the five stores draw the same thing. For the third, each answer is replayed through what Grafana does between the query and the screen, the Prometheus datasource’s own reshaping of a table, the panel’s transformations and its field overrides, and every pair of stores is compared on what a reader would see: a tile’s number, its unit and the words it shows for nothing, a table’s rows over the columns both stores draw, and a bar’s name and length. A table is held as well to one order of the columns both stores draw, which rows matched over those columns cannot show: on the 2.6.2 branch 28 Prometheus tables and 4 Graphite ones headed them in another order than the SQL stores, “Every bucket” with Most used last. The review of 2.6.1 found eleven differences by putting the dashboards side by side, among them a table that drew seven rows for one repository and stat tiles that had lost their units. Nine of them lived in the dashboards, and run against that release’s dashboards this fails on eight. The ninth was four columns Elasticsearch’s “Open the longest” went without, which no comparison of values can see. What a store can hold at all is its own description’s business, so another assertion holds every column the SQL stores draw in a table to being drawn by each other store that draws the table, or named in that store’s own words about the panel: on the 2.6.2 branch 33 tables of the other three stores lacked a column their descriptions did not name, and each now draws it or says why not.

Every stat and gauge is then asked about nothing, twice: about a repository no sweep wrote anything for, and over thirty days no point of any sweep falls in, four hundred days back or more. Over no rows a SQL count is 0 and anything else is null, which a tile draws as the words its panel gives a value that is not there, and every store is held to drawing the tiles the SQL stores draw, with the same words. The range is asked because the account’s own snapshots do not follow the repository picker: over a range four hundred days back Graphite drew three stat groups as panels with nothing in them, not even the names. A tile a store cannot draw over nothing is listed in tilesLeftOverNothing, in test/e2e/docker/tiles_over_nothing_test.go, under words of that store’s own description, and held to them as dashboardsDiffer is.

InfluxDB and PostgreSQL run one statement over the same rows, so they are also held to drawing the rows of a table or a bar chart in one order, which a comparison of the rows as a set cannot see: before the statements named their tie-breakers, twelve of the 49 panels both draw with more than one row put two of them the other way round in one run, and eleven in the next. TestEverySQLListOrdersItsRowsCompletely, in internal/dashboards, holds every list’s ORDER BY to naming what tells two of its rows apart, so that the next tie does not wait for a run that happens to draw it.

Two of the rules held of what is drawn are held of the specification as well, since a table the fixture leaves empty in both SQL stores is never drawn at all: every column the SQL stores draw that another store does not is named in that store’s own description, and a chart the SQL stores fold into other folds in each other store or says it does not (TestEveryColumnAStoreLacksIsNamedInItsDescription and TestEveryStoreFoldsTheRestIntoOtherOrSaysItDoesNot). And a Prometheus query that is wrong only for series the fixture never makes, a label value nobody wrote or a series that stood still over the range, is put to promtool inside the stack’s own Prometheus over series written for it (TestPrometheusAnswersSeriesTheFixtureHasNoneOfAsTheSQLStoresDo): until it was, the Signed commits of an account that never signs read “no commits” beside a count of 57.

A store that draws a panel differently on purpose says why in its own description of the panel, and the difference is listed in dashboardsDiffer, in test/e2e/docker/dashboards_agree_test.go, under those words. The test fails when the words are no longer in the description, when the stores have come to draw the panel alike, and when a store the entry names draws the panel as the stores it does not name, so an entry can neither outlive its reason nor excuse a store that needs no excuse. What the harness itself causes is absorbed where it arises rather than listed: the exporter’s half minute of history, values the collector computes from its own clock across three sweeps a minute apart, values a query computes from now() across stores asked seconds apart, and names Graphite holds as path nodes. A separate assertion fails on any panel, time series included, that draws a field under the name its datasource gave it, such as p50.0 seconds_to_merge.

Every layer runs against the fake GitHub, so nothing here notices GitHub changing a payload, retiring an endpoint or throttling differently. That is what ghchronicle -once against a real token is for.

They also prove nothing about a store the suite does not start. The answer covers InfluxDB 3 Core, PostgreSQL 18, Elasticsearch 9, Graphite 1.1, Prometheus 3, Loki 3, the OpenTelemetry collector and Telegraf, at the pinned versions. OpenSearch, TimescaleDB and anything behind the Telegraf or OTLP hop are still an inference from the format.

There are five more things, and none of them is a layer. Each is switched on by an environment variable and skipped when it is absent, so an ordinary go test ./... stays offline.

The store you actually run. test/live pushes a handful of points at a Loki or an OpenTelemetry collector named in GHC_LIVE_LOKI or GHC_LIVE_OTLP. It answers the one question containers cannot: whether your instance accepts them.

Terminal window
GHC_LIVE_LOKI=http://localhost:3100 go test ./test/live/

The real API, end to end. GHC_E2E_LIVE=1 runs TestLiveAPI against GitHub itself rather than the fake, with a real GITHUB_TOKEN, sweeping the account named in GHC_E2E_USER.

Terminal window
GHC_E2E_LIVE=1 GHC_E2E_USER=octocat GITHUB_TOKEN=ghp_... go test ./test/e2e/ -run TestLiveAPI

What a sweep costs in cache. GHC_LIVE_CONFIG points TestLiveSweepCacheFootprint at a configuration file and sweeps the account it names, reporting the entries and the bytes the conditional-request cache holds after each sweep. Those are the figures the 256 MB bound and the cost of a sweep rest on, and this is how to reproduce them for your own account. GHC_LIVE_DUMP=1 adds the per-URL list to standard output.

Terminal window
GHC_LIVE_CONFIG=config.yaml go test ./internal/ghapi/ -run TestLiveSweepCacheFootprint -v

One repository, one family. cmd/probe runs the collectors against a single repository and prints the line they would write, writing nothing anywhere. GHC_DUMP=<family> prints every point of that family in full, which is the fastest way to see what a collector actually produces.

Terminal window
go run ./cmd/probe owner/name
GHC_DUMP=actions go run ./cmd/probe owner/name

The pictures of the card. GHC_CARD_GALLERY names an existing directory and TestCardGallery renders one card per layout into it, from the fake GitHub rather than from anybody’s account. That is where the pictures on the layouts page come from, and a layout that changes shape is one command away from a set that agrees with it. The account is the base fixtures with test/e2e/testdata/gallery/ laid over them: a year of contributions, GitHub’s whole fourteen days of traffic, five repositories to rank and one of them in six languages, which the smaller account every other suite asserts on cannot give a picture. A fixture named <repo>~<fixture> there answers for that one repository, and any other repository borrows hello-world’s. Each layout comes out of one sweep under -card-theme both as two files, card-<layout>.svg in the light palette and card-<layout>_dark.svg in the dark one, which is what the site’s ThemeImage and the README’s <picture> read. The two layouts that loop come out a second time under -card-motion loop, as card-<layout>-loop.svg and its _dark twin. Only those two: on every other layout loop draws the same card as once, so a looping picture of one would be a second copy of the first under a name that promises something else. Which layouts they are is the registry’s Loops, and the gallery reads it rather than keeping its own list.

Terminal window
mkdir -p /tmp/cards
GHC_CARD_GALLERY=/tmp/cards go test ./test/e2e/ -run TestCardGallery

make check-gallery renders the gallery into a scratch directory and fails, naming every difference, if the committed set no longer matches it byte for byte; make gallery regenerates it in place. CI’s “Generated artifacts” job runs the check on every pull request.

Docker with the compose plugin, and room for the images. Then:

Terminal window
make test-e2e-docker

Up, run, down on every path including a failing assertion, and then a check that docker ps shows nothing of the project left. A suite that leaves nine containers behind on a failure is a suite nobody runs twice.

Boot, measured cold with the images already pulled: Elasticsearch 29 s, Loki 21 s, Grafana 13 s, the Graphite render API 10 s, InfluxDB 8 s, PostgreSQL 6 s, the rest 6 s. The stack is ready in 30 s; the target end to end, teardown included, is 53 s.

The reason to fail an assertion is to go and look at the store, and a suite that tore the store down first cannot be looked at. So the two halves are separate targets, and the harness reuses a stack it finds already running and leaves it running.

  1. Start the stores and leave them up. The command prints the port each service ended up on.

    Terminal window
    make e2e-docker-up
  2. Run the suite, or one test of it, as many times as it takes.

    Terminal window
    go test -count=1 -tags dockere2e -timeout 30m -v ./test/e2e/docker/

    GHCHRONICLE_E2E_KEEP=1 also stops the test binary tearing down a stack it started itself, which is what you want when a single -run is failing.

  3. Ask the store what it thinks, then tear it down.

    Terminal window
    make e2e-docker-logs SERVICE=influxdb
    make e2e-docker-down

With the ports from step 1:

Terminal window
# What InfluxDB thinks each column is. This is the answer to a 400 on write.
curl -s "http://127.0.0.1:<influx>/api/v3/query_sql?db=ghchronicle" \
--data-urlencode "q=SELECT * FROM information_schema.columns WHERE table_name = 'gh_repo'"
# The mapping Elasticsearch built for itself.
curl -s "http://127.0.0.1:<es>/ghchronicle-*/_mapping?pretty"
# What the SQL sink actually created.
psql "postgres://ghchronicle:ghchronicle@127.0.0.1:<pg>/ghchronicle" -c '\d+ gh_repo'
# The Graphite path, node by node.
curl -s "http://127.0.0.1:<graphite>/metrics/find?query=github.repo.*"

Grafana is at the port it published, with admin and admin, and every datasource is already provisioned, so a panel query can be pasted into Explore and run by hand.

Each of these silently produced a wrong answer before it was found, and each is in the compose file or its configuration with the measurement beside it:

  • Carbon drops a point older than its longest archive without saying so. A star dated 2020 vanished under a six year retention and the write was reported as accepted. The retention is 1d:12y for that reason.
  • Carbon’s default MAX_CREATES_PER_MINUTE is 50, fewer paths than one sweep creates, so most of a first sweep would be dropped.
  • Loki answers a push with 204 and will not serve it until the chunk is flushed. The test polls rather than asking once, and chunk_idle_period is 5 s.

A machine that adds the rule runs those assertions for the first time, which is when promNeedsHistory in dashboards_test.go starts to matter: the exporter is alive only for the length of the test, so every timeseries panel and every panel built on increase() is held to nothing and only the instant panels are asserted.

.github/workflows/e2e.yml runs make test-e2e-docker and has three ways in: manual dispatch with an optional ref, a weekly schedule on main, and workflow_call, so a release pipeline gates a tag on it with one line rather than a copy of the job that drifts from the original.

It is not a required check on a pull request: nine containers and around 10 GB of image is too much for every push, and L1 is what covers every push. The weekly run is the point of the schedule. Nothing else in the repository ever starts a container, so without it the suite would run only when somebody remembered it, which is how a suite ends up broken for weeks with nobody knowing.

The same suite also runs under the race detector, in .github/workflows/race.yml: weekly, at every release beside the E2E gate, and by hand. The harness builds the collector with -race and starts it with GORACE=halt_on_error=1, so a race inside the collector fails the test that started it, with the report. Locally it is make test-e2e-docker-race.

Written and maintained by
MIT licenceRelease history