Skip to content

Elasticsearch

The Elasticsearch sink sends ghchronicle’s dated points to Elasticsearch or OpenSearch in bulk, one document per point.

sinks:
elasticsearch:
url: http://elasticsearch:9200
prefix: ghchronicle
api_key: ${ES_API_KEY} # or username and password
batch: 1000

The _bulk API, which Elasticsearch and OpenSearch share, so the same sink serves both, and Kibana or the OpenSearch dashboards on top of either.

api_key: ${ES_API_KEY}

Sent as Authorization: ApiKey.

Not both. Configuring an API key alongside a username fails validation at start-up with sinks.elasticsearch: set either api_key or username and password, not both.

One index per measurement, named <prefix>-<measurement>: ghchronicle-gh_repo, ghchronicle-gh_traffic.

One document per point, with the time as @timestamp in RFC 3339, measurement, and every tag and field as a top-level key, so nothing has to be unnested before it can be filtered on.

{
"@timestamp": "2026-09-07T00:00:00Z",
"measurement": "gh_traffic",
"owner": "acme",
"repo": "telemetry",
"full_name": "acme/telemetry",
"kind": "views",
"count": 220,
"uniques": 131
}

The document id is the SHA-256 of the measurement, the tags that are set and the timestamp, and the action is index rather than create. So writing the same fourteen-day traffic window every six hours replaces fourteen documents instead of adding fourteen more, which is the same convergence InfluxDB gives for free.

The sink creates no index template. Dynamic mapping gives every string field a .keyword sub-field, which is what the dashboard aggregates on.

If you want explicit mappings, create the index templates before the first write. Nothing in the sink depends on them; only the panels’ choice of .keyword does. -migrate reads the mapping before it asks whose rows an index holds, so a template that maps strings as keyword, with no .keyword sub-field, is read from the field itself.

dashboards/ghchronicle-elasticsearch.json has the same 154 panels as the InfluxDB one, as Lucene filters and aggregations over one datasource pointing at <prefix>-*, because each target names its own index in its query.

Set the datasource’s time field to @timestamp. A per-item table there is the newest documents themselves, or the newest document of each item where an item has several; everything else is a bucket aggregation. OpenSearch works through the same plugin.

Two things a query cannot do there, and the panels that need them say so. It cannot join two indices, so Work elsewhere has no Stars column and the community profile keeps the API’s issue template flag rather than the count of templates. And it cannot ask about one window while it reads another: the Overview and Every repository, ever count an archived repository the default filter sets aside from its documents of the last seven days, where the SQL stores ask whether the collector still writes it, so a range that ended more than a week ago leaves those repositories out.

When a release changes what a measurement’s documents are keyed by, -migrate finds the documents of the old shape by counting the ones that carry the old tag, and applying the change sets the index aside in three steps. Measured against Elasticsearch 9.5.3:

  1. The index stops taking writes, through PUT <index>/_block/write.
  2. It is cloned to <index>-<instant>, the instant in UTC and in lower case, as an index name has to be, for example ghchronicle-gh_discussion_comment-20261001t091004, and the clone has to hold as many documents as the index before anything else happens.
  3. The index is deleted, and the sink’s next write creates it again, with a mapping of its own.

Before the change is applied, the plan asks whose rows the index holds and which repositories they belong to, with a terms aggregation on whatever field the mapping makes aggregatable: the .keyword sub-field dynamic mapping gives a string, or the field itself where a template maps it as keyword. The buckets have to add up to every document that holds the field, or the answer is not taken: measured on 9.5.3, an aggregation on user.keyword over an index whose template mapped user as keyword answered no buckets and no error, which read as an index holding nobody else’s rows, and a value past the sub-field’s ignore_above is in no bucket either. A document without the field is a row that names nobody, which the plan says as such.

A step that fails before the delete takes the write block off again, so a cluster that refused the clone goes on taking the sink’s writes, and the change stays pending with the cluster’s reason. A delete that fails is looked at again before anything is undone: measured through a proxy that answered 502 to a delete the cluster carried out, the index was gone, so the clone, the one copy of its documents, is kept and the change recorded as applied; only an index still there has its block taken off and its clone deleted. The API key needs the manage and delete_index privileges on the prefix’s indices; a key that can only write is refused that way.

The clone keeps the write block, and ghchronicle deletes it once it has been kept 24 hours: after a sweep of the service, at the next start of any run, or at the next -migrate -yes. Until then undoing the change is deleting the new index and cloning the copy back under its name. Every query of the shipped dashboard names its index whole, _index:ghchronicle-gh_discussion_comment, so no panel reads the clone: a panel’s query over ghchronicle-* counted the one new document and not the clone’s three. A Kibana data view over the same pattern without that filter does read it, for the day it is kept.

  • Choosing a store compares Elasticsearch with the others, and holds the write ledger every one of them shares.
  • The dashboards says which of the five is drawn against which store, and what a panel a store cannot answer becomes.
Written and maintained by
MIT licenceRelease history