Targets
The targets block decides which GitHub repositories a sweep collects: an account’s own, those of its organisations, and any named one.
targets: user: your-github-login orgs: [] repos: [] exclude: [] include_forks: false include_archived: false include_private: trueThe keys
Section titled “The keys”| Key | Default | What it does |
|---|---|---|
user | none | The account to collect. Its repositories are discovered automatically |
orgs | [] | Organisations to include as well |
repos | [] | Repositories to collect whatever the filters say |
exclude | [] | Shell globs, matched against owner/name |
include_ | false | Whether forks are collected |
include_ | false | Whether archived repositories are collected |
include_ | true | Whether private repositories are collected |
At least one of user, orgs or repos must be set, or start-up fails with
targets: set at least one of user, orgs or repos.
user is also what makes the account-wide families possible. A configuration
that names repositories alone has no login to hand to GraphQL, so the
contribution calendar, the event feed, notifications, billing, packages and the
outbound families are all skipped.
Naming a repository overrides every exclusion
Section titled “Naming a repository overrides every exclusion”repos is not another filter. It is a statement of intent, and it wins:
naming someone/thing collects it even if it is a fork, even if it is
archived, even if a glob in exclude would have matched it.
targets: user: acme repos: - someone-else/a-fork-i-actually-maintain exclude: - "acme/experiment-*"exclude takes shell globs, so someone/experiment-* drops a whole prefix in
one line.
Naming one of the account’s own repositories, a fork of yours for instance,
costs no request to discover it: the listing already returned it, with the
four things discovery wants to know about it. Only a name no listing returns,
someone else’s repository or one of an organisation not in orgs, is read on
its own each time the list is rebuilt.
Why forks and archived are off by default
Section titled “Why forks and archived are off by default”Both defaults exist for the same reason, which is the rate limit.
- A fork’s traffic is almost always zero. GitHub reports views and clones per repository, and for a fork that nobody visits, that is fourteen days of zeroes per sweep. It costs four calls per repository to learn nothing.
- An archived repository takes no more work. Its stars, forks and watchers can still move, and those are read anyway, as below. Nothing else can, and collecting it spends the budget on rows that never move again.
Turn either on when the assumption does not hold for you. A fork you actually
develop in is a real repository with real traffic, and repos is the way to
name that one without also collecting the forty bookmarks.
Off is not invisible, on two counts. An archived repository still gets two
rows from every totals sweep: gh_repo_archived, dated the instant it was
archived, and gh_repo_total, its lifetime counts as they stand, stars and
forks included, stamped at the sweep. The listing a sweep already pays for
says which repositories are archived, and the totals family asks about all
of them in one query per twenty five at its own cadence, so the Repositories
archived table exists from the first sweep, the account’s star and fork
totals count them, and the cost is a point per twenty five archived
repositories per totals sweep. And a backfill collects archived
repositories in full whatever this key says, because their history is the
account’s history and one walk of it is enough. Forks stay as configured in
both cases: an archived fork under include_forks: false has no row.
Private repositories
Section titled “Private repositories”include_private defaults to true, because a token that can see them was
given deliberately and the point of the tool is to keep the history of the
account, not of its public half.
What is stored is the same measurements as for a public repository: counts, durations and names. Two things are worth knowing about what that means:
- Repository and branch names appear as tag values, so they will be visible to anyone who can read the dashboard.
- Only the host of a webhook URL is stored, never the path, because the path usually carries a secret.
Set it to false to collect only public repositories.
The discovery cadence
Section titled “The discovery cadence”The repository list is rebuilt once an hour: at the first sweep an hour after
the last listing, to within half a tick, the margin a family’s
cadence
gets too. Repositories are created rarely and listing them costs a page per
hundred, so a new repository can take up to an hour to enter the sweep. Before
2.6.1 the list had no margin and could outlive the sweep an hour after it by a
few milliseconds, so at the quarter-hour tick a new repository waited up to an
hour and a quarter about half the time. Naming it in repos does not change
that; restarting the process does, and so does each -once, which lists them
afresh.