The shipped system profile binds status to 127.0.0.1:8085, disables guest access
and enables token_repo_restrictions. Start with the first-login instructions,
then add repositories. Schema defaults for omitted fields can
differ; upgrades preserve your existing configuration. Package-cache access is
separate from dashboard/status authentication.
repowatch's HTTP server (repowatch run / serve-status) exposes three
different kinds of access, each with its own credential:
- The dashboard and its administrative API — a single admin password.
status.jsononly (the global one and/status/<repo_id>.json), meant for automated clients (hosts polling beforepacman -Syu/apt update/apk update) — per-host bearer tokens. History, packages, and everything else the dashboard shows are administrative data, gated by the admin session in (1), not by a bearer token — a host token cannot read them (verified:/status.jsonreturns200,/status/<id>/historyand/api/repos/<id>/packagesreturn401for the same token).- An optional, fully anonymous read-only mode (
guest_read_only) — for internal networks where even a shared token is more friction than you want. This is what actually opens history/packages/warmed/bans for unauthenticated reading, independent of bearer tokens entirely — see Guest read-only mode below.
There's no framework underneath this — it's http.server plus a small
amount of hand-written session/token/CSRF logic (auth.py, web/access.py).
This document describes the resulting behavior; see those two files for the
implementation.
- Admin login
- Host tokens (
status.jsonclients) - Versioned status API (
/api/v1/...) - Guest read-only mode
- TLS and
allow_insecure_http - Reverse proxies and
trusted_proxies /metricsand/healthz- What's deliberately not there
- Physical cache completeness
- Repository and group cache sizes
There is exactly one administrator password, hashed with PBKDF2-HMAC-SHA256
(600,000 iterations, matching OWASP's current Password Storage Cheat Sheet
recommendation for that algorithm as of this writing, for newly generated
hashes) and stored as
admin_password_hash in config.yaml. The plaintext password itself is
never stored anywhere. Existing hashes retain their original iteration count
until the password is set again. Set it with:
repowatch hash-password # prints a hash to paste into config.yaml
repowatch set-password # changes it in place, also revokes existing sessionsUntil admin_password_hash is set, the dashboard and every administrative
route are closed (there's no "no password configured = wide open" mode).
Sessions. POST /api/auth/login with the password creates a 12-hour
session, stored in state_db, identified by an HttpOnly cookie
(repowatch_session; Secure is set whenever the connection is HTTPS —
see TLS). The cookie carries an opaque
random secret, not the password or a JWT — the server looks it up in SQLite
on every request and checks it against the current password's fingerprint,
so changing the password (or running set-password) invalidates every
existing session immediately, without needing to enumerate or delete them.
CSRF. Every state-changing (POST) admin request must also include an
X-CSRF-Token header matching a per-session token, obtained from
GET /api/auth/session. The cookie alone isn't enough to make a change —
this is what stops a malicious page the admin happens to have open in
another tab from silently POSTing to the dashboard using the ambient
session cookie. POST requests are additionally checked for same-origin
(Origin/Sec-Fetch-Site) as a second, independent layer.
Login itself requires HTTPS — either terminated by the server itself,
by a trusted_proxies-listed reverse proxy declaring X-Forwarded-Proto: https, or by allow_insecure_http as an explicit opt-out (see
TLS for all three) — with one standing
exception: a connection arriving directly on loopback (no forwarding
headers, not through a proxy) is always allowed to use credentials over
plain HTTP, on the reasoning that the server can be certain there's no
untrusted network hop between it and itself. That's what lets
curl http://127.0.0.1:.../api/auth/login work out of the box on the same
host repowatch runs on, without setting anything.
The machines that actually poll status.json before running their package
manager don't get the admin password — they get a separate, revocable
host token, issued from the dashboard (POST /api/tokens, admin-only)
or its API:
- The token itself (
rw_..., a random secret) is shown once, at issue time — only its SHA-256 hash is stored instate_db. If you lose it, revoke it and issue a new one. - Tokens can optionally expire (
expires_at) and can be named freely, so you can tell "the token ondb-primary" apart from "the token onweb-03" in the token list. - Revocation (
POST /api/tokens/<id>/revoke) is checked on every request against SQLite — never cached in a running process — so a revoked token stops working immediately, not "after the next restart". - A client authenticates with a standard
Authorization: Bearer <token>header againststatus.json(the global one and/status/<repo_id>.json) — that's the only thing a bearer token unlocks; see below.
Scoping tokens to specific repositories. By default a token can read
status for every repository. If you want a host to only be able to see
repositories it's actually supposed to use, turn on
status_server.token_repo_restrictions: true in config.yaml, then issue
new tokens with an explicit repo_ids list. This is opt-in and
non-retroactive: tokens issued before you turn the setting on (or issued
without a repo_ids list afterward) keep unrestricted access — turning the
setting on doesn't quietly narrow anything that already exists. A scoped
token gets a 403 for a repository outside its list, and its
GET /status.json response is filtered down to only the repositories it's
allowed to see, rather than erroring.
Note what host tokens don't grant: they unlock status.json specifically
and nothing else — not the per-repository history endpoint
(/status/<repo_id>/history), not /api/repos/<repo_id>/packages, not
warmed/bans, none of it. Those are administrative reads and require
either the admin session below or guest_read_only (see Guest read-only
mode) — a bearer token gets neither. Tokens also
cannot log into the dashboard, trigger a warm-up, add a repository, or
change any config — those all require the admin session and CSRF token
above.
Every route a host client actually needs — status.json, per-repository
status, per-repository history, /healthz, /metrics — is also reachable
under an /api/v1/ prefix, with identical behavior and the same
authentication rules described above:
GET /api/v1/status.json
GET /api/v1/status/<repo_id>.json
GET /api/v1/status/<repo_id>/history
GET /api/v1/healthz
GET /api/v1/metrics
Today /api/v1/... and the unversioned paths above are the same thing —
/api/v1/ is a stable alias, not a different implementation. The point of
having it is forward-looking: if this specific client-facing contract ever
needs an incompatible change, that change lands in /api/v2/... instead of
breaking what's already polling /api/v1/... (or the unversioned routes,
which keep working as a permanent alias of v1). New integrations should
prefer the /api/v1/ form for that reason, but nothing currently pointed at
the unversioned paths needs to change.
This versioning applies only to this client-facing status API. The
dashboard's own API (/api/repos, /api/config, /api/tokens, and so on)
is intentionally not versioned — it's an implementation detail of the
bundled dashboard, released and upgraded together with it, not a contract
promised to outside consumers.
status_server.guest_read_only: true (default false) opens the dashboard
and a fixed, explicit set of read routes — status, per-repo history,
packages, warmed-package lists, ban lists, request statistics, the safe
config fields — to anyone, with no login and no token at all. It's meant for
trusted internal networks where even distributing a token is unnecessary
friction, not for exposing repowatch to the open internet.
What it does not do:
- It's an allowlist of specific GET routes, not "every GET route" —
/api/tokens(the list of issued host tokens) stays admin-only regardless, since that list is itself sensitive. - Nothing that changes state is affected — every
POSTroute still requires a real admin session and CSRF token. A guest can watch the dashboard render but can't add a repository, trigger a warm-up, or edit anything; the write buttons are hidden in the UI, and the server enforces the same restriction independently if you bypass the UI and call the API directly. /metricshas its own separate IP-based rule (see below) — guest mode doesn't widen or narrow it.- Turning
guest_read_onlyback off closes anonymous access on the very next request; nothing is cached that would keep it open.
If you enable this, treat status.json itself as public information for
anyone who can reach the port — it's the explicit tradeoff you're making.
The built-in server can terminate TLS itself
(status_server.tls_cert_path/tls_key_path in config.yaml, see
configuration.md) or you can put a reverse proxy in
front of it and leave those unset. Either way, anything credentialed —
admin login/session, CSRF, and Bearer-token status requests — is refused
over a connection the server can't otherwise vouch for. Concretely, one of
the following three has to be true:
- The connection to repowatch itself uses TLS. For a proxy this requires an HTTPS backend connection to the built-in TLS server; terminating TLS at the proxy and forwarding plain HTTP instead uses rule 2.
- It arrives through a
trusted_proxies-listed peer that declaresX-Forwarded-Proto: https— this is real, live-checked support for a TLS-terminating reverse proxy, not something you needallow_insecure_httpfor. An UNLISTED peer'sX-Forwarded-Protoheader is never trusted for this (see the next section for why). - It arrives directly on loopback — no forwarding headers, not routed
through any proxy. The server can be certain there's no untrusted
network hop between it and itself in this specific case, so this is
always allowed regardless of
allow_insecure_http. This is what letscurl http://127.0.0.1:.../api/auth/loginwork on the same host out of the box, with nothing configured.
If none of those three hold — a REMOTE, un-proxied, plain-HTTP
connection — status_server.allow_insecure_http: true is the explicit
opt-out: set it if you deliberately want to run that way (e.g. a
network you already trust end-to-end without TLS). Passwords and tokens
travel unencrypted on the wire whenever this is what's actually in effect
— it's a conscious tradeoff you're opting into, not a default.
GET /login (the login page itself, not the dashboard) and GET /healthz
are reachable over plain HTTP unconditionally, from anywhere (there's
nothing secret in either). GET //GET /dashboard are NOT unconditionally
reachable, though: without an admin session (or a satisfied
guest_read_only), they respond with a 303 redirect to /login rather
than serving the dashboard's actual HTML — so "the dashboard's static HTML
shell is public" isn't quite accurate; what's public is the separate login
page you get redirected to.
Each repository in status.json includes pending_replacements, the number
of detected same-version replacements awaiting cache invalidation or warming.
Host tokens can read this counter under their normal repository scope.
A nonzero value means index freshness alone is not sufficient to conclude
that replacement bytes are available in the cache. /healthz is unchanged.
History adds modified_packages for these replacements; history still
requires an admin session or guest-read access, not a host token.
See replacement handling.
If you run repowatch behind a reverse proxy, the proxy's own address is what
the server sees as the "client" by default — which matters for
metrics_allowed_networks (below) and for what's logged as the requesting
IP. List your proxy's address(es) in status_server.trusted_proxies (see
configuration.md) to have X-Forwarded-For /
X-Forwarded-Proto / X-Forwarded-Host from that peer honored instead.
This is deliberately conservative: a peer that isn't in trusted_proxies
can't just claim to be forwarding for someone else by sending those headers
itself — they're only honored from peers you've explicitly named. An
unconfigured, non-loopback peer sending forwarding headers is treated as
suspicious input, not trusted data.
GET /metrics(Prometheus text format) is restricted by client network, not by password or token —status_server.metrics_allowed_networks(default: loopback only). Widen it to your monitoring subnet if Prometheus scrapes from somewhere else. This is deliberately a different mechanism from the admin/token/guest system above: a scraper is a different kind of client than a browser or a package-manager host, and tying it to the admin password or a host token would mean either sharing admin credentials with your monitoring stack or teaching every scraped target's token about metrics it has nothing to do with.GET /healthzis completely open (only{"healthy": true/false}, HTTP 200/503). It reports catalog readiness: every configured repository must have a successful check no older than three times its effective check interval. A newly added repository or one failing from its first check returns 503 until it succeeds. A successful empty catalog is valid; an installation with no configured repositories returns 200. An invalid current configuration returns 503, even if the scheduler continues with its last valid configuration./api/v1/healthzhas the same contract. This is not a process-liveness probe: restarting the service cannot repair an unavailable upstream. A green result does not guarantee that packages are cached or that every download will succeed.
- No per-IP or per-token rate limiting on reads. Read access (status API, dashboard, packages/history) is not rate-limited by client IP. A single address can legitimately represent an entire NAT'd network or reverse proxy fronting many real clients, so an IP-based quota would punish shared infrastructure, not abuse. If you need to defend against actual abusive traffic, do it at the reverse proxy / firewall layer, which has better visibility into real client identity than repowatch does.
- No OAuth/SSO/multi-user accounts. There's one administrator password, not a user database — this is a small self-hosted tool for a small operations team, not a multi-tenant service. If you need per-person audit trails for admin actions, that's out of scope today.
In the administrator dashboard, open Storage → Measure cache coverage and size.
The button calls GET /api/storage/completeness?usage=1, using the same
administrator session cookie. Omit usage=1 for the coverage-only response. Guest access and host bearer tokens do not grant
access to this scan. It is separate from the minimal status.json response and
is never run by dashboard polling.
Enable nginx.enable_cache_probe and apply the generated nginx configuration
first. The measurement compares the configured repositories' current index
entries with nginx's on-disk inventory. It includes packages outside warming
whitelists and packages fetched by clients without a warm record. Deduplicated
packages use the same accepted canonical rewrites as the nginx renderer; an old
copy under an unused key or key version does not count. Results assume nginx is
running the current generated routing/dedup configuration.
A successful response contains source: "cache_inventory", started_at,
finished_at, failed_leaves, unreadable_entries, and items. Each item has:
| Field | Meaning |
|---|---|
repo_id |
Configured repository id |
total_packages |
Number of index entries, or null before the first snapshot |
cached_packages |
Entries whose required cache files were observed |
missing_packages |
Entries absent from an otherwise complete inventory |
unknown_packages |
Entries without enough evidence to classify |
percent |
cached_packages / total_packages * 100, rounded to two decimals; null when empty or uncertain |
state |
complete, partial, empty, or unknown |
last_check |
Catalog check time, or null |
reason |
no_snapshot, source_changed, incomplete_evidence, or null |
For example, a ten-package catalog with seven observed files and a complete scan
reports cached_packages: 7, missing_packages: 3, unknown_packages: 0,
percent: 70, and state: "partial". If part of the scan failed, the three
unobserved packages become unknown and percent becomes null. Known hits are
still reported. An empty successfully checked catalog is empty, not 100%.
For Nix, a root narinfo alone does not establish coverage. The scan uses the recorded closure artifact set, including metadata and NAR files. Undiscovered closures are unknown. Artifact history retains old encodings for purge, so missing recorded files conservatively produce unknown rather than proving the current closure incomplete. No Nix evaluation or upstream discovery runs during this scan.
The endpoint returns 409 when probing is disabled, another completeness scan or dedup cleanup/preview is active in the same process, or configuration/catalog snapshots changed during the scan; 502 means the inventory endpoint failed or returned invalid data. Retry after correcting the cause. Large scans can take minutes; configure any fronting proxy's response timeout accordingly. This measures file presence over a time interval, not an atomic snapshot, freshness, signature validity or a guarantee of offline operation. Cache files can expire or change afterwards.
The usage=1 option adds storage_usage to the completeness response, using the
same inventory and observation timestamps. It does not perform a second scan.
Catalog rows are streamed after the network inventory completes; storage ownership
and availability maps retain only keys observed in that inventory.
Even with no repositories configured, this option scans existing cache files and
reports them as unattributed. The coverage-only request still skips an empty
configuration. The existing GET /api/stats?cache_dir=1 and repowatch stats --cache-dir continue to provide overall cache-directory statistics.
Two byte counts are shown for each repository and group:
- Physical bytes: each observed cache file is assigned once to its direct key's owner. Canonical dedup copies belong to their canonical repository. Repositories sharing the same direct route use the alphabetically first known repository id as the accounting owner. Group physical bytes are the sum of their repository owners' bytes.
- Available bytes: distinct observed keys referenced by current catalog entries, after canonical rewrites. Each repository using a shared copy counts it; each group counts it once across its members. These numbers overlap between repositories and groups and must not be summed as physical usage.
For example, repositories A and B in group G share a 100 MiB cache file owned by A. Their physical counts are 100 MiB and 0; their available counts are both 100 MiB. Group G reports 100 MiB physical and 100 MiB available. If A and B are in different groups, both groups have 100 MiB available, while the physical counts still total 100 MiB.
All sizes are cache-file lengths including nginx headers, not package payload
sizes or filesystem allocated blocks (du may differ). Available bytes are file
presence, not proof of freshness, integrity, a completed replacement or a complete
Nix closure. Nix accounting includes known artifacts for current roots; historical
encodings retained in artifact bookkeeping may also contribute.
Physical attribution includes known old warmed files and copies under both dedup key bases. The files must actually appear in the inventory: bookkeeping alone adds no bytes. Files outside known catalogs/artifact/warm history, metadata, unknown keys and old key versions remain unattributed. Repositories without a snapshot or with an outdated source identity have unavailable attribution, not a claimed zero-byte cache. Group rows show the number of such repositories.
storage_usage contains:
| Field | Meaning |
|---|---|
total_bytes, file_count |
Known bytes and files observed across the scan |
attributed_bytes |
Sum of repository physical bytes |
unattributed_bytes, unattributed_files |
Observed files not assigned to a repository |
unreadable_sizes, unreadable_keys, failed_leaves |
Gaps in the scan |
inventory_complete |
Whether those gap counters are all zero; not a guarantee of atomicity or complete attribution |
repositories |
Rows with repo_id, group, physical_bytes, physical_files, available_bytes, available_files, and nullable reason |
groups |
Rows with nullable group, both byte/file counts, and unknown_repositories |
Repository reason is no_snapshot, source_changed, or null; an ungrouped
repository has group: null. The invariant is
sum(repository.physical_bytes) + unattributed_bytes == total_bytes, including
partial scans. Unreadable sizes contribute no guessed bytes; failed leaves are
not counted. When scan gaps exist, the displayed byte totals are known subtotals
and may be undercounts. A complete inventory can still contain unattributed files.
Storage → Preview redundant copies scans the physical cache on demand and lists obsolete copies whose current canonical file is also present. It shows up to 1000 candidates and the observed byte total; unknown sizes contribute no bytes. Select copies and click Remove selected copies. The browser submits batches of at most 256 keys and keeps partial results if a batch stops. Scan again to refresh the list or continue beyond the preview limit.
The administrator-only API is GET /api/storage/dedup for preview and
POST /api/storage/dedup with the usual session cookie and CSRF token:
{"generation": "<generation returned by preview>", "keys": ["<key returned by preview>"]}A POST rederives candidates on the server; submitting an arbitrary key cannot
turn this into a general purge endpoint. The response contains per-key results,
observed_removed_bytes, and stopped (null if the batch finished). The counts
object separates examined, source_missing, canonical_missing, eligible
and purged; eligible counts pairs whose two copies were found before purge,
not necessarily successful removals. absent_keys lists sources observed absent
or successfully removed during this batch. These are observations, not permanent
absence guarantees: concurrent clients can populate a key again. A skipped key
no longer qualifies or lacks one of its copies. HTTP 409 means the evidence
is stale, the features are disabled, or another cleanup/preview is running.
A mid-batch change returns partial results with stopped; inspect this field
even on HTTP 200. The byte count is the pre-purge observed file size, including
nginx headers, not a filesystem allocation measurement.
The same cleanup runs automatically after hourly retention when nginx, dedup, purge and cache probing are enabled. It retains current canonical keys and warmed bookkeeping. It does not implement warm-list enforcement: client-fetched packages outside the warm list are still valid cached packages. See dedup cleanup limits for the safety checks and scheduling limits.