Skip to content

feat: group hosts, and act on a whole group at once - #16

Merged
stahnma merged 1 commit into
mainfrom
feat/host-groups
Oct 1, 2026
Merged

stahnma merged 1 commit into
mainfrom
feat/host-groups

Conversation

@stahnma

@stahnma stahnma commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Every action the dashboard offered was one host at a time: one row to expand, one button, one systems.commands.<verb>.<hostname> subject. Patching a dozen machines meant a dozen of each. A group is a named set of hosts that check-in, update and reboot can all be run against together.

Two kinds, differing only in where membership comes from:

  • Ordinary groups you make and fill yourself.
  • Derived groups the server computes from what hosts already report — os:fedora, pkg:rpm, arch:x86_64, state:needs-reboot. No maintenance: one exists only while something matches it.
All hosts 6  ◆arch:aarch64 2  ◆arch:x86_64 4  ◆os:debian 1  ◆os:fedora 2
◆pkg:deb 2  ◆pkg:rpm 3  ◆state:has-updates 3  ◆state:needs-reboot 2
prod 3  Ungrouped 2  + New group

Design notes

Membership lives in its own bbolt bucket, not on models.System. A client publishes a whole System on every check-in and the server stores what it sent, so anything server-owned kept there must be carried forward by hand — there are already three such fields, and the next person to add one would have no reason to suspect groups were among them. It also keeps groups out of SystemSummary, so the reflection tests pinning those together are untouched.

The per-host precondition ladder is now written once. All three single-host handlers ran the same sequence, and a group must skip exactly what they refuse. Had the two drifted, the dashboard would offer a button for one host and silently do nothing for that same host inside a group — a patch window gone by with a machine left behind. dispatch() holds the ladder, statusFor() the status mapping, and a test drives both paths from one setup requiring the statuses to match. The handlers keep their wording to the character, so their 620 lines of tests needed no edit; routes.go came out 120 lines shorter.

Members are asked all at once, and a host that cannot take the command is skipped rather than failing the batch. One machine declining should not stop the other eleven being patched. Each request is already bounded at 10s, so a group of fifty finishes in about as long as the slowest single host.

Status is 200 whenever the group exists and the feature is on, however the hosts answered. No code summarises a dozen different answers; 207 would be decoration (a WebDAV code whose body is defined XML, and fetch already treats it as ok). The body is what to read, and the dashboard renders it into a panel that stays up rather than an alert gone on dismissal.

Derived groups resolve at dispatch time, which is what makes state:needs-reboot mean what is true as you press it. They refuse every edit at every route, and the UI draws no Rename/Delete for one. : is reserved in ordinary names so a hand-made os:fedora cannot shadow the computed one.

Package family is inferred from the OS string the host already reports — the same PRETTY_NAME used to pick a distro icon. Nothing to configure, nothing to roll out, works for hosts on older clients. The cost is a matching table; a distribution missing from it joins no os:/pkg: group at all, keeping only its arch: group. That silence is deliberate — a gap is a one-line fix, a wrong guess hands a host a dnf command it cannot run.

The table stays flat, filtered by chips rather than divided into sections. One row per host is an invariant the file leans on — eight places reach for a hostname with document.querySelector — and a host in several groups drawn several times would leave all of them driving only the first copy.

Blast radius

The existing caveat that this is a convenience for a trusted network, not an authorization boundary, now matters more: one unauthenticated POST reboots twelve machines instead of one. Nothing changes about who may do what — a host that has not opted in is still untouchable — but every group action logs at WARN with the group, action, requesting address and counts. The README says this in as many words.

Testing

make test, make build and go vet all pass. New coverage: the storage layer (which had none before), the shared ladder, group CRUD, fan-out ordering and concurrency, mixed-outcome batches, derived derivation and every edit guard.

Verified end-to-end against an isolated server with real NATS and scripted hosts — every outcome exercised:

POST /api/groups/pkg:rpm/update
  3 requested / 1 accepted / 1 skipped / 1 failed
   db01   skipped   not_opted_in   This host has not opted into remote updates...
   web01  accepted
   web02  refused   host_refused   an update is already running on this host

Not verified: the browser behaviour of app.js. There is no JS engine on the machine this was built on, so the ~600 new lines have had only a structural check (balanced brackets, strings and template literals, validated against the pre-change file first). Worth a click-through before merge.

Every action the dashboard offered was one host at a time: one row to
expand, one button, one `systems.commands.<verb>.<hostname>` subject.
Patching a dozen machines meant a dozen of each, in a table that
re-rendered underneath you whenever any of them checked back in. A group
is a named set of hosts that check-in, update and reboot can all be run
against together.

There are two kinds, and only where the membership comes from differs.
Ordinary groups are made and filled by hand. Derived groups are computed
by the server from what the hosts already report -- os:fedora, pkg:rpm,
arch:x86_64, state:needs-reboot -- and need no maintenance at all,
because "every Fedora box" is a fact about the fleet rather than a
decision about it.

Membership lives in its own bbolt bucket rather than on models.System.
A client publishes a whole System on every check-in and the server
stores what it sent, so anything server-owned kept there has to be
carried forward by hand; there are already three such fields and the
next person to add one would have no reason to suspect groups were among
them. Keeping groups out of the system record also keeps them out of
SystemSummary, so the reflection tests that pin the two together are
untouched -- the dashboard reads /api/groups and inverts it itself.

Deleting a host purges it from every group in the same transaction. The
alternative was tempting, since the delete dialog says the row will
reappear when the host next checks in, but a name left behind in a group
is a member nothing can act on, quietly padding the count of every
action from then on.

The per-host precondition ladder is now written once. Each of the three
single-host handlers ran the same sequence -- look the host up, check
whichever opt-ins the action needs, send, read the ack -- and a group
has to skip exactly what those handlers refuse. Had the two drifted, the
dashboard would have offered a button for one host and silently done
nothing for that same host inside a group, which is a patch window gone
by with a machine left behind and nobody any the wiser. So dispatch()
holds the ladder, statusFor() holds the one outcome-to-status mapping,
and a test drives both paths from one setup and requires the statuses to
match. The handlers keep their wording to the character, which is why
their 620 lines of tests needed no edit; routes.go came out 120 lines
shorter.

Members are asked all at once. The fleet is tens of hosts rather than
thousands, each request is already bounded at ten seconds, and the hosts
rate-limit themselves anyway, so a group of fifty finishes in about as
long as the slowest single one and there is no deadline here to invent a
failure mode with. A host that cannot take the command is skipped rather
than failing the batch -- one machine declining should not stop the
other eleven being patched -- and every member comes back with its own
outcome, with skipped and refused kept apart because the first is what
the server knew before it sent anything and the second is what the host
said back.

The status is 200 whenever the group exists and the feature is on,
however the hosts answered. The request was "fan this out and tell me
what happened" and it did precisely that; no code can summarise a dozen
different answers, and 207 would have been decoration -- a WebDAV code
whose body is a defined XML document, which fetch already treats as ok.
The body is the thing to read, and the dashboard reads it into a panel
that stays up rather than an alert that is gone on dismissal, because a
reboot that skipped two hosts is something to keep reading while you go
and look at them.

Derived groups are recomputed on every read, so they are always current
and there is never a stale one to clean up; one exists only while
something matches it, and its members are resolved when the action runs
rather than when the page was drawn, which is what makes "reboot
everything that needs a reboot" mean what is true as you press it. They
refuse every edit at every route, and the dashboard draws no Rename or
Delete for one, since offering a control and then refusing it is worse
than not offering it. ":" is reserved in an ordinary group name so a
hand-made os:fedora can never shadow the computed one and leave an
action ambiguous about which hosts it meant.

The package family is inferred from the OS string the host already
reports, the same PRETTY_NAME the dashboard reads to pick a distro icon.
That asks nothing of the client and groups every host already checking
in, including ones on an older build. The cost is a matching table, and
a distribution missing from it joins no os: or pkg: group at all --
keeping its arch: group, which is reported rather than guessed. The
silence is deliberate: a gap in the list is a one-line fix, and a wrong
guess hands a host a dnf command it cannot run.

The table stays flat, filtered by the chips rather than divided into
sections. One row per host is an invariant the whole file leans on --
eight places reach for a hostname with document.querySelector -- and a
host in several groups drawn several times would have left every one of
them driving only the first copy. Row labels show only your own groups,
since os:fedora and arch:x86_64 beside every hostname would just restate
two columns that are already there. Accepted hosts are fed through the
existing pending maps, so each row shows the same spinners and timeouts
it would had you pressed every button yourself, which is all a group
action is.
@stahnma
stahnma merged commit 2a05da8 into main Oct 1, 2026
1 check passed
@stahnma
stahnma deleted the feat/host-groups branch October 1, 2026 19:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant