From db0e01305e32966e6d8b2af8a2e1e36f3bf2c8a4 Mon Sep 17 00:00:00 2001 From: Michael Stahnke Date: Thu, 1 Oct 2026 13:40:01 -0400 Subject: [PATCH] feat: group hosts, and act on a whole group at once Every action the dashboard offered was one host at a time: one row to expand, one button, one `systems.commands..` subject. Patching a dozen machines meant a dozen of each, in a table that re-rendered underneath you whenever any of them checked back in. A group is a named set of hosts that check-in, update and reboot can all be run against together. There are two kinds, and only where the membership comes from differs. Ordinary groups are made and filled by hand. Derived groups are computed by the server from what the hosts already report -- os:fedora, pkg:rpm, arch:x86_64, state:needs-reboot -- and need no maintenance at all, because "every Fedora box" is a fact about the fleet rather than a decision about it. Membership lives in its own bbolt bucket rather than on models.System. A client publishes a whole System on every check-in and the server stores what it sent, so anything server-owned kept there has to be carried forward by hand; there are already three such fields and the next person to add one would have no reason to suspect groups were among them. Keeping groups out of the system record also keeps them out of SystemSummary, so the reflection tests that pin the two together are untouched -- the dashboard reads /api/groups and inverts it itself. Deleting a host purges it from every group in the same transaction. The alternative was tempting, since the delete dialog says the row will reappear when the host next checks in, but a name left behind in a group is a member nothing can act on, quietly padding the count of every action from then on. The per-host precondition ladder is now written once. Each of the three single-host handlers ran the same sequence -- look the host up, check whichever opt-ins the action needs, send, read the ack -- and a group has to skip exactly what those handlers refuse. Had the two drifted, the dashboard would have offered a button for one host and silently done nothing for that same host inside a group, which is a patch window gone by with a machine left behind and nobody any the wiser. So dispatch() holds the ladder, statusFor() holds the one outcome-to-status mapping, and a test drives both paths from one setup and requires the statuses to match. The handlers keep their wording to the character, which is why their 620 lines of tests needed no edit; routes.go came out 120 lines shorter. Members are asked all at once. The fleet is tens of hosts rather than thousands, each request is already bounded at ten seconds, and the hosts rate-limit themselves anyway, so a group of fifty finishes in about as long as the slowest single one and there is no deadline here to invent a failure mode with. A host that cannot take the command is skipped rather than failing the batch -- one machine declining should not stop the other eleven being patched -- and every member comes back with its own outcome, with skipped and refused kept apart because the first is what the server knew before it sent anything and the second is what the host said back. The status is 200 whenever the group exists and the feature is on, however the hosts answered. The request was "fan this out and tell me what happened" and it did precisely that; no code can summarise a dozen different answers, and 207 would have been decoration -- a WebDAV code whose body is a defined XML document, which fetch already treats as ok. The body is the thing to read, and the dashboard reads it into a panel that stays up rather than an alert that is gone on dismissal, because a reboot that skipped two hosts is something to keep reading while you go and look at them. Derived groups are recomputed on every read, so they are always current and there is never a stale one to clean up; one exists only while something matches it, and its members are resolved when the action runs rather than when the page was drawn, which is what makes "reboot everything that needs a reboot" mean what is true as you press it. They refuse every edit at every route, and the dashboard draws no Rename or Delete for one, since offering a control and then refusing it is worse than not offering it. ":" is reserved in an ordinary group name so a hand-made os:fedora can never shadow the computed one and leave an action ambiguous about which hosts it meant. The package family is inferred from the OS string the host already reports, the same PRETTY_NAME the dashboard reads to pick a distro icon. That asks nothing of the client and groups every host already checking in, including ones on an older build. The cost is a matching table, and a distribution missing from it joins no os: or pkg: group at all -- keeping its arch: group, which is reported rather than guessed. The silence is deliberate: a gap in the list is a one-line fix, and a wrong guess hands a host a dnf command it cannot run. The table stays flat, filtered by the chips rather than divided into sections. One row per host is an invariant the whole file leans on -- eight places reach for a hostname with document.querySelector -- and a host in several groups drawn several times would have left every one of them driving only the first copy. Row labels show only your own groups, since os:fedora and arch:x86_64 beside every hostname would just restate two columns that are already there. Accepted hosts are fed through the existing pending maps, so each row shows the same spinners and timeouts it would had you pressed every button yourself, which is all a group action is. --- README.md | 218 ++++++++++ server/api/actions.go | 257 ++++++++++++ server/api/actions_test.go | 333 +++++++++++++++ server/api/groups.go | 464 +++++++++++++++++++++ server/api/groups_test.go | 669 +++++++++++++++++++++++++++++++ server/api/routes.go | 132 +----- server/api/routes_test.go | 175 ++++++++ server/models/derived.go | 194 +++++++++ server/models/derived_test.go | 215 ++++++++++ server/models/groups.go | 83 ++++ server/nats/subscriber_test.go | 12 + server/storage/bbolt.go | 10 +- server/storage/groups.go | 384 ++++++++++++++++++ server/storage/groups_test.go | 380 ++++++++++++++++++ server/storage/storage.go | 10 + server/web/server.go | 15 + server/web/static/apidoc.css | 6 + server/web/static/app.js | 596 ++++++++++++++++++++++++++- server/web/static/styles.css | 372 +++++++++++++++++ server/web/templates/apidoc.html | 408 +++++++++++++++++++ server/web/templates/index.html | 9 +- 21 files changed, 4822 insertions(+), 120 deletions(-) create mode 100644 server/api/actions.go create mode 100644 server/api/actions_test.go create mode 100644 server/api/groups.go create mode 100644 server/api/groups_test.go create mode 100644 server/models/derived.go create mode 100644 server/models/derived_test.go create mode 100644 server/models/groups.go create mode 100644 server/storage/groups.go create mode 100644 server/storage/groups_test.go diff --git a/README.md b/README.md index 6307eb3..00fbbe3 100644 --- a/README.md +++ b/README.md @@ -80,6 +80,8 @@ Configuration is powered by [Viper](https://github.com/spf13/viper). | `remote_updates` | `MUC_REMOTE_UPDATES` | `false` | Allow the dashboard to run updates on hosts that have opted in (see [Running updates from the dashboard](#running-updates-from-the-dashboard)) | | `remote_reboot` | `MUC_REMOTE_REBOOT` | `false` | Allow the dashboard to reboot hosts that have opted in (see [Rebooting a host from the dashboard](#rebooting-a-host-from-the-dashboard)) | +Grouping hosts needs no configuration at all — see [Grouping hosts](#grouping-hosts). The group routes are always available; a group update or reboot is gated by the same `remote_updates` / `remote_reboot` flags as its single-host counterpart. + **CLI Flags:** - `--dev`: Enable dev mode (debug logging enabled) - `--json`: Output logs in JSON format (default: text format) @@ -339,6 +341,7 @@ The web dashboard provides: - Tailnet status for hosts that use Tailscale (see below) - A check-in button on every host, to refresh a row now rather than at its next poll (see [Asking a host to check in](#asking-a-host-to-check-in)) - An update button for hosts that have opted in (see [Running updates from the dashboard](#running-updates-from-the-dashboard)) +- Group chips that filter the table, and run check-in, updates or a reboot across a whole group at once — groups you make yourself, plus derived ones like `os:fedora`, `pkg:rpm` and `state:needs-reboot` that the server works out for itself (see [Grouping hosts](#grouping-hosts)) ### Tailnet status @@ -660,6 +663,221 @@ failed** with the command's own error, and the control comes back. The request is logged on the host before the reboot, with the address it came from, so `journalctl -u muc-client -b -1` says who did it. +## Grouping hosts + +Patching a dozen machines one row at a time is the thing this dashboard was +worst at. A **group** is a named set of hosts that the three per-host +actions — check in, update, reboot — can be run against all at once. + +There are two kinds. **Ordinary groups** are the ones you make and fill +yourself: `prod`, `the noisy ones in the basement`. **Derived groups** are +computed by the server from what the hosts already report — every Fedora box, +everything rpm-based, everything needing a reboot — and need no maintenance at +all. They behave identically once they exist; only where their membership +comes from differs. + +Groups are kept by the server and managed from the dashboard. **Nothing changes +on a host**: there is no key in `/etc/muc/client.yml`, no client to restart, and +a host is never told which groups it is in. That is deliberate. A group is a +convenience for whoever is doing the patching, not a property of the machine, +and putting it in the client config would make it one more thing to deploy and +keep in step with a dashboard that can already see every host anyway. + +Nothing has to be configured on the server either. The group routes are always +on; what stays gated is what was already gated — a group update still needs +`remote_updates`, and a group reboot still needs `remote_reboot`. + +A host can be in as many groups as you like. `web01` can be in `prod` and `web` +and `rocky10` at once, and show up under each. + +### Using them + +The bar above the table holds a chip per group. Clicking one filters the table +to its members and reveals the actions for that group; clicking **All hosts** +puts it back. **Ungrouped** is there too, which is the quickest way to find a +host you forgot to file. + +The table itself stays flat — one row per host, however many groups it is in — +and each row carries its groups as small labels beside the hostname. + +To put a host in a group, expand its row: the **Groups** block in the details +holds a checkbox per group. New groups are made from **+ New group** on the bar, +which is a separate act on purpose: a typo that founds a group of one is much +harder to notice afterwards than a typo that is simply refused. + +### Derived groups + +Some groups are not worth maintaining by hand, because the server can already +see the answer. "Every Fedora box" and "everything rpm-based" are facts about +the fleet rather than decisions about it, so MUC derives them for you. + +They appear in the bar marked with a ◆ and a dashed outline, ahead of your own +groups: + +| Group | Members | +|---|---| +| `os:fedora`, `os:debian`, `os:rocky`, `os:ubuntu`, … | one per distribution actually present | +| `pkg:rpm`, `pkg:deb`, `pkg:pacman`, `pkg:nix`, `pkg:brew`, … | the package family, which is what `all rpm` and `all deb` mean | +| `arch:x86_64`, `arch:aarch64` | one per architecture present | +| `state:needs-reboot`, `state:has-updates` | what the fleet is currently reporting | + +A derived group exists only while something matches it. Boot a Debian machine +and `pkg:deb` appears; retire the last one and it is gone. There is nothing to +create and nothing to clean up, which is the point — a list of every +distribution MUC has heard of would be noise, so the bar shows the fleet you +actually have. + +Anything you can do to a group you can do to a derived one: + +```bash +curl -X POST http://muc-server:8080/api/groups/pkg:rpm/update +curl -X POST http://muc-server:8080/api/groups/state:needs-reboot/reboot +``` + +That second one is worth noticing. A derived group's members are worked out +when the action runs, not when the page was drawn, so "reboot everything that +needs a reboot" means what is true at the moment you press it. + +**They cannot be edited.** There is no renaming, no deleting, and no adding a +host — a machine joins `os:fedora` by being a Fedora box, and the dashboard +draws the controls accordingly rather than offering them and then refusing. The +expanded row shows a host's derived groups as plain labels beside its editable +ones. If you want a hand-picked set, make an ordinary group. + +The `:` in the name is what keeps the two kinds apart, so it is reserved: an +ordinary group name may not contain one. Without that, a hand-made `os:fedora` +could sit beside the computed one and a group action would be ambiguous about +which set of hosts it meant. + +The derived groups are also why **Ungrouped** means "in none of *your* groups". +Every host is in several derived groups, so counting those would make it +permanently empty and useless. + +#### Where the distribution comes from + +The package family is inferred from the OS string the host already reports — +its own `PRETTY_NAME`, the same string the dashboard uses to pick a distro +icon. Nothing has to be configured and nothing has to be upgraded: every host +already checking in is grouped, including ones running an older client. + +The trade is that it is a matching table, in `server/models/derived.go`, and a +distribution it does not recognise joins **no** `os:` or `pkg:` group at all — +it still gets its `arch:` group, since the architecture is reported rather than +guessed. That silence is deliberate. Guessing is how a host ends up in `all +rpm` and gets handed a `dnf` command it cannot run; a distribution missing from +the list is a one-line fix, and a wrong guess is an incident. + +SUSE is listed as `pkg:rpm`. It reaches rpm through zypper rather than dnf, but +"which hosts take an rpm" has one answer, and the `upd` script picks the right +manager per host regardless. + +### What a group action does + +All the members at once, not one after another. The fleet this serves is tens of +hosts rather than thousands; running a group check-in in sequence would make it +take minutes for no benefit, and the hosts enforce their own minimum gap between +commands regardless. A group of fifty finishes in about as long as the slowest +single host. + +**A member that cannot take the command is skipped, not an error.** One machine +that has not opted into remote updates should not stop the other eleven being +patched. Every member comes back with its own outcome, and the report under the +action bar lists them: + +| Outcome | What it means | +|---|---| +| **accepted** | The host took the command. It reports back separately, exactly as it does for the single-host button. | +| **skipped** | It was never a candidate: not opted in, no reboot pending, or not a host this server knows. | +| **refused** | The host itself said no, and said why — "an update is already running on this host", say. | +| **unreachable** | Nothing answered. The host is off, or its client no longer listens for that command. | +| **failed** | The request broke for some other reason. | + +*Skipped* and *refused* are worth telling apart: the first is what the server +knew before it sent anything, the second is what the host said back. + +A group reboot only reboots members that reported a pending reboot, which is +what makes it safe to press straight after a group update — it reboots what +needs it and leaves the rest alone. It uses the same **Confirm reboot** checkbox +as the per-host control, for the same reason and with no extra dialog. Check-in +and update have no checkbox; they are the ones you will press most often. + +Accepted hosts get the same ⏳ indicators in their rows as if you had pressed +each button yourself, because that is all a group action is. + +### Members the server does not know + +A group can name a host that has never checked in, and it keeps naming a host +whose row you delete. Both are intentional. Deleting a row is a tidying gesture +— the dashboard says as much, "it will reappear when it checks in again" — so +losing the grouping you did by hand would be a poor trade; and building a group +before its machines exist is a reasonable way to work. + +Such a member shows in the group's count as "not currently known", and a group +action skips it. To get rid of one for good: + +```bash +curl -X DELETE http://muc-server:8080/api/groups/prod/members/retired01 +``` + +### The same thing over the API + +```bash +# make a group and fill it +curl -X POST http://muc-server:8080/api/groups -d '{"name":"prod"}' +curl -X PUT http://muc-server:8080/api/groups/prod \ + -d '{"members":["web01","web02","db01"]}' + +# run updates across it +curl -X POST http://muc-server:8080/api/groups/prod/update +``` + +```json +{ + "group": "prod", + "action": "update", + "requested": 3, + "accepted": 2, + "skipped": 1, + "failed": 0, + "results": [ + { "hostname": "db01", "outcome": "skipped", "code": "not_opted_in", + "reason": "This host has not opted into remote updates (set allow_remote_updates: true in its client config)" }, + { "hostname": "web01", "outcome": "accepted", "id": "9f2c1b0a4d5e6f70" }, + { "hostname": "web02", "outcome": "accepted", "id": "1a2b3c4d5e6f7081" } + ] +} +``` + +The status is **200 whenever the group exists and the feature is enabled**, even +when every member was skipped. The call was "fan this out and tell me what +happened", and it did; no status code can summarise a dozen different answers, +so it does not try. Read the body. + +Group names ignore case — `prod` and `Prod` are one group, and cannot both +exist, because two chips that look alike and hold different hosts is the one +mistake here that goes unnoticed. + +### How much this is trusted + +The earlier sections say that remote updates and reboots are a convenience for a +trusted network, not an authorization boundary. That is still true and now it +matters more: **the blast radius of a group action is the whole group.** One +unauthenticated HTTP POST, from anything that can reach the dashboard, now +reboots twelve machines instead of one. + +Nothing here changes who is allowed to do what — a host that has not opted in is +still untouchable, and that remains the only real gate. But two things follow +from it. Keep a group no larger than the thing you actually want to act on. And +`remote_reboot` has no business being enabled anywhere the dashboard is +reachable by something you do not trust. + +Every group action is logged at WARN with the group, the action, the address it +came from and the counts, so "who rebooted production" is one `grep` away: + +```bash +journalctl -u muc-server | grep "Group action" +``` + ## Alternatives Instead of using this tool, you could run a cron job or systemd timer to auto-update. However, this approach has drawbacks: diff --git a/server/api/actions.go b/server/api/actions.go new file mode 100644 index 0000000..78e24f6 --- /dev/null +++ b/server/api/actions.go @@ -0,0 +1,257 @@ +package api + +import ( + "encoding/json" + "errors" + "log/slog" + "net/http" + "server/models" + "server/storage" + "sync" +) + +// action names the three things the dashboard can ask of a host. +// +// It exists so the precondition ladder below is written once instead of three +// times, and — the real point — so a group action is guaranteed to skip exactly +// what the single-host route would have refused. Those two paths drifting apart +// is the failure nobody would notice: the dashboard would offer a button for +// one host and quietly do nothing for the same host inside a group. +type action string + +const ( + actionUpdate action = "update" + actionCheckIn action = "checkin" + actionReboot action = "reboot" +) + +// The outcome of asking one host to do one thing. +// +// There are five rather than three because "skipped" and "refused" answer +// different questions, and an operator reading a group report needs to tell +// them apart: skipped is what the server knew before it sent anything, refused +// is what the host said back. +const ( + OutcomeAccepted = "accepted" + OutcomeSkipped = "skipped" + OutcomeRefused = "refused" + OutcomeUnreachable = "unreachable" + OutcomeFailed = "failed" +) + +// Machine-readable reasons, so the dashboard can style a result without +// matching on prose. +const ( + CodeUnknownHost = "unknown_host" + CodeNotOptedIn = "not_opted_in" + CodeNoRebootPending = "no_reboot_pending" + CodeHostRefused = "host_refused" + CodeNotListening = "not_listening" + CodeRequestFailed = "request_failed" +) + +// HostActionResult is what one host did when it was asked. Every member of a +// group gets exactly one, whether or not any command was actually sent to it. +// +// Reason is the same sentence the single-host route puts in its error body, so +// the two paths explain themselves identically. +type HostActionResult struct { + Hostname string `json:"hostname"` + Outcome string `json:"outcome"` + Code string `json:"code,omitempty"` + Reason string `json:"reason,omitempty"` + ID string `json:"id,omitempty"` + Command string `json:"command,omitempty"` +} + +// requesters bundles the three narrow interfaces so one value reaches dispatch. +// Any of them may be nil, but the nil check belongs to the caller: a feature is +// off once per request, not once per member, and the two routes that can be off +// say so differently. +type requesters struct { + update UpdateRequester + checkIn CheckInRequester + reboot RebootRequester +} + +// statusFor maps an outcome back to the HTTP status the single-host route has +// always returned for it. Keeping the mapping in one place is what lets a test +// assert that the group path skips precisely what the single-host path refuses. +func statusFor(res HostActionResult) int { + switch res.Outcome { + case OutcomeAccepted: + return http.StatusAccepted + case OutcomeSkipped: + if res.Code == CodeUnknownHost { + return http.StatusNotFound + } + return http.StatusConflict + case OutcomeRefused: + return http.StatusConflict + case OutcomeUnreachable: + return http.StatusServiceUnavailable + default: + return http.StatusBadGateway + } +} + +// dispatch runs the whole ladder for one host and one action: look the host up, +// check whichever opt-ins the action needs, send the command, read the answer. +// It never writes to the response — the single-host routes turn its result into +// a status code, the group routes turn it into one line of a result list. +func dispatch(store storage.Storage, act action, req requesters, hostname, requestedBy string) HostActionResult { + res := HostActionResult{Hostname: hostname} + + system, err := store.GetSystem(hostname) + if err != nil { + res.Outcome, res.Code, res.Reason = OutcomeSkipped, CodeUnknownHost, "System not found" + return res + } + + switch act { + case actionUpdate: + if !system.RemoteUpdatesEnabled { + res.Outcome, res.Code = OutcomeSkipped, CodeNotOptedIn + res.Reason = "This host has not opted into remote updates (set allow_remote_updates: true in its client config)" + return res + } + case actionReboot: + if !system.RemoteRebootEnabled { + res.Outcome, res.Code = OutcomeSkipped, CodeNotOptedIn + res.Reason = "This host has not opted into remote reboots (set allow_remote_reboot: true in its client config)" + return res + } + // The dashboard only enables its button when a reboot is pending, and + // the API should not be a way around that. The host checks again for + // itself when the command arrives. + if !system.RebootRequired { + res.Outcome, res.Code = OutcomeSkipped, CodeNoRebootPending + res.Reason = "This host did not report a pending reboot at its last check-in" + return res + } + } + + var ( + accepted bool + reason string + ) + switch act { + case actionUpdate: + ack, reqErr := req.update.RequestUpdate(hostname, requestedBy) + err, accepted, reason, res.ID, res.Command = reqErr, ack.Accepted, ack.Reason, ack.ID, ack.Command + case actionReboot: + ack, reqErr := req.reboot.RequestReboot(hostname, requestedBy) + err, accepted, reason, res.ID = reqErr, ack.Accepted, ack.Reason, ack.ID + case actionCheckIn: + ack, reqErr := req.checkIn.RequestCheckIn(hostname, requestedBy) + err, accepted, reason, res.ID = reqErr, ack.Accepted, ack.Reason, ack.ID + } + + switch { + case errors.Is(err, models.ErrHostNotListening): + // The stored opt-in said yes but nothing answered, so the host is down + // or its client has since been reconfigured. + res.Outcome, res.Code, res.Reason = OutcomeUnreachable, CodeNotListening, notListeningReason(act, hostname) + res.ID, res.Command = "", "" + return res + case err != nil: + slog.Error(string(act)+" request failed", "hostname", hostname, "error", err) + res.Outcome, res.Code = OutcomeFailed, CodeRequestFailed + res.Reason = failedPrefix(act) + ": " + err.Error() + res.ID, res.Command = "", "" + return res + } + + if !accepted { + if reason == "" { + reason = refusalReason(act) + } + res.Outcome, res.Code, res.Reason = OutcomeRefused, CodeHostRefused, reason + return res + } + + res.Outcome = OutcomeAccepted + return res +} + +// notListeningReason, failedPrefix and refusalReason keep the per-action +// wording the single-host routes have always used. The sentences differ by +// action on purpose — "too old to accept check-in requests" is true of a client +// from before check-ins existed, and would be wrong about an update. +func notListeningReason(act action, hostname string) string { + switch act { + case actionUpdate: + return "No response from " + hostname + ": it is offline, or its client is no longer accepting update commands" + case actionReboot: + return "No response from " + hostname + ": it is offline, or its client is no longer accepting reboot commands" + default: + return "No response from " + hostname + ": it is offline, or its client is too old to accept check-in requests" + } +} + +func failedPrefix(act action) string { + switch act { + case actionUpdate: + return "Update request failed" + case actionReboot: + return "Reboot request failed" + default: + return "Check-in request failed" + } +} + +func refusalReason(act action) string { + switch act { + case actionUpdate: + return "the host refused the update request" + case actionReboot: + return "the host refused the reboot request" + default: + return "the host refused the check-in request" + } +} + +// dispatchGroup fans one action out to every member at once and collects what +// each host said. +// +// Members are not paced. The point of a group is that it happens now; the fleet +// this serves is tens of hosts rather than thousands, serialising a group +// check-in would make it take minutes for no benefit, and the hosts enforce +// their own minimum gap between commands anyway. +// +// Results come back in the group's member order whatever order they finished +// in, so the list reads the same twice running. The whole call is bounded by +// one host's request timeout — ten seconds — rather than by the number of +// members, because they all wait at the same time. That is why there is no +// deadline here: one would only invent a failure mode. +// +// Nor is the request context threaded through. The three requester interfaces +// take no context and a published NATS request cannot be recalled, so +// cancelling the wait would end goroutines that were going to finish inside ten +// seconds anyway, while the commands it "cancelled" still ran. +func dispatchGroup(store storage.Storage, act action, req requesters, members []string, requestedBy string) []HostActionResult { + results := make([]HostActionResult, len(members)) + + var wg sync.WaitGroup + for i, hostname := range members { + wg.Add(1) + go func(i int, hostname string) { + defer wg.Done() + // Each goroutine owns one index of a pre-sized slice, so there is + // nothing to lock. + results[i] = dispatch(store, act, req, hostname, requestedBy) + }(i, hostname) + } + wg.Wait() + + return results +} + +// writeJSON is the success-path counterpart to writeAPIError. +func writeJSON(w http.ResponseWriter, status int, v any) { + w.Header().Set("Content-Type", "application/json") + w.WriteHeader(status) + if err := json.NewEncoder(w).Encode(v); err != nil { + slog.Error("Failed to write response", "error", err) + } +} diff --git a/server/api/actions_test.go b/server/api/actions_test.go new file mode 100644 index 0000000..035b80b --- /dev/null +++ b/server/api/actions_test.go @@ -0,0 +1,333 @@ +package api + +import ( + "net/http" + "server/models" + "slices" + "sync" + "testing" + "time" +) + +// optedInSystem returns a host that would accept any of the three commands, so +// a test only has to spell out the thing it is taking away. +func optedInSystem(hostname string) models.System { + return models.System{ + Hostname: hostname, + RemoteUpdatesEnabled: true, + RemoteRebootEnabled: true, + RebootRequired: true, + } +} + +// TestDispatchStatusMatchesSingleHostRoutes is the test this whole refactor +// exists for. +// +// The group routes classify a member by the status the single-host route would +// have returned for it. If the two ever disagree, the dashboard offers a button +// for one host and silently does nothing for that same host inside a group — +// a failure nobody would notice until a patch window went by with a machine +// left behind. Here the same setup is driven down both paths and the statuses +// are required to match. +func TestDispatchStatusMatchesSingleHostRoutes(t *testing.T) { + notListening := models.ErrHostNotListening + + for _, tc := range []struct { + name string + system models.System + ack any + err error + want int + }{ + {"accepted", optedInSystem("smallboi"), true, nil, http.StatusAccepted}, + {"host refused", optedInSystem("smallboi"), false, nil, http.StatusConflict}, + {"not listening", optedInSystem("smallboi"), true, notListening, http.StatusServiceUnavailable}, + {"request failed", optedInSystem("smallboi"), true, errNotFound{}, http.StatusBadGateway}, + } { + for _, act := range []action{actionUpdate, actionCheckIn, actionReboot} { + t.Run(string(act)+"/"+tc.name, func(t *testing.T) { + store := &fakeStorage{systems: []models.System{tc.system}} + accepted := tc.ack.(bool) + + var ( + got int + viaRoute int + ) + switch act { + case actionUpdate: + f := &fakeUpdater{ack: models.UpdateAck{Accepted: accepted, ID: "abc"}, err: tc.err} + got = statusFor(dispatch(store, act, requesters{update: f}, "smallboi", "tester")) + viaRoute = postUpdate(RunUpdateHandler(store, f), "smallboi").Code + case actionCheckIn: + f := &fakeCheckIner{ack: models.CheckInAck{Accepted: accepted, ID: "abc"}, err: tc.err} + got = statusFor(dispatch(store, act, requesters{checkIn: f}, "smallboi", "tester")) + viaRoute = postCheckIn(CheckInHandler(store, f), "smallboi").Code + case actionReboot: + f := &fakeRebooter{ack: models.RebootAck{Accepted: accepted, ID: "abc"}, err: tc.err} + got = statusFor(dispatch(store, act, requesters{reboot: f}, "smallboi", "tester")) + viaRoute = postReboot(RebootHandler(store, f), "smallboi").Code + } + + if got != tc.want { + t.Errorf("statusFor(dispatch) = %d, want %d", got, tc.want) + } + if viaRoute != got { + t.Errorf("the single-host route returned %d but dispatch classifies it as %d; "+ + "the group path and the per-host path have drifted apart", viaRoute, got) + } + }) + } + } +} + +// TestDispatchSkipsWhatTheRouteRefuses covers the rungs that are specific to +// one action: the opt-ins and the pending-reboot precondition. These are the +// ones a group has to skip rather than fail on. +func TestDispatchSkipsWhatTheRouteRefuses(t *testing.T) { + noUpdates := optedInSystem("smallboi") + noUpdates.RemoteUpdatesEnabled = false + + noReboots := optedInSystem("smallboi") + noReboots.RemoteRebootEnabled = false + + noPending := optedInSystem("smallboi") + noPending.RebootRequired = false + + for _, tc := range []struct { + name string + act action + system models.System + wantCode string + wantStatus int + }{ + {"update, host not opted in", actionUpdate, noUpdates, CodeNotOptedIn, http.StatusConflict}, + {"reboot, host not opted in", actionReboot, noReboots, CodeNotOptedIn, http.StatusConflict}, + {"reboot, nothing pending", actionReboot, noPending, CodeNoRebootPending, http.StatusConflict}, + } { + t.Run(tc.name, func(t *testing.T) { + store := &fakeStorage{systems: []models.System{tc.system}} + updater := &fakeUpdater{ack: models.UpdateAck{Accepted: true}} + rebooter := &fakeRebooter{ack: models.RebootAck{Accepted: true}} + + res := dispatch(store, tc.act, requesters{update: updater, reboot: rebooter}, "smallboi", "tester") + + if res.Outcome != OutcomeSkipped { + t.Errorf("Outcome = %q, want %q", res.Outcome, OutcomeSkipped) + } + if res.Code != tc.wantCode { + t.Errorf("Code = %q, want %q", res.Code, tc.wantCode) + } + if got := statusFor(res); got != tc.wantStatus { + t.Errorf("statusFor = %d, want %d", got, tc.wantStatus) + } + // Nothing was sent: a skip is decided before the host is bothered. + if updater.callCount()+rebooter.callCount() != 0 { + t.Error("a skipped host was still sent a command") + } + }) + } +} + +// TestDispatchUnknownHostIsSkippedNotFailed: a group may name a host the server +// has never seen — a group built before its hosts exist, or one whose host was +// deleted from the dashboard. That is a legitimate state, so it is a skip, and +// nothing is sent. +func TestDispatchUnknownHostIsSkipped(t *testing.T) { + store := &fakeStorage{} + updater := &fakeUpdater{ack: models.UpdateAck{Accepted: true}} + + res := dispatch(store, actionUpdate, requesters{update: updater}, "ghost", "tester") + + if res.Outcome != OutcomeSkipped || res.Code != CodeUnknownHost { + t.Errorf("got %q/%q, want %q/%q", res.Outcome, res.Code, OutcomeSkipped, CodeUnknownHost) + } + if statusFor(res) != http.StatusNotFound { + t.Errorf("statusFor = %d, want 404", statusFor(res)) + } + if updater.callCount() != 0 { + t.Error("a host the server does not know was still sent a command") + } +} + +// TestDispatchRefusalReasonFallsBackPerAction pins the default sentences. A +// host that refuses without saying why still has to produce something the +// operator can read, and the wording differs by action on purpose. +func TestDispatchRefusalReasonFallsBackPerAction(t *testing.T) { + store := &fakeStorage{systems: []models.System{optedInSystem("smallboi")}} + + for _, tc := range []struct { + act action + want string + }{ + {actionUpdate, "the host refused the update request"}, + {actionReboot, "the host refused the reboot request"}, + {actionCheckIn, "the host refused the check-in request"}, + } { + t.Run(string(tc.act), func(t *testing.T) { + req := requesters{ + update: &fakeUpdater{ack: models.UpdateAck{Accepted: false}}, + reboot: &fakeRebooter{ack: models.RebootAck{Accepted: false}}, + checkIn: &fakeCheckIner{ack: models.CheckInAck{Accepted: false}}, + } + res := dispatch(store, tc.act, req, "smallboi", "tester") + if res.Reason != tc.want { + t.Errorf("Reason = %q, want %q", res.Reason, tc.want) + } + }) + } +} + +// TestDispatchKeepsTheHostsOwnReason: when the host says why, that is what the +// operator sees — "an update is already running on this host" is far more use +// than a generic refusal. +func TestDispatchKeepsTheHostsOwnReason(t *testing.T) { + store := &fakeStorage{systems: []models.System{optedInSystem("smallboi")}} + updater := &fakeUpdater{ack: models.UpdateAck{Accepted: false, Reason: "an update is already running on this host"}} + + res := dispatch(store, actionUpdate, requesters{update: updater}, "smallboi", "tester") + + if res.Reason != "an update is already running on this host" { + t.Errorf("Reason = %q, want the host's own reason", res.Reason) + } +} + +// TestDispatchGroupPreservesMemberOrder: results are read as a list by a human, +// so they come back in the group's order however the goroutines finished. The +// scripted delays here are deliberately inverted against the member order. +func TestDispatchGroupPreservesMemberOrder(t *testing.T) { + members := []string{"a", "b", "c", "d"} + systems := make([]models.System, 0, len(members)) + for _, m := range members { + systems = append(systems, optedInSystem(m)) + } + store := &fakeStorage{systems: systems} + + delays := map[string]time.Duration{"a": 40 * time.Millisecond, "b": 30 * time.Millisecond, "c": 20 * time.Millisecond, "d": 0} + updater := &fakeUpdater{ + ack: models.UpdateAck{Accepted: true}, + hook: func(hostname string) { time.Sleep(delays[hostname]) }, + } + + results := dispatchGroup(store, actionUpdate, requesters{update: updater}, members, "tester") + + if len(results) != len(members) { + t.Fatalf("got %d results, want %d", len(results), len(members)) + } + for i, m := range members { + if results[i].Hostname != m { + t.Errorf("results[%d].Hostname = %q, want %q", i, results[i].Hostname, m) + } + } +} + +// TestDispatchGroupRunsConcurrently pins the "all at once" decision in a way +// that cannot silently regress: every member blocks until all of them have +// arrived. A serialised implementation never gets past the first one and the +// test times out rather than passing slowly. +func TestDispatchGroupRunsConcurrently(t *testing.T) { + members := []string{"a", "b", "c", "d", "e"} + systems := make([]models.System, 0, len(members)) + for _, m := range members { + systems = append(systems, optedInSystem(m)) + } + store := &fakeStorage{systems: systems} + + var ( + mu sync.Mutex + arrived int + ) + allHere := make(chan struct{}) + updater := &fakeUpdater{ + ack: models.UpdateAck{Accepted: true}, + hook: func(string) { + mu.Lock() + arrived++ + if arrived == len(members) { + close(allHere) + } + mu.Unlock() + <-allHere + }, + } + + done := make(chan []HostActionResult, 1) + go func() { + done <- dispatchGroup(store, actionUpdate, requesters{update: updater}, members, "tester") + }() + + select { + case results := <-done: + for _, res := range results { + if res.Outcome != OutcomeAccepted { + t.Errorf("%s: outcome = %q, want %q", res.Hostname, res.Outcome, OutcomeAccepted) + } + } + case <-time.After(5 * time.Second): + t.Fatal("dispatchGroup did not finish: the members are not being asked at the same time") + } +} + +// TestDispatchGroupMixedOutcomes is the case worth pinning — a group where +// every host agrees proves very little. One host accepts, one was never a +// candidate, one is offline, and the batch still runs for the rest. +func TestDispatchGroupMixedOutcomes(t *testing.T) { + optedOut := optedInSystem("db01") + optedOut.RemoteUpdatesEnabled = false + + store := &fakeStorage{systems: []models.System{ + optedInSystem("web01"), + optedOut, + optedInSystem("web02"), + }} + + updater := &fakeUpdater{ + perHost: map[string]fakeAnswer{ + "web01": {accepted: true, id: "run-1", command: "/usr/libexec/muc/upd"}, + "web02": {err: models.ErrHostNotListening}, + }, + } + + members := []string{"web01", "db01", "web02", "ghost01"} + results := dispatchGroup(store, actionUpdate, requesters{update: updater}, members, "tester") + + want := []struct { + outcome string + code string + }{ + {OutcomeAccepted, ""}, + {OutcomeSkipped, CodeNotOptedIn}, + {OutcomeUnreachable, CodeNotListening}, + {OutcomeSkipped, CodeUnknownHost}, + } + for i, w := range want { + if results[i].Outcome != w.outcome || results[i].Code != w.code { + t.Errorf("results[%d] (%s) = %q/%q, want %q/%q", + i, members[i], results[i].Outcome, results[i].Code, w.outcome, w.code) + } + } + + // Only the two candidates were actually asked. The opted-out host and the + // unknown one must not reach NATS at all. + seen := updater.hostsSeen() + slices.Sort(seen) + if !slices.Equal(seen, []string{"web01", "web02"}) { + t.Errorf("hosts asked = %v, want [web01 web02]", seen) + } +} + +// TestDispatchGroupOnNoMembers: an empty group is a legitimate state, not an +// error, and must produce an empty list rather than a nil one — the dashboard +// iterates it without checking. +func TestDispatchGroupOnNoMembers(t *testing.T) { + store := &fakeStorage{} + updater := &fakeUpdater{} + + results := dispatchGroup(store, actionUpdate, requesters{update: updater}, nil, "tester") + + if results == nil { + t.Fatal("results is nil, want an empty slice") + } + if len(results) != 0 { + t.Errorf("got %d results, want 0", len(results)) + } +} diff --git a/server/api/groups.go b/server/api/groups.go new file mode 100644 index 0000000..8d2585b --- /dev/null +++ b/server/api/groups.go @@ -0,0 +1,464 @@ +package api + +import ( + "encoding/json" + "fmt" + "log/slog" + "net/http" + "server/models" + "server/storage" + "slices" + "strings" + + "github.com/gorilla/mux" +) + +// GroupActionResponse is what a group action reports back. +// +// The status is 200 whenever the group exists and the feature is on, even when +// some members were skipped or could not be reached. The request was "fan this +// out and tell me what happened", and it did exactly that — nothing about the +// HTTP transaction failed. A status cannot summarise seven different answers, +// so it does not try: the body is the thing to read. (207 Multi-Status would be +// no better. It is a WebDAV code whose body is a defined XML document, no +// client treats it specially, and fetch's response.ok is already true for it.) +type GroupActionResponse struct { + Group string `json:"group"` + Action string `json:"action"` + Requested int `json:"requested"` + Accepted int `json:"accepted"` + Skipped int `json:"skipped"` + Failed int `json:"failed"` + Results []HostActionResult `json:"results"` +} + +// groupWriteRequest is the body of PUT /api/groups/{group}. +// +// Both fields are pointers so that leaving one out and sending it empty are +// different requests: omitting members means "rename only, leave membership +// alone", while "members": [] means "empty this group". +type groupWriteRequest struct { + Name *string `json:"name"` + Members *[]string `json:"members"` +} + +// groupCreateRequest is the body of POST /api/groups. +type groupCreateRequest struct { + Name string `json:"name"` + Members []string `json:"members"` +} + +// hostGroupsRequest is the body of PUT /api/systems/{hostname}/groups: the +// complete set of groups the host should now be in. +type hostGroupsRequest struct { + Groups []string `json:"groups"` +} + +// isNotFound reports whether a storage error means "no such thing". The bbolt +// layer spells it in its message, as DeleteSystemHandler has always relied on. +func isNotFound(err error) bool { + return err != nil && strings.Contains(err.Error(), "not found") +} + +// resolveGroup finds a group by name, stored or derived. +// +// Stored groups are checked first and the systems are only loaded when the +// name is in the derived namespace, so an ordinary group lookup stays a single +// key read. +func resolveGroup(store storage.Storage, name string) (models.Group, error) { + if !models.IsDerivedName(name) { + return store.GetGroup(name) + } + + systems, err := store.GetAllSystems() + if err != nil { + return models.Group{}, err + } + group, ok := models.FindDerivedGroup(systems, name) + if !ok { + // A derived group with no members does not exist, which is why this + // reads as "not found" rather than as an empty group: nothing is + // Fedora, so there is no os:fedora to act on. + return models.Group{}, fmt.Errorf("group '%s' not found", name) + } + return group, nil +} + +// refuseIfDerived writes the explanation and reports true when a route that +// changes a group has been pointed at one that is computed. +func refuseIfDerived(w http.ResponseWriter, name string) bool { + if !models.IsDerivedName(name) { + return false + } + writeAPIError(w, http.StatusConflict, + name+" is a derived group: its members come from what the hosts report, so it cannot be edited. "+ + "Change the host, or make an ordinary group instead.") + return true +} + +// ListGroupsHandler returns every group. The dashboard asks for this on load +// and after every edit, and inverts it into a per-host view itself — which is +// why groups are not folded into the /api/systems payload. +func ListGroupsHandler(store storage.Storage) http.HandlerFunc { + return func(w http.ResponseWriter, r *http.Request) { + groups, err := store.GetAllGroups() + if err != nil { + slog.Error("Failed to list groups", "error", err) + writeAPIError(w, http.StatusInternalServerError, "Failed to fetch groups") + return + } + if groups == nil { + groups = []models.Group{} + } + + // Derived groups are computed on every read rather than stored, so + // they are always current and there is never a stale one to clean up. + // A failure here is not fatal: the hand-made groups are still worth + // returning. + if systems, err := store.GetAllSystems(); err != nil { + slog.Error("Failed to read systems for derived groups", "error", err) + } else { + groups = append(groups, models.DeriveGroups(systems)...) + } + + writeJSON(w, http.StatusOK, groups) + } +} + +// GetGroupHandler returns one group, matched without regard to case. +func GetGroupHandler(store storage.Storage) http.HandlerFunc { + return func(w http.ResponseWriter, r *http.Request) { + group, err := resolveGroup(store, mux.Vars(r)["group"]) + if err != nil { + writeAPIError(w, http.StatusNotFound, "Group not found") + return + } + writeJSON(w, http.StatusOK, group) + } +} + +// CreateGroupHandler makes a new, usually empty, group. +// +// Creating a group is deliberately its own act rather than something that +// happens the first time a name is typed into the host editor: a typo that +// founds a group of one is far harder to notice than a typo that is refused. +func CreateGroupHandler(store storage.Storage) http.HandlerFunc { + return func(w http.ResponseWriter, r *http.Request) { + var req groupCreateRequest + if err := json.NewDecoder(r.Body).Decode(&req); err != nil { + writeAPIError(w, http.StatusBadRequest, "Invalid request body") + return + } + + name := models.NormalizeGroupName(req.Name) + if err := models.ValidateGroupName(name); err != nil { + writeAPIError(w, http.StatusBadRequest, err.Error()) + return + } + + if existing, err := store.GetGroup(name); err == nil { + writeAPIError(w, http.StatusConflict, + "A group called "+existing.Name+" already exists") + return + } + + group := models.Group{Name: name, Members: req.Members} + if err := store.SaveGroup(group); err != nil { + slog.Error("Failed to create group", "group", name, "error", err) + writeAPIError(w, http.StatusInternalServerError, "Failed to create group") + return + } + + slog.Info("Group created", "group", name, "members", len(req.Members), "requested_by", requesterAddress(r)) + + saved, err := store.GetGroup(name) + if err != nil { + saved = group + } + writeJSON(w, http.StatusCreated, saved) + } +} + +// UpdateGroupHandler renames a group, replaces its membership, or both. +func UpdateGroupHandler(store storage.Storage) http.HandlerFunc { + return func(w http.ResponseWriter, r *http.Request) { + name := mux.Vars(r)["group"] + if refuseIfDerived(w, name) { + return + } + + var req groupWriteRequest + if err := json.NewDecoder(r.Body).Decode(&req); err != nil { + writeAPIError(w, http.StatusBadRequest, "Invalid request body") + return + } + + group, err := store.GetGroup(name) + if err != nil { + writeAPIError(w, http.StatusNotFound, "Group not found") + return + } + + if req.Members != nil { + group.Members = *req.Members + if err := store.SaveGroup(group); err != nil { + slog.Error("Failed to set group members", "group", group.Name, "error", err) + writeAPIError(w, http.StatusInternalServerError, "Failed to update group") + return + } + } + + if req.Name != nil { + newName := models.NormalizeGroupName(*req.Name) + if err := models.ValidateGroupName(newName); err != nil { + writeAPIError(w, http.StatusBadRequest, err.Error()) + return + } + // A change of capitalisation alone is a relabel in place rather + // than a move, and the storage layer already tells the two apart. + if newName != group.Name { + switch err := store.RenameGroup(group.Name, newName); { + case err == nil: + case isNotFound(err): + writeAPIError(w, http.StatusNotFound, "Group not found") + return + case strings.Contains(err.Error(), "already exists"): + writeAPIError(w, http.StatusConflict, + "A group called "+newName+" already exists") + return + default: + slog.Error("Failed to rename group", "group", group.Name, "error", err) + writeAPIError(w, http.StatusInternalServerError, "Failed to rename group") + return + } + group.Name = newName + slog.Info("Group renamed", "group", name, "to", newName, "requested_by", requesterAddress(r)) + } + } + + saved, err := store.GetGroup(group.Name) + if err != nil { + saved = group + } + writeJSON(w, http.StatusOK, saved) + } +} + +// DeleteGroupHandler removes a group. The hosts in it are untouched — a group +// is a label, and dropping the label is not dropping the machines. +func DeleteGroupHandler(store storage.Storage) http.HandlerFunc { + return func(w http.ResponseWriter, r *http.Request) { + name := mux.Vars(r)["group"] + if refuseIfDerived(w, name) { + return + } + + if err := store.DeleteGroup(name); err != nil { + if isNotFound(err) { + writeAPIError(w, http.StatusNotFound, "Group not found") + return + } + slog.Error("Failed to delete group", "group", name, "error", err) + writeAPIError(w, http.StatusInternalServerError, "Failed to delete group") + return + } + + slog.Info("Group deleted", "group", name, "requested_by", requesterAddress(r)) + writeJSON(w, http.StatusOK, map[string]string{ + "status": "success", + "message": "Group deleted successfully", + "group": name, + }) + } +} + +// RemoveGroupMemberHandler drops one host from one group. +// +// It is how a member the server no longer knows about gets evicted — a host +// that was decommissioned, or a name that was mistyped into the list. The +// per-host editor cannot do it, because that host has no row to expand. +func RemoveGroupMemberHandler(store storage.Storage) http.HandlerFunc { + return func(w http.ResponseWriter, r *http.Request) { + vars := mux.Vars(r) + hostname := strings.TrimSpace(vars["hostname"]) + if refuseIfDerived(w, vars["group"]) { + return + } + + group, err := store.GetGroup(vars["group"]) + if err != nil { + writeAPIError(w, http.StatusNotFound, "Group not found") + return + } + + if !slices.Contains(group.Members, hostname) { + writeAPIError(w, http.StatusNotFound, hostname+" is not a member of "+group.Name) + return + } + + group.Members = slices.DeleteFunc(group.Members, func(m string) bool { return m == hostname }) + if err := store.SaveGroup(group); err != nil { + slog.Error("Failed to remove group member", "group", group.Name, "hostname", hostname, "error", err) + writeAPIError(w, http.StatusInternalServerError, "Failed to update group") + return + } + + saved, err := store.GetGroup(group.Name) + if err != nil { + saved = group + } + writeJSON(w, http.StatusOK, saved) + } +} + +// SetSystemGroupsHandler replaces one host's memberships with exactly the list +// it is given, so a group left out is a group the host leaves. +// +// The host need not have checked in: putting a machine into its groups before +// it is built is a reasonable thing to want, and membership that outlives the +// system record is already how groups behave. +func SetSystemGroupsHandler(store storage.Storage) http.HandlerFunc { + return func(w http.ResponseWriter, r *http.Request) { + hostname := strings.TrimSpace(mux.Vars(r)["hostname"]) + if hostname == "" { + writeAPIError(w, http.StatusBadRequest, "Hostname is required") + return + } + + var req hostGroupsRequest + if err := json.NewDecoder(r.Body).Decode(&req); err != nil { + writeAPIError(w, http.StatusBadRequest, "Invalid request body") + return + } + + for _, name := range req.Groups { + if models.IsDerivedName(name) { + writeAPIError(w, http.StatusConflict, + name+" is a derived group: a host joins it by being a Fedora box or needing a reboot, not by being put in it.") + return + } + } + + if err := store.SetHostGroups(hostname, req.Groups); err != nil { + if isNotFound(err) { + // A name that is not already a group. Refusing is the point: + // joining nothing is recoverable, quietly founding a group of + // one is not. + writeAPIError(w, http.StatusBadRequest, + "No such group — create it before adding hosts to it ("+err.Error()+")") + return + } + slog.Error("Failed to set host groups", "hostname", hostname, "error", err) + writeAPIError(w, http.StatusInternalServerError, "Failed to update groups") + return + } + + groups, err := store.GetAllGroups() + if err != nil { + slog.Error("Failed to re-read groups", "error", err) + writeAPIError(w, http.StatusInternalServerError, "Failed to update groups") + return + } + + joined := []string{} + for _, group := range groups { + if slices.Contains(group.Members, hostname) { + joined = append(joined, group.Name) + } + } + + slog.Info("Host groups set", "hostname", hostname, "groups", joined, "requested_by", requesterAddress(r)) + writeJSON(w, http.StatusOK, map[string]any{ + "hostname": hostname, + "groups": joined, + }) + } +} + +// GroupUpdateHandler asks every member of a group to install its pending +// packages. GroupCheckInHandler and GroupRebootHandler are the same shape. +// +// Gated exactly like the single-host route: the server flag decides whether the +// route works at all, and each host's own opt-in decides whether it is a +// candidate. A member that has not opted in is skipped, not an error — one +// machine that declines should not stop the other eleven being patched. +func GroupUpdateHandler(store storage.Storage, updater UpdateRequester) http.HandlerFunc { + return groupActionHandler(store, actionUpdate, func() (requesters, bool, int, string) { + return requesters{update: updater}, updater != nil, http.StatusForbidden, + "Remote updates are disabled on this server (set remote_updates: true to enable them)" + }) +} + +// GroupRebootHandler reboots every member of a group that reports a pending +// reboot. Members with nothing pending are skipped, which is also what makes +// the button safe to press after a group update: it reboots what needs it. +func GroupRebootHandler(store storage.Storage, rebooter RebootRequester) http.HandlerFunc { + return groupActionHandler(store, actionReboot, func() (requesters, bool, int, string) { + return requesters{reboot: rebooter}, rebooter != nil, http.StatusForbidden, + "Remote reboots are disabled on this server (set remote_reboot: true to enable them)" + }) +} + +// GroupCheckInHandler asks every member of a group to publish a fresh check-in. +// +// Nothing configured gates it, as with the single-host route — a check-in +// installs nothing. Its unavailability is an infrastructure fact rather than a +// policy one, which is why it answers 503 where the other two answer 403. +func GroupCheckInHandler(store storage.Storage, checkins CheckInRequester) http.HandlerFunc { + return groupActionHandler(store, actionCheckIn, func() (requesters, bool, int, string) { + return requesters{checkIn: checkins}, checkins != nil, http.StatusServiceUnavailable, + "Check-in requests are unavailable: the server has no NATS connection" + }) +} + +// groupActionHandler is the body all three group actions share: check the +// feature is on, look the group up, fan out, count the answers. +func groupActionHandler(store storage.Storage, act action, gate func() (requesters, bool, int, string)) http.HandlerFunc { + return func(w http.ResponseWriter, r *http.Request) { + req, enabled, status, message := gate() + if !enabled { + writeAPIError(w, status, message) + return + } + + // Resolved when the action runs, not when the page was drawn. For a + // derived group that is the point: "reboot everything that needs a + // reboot" should mean what is true now. + group, err := resolveGroup(store, mux.Vars(r)["group"]) + if err != nil { + writeAPIError(w, http.StatusNotFound, "Group not found") + return + } + + results := dispatchGroup(store, act, req, group.Members, requesterAddress(r)) + + resp := GroupActionResponse{ + Group: group.Name, + Action: string(act), + Requested: len(results), + Results: results, + } + for _, res := range results { + switch res.Outcome { + case OutcomeAccepted: + resp.Accepted++ + case OutcomeSkipped: + resp.Skipped++ + default: + // refused, unreachable and failed all read as "failed" in the + // summary line; the per-host list keeps them apart. + resp.Failed++ + } + } + + // A group action is the most consequential thing this server does, and + // nothing authenticates the person who pressed the button. One line per + // action means "who rebooted production" is one grep away. + slog.Warn("Group action dispatched", + "group", group.Name, "action", string(act), "requested_by", requesterAddress(r), + "requested", resp.Requested, "accepted", resp.Accepted, + "skipped", resp.Skipped, "failed", resp.Failed) + + writeJSON(w, http.StatusOK, resp) + } +} diff --git a/server/api/groups_test.go b/server/api/groups_test.go new file mode 100644 index 0000000..0ac1467 --- /dev/null +++ b/server/api/groups_test.go @@ -0,0 +1,669 @@ +package api + +import ( + "encoding/json" + "net/http" + "net/http/httptest" + "server/models" + "slices" + "strings" + "testing" + + "github.com/gorilla/mux" +) + +// storeWithGroups builds a fakeStorage holding the given groups, keyed and +// sorted the way the bbolt layer keys and sorts them — a fake that is tidier +// than the real store would hide an ordering bug rather than catch one. +func storeWithGroups(systems []models.System, groups ...models.Group) *fakeStorage { + store := &fakeStorage{systems: systems, groups: map[string]models.Group{}} + for _, g := range groups { + slices.Sort(g.Members) + store.groups[strings.ToLower(g.Name)] = g + } + return store +} + +func doJSON(handler http.HandlerFunc, method, path, body string, vars map[string]string) *httptest.ResponseRecorder { + var reader *strings.Reader + if body == "" { + reader = strings.NewReader("{}") + } else { + reader = strings.NewReader(body) + } + req := httptest.NewRequest(method, path, reader) + if vars != nil { + req = mux.SetURLVars(req, vars) + } + rec := httptest.NewRecorder() + handler(rec, req) + return rec +} + +func decodeGroup(t *testing.T, rec *httptest.ResponseRecorder) models.Group { + t.Helper() + var group models.Group + if err := json.NewDecoder(rec.Body).Decode(&group); err != nil { + t.Fatalf("decoding group: %v", err) + } + return group +} + +func decodeAction(t *testing.T, rec *httptest.ResponseRecorder) GroupActionResponse { + t.Helper() + var resp GroupActionResponse + if err := json.NewDecoder(rec.Body).Decode(&resp); err != nil { + t.Fatalf("decoding group action response: %v", err) + } + return resp +} + +// TestListGroupsOnEmptyServerReturnsAnArray: the dashboard iterates the +// response without checking it first, and a JSON null would throw. An empty +// fleet is the state every new install starts in. +func TestListGroupsOnEmptyServerReturnsAnArray(t *testing.T) { + rec := doJSON(ListGroupsHandler(&fakeStorage{}), http.MethodGet, "/api/groups", "", nil) + + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200", rec.Code) + } + if body := strings.TrimSpace(rec.Body.String()); body != "[]" { + t.Errorf("body = %s, want []", body) + } +} + +func TestCreateGroup(t *testing.T) { + store := &fakeStorage{} + + rec := doJSON(CreateGroupHandler(store), http.MethodPost, "/api/groups", `{"name":"prod"}`, nil) + + if rec.Code != http.StatusCreated { + t.Fatalf("status = %d, want 201: %s", rec.Code, rec.Body.String()) + } + if group := decodeGroup(t, rec); group.Name != "prod" { + t.Errorf("Name = %q, want %q", group.Name, "prod") + } +} + +// TestCreateGroupRefusesCaseDuplicate is the one name mistake worth designing +// against: "prod" and "Prod" both look right in the chip bar, and an action on +// either silently misses the hosts in the other. +func TestCreateGroupRefusesCaseDuplicate(t *testing.T) { + store := storeWithGroups(nil, models.Group{Name: "Prod", Members: []string{"web01"}}) + + rec := doJSON(CreateGroupHandler(store), http.MethodPost, "/api/groups", `{"name":"prod"}`, nil) + + if rec.Code != http.StatusConflict { + t.Fatalf("status = %d, want 409", rec.Code) + } + // The message names the group as it is actually spelled, which is the + // thing the operator needs to go and look at. + if !strings.Contains(rec.Body.String(), "Prod") { + t.Errorf("the refusal does not name the existing group: %s", rec.Body.String()) + } +} + +func TestCreateGroupRejectsBadNames(t *testing.T) { + for _, tc := range []struct{ name, body string }{ + {"empty", `{"name":""}`}, + {"whitespace only", `{"name":" "}`}, + {"path separator", `{"name":"prod/web"}`}, + {"too long", `{"name":"` + strings.Repeat("a", 65) + `"}`}, + } { + t.Run(tc.name, func(t *testing.T) { + rec := doJSON(CreateGroupHandler(&fakeStorage{}), http.MethodPost, "/api/groups", tc.body, nil) + if rec.Code != http.StatusBadRequest { + t.Errorf("status = %d, want 400", rec.Code) + } + }) + } +} + +func TestGetGroupNotFound(t *testing.T) { + rec := doJSON(GetGroupHandler(&fakeStorage{}), http.MethodGet, "/api/groups/nope", "", + map[string]string{"group": "nope"}) + + if rec.Code != http.StatusNotFound { + t.Errorf("status = %d, want 404", rec.Code) + } +} + +// TestUpdateGroupOmittingMembersLeavesThemAlone is why the request struct uses +// pointers: a rename must not be a way to accidentally empty a group. +func TestUpdateGroupOmittingMembersLeavesThemAlone(t *testing.T) { + store := storeWithGroups(nil, models.Group{Name: "prod", Members: []string{"web01", "db01"}}) + + rec := doJSON(UpdateGroupHandler(store), http.MethodPut, "/api/groups/prod", + `{"name":"production"}`, map[string]string{"group": "prod"}) + + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200: %s", rec.Code, rec.Body.String()) + } + group := decodeGroup(t, rec) + if group.Name != "production" { + t.Errorf("Name = %q, want %q", group.Name, "production") + } + if !slices.Equal(group.Members, []string{"db01", "web01"}) { + t.Errorf("a rename lost the members: %v", group.Members) + } +} + +// TestUpdateGroupWithEmptyMembersEmptiesIt is the other half of the pointer: +// sending an explicit empty list must mean what it says. +func TestUpdateGroupWithEmptyMembersEmptiesIt(t *testing.T) { + store := storeWithGroups(nil, models.Group{Name: "prod", Members: []string{"web01"}}) + + rec := doJSON(UpdateGroupHandler(store), http.MethodPut, "/api/groups/prod", + `{"members":[]}`, map[string]string{"group": "prod"}) + + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200", rec.Code) + } + if group := decodeGroup(t, rec); len(group.Members) != 0 { + t.Errorf("Members = %v, want none", group.Members) + } +} + +func TestUpdateGroupRenameCollision(t *testing.T) { + store := storeWithGroups(nil, + models.Group{Name: "prod", Members: []string{"web01"}}, + models.Group{Name: "staging", Members: []string{"web02"}}, + ) + + rec := doJSON(UpdateGroupHandler(store), http.MethodPut, "/api/groups/prod", + `{"name":"STAGING"}`, map[string]string{"group": "prod"}) + + if rec.Code != http.StatusConflict { + t.Fatalf("status = %d, want 409: %s", rec.Code, rec.Body.String()) + } +} + +func TestDeleteGroup(t *testing.T) { + store := storeWithGroups(nil, models.Group{Name: "prod"}) + + rec := doJSON(DeleteGroupHandler(store), http.MethodDelete, "/api/groups/prod", "", + map[string]string{"group": "prod"}) + + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200", rec.Code) + } + if len(store.groups) != 0 { + t.Errorf("the group is still in storage: %v", store.groups) + } +} + +func TestDeleteGroupNotFound(t *testing.T) { + rec := doJSON(DeleteGroupHandler(&fakeStorage{groups: map[string]models.Group{}}), + http.MethodDelete, "/api/groups/nope", "", map[string]string{"group": "nope"}) + + if rec.Code != http.StatusNotFound { + t.Errorf("status = %d, want 404", rec.Code) + } +} + +// TestRemoveGroupMemberEvictsAHostTheServerDoesNotKnow is the whole point of +// the route: a decommissioned machine has no row to expand, so the per-host +// editor cannot reach it. +func TestRemoveGroupMemberEvictsAHostTheServerDoesNotKnow(t *testing.T) { + store := storeWithGroups(nil, models.Group{Name: "prod", Members: []string{"web01", "retired01"}}) + + rec := doJSON(RemoveGroupMemberHandler(store), http.MethodDelete, + "/api/groups/prod/members/retired01", "", + map[string]string{"group": "prod", "hostname": "retired01"}) + + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200: %s", rec.Code, rec.Body.String()) + } + if group := decodeGroup(t, rec); !slices.Equal(group.Members, []string{"web01"}) { + t.Errorf("Members = %v, want [web01]", group.Members) + } +} + +func TestRemoveGroupMemberNotAMember(t *testing.T) { + store := storeWithGroups(nil, models.Group{Name: "prod", Members: []string{"web01"}}) + + rec := doJSON(RemoveGroupMemberHandler(store), http.MethodDelete, + "/api/groups/prod/members/db01", "", + map[string]string{"group": "prod", "hostname": "db01"}) + + if rec.Code != http.StatusNotFound { + t.Errorf("status = %d, want 404", rec.Code) + } +} + +func TestSetSystemGroups(t *testing.T) { + store := storeWithGroups(nil, + models.Group{Name: "prod", Members: []string{"web01"}}, + models.Group{Name: "web"}, + ) + + rec := doJSON(SetSystemGroupsHandler(store), http.MethodPut, "/api/systems/web01/groups", + `{"groups":["web"]}`, map[string]string{"hostname": "web01"}) + + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200: %s", rec.Code, rec.Body.String()) + } + + var got struct { + Hostname string `json:"hostname"` + Groups []string `json:"groups"` + } + if err := json.NewDecoder(rec.Body).Decode(&got); err != nil { + t.Fatalf("decoding: %v", err) + } + if !slices.Equal(got.Groups, []string{"web"}) { + t.Errorf("Groups = %v, want [web] — the host should have left prod", got.Groups) + } +} + +// TestSetSystemGroupsRefusesUnknownGroup: a typo must join nothing rather than +// found a group of one, which is far harder to spot afterwards. +func TestSetSystemGroupsRefusesUnknownGroup(t *testing.T) { + store := storeWithGroups(nil, models.Group{Name: "prod"}) + + rec := doJSON(SetSystemGroupsHandler(store), http.MethodPut, "/api/systems/web01/groups", + `{"groups":["prodd"]}`, map[string]string{"hostname": "web01"}) + + if rec.Code != http.StatusBadRequest { + t.Fatalf("status = %d, want 400: %s", rec.Code, rec.Body.String()) + } + if group := store.groups["prod"]; slices.Contains(group.Members, "web01") { + t.Error("the valid half of a refused list was still applied") + } +} + +// TestSetSystemGroupsWorksForAHostThatHasNeverCheckedIn: pre-staging a machine +// into its groups before it is built is a reasonable thing to want, and +// membership already outlives the system record. +func TestSetSystemGroupsWorksForAHostThatHasNeverCheckedIn(t *testing.T) { + store := storeWithGroups(nil, models.Group{Name: "prod"}) + + rec := doJSON(SetSystemGroupsHandler(store), http.MethodPut, "/api/systems/not-yet-built/groups", + `{"groups":["prod"]}`, map[string]string{"hostname": "not-yet-built"}) + + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200: %s", rec.Code, rec.Body.String()) + } + if group := store.groups["prod"]; !slices.Contains(group.Members, "not-yet-built") { + t.Errorf("the host did not join: %v", group.Members) + } +} + +// TestGroupActionDisabled pins the server-side half of each opt-in, and the +// deliberate difference between them: update and reboot are policy decisions +// and answer 403, while an unavailable check-in is an infrastructure fact and +// answers 503. +func TestGroupActionDisabled(t *testing.T) { + store := storeWithGroups(nil, models.Group{Name: "prod", Members: []string{"web01"}}) + vars := map[string]string{"group": "prod"} + + for _, tc := range []struct { + name string + handler http.HandlerFunc + want int + }{ + {"update", GroupUpdateHandler(store, nil), http.StatusForbidden}, + {"reboot", GroupRebootHandler(store, nil), http.StatusForbidden}, + {"checkin", GroupCheckInHandler(store, nil), http.StatusServiceUnavailable}, + } { + t.Run(tc.name, func(t *testing.T) { + rec := doJSON(tc.handler, http.MethodPost, "/api/groups/prod/"+tc.name, "", vars) + if rec.Code != tc.want { + t.Errorf("status = %d, want %d", rec.Code, tc.want) + } + }) + } +} + +func TestGroupActionUnknownGroup(t *testing.T) { + store := &fakeStorage{groups: map[string]models.Group{}} + updater := &fakeUpdater{ack: models.UpdateAck{Accepted: true}} + + rec := doJSON(GroupUpdateHandler(store, updater), http.MethodPost, "/api/groups/nope/update", "", + map[string]string{"group": "nope"}) + + if rec.Code != http.StatusNotFound { + t.Fatalf("status = %d, want 404", rec.Code) + } + if updater.callCount() != 0 { + t.Error("an unknown group still dispatched commands") + } +} + +// TestGroupActionOnEmptyGroup: an empty group is a legitimate state, not an +// error. Keeping the client on one code path is worth more than the +// distinction, so it answers 200 with nothing requested. +func TestGroupActionOnEmptyGroup(t *testing.T) { + store := storeWithGroups(nil, models.Group{Name: "prod"}) + updater := &fakeUpdater{ack: models.UpdateAck{Accepted: true}} + + rec := doJSON(GroupUpdateHandler(store, updater), http.MethodPost, "/api/groups/prod/update", "", + map[string]string{"group": "prod"}) + + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200", rec.Code) + } + resp := decodeAction(t, rec) + if resp.Requested != 0 { + t.Errorf("Requested = %d, want 0", resp.Requested) + } + if resp.Results == nil { + t.Error("Results is null; the dashboard iterates it without checking") + } +} + +// TestGroupUpdateMixedOutcomes is the case the whole feature turns on: one +// machine declining must not stop the rest being patched, and the operator has +// to be able to see which was which. +func TestGroupUpdateMixedOutcomes(t *testing.T) { + optedOut := optedInSystem("db01") + optedOut.RemoteUpdatesEnabled = false + + store := storeWithGroups( + []models.System{optedInSystem("web01"), optedOut, optedInSystem("web02")}, + models.Group{Name: "prod", Members: []string{"db01", "ghost01", "web01", "web02"}}, + ) + updater := &fakeUpdater{perHost: map[string]fakeAnswer{ + "web01": {accepted: true, id: "run-1"}, + "web02": {err: models.ErrHostNotListening}, + }} + + rec := doJSON(GroupUpdateHandler(store, updater), http.MethodPost, "/api/groups/prod/update", "", + map[string]string{"group": "prod"}) + + // 200, not 207 and not an error: the fan-out ran and the body says what + // happened to each host. + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200: %s", rec.Code, rec.Body.String()) + } + + resp := decodeAction(t, rec) + if resp.Requested != 4 || resp.Accepted != 1 || resp.Skipped != 2 || resp.Failed != 1 { + t.Errorf("counts = requested %d / accepted %d / skipped %d / failed %d; want 4/1/2/1", + resp.Requested, resp.Accepted, resp.Skipped, resp.Failed) + } + if resp.Group != "prod" || resp.Action != "update" { + t.Errorf("Group/Action = %q/%q, want prod/update", resp.Group, resp.Action) + } + + // One result per member, in the group's order, so the list reads the same + // twice running. + wantHosts := []string{"db01", "ghost01", "web01", "web02"} + for i, want := range wantHosts { + if resp.Results[i].Hostname != want { + t.Errorf("Results[%d].Hostname = %q, want %q", i, resp.Results[i].Hostname, want) + } + } +} + +// TestGroupRebootSkipsHostsWithNothingPending is what makes the button safe to +// press after a group update: it reboots what needs it and leaves the rest. +func TestGroupRebootSkipsHostsWithNothingPending(t *testing.T) { + noPending := optedInSystem("web02") + noPending.RebootRequired = false + + store := storeWithGroups( + []models.System{optedInSystem("web01"), noPending}, + models.Group{Name: "prod", Members: []string{"web01", "web02"}}, + ) + rebooter := &fakeRebooter{ack: models.RebootAck{Accepted: true, ID: "r1"}} + + rec := doJSON(GroupRebootHandler(store, rebooter), http.MethodPost, "/api/groups/prod/reboot", "", + map[string]string{"group": "prod"}) + + resp := decodeAction(t, rec) + if resp.Accepted != 1 || resp.Skipped != 1 { + t.Errorf("accepted %d / skipped %d, want 1/1", resp.Accepted, resp.Skipped) + } + if resp.Results[1].Code != CodeNoRebootPending { + t.Errorf("Results[1].Code = %q, want %q", resp.Results[1].Code, CodeNoRebootPending) + } + if rebooter.callCount() != 1 { + t.Errorf("%d hosts were sent a reboot, want 1", rebooter.callCount()) + } +} + +// TestGroupActionIsCaseInsensitive: a chip the dashboard drew from the stored +// spelling has to work whatever case the URL carries. +func TestGroupActionIsCaseInsensitive(t *testing.T) { + store := storeWithGroups( + []models.System{optedInSystem("web01")}, + models.Group{Name: "Prod", Members: []string{"web01"}}, + ) + checkins := &fakeCheckIner{ack: models.CheckInAck{Accepted: true, ID: "c1"}} + + rec := doJSON(GroupCheckInHandler(store, checkins), http.MethodPost, "/api/groups/PROD/checkin", "", + map[string]string{"group": "PROD"}) + + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200", rec.Code) + } + if resp := decodeAction(t, rec); resp.Group != "Prod" { + t.Errorf("Group = %q, want the stored spelling %q", resp.Group, "Prod") + } +} + +// --- derived groups ------------------------------------------------------- + +// fedoraFleet is a set of hosts whose facts produce derived groups: two Fedora +// boxes and one Ubuntu one, with one of each opted into remote updates. +func fedoraFleet() []models.System { + web01 := optedInSystem("web01") + web01.OS = "Fedora Linux 42" + web01.Architecture = "x86_64" + + web02 := optedInSystem("web02") + web02.OS = "Fedora Linux 42" + web02.Architecture = "x86_64" + web02.RemoteUpdatesEnabled = false + + build01 := optedInSystem("build01") + build01.OS = "Ubuntu 24.04.1 LTS" + build01.Architecture = "x86_64" + + return []models.System{web01, web02, build01} +} + +// TestListGroupsIncludesDerived: the dashboard reads one route and gets both +// kinds, so a chip bar needs no second request to be complete. +func TestListGroupsIncludesDerived(t *testing.T) { + store := storeWithGroups(fedoraFleet(), models.Group{Name: "prod", Members: []string{"web01"}}) + + rec := doJSON(ListGroupsHandler(store), http.MethodGet, "/api/groups", "", nil) + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200", rec.Code) + } + + var groups []models.Group + if err := json.NewDecoder(rec.Body).Decode(&groups); err != nil { + t.Fatalf("decoding: %v", err) + } + + byName := map[string]models.Group{} + for _, g := range groups { + byName[g.Name] = g + } + + if g, ok := byName["prod"]; !ok { + t.Error("the stored group is missing") + } else if g.Derived { + t.Error("a stored group is marked derived") + } + + for _, name := range []string{"os:fedora", "os:ubuntu", "pkg:rpm", "pkg:deb", "arch:x86_64"} { + g, ok := byName[name] + if !ok { + t.Errorf("derived group %q is missing", name) + continue + } + if !g.Derived { + t.Errorf("%q is not marked derived, so the dashboard cannot style it apart", name) + } + } + + if !slices.Equal(byName["pkg:rpm"].Members, []string{"web01", "web02"}) { + t.Errorf("pkg:rpm = %v, want [web01 web02]", byName["pkg:rpm"].Members) + } +} + +// TestGroupActionOnDerivedGroup is the point of the whole feature: "update +// every Fedora box" without anyone maintaining a list of which ones they are. +func TestGroupActionOnDerivedGroup(t *testing.T) { + store := storeWithGroups(fedoraFleet()) + updater := &fakeUpdater{perHost: map[string]fakeAnswer{ + "web01": {accepted: true, id: "run-1"}, + }} + + rec := doJSON(GroupUpdateHandler(store, updater), http.MethodPost, "/api/groups/os:fedora/update", "", + map[string]string{"group": "os:fedora"}) + + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200: %s", rec.Code, rec.Body.String()) + } + + resp := decodeAction(t, rec) + if resp.Group != "os:fedora" { + t.Errorf("Group = %q, want os:fedora", resp.Group) + } + // Both Fedora hosts were considered; the Ubuntu one was never in the group. + if resp.Requested != 2 { + t.Errorf("Requested = %d, want 2 (the two Fedora hosts)", resp.Requested) + } + if resp.Accepted != 1 || resp.Skipped != 1 { + t.Errorf("accepted %d / skipped %d, want 1/1 (web02 has not opted in)", resp.Accepted, resp.Skipped) + } + if slices.Contains(updater.hostsSeen(), "build01") { + t.Error("the Ubuntu host was asked to update as part of os:fedora") + } +} + +// TestGroupActionOnPackageFamily is the "all rpm" case, which is the one that +// spans distributions. +func TestGroupActionOnPackageFamily(t *testing.T) { + fleet := fedoraFleet() + rocky := optedInSystem("db01") + rocky.OS = "Rocky Linux 10.2" + rocky.Architecture = "aarch64" + fleet = append(fleet, rocky) + + store := storeWithGroups(fleet) + checkins := &fakeCheckIner{ack: models.CheckInAck{Accepted: true, ID: "c1"}} + + rec := doJSON(GroupCheckInHandler(store, checkins), http.MethodPost, "/api/groups/pkg:rpm/checkin", "", + map[string]string{"group": "pkg:rpm"}) + + resp := decodeAction(t, rec) + if resp.Requested != 3 { + t.Errorf("Requested = %d, want 3 (two Fedora and one Rocky)", resp.Requested) + } + hosts := []string{} + for _, res := range resp.Results { + hosts = append(hosts, res.Hostname) + } + if !slices.Equal(hosts, []string{"db01", "web01", "web02"}) { + t.Errorf("members = %v, want [db01 web01 web02]", hosts) + } +} + +// TestDerivedGroupResolvesWhenTheActionRuns: a state group has to mean what is +// true now, not what was true when the page was drawn. This is what makes +// "reboot everything that needs a reboot" correct rather than approximate. +func TestDerivedGroupResolvesAtDispatchTime(t *testing.T) { + needsReboot := optedInSystem("web01") + needsReboot.OS = "Fedora Linux 42" + noReboot := optedInSystem("web02") + noReboot.OS = "Fedora Linux 42" + noReboot.RebootRequired = false + + store := storeWithGroups([]models.System{needsReboot, noReboot}) + rebooter := &fakeRebooter{ack: models.RebootAck{Accepted: true, ID: "r1"}} + + rec := doJSON(GroupRebootHandler(store, rebooter), http.MethodPost, "/api/groups/state:needs-reboot/reboot", "", + map[string]string{"group": "state:needs-reboot"}) + + resp := decodeAction(t, rec) + if resp.Requested != 1 || resp.Accepted != 1 { + t.Errorf("requested %d / accepted %d, want 1/1", resp.Requested, resp.Accepted) + } + if resp.Results[0].Hostname != "web01" { + t.Errorf("rebooted %q, want web01", resp.Results[0].Hostname) + } +} + +// TestDerivedGroupWithNoMembersIsNotFound: nothing is Debian, so there is no +// pkg:deb to act on. An empty group here would invite a button for a set that +// does not exist. +func TestDerivedGroupWithNoMembersIsNotFound(t *testing.T) { + store := storeWithGroups(fedoraFleet()) + rebooter := &fakeRebooter{ack: models.RebootAck{Accepted: true}} + + rec := doJSON(GroupRebootHandler(store, rebooter), http.MethodPost, "/api/groups/os:debian/reboot", "", + map[string]string{"group": "os:debian"}) + + if rec.Code != http.StatusNotFound { + t.Errorf("status = %d, want 404", rec.Code) + } + if rebooter.callCount() != 0 { + t.Error("a group with no members still dispatched commands") + } +} + +// TestDerivedGroupsCannotBeEdited. Every route that changes a group has to +// refuse, or membership would appear editable and then silently revert on the +// next read, which is worse than refusing. +func TestDerivedGroupsCannotBeEdited(t *testing.T) { + store := storeWithGroups(fedoraFleet()) + vars := map[string]string{"group": "os:fedora"} + + for _, tc := range []struct { + name string + handler http.HandlerFunc + method string + body string + vars map[string]string + }{ + {"rename", UpdateGroupHandler(store), http.MethodPut, `{"name":"fedora"}`, vars}, + {"set members", UpdateGroupHandler(store), http.MethodPut, `{"members":["web01"]}`, vars}, + {"delete", DeleteGroupHandler(store), http.MethodDelete, "", vars}, + {"remove member", RemoveGroupMemberHandler(store), http.MethodDelete, "", + map[string]string{"group": "os:fedora", "hostname": "web01"}}, + } { + t.Run(tc.name, func(t *testing.T) { + rec := doJSON(tc.handler, tc.method, "/api/groups/os:fedora", tc.body, tc.vars) + if rec.Code != http.StatusConflict { + t.Errorf("status = %d, want 409: %s", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), "derived") { + t.Errorf("the refusal does not explain why: %s", rec.Body.String()) + } + }) + } +} + +// TestSetSystemGroupsRefusesDerived: a host joins os:fedora by being a Fedora +// box, not by being put there. +func TestSetSystemGroupsRefusesDerived(t *testing.T) { + store := storeWithGroups(fedoraFleet(), models.Group{Name: "prod"}) + + rec := doJSON(SetSystemGroupsHandler(store), http.MethodPut, "/api/systems/web01/groups", + `{"groups":["prod","os:fedora"]}`, map[string]string{"hostname": "web01"}) + + if rec.Code != http.StatusConflict { + t.Fatalf("status = %d, want 409: %s", rec.Code, rec.Body.String()) + } + // Nothing was applied: the valid half of the list must not have landed. + if group := store.groups["prod"]; slices.Contains(group.Members, "web01") { + t.Error("the valid half of a refused list was still applied") + } +} + +// TestCreateGroupRefusesDerivedNamespace. The namespace only works if nothing +// can be hand-made inside it. +func TestCreateGroupRefusesDerivedNamespace(t *testing.T) { + rec := doJSON(CreateGroupHandler(&fakeStorage{}), http.MethodPost, "/api/groups", + `{"name":"os:fedora"}`, nil) + + if rec.Code != http.StatusBadRequest { + t.Errorf("status = %d, want 400: %s", rec.Code, rec.Body.String()) + } +} diff --git a/server/api/routes.go b/server/api/routes.go index f51b408..34fa495 100644 --- a/server/api/routes.go +++ b/server/api/routes.go @@ -3,7 +3,6 @@ package api import ( "bytes" "encoding/json" - "errors" "log/slog" "net" "net/http" @@ -223,51 +222,19 @@ func RunUpdateHandler(store storage.Storage, updater UpdateRequester) http.Handl return } - system, err := store.GetSystem(hostname) - if err != nil { - writeAPIError(w, http.StatusNotFound, "System not found") - return - } - if !system.RemoteUpdatesEnabled { - writeAPIError(w, http.StatusConflict, - "This host has not opted into remote updates (set allow_remote_updates: true in its client config)") - return - } - - ack, err := updater.RequestUpdate(hostname, requesterAddress(r)) - switch { - case errors.Is(err, models.ErrHostNotListening): - // The stored opt-in said yes but nothing answered, so the host is - // down or its client has since been reconfigured. - writeAPIError(w, http.StatusServiceUnavailable, - "No response from "+hostname+": it is offline, or its client is no longer accepting update commands") - return - case err != nil: - slog.Error("Update request failed", "hostname", hostname, "error", err) - writeAPIError(w, http.StatusBadGateway, "Update request failed: "+err.Error()) - return - } - - if !ack.Accepted { - reason := ack.Reason - if reason == "" { - reason = "the host refused the update request" - } - writeAPIError(w, http.StatusConflict, reason) + res := dispatch(store, actionUpdate, requesters{update: updater}, hostname, requesterAddress(r)) + if res.Outcome != OutcomeAccepted { + writeAPIError(w, statusFor(res), res.Reason) return } - w.Header().Set("Content-Type", "application/json") - w.WriteHeader(http.StatusAccepted) - if err := json.NewEncoder(w).Encode(map[string]string{ + writeJSON(w, http.StatusAccepted, map[string]string{ "status": "accepted", "hostname": hostname, - "id": ack.ID, - "command": ack.Command, + "id": res.ID, + "command": res.Command, "message": "Update started on " + hostname, - }); err != nil { - slog.Error("Failed to write update response", "error", err) - } + }) } } @@ -297,53 +264,18 @@ func RebootHandler(store storage.Storage, rebooter RebootRequester) http.Handler return } - system, err := store.GetSystem(hostname) - if err != nil { - writeAPIError(w, http.StatusNotFound, "System not found") - return - } - if !system.RemoteRebootEnabled { - writeAPIError(w, http.StatusConflict, - "This host has not opted into remote reboots (set allow_remote_reboot: true in its client config)") - return - } - if !system.RebootRequired { - writeAPIError(w, http.StatusConflict, - "This host did not report a pending reboot at its last check-in") + res := dispatch(store, actionReboot, requesters{reboot: rebooter}, hostname, requesterAddress(r)) + if res.Outcome != OutcomeAccepted { + writeAPIError(w, statusFor(res), res.Reason) return } - ack, err := rebooter.RequestReboot(hostname, requesterAddress(r)) - switch { - case errors.Is(err, models.ErrHostNotListening): - writeAPIError(w, http.StatusServiceUnavailable, - "No response from "+hostname+": it is offline, or its client is no longer accepting reboot commands") - return - case err != nil: - slog.Error("Reboot request failed", "hostname", hostname, "error", err) - writeAPIError(w, http.StatusBadGateway, "Reboot request failed: "+err.Error()) - return - } - - if !ack.Accepted { - reason := ack.Reason - if reason == "" { - reason = "the host refused the reboot request" - } - writeAPIError(w, http.StatusConflict, reason) - return - } - - w.Header().Set("Content-Type", "application/json") - w.WriteHeader(http.StatusAccepted) - if err := json.NewEncoder(w).Encode(map[string]string{ + writeJSON(w, http.StatusAccepted, map[string]string{ "status": "accepted", "hostname": hostname, - "id": ack.ID, + "id": res.ID, "message": hostname + " is rebooting", - }); err != nil { - slog.Error("Failed to write reboot response", "error", err) - } + }) } } @@ -374,44 +306,18 @@ func CheckInHandler(store storage.Storage, requester CheckInRequester) http.Hand return } - if _, err := store.GetSystem(hostname); err != nil { - writeAPIError(w, http.StatusNotFound, "System not found") - return - } - - ack, err := requester.RequestCheckIn(hostname, requesterAddress(r)) - switch { - case errors.Is(err, models.ErrHostNotListening): - // Every current client subscribes, so this is a host that is down - // or one running a client from before the command existed. - writeAPIError(w, http.StatusServiceUnavailable, - "No response from "+hostname+": it is offline, or its client is too old to accept check-in requests") - return - case err != nil: - slog.Error("Check-in request failed", "hostname", hostname, "error", err) - writeAPIError(w, http.StatusBadGateway, "Check-in request failed: "+err.Error()) - return - } - - if !ack.Accepted { - reason := ack.Reason - if reason == "" { - reason = "the host refused the check-in request" - } - writeAPIError(w, http.StatusConflict, reason) + res := dispatch(store, actionCheckIn, requesters{checkIn: requester}, hostname, requesterAddress(r)) + if res.Outcome != OutcomeAccepted { + writeAPIError(w, statusFor(res), res.Reason) return } - w.Header().Set("Content-Type", "application/json") - w.WriteHeader(http.StatusAccepted) - if err := json.NewEncoder(w).Encode(map[string]string{ + writeJSON(w, http.StatusAccepted, map[string]string{ "status": "accepted", "hostname": hostname, - "id": ack.ID, + "id": res.ID, "message": hostname + " is checking in", - }); err != nil { - slog.Error("Failed to write check-in response", "error", err) - } + }) } } diff --git a/server/api/routes_test.go b/server/api/routes_test.go index 8b7c6b7..d544183 100644 --- a/server/api/routes_test.go +++ b/server/api/routes_test.go @@ -2,11 +2,15 @@ package api import ( "encoding/json" + "errors" "net/http" "net/http/httptest" "reflect" "server/models" + "slices" + "sort" "strings" + "sync" "testing" "github.com/gorilla/mux" @@ -16,6 +20,13 @@ import ( type fakeStorage struct { systems []models.System err error + + // groups is keyed by the group's name folded to lower case, mirroring what + // the bbolt store does with its keys. + groups map[string]models.Group + // groupErr, when set, fails every group read and write, for the tests that + // need the storage layer to break. + groupErr error } func (f *fakeStorage) SaveSystem(hostname string, system models.System) error { return nil } @@ -38,6 +49,100 @@ func (f *fakeStorage) SubscribeToUpdates() <-chan models.System { return make(chan models.System) } +func (f *fakeStorage) GetAllGroups() ([]models.Group, error) { + if f.groupErr != nil { + return nil, f.groupErr + } + groups := []models.Group{} + for _, g := range f.groups { + groups = append(groups, g) + } + sort.Slice(groups, func(i, j int) bool { + return strings.ToLower(groups[i].Name) < strings.ToLower(groups[j].Name) + }) + return groups, nil +} + +func (f *fakeStorage) GetGroup(name string) (models.Group, error) { + if f.groupErr != nil { + return models.Group{}, f.groupErr + } + group, ok := f.groups[strings.ToLower(strings.TrimSpace(name))] + if !ok { + return models.Group{}, errNotFound{} + } + return group, nil +} + +func (f *fakeStorage) SaveGroup(group models.Group) error { + if f.groupErr != nil { + return f.groupErr + } + if f.groups == nil { + f.groups = map[string]models.Group{} + } + sort.Strings(group.Members) + f.groups[strings.ToLower(strings.TrimSpace(group.Name))] = group + return nil +} + +func (f *fakeStorage) DeleteGroup(name string) error { + if f.groupErr != nil { + return f.groupErr + } + key := strings.ToLower(strings.TrimSpace(name)) + if _, ok := f.groups[key]; !ok { + return errNotFound{} + } + delete(f.groups, key) + return nil +} + +func (f *fakeStorage) RenameGroup(oldName, newName string) error { + if f.groupErr != nil { + return f.groupErr + } + oldKey := strings.ToLower(strings.TrimSpace(oldName)) + newKey := strings.ToLower(strings.TrimSpace(newName)) + group, ok := f.groups[oldKey] + if !ok { + return errNotFound{} + } + if oldKey != newKey { + if _, taken := f.groups[newKey]; taken { + return errors.New("group already exists") + } + delete(f.groups, oldKey) + } + group.Name = newName + f.groups[newKey] = group + return nil +} + +func (f *fakeStorage) SetHostGroups(hostname string, groups []string) error { + if f.groupErr != nil { + return f.groupErr + } + wanted := map[string]bool{} + for _, name := range groups { + key := strings.ToLower(strings.TrimSpace(name)) + if _, ok := f.groups[key]; !ok { + return errNotFound{} + } + wanted[key] = true + } + for key, group := range f.groups { + members := slices.DeleteFunc(slices.Clone(group.Members), func(m string) bool { return m == hostname }) + if wanted[key] { + members = append(members, hostname) + } + sort.Strings(members) + group.Members = members + f.groups[key] = group + } + return nil +} + type errNotFound struct{} func (errNotFound) Error() string { return "not found" } @@ -215,54 +320,124 @@ func TestSystemSummaryCoversFreshnessFields(t *testing.T) { } } +// The three fakes below stand in for the NATS connection. +// +// They are mutex-guarded because a group action calls them from one goroutine +// per member at once; without it `go test -race` fails on the call counter +// rather than on anything the production code did wrong. +// +// perHost lets one test produce accepted, skipped and failed answers in a +// single batch, which is the case worth pinning — a group where every host +// agrees proves very little. + // fakeUpdater stands in for the NATS connection in the update-request tests. type fakeUpdater struct { + mu sync.Mutex ack models.UpdateAck err error + perHost map[string]fakeAnswer hostname string // recorded from the last call by string calls int + seen []string + hook func(hostname string) +} + +// fakeAnswer is one host's scripted reply, for the per-host maps. +type fakeAnswer struct { + accepted bool + reason string + id string + command string + err error } func (f *fakeUpdater) RequestUpdate(hostname, requestedBy string) (models.UpdateAck, error) { + if f.hook != nil { + f.hook(hostname) + } + f.mu.Lock() + defer f.mu.Unlock() f.calls++ f.hostname = hostname f.by = requestedBy + f.seen = append(f.seen, hostname) + if a, ok := f.perHost[hostname]; ok { + return models.UpdateAck{ID: a.id, Hostname: hostname, Accepted: a.accepted, Reason: a.reason, Command: a.command}, a.err + } return f.ack, f.err } +func (f *fakeUpdater) hostsSeen() []string { + f.mu.Lock() + defer f.mu.Unlock() + return slices.Clone(f.seen) +} + +func (f *fakeUpdater) callCount() int { + f.mu.Lock() + defer f.mu.Unlock() + return f.calls +} + // fakeCheckIner stands in for the NATS connection in the check-in tests. type fakeCheckIner struct { + mu sync.Mutex ack models.CheckInAck err error + perHost map[string]fakeAnswer hostname string // recorded from the last call by string calls int } func (f *fakeCheckIner) RequestCheckIn(hostname, requestedBy string) (models.CheckInAck, error) { + f.mu.Lock() + defer f.mu.Unlock() f.calls++ f.hostname = hostname f.by = requestedBy + if a, ok := f.perHost[hostname]; ok { + return models.CheckInAck{ID: a.id, Hostname: hostname, Accepted: a.accepted, Reason: a.reason}, a.err + } return f.ack, f.err } +func (f *fakeCheckIner) callCount() int { + f.mu.Lock() + defer f.mu.Unlock() + return f.calls +} + // fakeRebooter stands in for the NATS connection in the reboot tests. type fakeRebooter struct { + mu sync.Mutex ack models.RebootAck err error + perHost map[string]fakeAnswer hostname string // recorded from the last call by string calls int } func (f *fakeRebooter) RequestReboot(hostname, requestedBy string) (models.RebootAck, error) { + f.mu.Lock() + defer f.mu.Unlock() f.calls++ f.hostname = hostname f.by = requestedBy + if a, ok := f.perHost[hostname]; ok { + return models.RebootAck{ID: a.id, Hostname: hostname, Accepted: a.accepted, Reason: a.reason}, a.err + } return f.ack, f.err } +func (f *fakeRebooter) callCount() int { + f.mu.Lock() + defer f.mu.Unlock() + return f.calls +} + func postReboot(handler http.HandlerFunc, hostname string) *httptest.ResponseRecorder { req := httptest.NewRequest(http.MethodPost, "/api/systems/"+hostname+"/reboot", nil) req = mux.SetURLVars(req, map[string]string{"hostname": hostname}) diff --git a/server/models/derived.go b/server/models/derived.go new file mode 100644 index 0000000..d61e4d6 --- /dev/null +++ b/server/models/derived.go @@ -0,0 +1,194 @@ +package models + +import ( + "sort" + "strings" +) + +// Derived groups are computed from what hosts already report, rather than +// curated by hand. "Every Fedora box" and "everything rpm-based" are facts +// about the fleet, not decisions about it, so asking someone to maintain them +// as membership lists would be asking them to keep re-deriving something the +// server can see for itself. +// +// They are namespaced by a prefix so a derived group can never collide with a +// hand-made one; ValidateGroupName rejects ":" in a name for the same reason. +// A derived group exists only while at least one host matches it: boot a +// Debian machine and pkg:deb appears, retire it and the group goes away. There +// is nothing to clean up, which is the whole point. +const ( + DerivedPrefixOS = "os:" + DerivedPrefixPkg = "pkg:" + DerivedPrefixArch = "arch:" + DerivedPrefixState = "state:" +) + +// derivedPrefixes is the set used to tell a derived name from a stored one. +var derivedPrefixes = []string{ + DerivedPrefixOS, DerivedPrefixPkg, DerivedPrefixArch, DerivedPrefixState, +} + +// IsDerivedName reports whether a name belongs to the derived namespace, and +// therefore names a group that is computed rather than stored. Nothing can +// create, rename, delete or edit the membership of such a group. +func IsDerivedName(name string) bool { + lower := strings.ToLower(strings.TrimSpace(name)) + for _, prefix := range derivedPrefixes { + if strings.HasPrefix(lower, prefix) { + return true + } + } + return false +} + +// osFamilyRule maps a substring of the OS string to a family slug. +// +// The OS string is the distribution's own PRETTY_NAME, so this is a matching +// table rather than an enumeration, and order matters: Pop!_OS has to be +// tested before Ubuntu because it reports "Pop!_OS 22.04 LTS" and is an Ubuntu +// derivative, and the same holds for the RHEL rebuilds. +// +// A distribution not listed here simply joins no os: or pkg: group. That is +// the honest outcome — guessing a package family from an unrecognised name is +// how a host ends up in "all rpm" and gets handed a dnf command it cannot run. +type osFamilyRule struct { + match []string + family string + pkg string +} + +var osFamilyRules = []osFamilyRule{ + // macOS first: "darwin" collides with nothing, but brew is not a Linux + // package family and should not fall through to one. + {[]string{"darwin", "macos", "mac os"}, "macos", "brew"}, + + // Debian and its derivatives. Pop!_OS and Mint precede Ubuntu, and Ubuntu + // precedes Debian, because each reports its own name and is a derivative + // of the next. + {[]string{"pop!_os", "pop_os", "pop-os", "pop os"}, "pop-os", "deb"}, + {[]string{"linux mint", "linuxmint"}, "mint", "deb"}, + {[]string{"raspbian", "raspberry pi os"}, "raspbian", "deb"}, + {[]string{"kali"}, "kali", "deb"}, + {[]string{"devuan"}, "devuan", "deb"}, + {[]string{"ubuntu"}, "ubuntu", "deb"}, + {[]string{"debian"}, "debian", "deb"}, + + // The RHEL family. The rebuilds name themselves, so they are tested before + // "red hat" to keep each one its own os: group while they share pkg:rpm. + {[]string{"rocky"}, "rocky", "rpm"}, + {[]string{"almalinux", "alma linux"}, "alma", "rpm"}, + {[]string{"centos"}, "centos", "rpm"}, + {[]string{"oracle linux"}, "oracle", "rpm"}, + {[]string{"amazon linux"}, "amazon", "rpm"}, + {[]string{"scientific linux"}, "scientific", "rpm"}, + {[]string{"red hat", "redhat", "rhel"}, "rhel", "rpm"}, + {[]string{"fedora"}, "fedora", "rpm"}, + + // SUSE is rpm-based but reaches it through zypper. It is still pkg:rpm: + // the question "which hosts take an rpm" has one answer, and the update + // script picks the manager per host anyway. + {[]string{"opensuse", "suse"}, "suse", "rpm"}, + + // The rest, each its own package family. + {[]string{"nixos"}, "nixos", "nix"}, + {[]string{"manjaro"}, "manjaro", "pacman"}, + {[]string{"endeavouros"}, "endeavouros", "pacman"}, + {[]string{"arch linux", "archlinux"}, "arch", "pacman"}, + {[]string{"gentoo"}, "gentoo", "portage"}, + {[]string{"alpine"}, "alpine", "apk"}, + {[]string{"void linux"}, "void", "xbps"}, +} + +// OSFamily returns the distribution slug for an OS string, or "" when the +// distribution is not one this table knows. +func OSFamily(os string) string { + lower := strings.ToLower(os) + for _, rule := range osFamilyRules { + for _, needle := range rule.match { + if strings.Contains(lower, needle) { + return rule.family + } + } + } + return "" +} + +// PackageFamily returns the package family slug for an OS string ("rpm", +// "deb", ...), or "" when it cannot be determined. +// +// It is inferred from the distribution rather than reported by the host. The +// client knows exactly which package manager it used, so reporting it would be +// more precise; inferring keeps this entirely server-side and works for every +// host already checking in, including ones running an older client. +func PackageFamily(os string) string { + lower := strings.ToLower(os) + for _, rule := range osFamilyRules { + for _, needle := range rule.match { + if strings.Contains(lower, needle) { + return rule.pkg + } + } + } + return "" +} + +// DeriveGroups computes every derived group that currently has a member. +// +// Groups with no members are not returned at all, so the dashboard's chip bar +// shows the fleet that exists rather than a list of everything it might +// contain. +func DeriveGroups(systems []System) []Group { + members := map[string][]string{} + add := func(name, hostname string) { + members[name] = append(members[name], hostname) + } + + for _, system := range systems { + host := strings.TrimSpace(system.Hostname) + if host == "" { + continue + } + + if family := OSFamily(system.OS); family != "" { + add(DerivedPrefixOS+family, host) + } + if pkg := PackageFamily(system.OS); pkg != "" { + add(DerivedPrefixPkg+pkg, host) + } + if arch := strings.TrimSpace(system.Architecture); arch != "" { + add(DerivedPrefixArch+strings.ToLower(arch), host) + } + + // State groups move under you as hosts check in, unlike the three + // above. That is what makes them useful — "everything that needs a + // reboot" is exactly the set you want to reboot — and a group action + // resolves its members when it runs, not when the page was drawn. + if system.RebootRequired { + add(DerivedPrefixState+"needs-reboot", host) + } + if system.UpdatesAvailable { + add(DerivedPrefixState+"has-updates", host) + } + } + + groups := make([]Group, 0, len(members)) + for name, hosts := range members { + sort.Strings(hosts) + groups = append(groups, Group{Name: name, Members: hosts, Derived: true}) + } + sort.Slice(groups, func(i, j int) bool { return groups[i].Name < groups[j].Name }) + + return groups +} + +// FindDerivedGroup returns the derived group of that name, if it currently has +// any members. The lookup is case-insensitive, like a stored group's. +func FindDerivedGroup(systems []System, name string) (Group, bool) { + name = strings.TrimSpace(name) + for _, group := range DeriveGroups(systems) { + if SameGroupName(group.Name, name) { + return group, true + } + } + return Group{}, false +} diff --git a/server/models/derived_test.go b/server/models/derived_test.go new file mode 100644 index 0000000..9553fb0 --- /dev/null +++ b/server/models/derived_test.go @@ -0,0 +1,215 @@ +package models + +import ( + "slices" + "testing" +) + +// TestOSFamilyAndPackageFamily is the matching table's guard. The OS string is +// whatever the distribution put in its own PRETTY_NAME, so these are the real +// shapes rather than invented ones. +func TestOSFamilyAndPackageFamily(t *testing.T) { + for _, tc := range []struct { + os string + family string + pkg string + }{ + {"Fedora Linux 42 (Workstation Edition)", "fedora", "rpm"}, + {"Rocky Linux 10.2 (Red Quartz)", "rocky", "rpm"}, + {"AlmaLinux 9.4 (Seafoam Ocelot)", "alma", "rpm"}, + {"CentOS Stream 10", "centos", "rpm"}, + {"Red Hat Enterprise Linux 9.4 (Plow)", "rhel", "rpm"}, + {"openSUSE Tumbleweed", "suse", "rpm"}, + {"Debian GNU/Linux 12 (bookworm)", "debian", "deb"}, + {"Ubuntu 24.04.1 LTS", "ubuntu", "deb"}, + {"Linux Mint 22", "mint", "deb"}, + {"Raspbian GNU/Linux 11 (bullseye)", "raspbian", "deb"}, + {"NixOS 24.11 (Vicuna)", "nixos", "nix"}, + {"Arch Linux", "arch", "pacman"}, + {"Manjaro Linux", "manjaro", "pacman"}, + {"Gentoo Linux", "gentoo", "portage"}, + {"Alpine Linux v3.20", "alpine", "apk"}, + {"macOS 15.1", "macos", "brew"}, + + // Not in the table: joins no os: or pkg: group rather than being + // guessed at. Guessing is how a host lands in "all rpm" and gets + // handed a dnf command it cannot run. + {"Slackware 15.0", "", ""}, + {"", "", ""}, + } { + t.Run(tc.os, func(t *testing.T) { + if got := OSFamily(tc.os); got != tc.family { + t.Errorf("OSFamily(%q) = %q, want %q", tc.os, got, tc.family) + } + if got := PackageFamily(tc.os); got != tc.pkg { + t.Errorf("PackageFamily(%q) = %q, want %q", tc.os, got, tc.pkg) + } + }) + } +} + +// TestOSFamilyChecksDerivativesFirst pins the ordering that the table depends +// on. Pop!_OS and Mint are Ubuntu derivatives and Ubuntu is a Debian one, so a +// table tested in the wrong order collapses all three into "debian" and the +// os: groups stop telling them apart. The package family is shared either way, +// which is exactly what makes the bug quiet. +func TestOSFamilyChecksDerivativesFirst(t *testing.T) { + for _, tc := range []struct{ os, want string }{ + {"Pop!_OS 22.04 LTS", "pop-os"}, + {"Linux Mint 22 (Wilma)", "mint"}, + {"Ubuntu 24.04.1 LTS", "ubuntu"}, + {"Rocky Linux 10.2", "rocky"}, + {"AlmaLinux 9.4", "alma"}, + {"Red Hat Enterprise Linux 9.4", "rhel"}, + } { + if got := OSFamily(tc.os); got != tc.want { + t.Errorf("OSFamily(%q) = %q, want %q", tc.os, got, tc.want) + } + } +} + +func TestIsDerivedName(t *testing.T) { + for _, tc := range []struct { + name string + want bool + }{ + {"os:fedora", true}, + {"pkg:rpm", true}, + {"arch:x86_64", true}, + {"state:needs-reboot", true}, + {"OS:Fedora", true}, + {" os:fedora ", true}, + {"prod", false}, + {"web servers", false}, + {"", false}, + } { + if got := IsDerivedName(tc.name); got != tc.want { + t.Errorf("IsDerivedName(%q) = %v, want %v", tc.name, got, tc.want) + } + } +} + +// TestValidateGroupNameReservesColon: the derived namespace only works as a +// namespace if nothing can be hand-made inside it. A stored "os:fedora" +// sitting beside a computed one would make a group action ambiguous about +// which set of hosts it was acting on. +func TestValidateGroupNameReservesColon(t *testing.T) { + if err := ValidateGroupName("os:fedora"); err == nil { + t.Error("a name in the derived namespace was accepted") + } + if err := ValidateGroupName("prod"); err != nil { + t.Errorf("an ordinary name was rejected: %v", err) + } +} + +func fleet() []System { + return []System{ + {Hostname: "web01", OS: "Fedora Linux 42", Architecture: "x86_64", RebootRequired: true, UpdatesAvailable: true}, + {Hostname: "web02", OS: "Fedora Linux 42", Architecture: "x86_64"}, + {Hostname: "db01", OS: "Rocky Linux 10.2", Architecture: "aarch64", UpdatesAvailable: true}, + {Hostname: "build01", OS: "Ubuntu 24.04.1 LTS", Architecture: "x86_64", RebootRequired: true}, + } +} + +func groupNames(groups []Group) []string { + names := make([]string, 0, len(groups)) + for _, g := range groups { + names = append(names, g.Name) + } + return names +} + +func memberList(t *testing.T, groups []Group, name string) []string { + t.Helper() + for _, g := range groups { + if g.Name == name { + return g.Members + } + } + t.Fatalf("no derived group %q in %v", name, groupNames(groups)) + return nil +} + +func TestDeriveGroups(t *testing.T) { + groups := DeriveGroups(fleet()) + + want := []string{ + "arch:aarch64", "arch:x86_64", + "os:fedora", "os:rocky", "os:ubuntu", + "pkg:deb", "pkg:rpm", + "state:has-updates", "state:needs-reboot", + } + if got := groupNames(groups); !slices.Equal(got, want) { + t.Errorf("groups = %v\nwant %v", got, want) + } + + // The whole point of pkg: — Fedora and Rocky are different distributions + // and the same package family. + if got := memberList(t, groups, "pkg:rpm"); !slices.Equal(got, []string{"db01", "web01", "web02"}) { + t.Errorf("pkg:rpm = %v, want [db01 web01 web02]", got) + } + if got := memberList(t, groups, "pkg:deb"); !slices.Equal(got, []string{"build01"}) { + t.Errorf("pkg:deb = %v, want [build01]", got) + } + if got := memberList(t, groups, "os:fedora"); !slices.Equal(got, []string{"web01", "web02"}) { + t.Errorf("os:fedora = %v, want [web01 web02]", got) + } + if got := memberList(t, groups, "state:needs-reboot"); !slices.Equal(got, []string{"build01", "web01"}) { + t.Errorf("state:needs-reboot = %v, want [build01 web01]", got) + } + + for _, g := range groups { + if !g.Derived { + t.Errorf("%s is not marked derived", g.Name) + } + } +} + +// TestDeriveGroupsOmitsEmptyOnes: the chip bar should show the fleet that +// exists, not a catalogue of every distribution this server has heard of. +func TestDeriveGroupsOmitsEmptyOnes(t *testing.T) { + groups := DeriveGroups([]System{ + {Hostname: "web01", OS: "Fedora Linux 42", Architecture: "x86_64"}, + }) + + for _, name := range []string{"pkg:deb", "os:debian", "state:needs-reboot", "state:has-updates"} { + if slices.Contains(groupNames(groups), name) { + t.Errorf("%s was derived with no members", name) + } + } +} + +// TestDeriveGroupsSkipsUnknownDistros: an unrecognised OS still gets its arch +// group, because the architecture is reported rather than inferred. It just +// joins no os: or pkg: group. +func TestDeriveGroupsSkipsUnknownDistros(t *testing.T) { + groups := DeriveGroups([]System{ + {Hostname: "odd01", OS: "Slackware 15.0", Architecture: "x86_64"}, + }) + + if got := groupNames(groups); !slices.Equal(got, []string{"arch:x86_64"}) { + t.Errorf("groups = %v, want [arch:x86_64]", got) + } +} + +func TestDeriveGroupsOnEmptyFleet(t *testing.T) { + if groups := DeriveGroups(nil); len(groups) != 0 { + t.Errorf("got %d groups from an empty fleet, want 0", len(groups)) + } +} + +func TestFindDerivedGroup(t *testing.T) { + systems := fleet() + + group, ok := FindDerivedGroup(systems, "PKG:RPM") + if !ok { + t.Fatal("a derived group was not found through a differently-cased name") + } + if group.Name != "pkg:rpm" || !group.Derived { + t.Errorf("got %+v, want the derived pkg:rpm", group) + } + + if _, ok := FindDerivedGroup(systems, "pkg:apk"); ok { + t.Error("a derived group with no members was reported as existing") + } +} diff --git a/server/models/groups.go b/server/models/groups.go new file mode 100644 index 0000000..be1c1e6 --- /dev/null +++ b/server/models/groups.go @@ -0,0 +1,83 @@ +package models + +import ( + "fmt" + "strings" + "unicode" +) + +// MaxGroupNameLength bounds a group name. Nothing technical needs the limit — +// it is there so a pasted accident cannot become a chip nobody can read or a +// bbolt key nobody can find again. +const MaxGroupNameLength = 64 + +// Group is a named set of hosts, managed on the server rather than declared by +// the hosts themselves. Nothing has to be configured on a host to put it in a +// group, and a host cannot put itself in one. +// +// Membership lives in its own bucket rather than on System because a client +// publishes a whole System on every check-in and the server stores what it +// sent. Anything server-owned kept there has to be carried forward by hand on +// every check-in — there are already three such fields — and the next person to +// add one would have no reason to suspect groups were among them. +// +// Members may name a host the server has never seen. That is deliberate: a host +// deleted from the dashboard and re-added should not silently lose its groups, +// and a group can be built before the hosts in it are provisioned. The +// dashboard shows such a member greyed out, and a group action skips it. +type Group struct { + Name string `json:"name"` + Members []string `json:"members"` + // Derived marks a group whose membership is computed from what hosts + // report rather than stored. It is never true for anything in the groups + // bucket, and a derived group cannot be created, renamed, deleted, or + // have its membership edited. See derived.go. + Derived bool `json:"derived,omitempty"` +} + +// ValidateGroupName reports why a name cannot be used, or nil if it can. The +// name is assumed already trimmed by NormalizeGroupName. +// +// A group name is a URL path segment and a bbolt key, so it rules out the +// separator and anything unprintable; the rest is just keeping it legible. +func ValidateGroupName(name string) error { + switch { + case name == "": + return fmt.Errorf("a group name is required") + case len(name) > MaxGroupNameLength: + return fmt.Errorf("a group name may be at most %d characters", MaxGroupNameLength) + case strings.ContainsRune(name, '/'): + return fmt.Errorf("a group name may not contain %q", "/") + case strings.ContainsRune(name, ':'): + // ":" is how a derived group is told from a stored one. Reserving it + // means a hand-made group can never shadow "os:fedora" or be shadowed + // by it, which would otherwise make a group action ambiguous about + // which set of hosts it was acting on. + return fmt.Errorf("a group name may not contain %q, which is reserved for derived groups like %q", ":", "os:fedora") + } + + for _, r := range name { + if !unicode.IsPrint(r) { + return fmt.Errorf("a group name may not contain control characters") + } + } + + return nil +} + +// NormalizeGroupName trims the surrounding whitespace a text input collects. +// Case is left alone: a group is stored as it was typed, and only compared +// case-insensitively (see SameGroupName). +func NormalizeGroupName(name string) string { + return strings.TrimSpace(name) +} + +// SameGroupName reports whether two names refer to the same group. +// +// Comparing case-insensitively is what stops "prod" and "Prod" from both +// existing. That pair is the one mistake worth designing against here: both +// chips look right, and a group action on either one silently misses half the +// fleet rather than failing. +func SameGroupName(a, b string) bool { + return strings.EqualFold(a, b) +} diff --git a/server/nats/subscriber_test.go b/server/nats/subscriber_test.go index 5b24975..b0da24c 100644 --- a/server/nats/subscriber_test.go +++ b/server/nats/subscriber_test.go @@ -2,6 +2,7 @@ package nats import ( "encoding/json" + "errors" "server/models" "testing" "time" @@ -53,6 +54,17 @@ func (s *memStore) DeleteSystem(hostname string) error { func (s *memStore) SubscribeToUpdates() <-chan models.System { return make(chan models.System) } +// Groups play no part in the NATS handlers — nothing a client publishes can +// change a group — so memStore satisfies the interface and does nothing. +func (s *memStore) GetAllGroups() ([]models.Group, error) { return nil, nil } +func (s *memStore) GetGroup(name string) (models.Group, error) { + return models.Group{}, errors.New("group not found") +} +func (s *memStore) SaveGroup(group models.Group) error { return nil } +func (s *memStore) DeleteGroup(name string) error { return nil } +func (s *memStore) RenameGroup(oldName, newName string) error { return nil } +func (s *memStore) SetHostGroups(hostname string, groups []string) error { return nil } + type errMissing struct{} func (errMissing) Error() string { return "not found" } diff --git a/server/storage/bbolt.go b/server/storage/bbolt.go index 5b7c576..09c9f18 100644 --- a/server/storage/bbolt.go +++ b/server/storage/bbolt.go @@ -187,7 +187,15 @@ func (s *BboltStorage) DeleteSystem(hostname string) error { return fmt.Errorf("system with hostname '%s' not found", hostname) } - return bucket.Delete([]byte(hostname)) + if err := bucket.Delete([]byte(hostname)); err != nil { + return err + } + + // Drop the host from every group it was in, in the same transaction. + // A deleted host that stayed in its groups would come back as a member + // nothing can act on, and would quietly pad the member count on every + // group action from then on. + return removeHostFromGroups(tx, hostname) }) if err != nil { diff --git a/server/storage/groups.go b/server/storage/groups.go new file mode 100644 index 0000000..941cb57 --- /dev/null +++ b/server/storage/groups.go @@ -0,0 +1,384 @@ +package storage + +import ( + "encoding/json" + "fmt" + "log/slog" + "server/metrics" + "server/models" + "slices" + "sort" + "strings" + "time" + + bolt "go.etcd.io/bbolt" +) + +// groupsBucket holds one record per group: key is the group's name folded to +// lower case, value is the JSON models.Group carrying the name as it was typed. +// +// Keying on the folded name is what makes "prod" and "Prod" the same group +// rather than two that look alike — the uniqueness check is the key itself +// rather than a scan that someone later forgets to do — and it lets a URL name +// a group in whatever case the operator typed. +const groupsBucket = "groups" + +// groupKey is the bucket key for a group name. +func groupKey(name string) []byte { + return []byte(strings.ToLower(strings.TrimSpace(name))) +} + +// normalizeMembers sorts and de-duplicates a member list so a stored group +// reads the same way every time and the dashboard never has to sort it. +func normalizeMembers(members []string) []string { + seen := make(map[string]bool, len(members)) + out := make([]string, 0, len(members)) + for _, m := range members { + m = strings.TrimSpace(m) + if m == "" || seen[m] { + continue + } + seen[m] = true + out = append(out, m) + } + sort.Strings(out) + return out +} + +// readGroup pulls one group out of an open transaction. ok is false when the +// bucket or the key is absent, which are both "no such group" rather than +// errors — the bucket does not exist until the first group is created. +func readGroup(tx *bolt.Tx, name string) (models.Group, bool, error) { + bucket := tx.Bucket([]byte(groupsBucket)) + if bucket == nil { + return models.Group{}, false, nil + } + + data := bucket.Get(groupKey(name)) + if data == nil { + return models.Group{}, false, nil + } + + var group models.Group + if err := json.Unmarshal(data, &group); err != nil { + return models.Group{}, false, fmt.Errorf("failed to unmarshal group %q: %w", name, err) + } + + return group, true, nil +} + +// writeGroup stores a group, creating the bucket if this is the first one. +func writeGroup(tx *bolt.Tx, group models.Group) error { + bucket, err := tx.CreateBucketIfNotExists([]byte(groupsBucket)) + if err != nil { + return err + } + + group.Members = normalizeMembers(group.Members) + data, err := json.Marshal(group) + if err != nil { + return err + } + + return bucket.Put(groupKey(group.Name), data) +} + +// GetAllGroups returns every group, ordered by name. A database with no groups +// yet returns an empty slice rather than an error. +func (s *BboltStorage) GetAllGroups() ([]models.Group, error) { + start := time.Now() + defer func() { + metrics.StorageOperationDuration.WithLabelValues("get_all_groups").Observe(time.Since(start).Seconds()) + }() + + groups := []models.Group{} + + err := s.db.View(func(tx *bolt.Tx) error { + bucket := tx.Bucket([]byte(groupsBucket)) + if bucket == nil { + return nil + } + + return bucket.ForEach(func(k, v []byte) error { + var group models.Group + if err := json.Unmarshal(v, &group); err != nil { + // Same tolerance GetAllSystems shows: one unreadable record + // should not hide every other group. + slog.Error("Failed to unmarshal group, skipping", "key", string(k), "error", err) + return nil + } + groups = append(groups, group) + return nil + }) + }) + if err != nil { + metrics.StorageOperationErrors.WithLabelValues("get_all_groups").Inc() + return nil, err + } + + sort.Slice(groups, func(i, j int) bool { + return strings.ToLower(groups[i].Name) < strings.ToLower(groups[j].Name) + }) + + return groups, nil +} + +// GetGroup returns one group by name, matched without regard to case. +func (s *BboltStorage) GetGroup(name string) (models.Group, error) { + start := time.Now() + defer func() { + metrics.StorageOperationDuration.WithLabelValues("get_group").Observe(time.Since(start).Seconds()) + }() + + var group models.Group + + err := s.db.View(func(tx *bolt.Tx) error { + found, ok, err := readGroup(tx, name) + if err != nil { + return err + } + if !ok { + return fmt.Errorf("group '%s' not found", name) + } + group = found + return nil + }) + if err != nil { + metrics.StorageOperationErrors.WithLabelValues("get_group").Inc() + return models.Group{}, err + } + + return group, nil +} + +// SaveGroup creates a group or replaces it wholesale, members and all. A group +// that already exists under a different spelling of the same name keeps the +// spelling passed here. +func (s *BboltStorage) SaveGroup(group models.Group) error { + start := time.Now() + defer func() { + metrics.StorageOperationDuration.WithLabelValues("save_group").Observe(time.Since(start).Seconds()) + }() + + s.Lock() + defer s.Unlock() + + err := s.db.Update(func(tx *bolt.Tx) error { + return writeGroup(tx, group) + }) + if err != nil { + metrics.StorageOperationErrors.WithLabelValues("save_group").Inc() + slog.Error("Failed to save group", "group", group.Name, "error", err) + return err + } + + return nil +} + +// DeleteGroup removes a group. The hosts in it are untouched — a group is a +// label, and dropping the label is not dropping the machines. +func (s *BboltStorage) DeleteGroup(name string) error { + start := time.Now() + defer func() { + metrics.StorageOperationDuration.WithLabelValues("delete_group").Observe(time.Since(start).Seconds()) + }() + + s.Lock() + defer s.Unlock() + + err := s.db.Update(func(tx *bolt.Tx) error { + if _, ok, err := readGroup(tx, name); err != nil { + return err + } else if !ok { + return fmt.Errorf("group '%s' not found", name) + } + return tx.Bucket([]byte(groupsBucket)).Delete(groupKey(name)) + }) + if err != nil { + metrics.StorageOperationErrors.WithLabelValues("delete_group").Inc() + slog.Error("Failed to delete group", "group", name, "error", err) + return err + } + + slog.Info("Group deleted", "group", name) + return nil +} + +// RenameGroup renames a group, keeping its members. +// +// It is one transaction rather than a delete and a create because the halfway +// state — the old group gone, the new one not yet written — would lose every +// member if the process died between them. +func (s *BboltStorage) RenameGroup(oldName, newName string) error { + start := time.Now() + defer func() { + metrics.StorageOperationDuration.WithLabelValues("rename_group").Observe(time.Since(start).Seconds()) + }() + + s.Lock() + defer s.Unlock() + + err := s.db.Update(func(tx *bolt.Tx) error { + group, ok, err := readGroup(tx, oldName) + if err != nil { + return err + } + if !ok { + return fmt.Errorf("group '%s' not found", oldName) + } + + // A rename that only changes case keeps the same key, so it is a + // relabel in place rather than a collision with itself. + if !models.SameGroupName(oldName, newName) { + if _, taken, err := readGroup(tx, newName); err != nil { + return err + } else if taken { + return fmt.Errorf("group '%s' already exists", newName) + } + } + + group.Name = newName + if err := writeGroup(tx, group); err != nil { + return err + } + + if !models.SameGroupName(oldName, newName) { + return tx.Bucket([]byte(groupsBucket)).Delete(groupKey(oldName)) + } + return nil + }) + if err != nil { + metrics.StorageOperationErrors.WithLabelValues("rename_group").Inc() + slog.Error("Failed to rename group", "group", oldName, "to", newName, "error", err) + return err + } + + slog.Info("Group renamed", "group", oldName, "to", newName) + return nil +} + +// SetHostGroups replaces one host's memberships: it joins every group named +// here and leaves every other one. +// +// This is a single transaction because it is a single edit as far as the +// operator is concerned — they ticked boxes on one host and pressed save. Doing +// it as one write per group would leave the host in a mixture of its old and +// new groups if any of them failed. +// +// Every named group must already exist; creating one is a separate, deliberate +// act, so that a typo in a group name joins nothing rather than quietly +// founding a group of one. +func (s *BboltStorage) SetHostGroups(hostname string, groups []string) error { + start := time.Now() + defer func() { + metrics.StorageOperationDuration.WithLabelValues("set_host_groups").Observe(time.Since(start).Seconds()) + }() + + hostname = strings.TrimSpace(hostname) + if hostname == "" { + return fmt.Errorf("a hostname is required") + } + + s.Lock() + defer s.Unlock() + + err := s.db.Update(func(tx *bolt.Tx) error { + wanted := make(map[string]bool, len(groups)) + for _, name := range groups { + group, ok, err := readGroup(tx, name) + if err != nil { + return err + } + if !ok { + return fmt.Errorf("group '%s' not found", name) + } + wanted[string(groupKey(group.Name))] = true + } + + bucket := tx.Bucket([]byte(groupsBucket)) + if bucket == nil { + // No groups exist, so there is nothing to join and nothing to + // leave. Asking for none of them is the only coherent request. + if len(wanted) > 0 { + return fmt.Errorf("no groups exist") + } + return nil + } + + // Walk every group once, adding or removing this host as the wanted + // set says, so a membership the caller did not mention is dropped. + var changed []models.Group + err := bucket.ForEach(func(k, v []byte) error { + var group models.Group + if err := json.Unmarshal(v, &group); err != nil { + slog.Error("Failed to unmarshal group, skipping", "key", string(k), "error", err) + return nil + } + + member := slices.Contains(group.Members, hostname) + switch { + case wanted[string(k)] && !member: + group.Members = append(group.Members, hostname) + case !wanted[string(k)] && member: + group.Members = slices.DeleteFunc(group.Members, func(m string) bool { return m == hostname }) + default: + return nil + } + + changed = append(changed, group) + return nil + }) + if err != nil { + return err + } + + for _, group := range changed { + if err := writeGroup(tx, group); err != nil { + return err + } + } + return nil + }) + if err != nil { + metrics.StorageOperationErrors.WithLabelValues("set_host_groups").Inc() + slog.Error("Failed to set host groups", "hostname", hostname, "error", err) + return err + } + + return nil +} + +// removeHostFromGroups drops a hostname from every group inside an open +// transaction. DeleteSystem calls it so that deleting a host from the dashboard +// does not leave its name behind in groups it used to be in. +func removeHostFromGroups(tx *bolt.Tx, hostname string) error { + bucket := tx.Bucket([]byte(groupsBucket)) + if bucket == nil { + return nil + } + + var changed []models.Group + err := bucket.ForEach(func(k, v []byte) error { + var group models.Group + if err := json.Unmarshal(v, &group); err != nil { + slog.Error("Failed to unmarshal group, skipping", "key", string(k), "error", err) + return nil + } + if !slices.Contains(group.Members, hostname) { + return nil + } + group.Members = slices.DeleteFunc(group.Members, func(m string) bool { return m == hostname }) + changed = append(changed, group) + return nil + }) + if err != nil { + return err + } + + for _, group := range changed { + if err := writeGroup(tx, group); err != nil { + return err + } + } + return nil +} diff --git a/server/storage/groups_test.go b/server/storage/groups_test.go new file mode 100644 index 0000000..4332bb4 --- /dev/null +++ b/server/storage/groups_test.go @@ -0,0 +1,380 @@ +package storage + +import ( + "path/filepath" + "server/models" + "slices" + "testing" +) + +// newTestStore opens a database in a directory the test owns, so nothing here +// can reach the real systems.db. +func newTestStore(t *testing.T) *BboltStorage { + t.Helper() + + store, err := NewBboltStorage(filepath.Join(t.TempDir(), "test.db")) + if err != nil { + t.Fatalf("opening test database: %v", err) + } + t.Cleanup(func() { store.Close() }) + return store +} + +func mustSaveGroup(t *testing.T, store *BboltStorage, name string, members ...string) { + t.Helper() + if err := store.SaveGroup(models.Group{Name: name, Members: members}); err != nil { + t.Fatalf("saving group %q: %v", name, err) + } +} + +func mustGetGroup(t *testing.T, store *BboltStorage, name string) models.Group { + t.Helper() + group, err := store.GetGroup(name) + if err != nil { + t.Fatalf("getting group %q: %v", name, err) + } + return group +} + +// TestGroupRoundTrip is the baseline: what goes in comes back out, with the +// name spelled as it was typed. +func TestGroupRoundTrip(t *testing.T) { + store := newTestStore(t) + + mustSaveGroup(t, store, "Production", "web01", "db01") + + group := mustGetGroup(t, store, "Production") + if group.Name != "Production" { + t.Errorf("Name = %q, want %q", group.Name, "Production") + } + if !slices.Equal(group.Members, []string{"db01", "web01"}) { + t.Errorf("Members = %v, want [db01 web01]", group.Members) + } +} + +// TestGetGroupIgnoresCase pins the lookup half of the case rule. A URL that +// names "prod" must find the group someone created as "Prod", or the dashboard +// and the API would disagree about which groups exist. +func TestGetGroupIgnoresCase(t *testing.T) { + store := newTestStore(t) + mustSaveGroup(t, store, "Prod", "web01") + + if _, err := store.GetGroup("PROD"); err != nil { + t.Errorf("GetGroup(%q) = %v, want the group created as %q", "PROD", err, "Prod") + } +} + +// TestSaveGroupDoesNotForkOnCase is the other half, and the one that matters: +// saving "prod" over "Prod" must replace it rather than create a second group. +// Two groups whose names differ only in case both look right in the chip bar, +// and an action on either one silently misses the hosts in the other. +func TestSaveGroupDoesNotForkOnCase(t *testing.T) { + store := newTestStore(t) + + mustSaveGroup(t, store, "Prod", "web01") + mustSaveGroup(t, store, "prod", "web01", "db01") + + groups, err := store.GetAllGroups() + if err != nil { + t.Fatalf("GetAllGroups: %v", err) + } + if len(groups) != 1 { + t.Fatalf("got %d groups, want 1: %+v", len(groups), groups) + } + if groups[0].Name != "prod" { + t.Errorf("Name = %q, want the spelling from the later save, %q", groups[0].Name, "prod") + } +} + +// TestSaveGroupNormalizesMembers checks a member list is stored sorted and +// de-duplicated, so the dashboard never has to sort and a double-add is not +// visible as two rows. +func TestSaveGroupNormalizesMembers(t *testing.T) { + store := newTestStore(t) + + mustSaveGroup(t, store, "prod", "web02", "db01", "web02", " ", "web01") + + group := mustGetGroup(t, store, "prod") + if !slices.Equal(group.Members, []string{"db01", "web01", "web02"}) { + t.Errorf("Members = %v, want [db01 web01 web02]", group.Members) + } +} + +// TestGetAllGroupsOnEmptyDatabase pins that a database where no group has ever +// been created reads as no groups rather than as an error. The bucket does not +// exist until the first write, and the dashboard asks for this on every load. +func TestGetAllGroupsOnEmptyDatabase(t *testing.T) { + store := newTestStore(t) + + groups, err := store.GetAllGroups() + if err != nil { + t.Fatalf("GetAllGroups on an empty database: %v", err) + } + if len(groups) != 0 { + t.Errorf("got %d groups, want 0", len(groups)) + } +} + +func TestGetGroupNotFound(t *testing.T) { + store := newTestStore(t) + mustSaveGroup(t, store, "prod") + + if _, err := store.GetGroup("staging"); err == nil { + t.Error("GetGroup on a group that does not exist returned no error") + } +} + +func TestDeleteGroup(t *testing.T) { + store := newTestStore(t) + mustSaveGroup(t, store, "prod", "web01") + + if err := store.DeleteGroup("PROD"); err != nil { + t.Fatalf("DeleteGroup: %v", err) + } + if _, err := store.GetGroup("prod"); err == nil { + t.Error("the group is still there after DeleteGroup") + } + if err := store.DeleteGroup("prod"); err == nil { + t.Error("deleting a group twice returned no error the second time") + } +} + +// TestDeleteGroupLeavesSystemsAlone: a group is a label. Dropping the label is +// not dropping the machines. +func TestDeleteGroupLeavesSystemsAlone(t *testing.T) { + store := newTestStore(t) + if err := store.SaveSystem("web01", models.System{Hostname: "web01"}); err != nil { + t.Fatalf("SaveSystem: %v", err) + } + mustSaveGroup(t, store, "prod", "web01") + + if err := store.DeleteGroup("prod"); err != nil { + t.Fatalf("DeleteGroup: %v", err) + } + if _, err := store.GetSystem("web01"); err != nil { + t.Errorf("deleting a group took its member with it: %v", err) + } +} + +// TestRenameGroupKeepsMembers is the point of having a rename at all: a delete +// and a create would drop everyone in the group. +func TestRenameGroupKeepsMembers(t *testing.T) { + store := newTestStore(t) + mustSaveGroup(t, store, "prod", "web01", "db01") + + if err := store.RenameGroup("prod", "production"); err != nil { + t.Fatalf("RenameGroup: %v", err) + } + + if _, err := store.GetGroup("prod"); err == nil { + t.Error("the old name still resolves after a rename") + } + group := mustGetGroup(t, store, "production") + if !slices.Equal(group.Members, []string{"db01", "web01"}) { + t.Errorf("Members = %v, want [db01 web01]", group.Members) + } +} + +// TestRenameGroupToDifferentCase is a relabel in place, not a collision with +// itself — the key does not change, so the group must survive. +func TestRenameGroupToDifferentCase(t *testing.T) { + store := newTestStore(t) + mustSaveGroup(t, store, "prod", "web01") + + if err := store.RenameGroup("prod", "Prod"); err != nil { + t.Fatalf("RenameGroup to a different case: %v", err) + } + + group := mustGetGroup(t, store, "prod") + if group.Name != "Prod" { + t.Errorf("Name = %q, want %q", group.Name, "Prod") + } + if !slices.Equal(group.Members, []string{"web01"}) { + t.Errorf("Members = %v, want [web01]", group.Members) + } +} + +func TestRenameGroupRefusesCollision(t *testing.T) { + store := newTestStore(t) + mustSaveGroup(t, store, "prod", "web01") + mustSaveGroup(t, store, "staging", "web02") + + if err := store.RenameGroup("prod", "STAGING"); err == nil { + t.Fatal("renaming onto an existing group returned no error") + } + + // And neither group was damaged by the attempt. + if group := mustGetGroup(t, store, "prod"); !slices.Equal(group.Members, []string{"web01"}) { + t.Errorf("prod members = %v, want [web01]", group.Members) + } + if group := mustGetGroup(t, store, "staging"); !slices.Equal(group.Members, []string{"web02"}) { + t.Errorf("staging members = %v, want [web02]", group.Members) + } +} + +func TestRenameGroupNotFound(t *testing.T) { + store := newTestStore(t) + + if err := store.RenameGroup("nope", "also-nope"); err == nil { + t.Error("renaming a group that does not exist returned no error") + } +} + +// TestSetHostGroupsJoinsAndLeaves is the shape of the dashboard's edit: the +// caller sends the complete set of groups the host should be in, so a group +// left out of the list is a group the host leaves. +func TestSetHostGroupsJoinsAndLeaves(t *testing.T) { + store := newTestStore(t) + mustSaveGroup(t, store, "prod", "web01", "db01") + mustSaveGroup(t, store, "web") + mustSaveGroup(t, store, "staging") + + if err := store.SetHostGroups("web01", []string{"web", "staging"}); err != nil { + t.Fatalf("SetHostGroups: %v", err) + } + + if group := mustGetGroup(t, store, "prod"); slices.Contains(group.Members, "web01") { + t.Errorf("web01 is still in prod, which was not in the list: %v", group.Members) + } + if group := mustGetGroup(t, store, "web"); !slices.Contains(group.Members, "web01") { + t.Errorf("web01 did not join web: %v", group.Members) + } + if group := mustGetGroup(t, store, "staging"); !slices.Contains(group.Members, "web01") { + t.Errorf("web01 did not join staging: %v", group.Members) + } + + // The other member of prod was not collateral damage. + if group := mustGetGroup(t, store, "prod"); !slices.Equal(group.Members, []string{"db01"}) { + t.Errorf("prod members = %v, want [db01]", group.Members) + } +} + +// TestSetHostGroupsToNothing: an empty list is how a host leaves every group, +// and must not be mistaken for "change nothing". +func TestSetHostGroupsToNothing(t *testing.T) { + store := newTestStore(t) + mustSaveGroup(t, store, "prod", "web01") + mustSaveGroup(t, store, "web", "web01") + + if err := store.SetHostGroups("web01", nil); err != nil { + t.Fatalf("SetHostGroups with no groups: %v", err) + } + + for _, name := range []string{"prod", "web"} { + if group := mustGetGroup(t, store, name); len(group.Members) != 0 { + t.Errorf("%s members = %v, want none", name, group.Members) + } + } +} + +// TestSetHostGroupsRefusesUnknownGroup pins that a typo joins nothing rather +// than quietly founding a group of one. Creating a group is a separate, +// deliberate act. +func TestSetHostGroupsRefusesUnknownGroup(t *testing.T) { + store := newTestStore(t) + mustSaveGroup(t, store, "prod") + mustSaveGroup(t, store, "web") + + if err := store.SetHostGroups("web01", []string{"web", "prodd"}); err == nil { + t.Fatal("naming a group that does not exist returned no error") + } + + // Nothing was written: the valid half of the list must not have landed. + if group := mustGetGroup(t, store, "web"); slices.Contains(group.Members, "web01") { + t.Errorf("web01 joined web even though the call failed: %v", group.Members) + } +} + +// TestSetHostGroupsIgnoresCase: the dashboard sends back whatever spelling it +// was given, so matching has to be as forgiving as GetGroup. +func TestSetHostGroupsIgnoresCase(t *testing.T) { + store := newTestStore(t) + mustSaveGroup(t, store, "Prod") + + if err := store.SetHostGroups("web01", []string{"PROD"}); err != nil { + t.Fatalf("SetHostGroups: %v", err) + } + if group := mustGetGroup(t, store, "prod"); !slices.Contains(group.Members, "web01") { + t.Errorf("web01 did not join Prod: %v", group.Members) + } +} + +// TestDeleteSystemRemovesItFromEveryGroup is the one that would otherwise rot +// quietly. A deleted host left behind in its groups is a member nothing can act +// on, padding the member count of every group action from then on. +func TestDeleteSystemRemovesItFromEveryGroup(t *testing.T) { + store := newTestStore(t) + for _, host := range []string{"web01", "web02"} { + if err := store.SaveSystem(host, models.System{Hostname: host}); err != nil { + t.Fatalf("SaveSystem %s: %v", host, err) + } + } + mustSaveGroup(t, store, "prod", "web01", "web02") + mustSaveGroup(t, store, "web", "web01") + mustSaveGroup(t, store, "staging", "web02") + + if err := store.DeleteSystem("web01"); err != nil { + t.Fatalf("DeleteSystem: %v", err) + } + + if group := mustGetGroup(t, store, "prod"); !slices.Equal(group.Members, []string{"web02"}) { + t.Errorf("prod members = %v, want [web02]", group.Members) + } + if group := mustGetGroup(t, store, "web"); len(group.Members) != 0 { + t.Errorf("web members = %v, want none", group.Members) + } + // A group the host was never in is untouched. + if group := mustGetGroup(t, store, "staging"); !slices.Equal(group.Members, []string{"web02"}) { + t.Errorf("staging members = %v, want [web02]", group.Members) + } +} + +// TestDeleteSystemWithNoGroups pins that the purge is harmless on a database +// where no group has ever been created — the groups bucket does not exist yet, +// and deleting a host must not start failing because of it. +func TestDeleteSystemWithNoGroups(t *testing.T) { + store := newTestStore(t) + if err := store.SaveSystem("web01", models.System{Hostname: "web01"}); err != nil { + t.Fatalf("SaveSystem: %v", err) + } + + if err := store.DeleteSystem("web01"); err != nil { + t.Errorf("DeleteSystem on a database with no groups: %v", err) + } +} + +// TestGroupMembersMayNameUnknownHosts: membership is not a foreign key. A group +// built before its hosts are provisioned, or one holding a host that was +// deleted and will be re-added, is a legitimate state. +func TestGroupMembersMayNameUnknownHosts(t *testing.T) { + store := newTestStore(t) + + mustSaveGroup(t, store, "prod", "not-yet-built") + + if group := mustGetGroup(t, store, "prod"); !slices.Equal(group.Members, []string{"not-yet-built"}) { + t.Errorf("Members = %v, want [not-yet-built]", group.Members) + } +} + +func TestValidateGroupName(t *testing.T) { + for _, tc := range []struct { + name string + input string + wantErr bool + }{ + {"ordinary", "prod", false}, + {"spaces inside", "web servers", false}, + {"empty", "", true}, + {"path separator", "prod/web", true}, + {"control character", "prod\nweb", true}, + {"at the limit", string(make([]byte, 0, 64)) + "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", false}, + {"over the limit", "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", true}, + } { + t.Run(tc.name, func(t *testing.T) { + err := models.ValidateGroupName(tc.input) + if (err != nil) != tc.wantErr { + t.Errorf("ValidateGroupName(%q) = %v, wantErr %v", tc.input, err, tc.wantErr) + } + }) + } +} diff --git a/server/storage/storage.go b/server/storage/storage.go index 104ce42..8386ce3 100644 --- a/server/storage/storage.go +++ b/server/storage/storage.go @@ -8,4 +8,14 @@ type Storage interface { GetAllSystems() ([]models.System, error) DeleteSystem(hostname string) error SubscribeToUpdates() <-chan models.System // This should be declared + + // Groups are server-owned: a host neither declares nor sees them. They are + // kept apart from System because a client overwrites its whole System + // record on every check-in. + GetAllGroups() ([]models.Group, error) + GetGroup(name string) (models.Group, error) + SaveGroup(group models.Group) error + DeleteGroup(name string) error + RenameGroup(oldName, newName string) error + SetHostGroups(hostname string, groups []string) error } diff --git a/server/web/server.go b/server/web/server.go index d916e12..417fa75 100644 --- a/server/web/server.go +++ b/server/web/server.go @@ -184,6 +184,21 @@ func StartWebServer(store storage.Storage, port string, version string, updater r.HandleFunc("/api/systems/{hostname}/update/output", api.UpdateOutputHandler(runs)).Methods("GET") r.HandleFunc("/api/features", api.FeaturesHandler(updater, rebooter)).Methods("GET") + // Groups. They need no configuration of their own: a group is server-side + // bookkeeping, so the CRUD routes are always on, and the two group actions + // that can do something to a host are gated by exactly the flags their + // single-host counterparts are. + r.HandleFunc("/api/groups", api.ListGroupsHandler(store)).Methods("GET") + r.HandleFunc("/api/groups", api.CreateGroupHandler(store)).Methods("POST") + r.HandleFunc("/api/groups/{group}", api.GetGroupHandler(store)).Methods("GET") + r.HandleFunc("/api/groups/{group}", api.UpdateGroupHandler(store)).Methods("PUT") + r.HandleFunc("/api/groups/{group}", api.DeleteGroupHandler(store)).Methods("DELETE") + r.HandleFunc("/api/groups/{group}/members/{hostname}", api.RemoveGroupMemberHandler(store)).Methods("DELETE") + r.HandleFunc("/api/groups/{group}/update", api.GroupUpdateHandler(store, updater)).Methods("POST") + r.HandleFunc("/api/groups/{group}/checkin", api.GroupCheckInHandler(store, checkins)).Methods("POST") + r.HandleFunc("/api/groups/{group}/reboot", api.GroupRebootHandler(store, rebooter)).Methods("POST") + r.HandleFunc("/api/systems/{hostname}/groups", api.SetSystemGroupsHandler(store)).Methods("PUT") + // API documentation endpoint r.HandleFunc("/apidoc", apiDocsHandler(version)) diff --git a/server/web/static/apidoc.css b/server/web/static/apidoc.css index 7c2c839..49899f4 100644 --- a/server/web/static/apidoc.css +++ b/server/web/static/apidoc.css @@ -130,6 +130,12 @@ h1 { border: 1px solid rgba(248, 81, 73, 0.3); } +.method.put { + background: rgba(210, 153, 34, 0.15); + color: var(--accent-yellow); + border: 1px solid rgba(210, 153, 34, 0.3); +} + .path { font-family: 'Monaco', 'Menlo', 'Ubuntu Mono', monospace; font-size: 16px; diff --git a/server/web/static/app.js b/server/web/static/app.js index c6b9c12..d10e94e 100644 --- a/server/web/static/app.js +++ b/server/web/static/app.js @@ -308,6 +308,555 @@ document.addEventListener("DOMContentLoaded", () => { }); } + // --------------------------------------------------------------------- + // Groups + // + // A group is server-side bookkeeping: a named set of hostnames the + // dashboard can act on at once. Nothing about it reaches a host, and a + // host does not know which groups it is in. + // + // Membership is deliberately not part of the /api/systems payload. It is + // server-owned and changes by hand a few times a year, while a system + // record is overwritten by its host every few minutes; keeping them apart + // means a check-in can never clobber a group. The dashboard reads + // /api/groups once at load and inverts it here. + // --------------------------------------------------------------------- + + // The sentinel for "hosts in no group at all". A real group name can never + // collide with it, because the API rejects "/" in a name. + const UNGROUPED = "/ungrouped"; + + let groupsData = []; + let groupFilter = ""; + let lastGroupReport = null; + // Which group's reboot checkbox is ticked. Held outside the DOM for the + // same reason armedReboots is: the bar is re-rendered whenever a host + // checks in, and a tick that only lived in the markup would be lost. + const armedGroupReboots = new Set(); + // A host's part-edited group selection, held across the re-renders that an + // unrelated check-in causes. Without this, ticking two boxes and having a + // payload land in between silently discards the first tick. + const pendingGroupEdits = new Map(); + + const GROUP_CHECK_IN_LABEL = "🔄 Check in all"; + const GROUP_UPDATE_LABEL = "⬇️ Run updates"; + const GROUP_REBOOT_LABEL = "⏻ Reboot group"; + + // Ask the server for the groups. Like fetchFeatures, a failure here is not + // fatal: the dashboard renders as a plain ungrouped list. + function fetchGroups() { + return fetch("/api/groups") + .then((response) => { + if (!response.ok) { + throw new Error(`Failed to fetch groups: ${response.status}`); + } + return response.json(); + }) + .then((data) => { + groupsData = Array.isArray(data) ? data : []; + // A group that has gone away must not keep filtering the table + // down to nothing. + if (groupFilter && groupFilter !== UNGROUPED && !findGroup(groupFilter)) { + groupFilter = ""; + } + }) + .catch((error) => { + console.warn("Could not read groups:", error); + groupsData = []; + }); + } + + // Group names are compared without regard to case, exactly as the server + // compares them, so a chip always finds the group it was drawn from. + function sameGroupName(a, b) { + return String(a).toLowerCase() === String(b).toLowerCase(); + } + + function findGroup(name) { + return groupsData.find((g) => g && sameGroupName(g.name, name)) || null; + } + + function groupMembers(name) { + const group = findGroup(name); + return (group && Array.isArray(group.members)) ? group.members : []; + } + + // The reverse index, computed rather than stored. At this scale it is a + // handful of array scans, and a second persisted index would be one more + // thing that can disagree with the first. + function groupsForHost(hostname) { + return groupsData + .filter((g) => g && Array.isArray(g.members) && g.members.includes(hostname)) + .map((g) => g.name); + } + + // Derived groups are computed by the server from what hosts report — + // os:fedora, pkg:rpm, arch:x86_64, state:needs-reboot. They are told apart + // by the flag rather than by the prefix, so the naming stays the server's + // business. + function isDerived(group) { + return !!(group && group.derived); + } + + function manualGroups() { + return groupsData.filter((g) => !isDerived(g)); + } + + function derivedGroups() { + return groupsData.filter(isDerived); + } + + // Only the hand-made groups. Derived membership is not a thing anyone + // chose, so it is not shown as a property of the host in the table, and + // "Ungrouped" has to mean "not filed anywhere by hand" — otherwise it + // would always be empty, since every host is in several derived groups. + function manualGroupsForHost(hostname) { + return manualGroups() + .filter((g) => Array.isArray(g.members) && g.members.includes(hostname)) + .map((g) => g.name); + } + + function filterByGroup(systems) { + if (!groupFilter) return systems; + if (groupFilter === UNGROUPED) { + return systems.filter((s) => s && manualGroupsForHost(s.hostname).length === 0); + } + const members = groupMembers(groupFilter); + return systems.filter((s) => s && members.includes(s.hostname)); + } + + // Members of the selected group that have no row in the table: a host that + // has never checked in, or one whose row was deleted. They are named + // rather than silently dropped, because "7 members, 5 rows" with nothing + // to explain it is the kind of thing that costs an afternoon. + function unknownMembers(name) { + if (!name || name === UNGROUPED) return []; + const known = new Set(systemsData.map((s) => s && s.hostname)); + return groupMembers(name).filter((m) => !known.has(m)); + } + + // The small group labels drawn beside a hostname. They go inside the + // existing cell on purpose: the table's colspan is hardcoded in three + // places and its header is hand-written, so a new column is a much larger + // change than it looks. + function groupChips(hostname) { + // Deliberately not the derived ones: os:fedora and arch:x86_64 beside + // every hostname would restate the OS and Architecture columns on + // every row. The derived groups are useful as filters, not as labels. + const names = manualGroupsForHost(hostname); + if (!names.length) return ''; + return ' ' + names + .map((n) => `${escapeHtml(n)}`) + .join(''); + } + + function renderGroupBar() { + const bar = document.getElementById("group-bar"); + if (!bar) return; + + const ungroupedCount = systemsData.filter((s) => s && manualGroupsForHost(s.hostname).length === 0).length; + + const chip = (group) => { + const active = sameGroupName(groupFilter, group.name); + const derived = isDerived(group); + return ``; + }; + + const chips = [ + ``, + ]; + // Derived first, then the hand-made ones: the derived set is a fixed + // reading of the fleet, while the groups below it are the ones someone + // decided on. + derivedGroups().forEach((group) => chips.push(chip(group))); + manualGroups().forEach((group) => chips.push(chip(group))); + if (ungroupedCount > 0) { + chips.push( + `` + ); + } + chips.push(``); + + bar.innerHTML = chips.join(''); + renderGroupActions(); + } + + function renderGroupActions() { + const el = document.getElementById("group-actions"); + if (!el) return; + + const group = (groupFilter && groupFilter !== UNGROUPED) ? findGroup(groupFilter) : null; + if (!group) { + el.innerHTML = ''; + el.style.display = 'none'; + renderGroupReport(); + return; + } + + const members = group.members || []; + const ghosts = unknownMembers(group.name); + const ghostNote = ghosts.length + ? ` ${ghosts.length} not currently known` + : ''; + + const updateBtn = features.remote_updates + ? `` + : ''; + + // The same arm-checkbox the per-host reboot uses, and it looks + // identical on purpose: a group reboot should feel like the control + // the operator already knows rather than a new one to learn. + const armed = armedGroupReboots.has(group.name.toLowerCase()); + const rebootControls = features.remote_reboot + ? ` + + + ` + : ''; + + // A derived group has nothing to rename or delete: it exists for as + // long as a host matches it and not a moment longer. Offering the + // controls and refusing them would be worse than not offering them. + const derived = isDerived(group); + const editButtons = derived + ? 'derived from host facts' + : ` + `; + + el.style.display = ''; + el.innerHTML = ` +
+ ${derived ? '◆' : ''}${escapeHtml(group.name)} + ${members.length} ${members.length === 1 ? 'host' : 'hosts'}${ghostNote} +
+
+ + ${updateBtn} + ${rebootControls} + ${editButtons} +
+ `; + renderGroupReport(); + } + + // The per-host outcomes of the last group action. + // + // Not an alert(): a reboot that skipped two hosts is something the operator + // needs to keep reading while they go and look, and a modal is gone the + // moment it is dismissed. The existing alerts are for a single host, where + // there is exactly one sentence to say. + function renderGroupReport() { + const el = document.getElementById("group-action-report"); + if (!el) return; + + if (!lastGroupReport) { + el.innerHTML = ''; + el.style.display = 'none'; + return; + } + + const r = lastGroupReport; + const verb = r.action === 'checkin' ? 'Check-in' : r.action === 'update' ? 'Update' : 'Reboot'; + const parts = []; + if (r.accepted) parts.push(`${r.accepted} accepted`); + if (r.skipped) parts.push(`${r.skipped} skipped`); + if (r.failed) parts.push(`${r.failed} failed`); + const summary = parts.length ? parts.join(' · ') : 'nothing to do'; + + const lines = (r.results || []).map((res) => { + const ghost = res.code === 'unknown_host' ? ' ghost' : ''; + return `
  • + ${escapeHtml(res.outcome)} + ${escapeHtml(res.hostname)} + ${res.reason ? `${escapeHtml(res.reason)}` : ''} +
  • `; + }).join(''); + + el.style.display = ''; + el.innerHTML = ` +
    + ${verb} on ${escapeHtml(r.group)} — ${escapeHtml(summary)} + +
    + ${lines ? `
      ${lines}
    ` : ''} + `; + } + + function handleGroupAction(action) { + const group = findGroup(groupFilter); + if (!group) return; + + if (action === 'reboot' && !armedGroupReboots.has(group.name.toLowerCase())) { + // The button should have been disabled; treat it as if it were. + return; + } + + document.querySelectorAll('#group-actions .group-action-btn').forEach((b) => { b.disabled = true; }); + + fetch(`/api/groups/${encodeURIComponent(group.name)}/${action}`, { method: 'POST' }) + .then((response) => response.json().catch(() => ({})).then((body) => ({ response, body }))) + .then(({ response, body }) => { + if (!response.ok) { + throw new Error(body.error || `Group ${action} failed: ${response.status}`); + } + lastGroupReport = body; + armedGroupReboots.delete(group.name.toLowerCase()); + + // Seed the per-host pending state for everything that was + // accepted, so each row shows the same spinner, timeout and + // WebSocket-driven release it would for a single-host action. + // A group action is N button presses and should look like it. + (body.results || []).forEach((res) => { + if (res.outcome !== 'accepted') return; + const system = systemsData.find((s) => s && s.hostname === res.hostname); + if (action === 'checkin') { + markCheckInPending(res.hostname, (system && system.updates_checked_at) || ''); + } else if (action === 'reboot') { + markRebootPending(res.hostname, (system && system.last_reboot && system.last_reboot.id) || ''); + } + }); + + renderSystems(Array.from(systemsData)); + }) + .catch((error) => { + console.error(`Group ${action} failed:`, error); + lastGroupReport = null; + renderGroupActions(); + alert(`Could not run ${action} on ${group.name}:\n\n${error.message}`); + }); + } + + function handleNewGroup() { + const name = prompt('Name for the new group:'); + if (name === null || !name.trim()) return; + + fetch('/api/groups', { + method: 'POST', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify({ name: name.trim() }), + }) + .then((response) => response.json().catch(() => ({})).then((body) => ({ response, body }))) + .then(({ response, body }) => { + if (!response.ok) { + throw new Error(body.error || `Could not create the group: ${response.status}`); + } + return fetchGroups(); + }) + .then(() => renderSystems(Array.from(systemsData))) + .catch((error) => alert(`Could not create the group:\n\n${error.message}`)); + } + + function handleRenameGroup() { + const group = findGroup(groupFilter); + if (!group) return; + const name = prompt(`Rename "${group.name}" to:`, group.name); + if (name === null || !name.trim() || name.trim() === group.name) return; + + fetch(`/api/groups/${encodeURIComponent(group.name)}`, { + method: 'PUT', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify({ name: name.trim() }), + }) + .then((response) => response.json().catch(() => ({})).then((body) => ({ response, body }))) + .then(({ response, body }) => { + if (!response.ok) { + throw new Error(body.error || `Rename failed: ${response.status}`); + } + groupFilter = body.name || name.trim(); + return fetchGroups(); + }) + .then(() => renderSystems(Array.from(systemsData))) + .catch((error) => alert(`Could not rename the group:\n\n${error.message}`)); + } + + function handleDeleteGroup() { + const group = findGroup(groupFilter); + if (!group) return; + const count = (group.members || []).length; + // A group is a label. Say so, so nobody reads this as deleting hosts. + if (!confirm(`Delete the group "${group.name}"?\n\n` + + `${count} ${count === 1 ? 'host is' : 'hosts are'} in it. They are not deleted, only the group is.`)) { + return; + } + + fetch(`/api/groups/${encodeURIComponent(group.name)}`, { method: 'DELETE' }) + .then((response) => response.json().catch(() => ({})).then((body) => ({ response, body }))) + .then(({ response, body }) => { + if (!response.ok) { + throw new Error(body.error || `Delete failed: ${response.status}`); + } + groupFilter = ''; + lastGroupReport = null; + return fetchGroups(); + }) + .then(() => renderSystems(Array.from(systemsData))) + .catch((error) => alert(`Could not delete the group:\n\n${error.message}`)); + } + + // The per-host editor in the expanded row. Groups are a property of the + // host rather than an action on it, so this sits with the system + // information and not in the actions footer. + function groupEditorHTML(hostname) { + const manual = manualGroups(); + const selected = pendingGroupEdits.get(hostname) || new Set(manualGroupsForHost(hostname)); + const dirty = pendingGroupEdits.has(hostname); + + // The derived memberships are shown but not offered as checkboxes: a + // host joins os:fedora by being a Fedora box, and a tickbox that + // reverted on the next read would be worse than no tickbox. + const derivedNames = derivedGroups() + .filter((g) => Array.isArray(g.members) && g.members.includes(hostname)) + .map((g) => g.name); + const derivedHTML = derivedNames.length + ? `
    + Derived + ${derivedNames.map((n) => `◆${escapeHtml(n)}`).join('')} +
    ` + : ''; + + const boxes = manual.length + ? manual.map((group) => ` + `).join('') + : 'No groups of your own yet. Create one from the bar above the table.'; + + return `
    +

    Groups

    +
    ${boxes}
    + ${manual.length ? `
    + + ${dirty ? 'unsaved' : ''} +
    ` : ''} + ${derivedHTML} +
    `; + } + + function handleGroupMemberToggle(event) { + const checkbox = event.currentTarget; + const hostname = checkbox.dataset.hostname; + const group = checkbox.dataset.group; + if (!hostname || !group) return; + + const selected = pendingGroupEdits.get(hostname) || new Set(manualGroupsForHost(hostname)); + if (checkbox.checked) { + selected.add(group); + } else { + selected.delete(group); + } + pendingGroupEdits.set(hostname, selected); + + const editor = checkbox.closest('.group-editor'); + const save = editor && editor.querySelector('.save-groups-btn'); + if (save) save.disabled = false; + } + + function handleSaveGroups(event) { + const button = event.currentTarget; + const hostname = button.dataset.hostname; + if (!hostname) return; + + const selected = pendingGroupEdits.get(hostname); + if (!selected) return; + + button.disabled = true; + fetch(`/api/systems/${encodeURIComponent(hostname)}/groups`, { + method: 'PUT', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify({ groups: Array.from(selected) }), + }) + .then((response) => response.json().catch(() => ({})).then((body) => ({ response, body }))) + .then(({ response, body }) => { + if (!response.ok) { + throw new Error(body.error || `Could not save groups: ${response.status}`); + } + pendingGroupEdits.delete(hostname); + return fetchGroups(); + }) + .then(() => renderSystems(Array.from(systemsData))) + .catch((error) => { + button.disabled = false; + console.error(`Failed to save groups for ${hostname}:`, error); + alert(`Could not save groups for ${hostname}:\n\n${error.message}`); + }); + } + + // A collapsed row drops its half-finished edit, for the same reason it + // drops a ticked reboot confirmation: what is no longer on screen should + // not still be pending. + function discardGroupEdit(hostname) { + pendingGroupEdits.delete(hostname); + } + + // Delegated listeners on the bar, the action strip and the report. They + // are attached once at boot rather than on every render, because all three + // are rebuilt whenever any host checks in. + function initGroupControls() { + const bar = document.getElementById("group-bar"); + if (bar) { + bar.addEventListener('click', (event) => { + if (event.target.closest('#new-group-btn')) { + handleNewGroup(); + return; + } + const chip = event.target.closest('.group-chip'); + if (!chip || chip.dataset.filter === undefined) return; + groupFilter = chip.dataset.filter; + lastGroupReport = null; + armedGroupReboots.clear(); + renderSystems(Array.from(systemsData)); + }); + } + + const actions = document.getElementById("group-actions"); + if (actions) { + actions.addEventListener('click', (event) => { + const edit = event.target.closest('[data-group-edit]'); + if (edit) { + if (edit.dataset.groupEdit === 'rename') handleRenameGroup(); + else handleDeleteGroup(); + return; + } + const btn = event.target.closest('[data-group-action]'); + if (btn && !btn.disabled) handleGroupAction(btn.dataset.groupAction); + }); + actions.addEventListener('change', (event) => { + const checkbox = event.target.closest('.group-reboot-arm-checkbox'); + if (!checkbox) return; + const group = findGroup(groupFilter); + if (!group) return; + if (checkbox.checked) { + armedGroupReboots.add(group.name.toLowerCase()); + } else { + armedGroupReboots.delete(group.name.toLowerCase()); + } + const button = actions.querySelector('[data-group-action="reboot"]'); + if (button) button.disabled = !checkbox.checked; + }); + } + + const report = document.getElementById("group-action-report"); + if (report) { + report.addEventListener('click', (event) => { + if (!event.target.closest('.group-report-dismiss')) return; + lastGroupReport = null; + renderGroupReport(); + }); + } + } + // Fetch and render systems list function fetchSystems() { fetch("/api/systems") @@ -731,6 +1280,34 @@ document.addEventListener("DOMContentLoaded", () => { systems = []; } + // The chip bar counts the whole fleet, so it is drawn from the + // unfiltered list before the filter is applied below. + renderGroupBar(); + + // One row per host, always: the table is filtered rather than divided + // into sections. A host in several groups would otherwise be drawn + // several times, and every document.querySelector('[data-hostname=...]') + // in this file would then drive only the first copy. + const unfilteredCount = systems.length; + systems = filterByGroup(systems); + + // An empty group is a different thing from an empty fleet, and saying + // "no systems have checked in yet" under a group whose members simply + // have no rows yet would be a lie. + if (systems.length === 0 && unfilteredCount > 0) { + const label = groupFilter === UNGROUPED ? 'Ungrouped' : groupFilter; + systemsTable.innerHTML = ` + + +
    \u{1F50D}
    +
    No hosts to show in ${escapeHtml(label)}
    +
    Its members may not have checked in yet
    + + + `; + return; + } + // Show message if no systems exist yet if (systems.length === 0) { systemsTable.innerHTML = ` @@ -788,7 +1365,7 @@ document.addEventListener("DOMContentLoaded", () => { return ` ▶ - ${tailnetIndicator(system, showTailnetSlot)}${escapeHtml(system.hostname)}${rebootIndicator(system)} + ${tailnetIndicator(system, showTailnetSlot)}${escapeHtml(system.hostname)}${rebootIndicator(system)}${groupChips(system.hostname)} ${getOSIcon(system.os)} ${escapeHtml(system.os || '')} ${escapeHtml(system.os_version || '')} ${escapeHtml(system.architecture || '')} ${escapeHtml(system.ip || '')} @@ -1491,14 +2068,14 @@ document.addEventListener("DOMContentLoaded", () => {
      ${warnings.map((w) => `
    • ${escapeHtml(w)}
    • `).join('')}
    ` : ''; - const systemInfoHTML = infoRows.length + const systemInfoHTML = (infoRows.length ? `

    System information

    ${infoRows.map(([label, value]) => `
    ${label}
    ${value}
    `).join('')}
    ` - : ''; + : '') + groupEditorHTML(hostname); if (data.pending_updates && data.pending_updates.length > 0) { const updatesList = data.pending_updates @@ -1580,6 +2157,15 @@ document.addEventListener("DOMContentLoaded", () => { rebootBtn.addEventListener('click', handleReboot); } + // Attach the group editor + detailsContent.querySelectorAll('.group-member-checkbox').forEach((box) => { + box.addEventListener('change', handleGroupMemberToggle); + }); + const saveGroupsBtn = detailsContent.querySelector('.save-groups-btn'); + if (saveGroupsBtn) { + saveGroupsBtn.addEventListener('click', handleSaveGroups); + } + restoreOutputView(hostname, detailsContent, data.last_update_run); seedLiveOutput(hostname, data.last_update_run); } @@ -1641,6 +2227,7 @@ document.addEventListener("DOMContentLoaded", () => { chevron.textContent = "▶"; expandedSystems.delete(hostname); // Remove from expanded set disarmReboot(hostname); // A closed row is not a confirmed one + discardGroupEdit(hostname); // nor does it hold an unsaved edit } }); @@ -1708,7 +2295,8 @@ document.addEventListener("DOMContentLoaded", () => { // Initial fetch and WebSocket connection. Features first, so the first // render of an expanded row already knows whether to offer the update // button; the fetch is not allowed to hold up the systems list for long. - fetchFeatures().finally(fetchSystems); + initGroupControls(); + Promise.all([fetchFeatures(), fetchGroups()]).finally(fetchSystems); initWebSocket(); // Set up periodic update of relative timestamps (every 30 seconds) diff --git a/server/web/static/styles.css b/server/web/static/styles.css index f2157ab..4b3a6ec 100644 --- a/server/web/static/styles.css +++ b/server/web/static/styles.css @@ -1027,3 +1027,375 @@ th.sortable.descending::after { border-bottom: none; } + +/* ------------------------------------------------------------------ */ +/* Groups */ +/* */ +/* Everything here is built from the theme's custom properties, so the */ +/* light theme needs no overrides of its own: the variables are */ +/* already redefined under [data-theme="light"] above. The two */ +/* literal rgba() colours below are the outcome pills, which need a */ +/* translucent wash over whatever the background happens to be. */ +/* ------------------------------------------------------------------ */ + +.group-bar { + display: flex; + flex-wrap: wrap; + align-items: center; + gap: 8px; + max-width: 1400px; + margin: 0 auto 12px auto; + padding: 0 20px; +} + +.group-chip { + display: inline-flex; + align-items: center; + gap: 6px; + padding: 5px 12px; + font-size: 13px; + font-family: inherit; + border-radius: 14px; + border: 1px solid var(--border-color); + background: var(--bg-secondary); + color: var(--text-secondary); + cursor: pointer; + transition: all 0.15s ease; +} + +.group-chip:hover { + background: var(--bg-hover); + color: var(--text-primary); + border-color: var(--accent-blue); +} + +.group-chip.active { + background: var(--bg-tertiary); + border-color: var(--accent-blue); + color: var(--accent-blue); + font-weight: 500; +} + +.group-count { + font-size: 11px; + color: var(--text-muted); + background: var(--bg-primary); + border-radius: 9px; + padding: 1px 6px; +} + +.group-chip.active .group-count { + color: var(--accent-blue); +} + +.group-chip.new-group-btn { + border-style: dashed; + color: var(--text-muted); +} + +/* The strip of things that act on the whole group. It is deliberately + set apart from the table: these buttons have a blast radius the + per-row ones do not. */ +.group-actions { + display: flex; + flex-wrap: wrap; + align-items: center; + justify-content: space-between; + gap: 12px; + max-width: 1400px; + margin: 0 auto 12px auto; + padding: 10px 20px; + background: var(--bg-secondary); + border: 1px solid var(--border-color); + border-radius: 6px; +} + +.group-action-summary { + display: inline-flex; + align-items: baseline; + gap: 10px; + color: var(--text-primary); + font-size: 14px; +} + +.group-member-count, +.group-ghost-note { + font-size: 12px; + color: var(--text-secondary); +} + +/* Members the server has no row for. Named rather than silently + dropped: "7 members, 5 rows" with no explanation costs an afternoon. */ +.group-ghost-note { + color: var(--accent-yellow); + cursor: help; +} + +.group-action-buttons { + display: inline-flex; + flex-wrap: wrap; + align-items: center; + gap: 8px; +} + +.group-action-btn, +.group-edit-btn { + font-family: inherit; + font-size: 13px; + padding: 6px 12px; + border-radius: 6px; + border: 1px solid var(--border-color); + background: var(--bg-tertiary); + color: var(--text-primary); + cursor: pointer; + transition: all 0.15s ease; +} + +.group-action-btn:hover:not(:disabled), +.group-edit-btn:hover:not(:disabled) { + background: var(--bg-hover); + border-color: var(--accent-blue); +} + +.group-action-btn:disabled, +.group-edit-btn:disabled { + color: var(--text-muted); + cursor: not-allowed; + opacity: 0.5; +} + +.group-edit-btn { + font-size: 12px; + color: var(--text-secondary); +} + +/* The report of what a group action actually did. It stays on screen + until dismissed, because a reboot that skipped two hosts is something + to keep reading while you go and look at them. */ +.group-action-report { + max-width: 1400px; + margin: 0 auto 12px auto; + padding: 10px 20px; + background: var(--bg-secondary); + border: 1px solid var(--border-color); + border-radius: 6px; + font-size: 13px; +} + +.group-report-head { + display: flex; + align-items: center; + justify-content: space-between; + gap: 12px; + color: var(--text-primary); +} + +.group-report-dismiss { + background: none; + border: none; + color: var(--text-muted); + font-size: 14px; + cursor: pointer; + padding: 0 4px; +} + +.group-report-dismiss:hover { + color: var(--text-primary); +} + +.group-report-list { + list-style: none; + margin: 10px 0 0 0; + padding: 0; + display: flex; + flex-direction: column; + gap: 5px; +} + +.group-report-list li { + display: flex; + align-items: baseline; + gap: 8px; +} + +.outcome-pill { + flex: 0 0 auto; + min-width: 88px; + text-align: center; + font-size: 11px; + text-transform: uppercase; + letter-spacing: 0.4px; + padding: 2px 8px; + border-radius: 10px; + border: 1px solid transparent; +} + +.outcome-pill.accepted { + background: rgba(63, 185, 80, 0.15); + border-color: var(--accent-green); + color: var(--accent-green); +} + +.outcome-pill.skipped { + background: var(--bg-tertiary); + border-color: var(--border-color); + color: var(--text-secondary); +} + +.outcome-pill.refused, +.outcome-pill.unreachable, +.outcome-pill.failed { + background: rgba(248, 81, 73, 0.15); + border-color: var(--accent-red); + color: var(--accent-red); +} + +.outcome-host { + font-weight: 500; + color: var(--text-primary); +} + +.outcome-host.ghost { + color: var(--text-muted); + font-style: italic; +} + +.outcome-reason { + color: var(--text-secondary); +} + +/* The group labels beside a hostname in the table. They sit inside the + hostname cell rather than in a column of their own: the table's + colspan is hardcoded in three places and its header is hand-written. */ +.host-group-chip { + display: inline-block; + margin-left: 5px; + padding: 1px 7px; + font-size: 11px; + border-radius: 9px; + background: var(--bg-tertiary); + border: 1px solid var(--border-color); + color: var(--text-secondary); + vertical-align: middle; +} + +/* The per-host editor in the expanded row. Groups are a property of the + host, so this sits with the system information, not in the actions. */ +.group-editor { + margin-top: 16px; +} + +.group-editor h4 { + margin: 0 0 8px 0; + font-size: 13px; + color: var(--text-secondary); + text-transform: uppercase; + letter-spacing: 0.5px; +} + +.group-editor-options { + display: flex; + flex-wrap: wrap; + gap: 12px; +} + +.group-editor-option { + display: inline-flex; + align-items: center; + gap: 5px; + font-size: 13px; + color: var(--text-primary); + cursor: pointer; + user-select: none; +} + +.group-editor-option input { + margin: 0; + accent-color: var(--accent-blue); + cursor: pointer; +} + +.group-editor-empty { + font-size: 13px; + color: var(--text-muted); +} + +.group-editor-actions { + display: flex; + align-items: center; + gap: 10px; + margin-top: 10px; +} + +.save-groups-btn { + font-family: inherit; + font-size: 13px; + padding: 5px 12px; + border-radius: 6px; + border: 1px solid var(--border-color); + background: var(--bg-tertiary); + color: var(--text-primary); + cursor: pointer; +} + +.save-groups-btn:hover:not(:disabled) { + background: var(--bg-hover); + border-color: var(--accent-blue); +} + +.save-groups-btn:disabled { + color: var(--text-muted); + cursor: not-allowed; + opacity: 0.5; +} + +.group-editor-dirty { + font-size: 12px; + color: var(--accent-yellow); +} + +/* Derived groups: computed by the server from what the hosts report, rather + than curated. Marked with a diamond and a dimmer, dashed outline so the bar + reads as two kinds of thing — facts about the fleet, and decisions about + it — without needing a legend. */ +.group-chip.derived { + border-style: dashed; + color: var(--text-muted); +} + +.group-chip.derived:hover { + color: var(--text-secondary); +} + +.group-chip.derived.active { + border-style: solid; + color: var(--accent-blue); +} + +.group-derived-note { + font-size: 12px; + font-style: italic; + color: var(--text-muted); + cursor: help; +} + +.host-group-chip.derived { + border-style: dashed; + color: var(--text-muted); +} + +.group-editor-derived { + display: flex; + flex-wrap: wrap; + align-items: center; + gap: 6px; + margin-top: 10px; +} + +.group-editor-derived-label { + font-size: 11px; + text-transform: uppercase; + letter-spacing: 0.5px; + color: var(--text-muted); + margin-right: 2px; +} diff --git a/server/web/templates/apidoc.html b/server/web/templates/apidoc.html index 97af6ac..b654696 100644 --- a/server/web/templates/apidoc.html +++ b/server/web/templates/apidoc.html @@ -410,6 +410,414 @@

    API Documentation

    + +
    +
    + GET + /api/groups +
    +
    + Lists every group with its members. A group is a named set of hostnames kept by the server so the + dashboard can act on several hosts at once. +

    + Two kinds come back from this one route. Stored groups are the ones you create and + fill. Derived groups are computed from what the hosts already report, are namespaced + by a prefix — os:, pkg:, arch:, state: — and carry + "derived": true. They are recomputed on every read, so they are always current and there + is never a stale one to clean up. +

    + Groups are deliberately not part of the /api/systems payload. They are server-owned and + change by hand, while a system record is overwritten by its host every few minutes; keeping them + apart means a check-in can never clobber a group. The dashboard reads this route and inverts it to + work out which groups a host is in. +

    + Nothing has to be configured for groups to work, on the server or on a host, and a host is never told + which groups it is in. +
    + +
    +
    Response
    +
    + 200 OK + Every group, ordered by name. An empty server returns []. +
    +
    [ + { + "name": "prod", + "members": ["db01", "web01", "web02"] + } +]
    +
    + +
    +
    Fields
    +
    +
    name
    +
    The group's name as it was typed. Names are compared without regard to + case, so prod and Prod are the same group and cannot both exist.
    +
    +
    +
    members
    +
    Hostnames, sorted and de-duplicated. A member may name a host the server + has never seen — one that has not checked in yet, or whose row was deleted. That is a + legitimate state, not an error; a group action skips such a member.
    +
    +
    +
    derived
    +
    Present and true only on a computed group. A derived group + exists for exactly as long as something matches it, and cannot be created, renamed, deleted, + or have its membership edited — every such route answers 409. The derived + groups are os:<distro>, pkg:<family> (rpm, deb, pacman, + nix, brew, portage, apk, xbps), arch:<arch>, + state:needs-reboot and state:has-updates.
    +
    +
    +
    + +
    +
    + POST + /api/groups +
    +
    + Creates a group. Creating one is deliberately its own act rather than something that happens the + first time a name is typed somewhere else: a typo that founds a group of one is far harder to notice + than a typo that is refused. +
    + +
    +
    Request Body
    +
    { + "name": "prod", + "members": ["web01"] +}
    +
    +
    name
    +
    Required. Trimmed; at most 64 characters; may not be empty, contain + / or :, or contain control characters. : is reserved + for the derived namespace, so a stored group can never shadow os:fedora and + make a group action ambiguous about which hosts it meant.
    +
    +
    +
    members
    +
    Optional. Defaults to an empty group.
    +
    +
    + +
    +
    Response
    +
    + 201 Created + The new group +
    +
    + 400 Bad Request + The name is empty, too long, or contains a forbidden character +
    +
    + 409 Conflict + A group of that name already exists, ignoring case +
    +
    +
    + +
    +
    + GET + /api/groups/{group} +
    +
    + Returns one group. The name in the path is matched without regard to case. +
    + +
    +
    Response
    +
    + 200 OK + The group +
    +
    + 404 Not Found + No such group +
    +
    +
    + +
    +
    + PUT + /api/groups/{group} +
    +
    + Renames a group, replaces its membership, or both. Derived groups answer 409: their + members come from what the hosts report, so there is nothing here to change. +

    + Both fields are optional, and leaving one out is different from sending it empty: omitting + members leaves membership alone, so a rename cannot accidentally empty a group, while + "members": [] empties it on purpose. +
    + +
    +
    Request Body
    +
    { + "name": "production", + "members": ["db01", "web01"] +}
    +
    + +
    +
    Response
    +
    + 200 OK + The group as it now stands +
    +
    + 400 Bad Request + The new name is not a valid group name +
    +
    + 404 Not Found + No such group +
    +
    + 409 Conflict + The new name is already taken by another group +
    +
    +
    + +
    +
    + DELETE + /api/groups/{group} +
    +
    + Deletes a group. The hosts in it are untouched — a group is a label, and dropping the label is not + dropping the machines. +

    + Derived groups answer 409. One disappears on its own as soon as nothing matches it, + which is the only way it can go. +
    + +
    +
    Response
    +
    + 200 OK + The group was deleted +
    +
    { + "status": "success", + "message": "Group deleted successfully", + "group": "prod" +}
    +
    + 404 Not Found + No such group +
    +
    +
    + +
    +
    + DELETE + /api/groups/{group}/members/{hostname} +
    +
    + Removes one host from one group. +

    + This is how a member the server no longer knows about gets evicted — a machine that was + decommissioned, or a name that was mistyped into the list. Such a host has no row in the dashboard, + so the per-host editor cannot reach it. +
    + +
    +
    Response
    +
    + 200 OK + The group as it now stands +
    +
    + 404 Not Found + No such group, or that host is not a member of it +
    +
    +
    + +
    +
    + PUT + /api/systems/{hostname}/groups +
    +
    + Replaces one host's memberships with exactly the list given, so a group left out of the list is a + group the host leaves. This is what the group editor in the dashboard's expanded row sends. +

    + Every group named must already exist. A name that is not a group is refused rather than created, so a + typo joins nothing instead of quietly founding a group of one — and nothing in the list is applied + when any of it is refused. Naming a derived group answers 409: a host joins + os:fedora by being a Fedora box, not by being put there. +

    + The host need not have checked in. Putting a machine into its groups before it is built is a + reasonable thing to want, and membership already outlives the system record. +
    + +
    +
    Request Body
    +
    { + "groups": ["prod", "web"] +}
    +
    + +
    +
    Response
    +
    + 200 OK + The groups the host is now in +
    +
    { + "hostname": "web01", + "groups": ["prod", "web"] +}
    +
    + 400 Bad Request + The body is not valid JSON, or it names a group that does not exist +
    +
    +
    + +
    +
    + POST + /api/groups/{group}/checkin +
    +
    + Asks every member of the group to check in. The members are asked all at once rather than in + sequence, so the whole call takes about as long as the slowest single host. +

    + The status is 200 even when members were skipped or could not be reached. The body is the + thing to read. The request was "fan this out and tell me what happened", and it did exactly + that; a status code cannot summarise a dozen different answers, so it does not try. See + /api/groups/{group}/update below for the response shape, which all three group actions + share. +

    + Like the single-host route, nothing configured gates this, and it answers + 503 rather than 403 when it is unavailable: a check-in has no feature flag, + so its unavailability is an infrastructure fact rather than a policy one. +
    +
    + +
    +
    + POST + /api/groups/{group}/update +
    +
    + Asks every member of the group to install its pending packages, all at once. +

    + Works on a derived group as readily as a stored one — /api/groups/pkg:rpm/update is the + "update everything rpm-based" case. A derived group's members are resolved when the action runs + rather than when the page was drawn, which is what makes + /api/groups/state:needs-reboot/reboot mean what is true at the moment you call it. A + derived group with no members does not exist, so it answers 404 rather than running + against nobody. +

    + Gated exactly like the single-host route: remote_updates on the server decides whether + the route works at all, and each host's own allow_remote_updates decides whether it is a + candidate. A member that has not opted in is skipped, not an error — one machine declining + should not stop the other eleven being patched. +

    + The status is 200 whenever the group exists and the feature is on, however the individual + hosts answered. +
    + +
    +
    Response
    +
    + 200 OK + The fan-out ran; every member has a result +
    +
    { + "group": "prod", + "action": "update", + "requested": 4, + "accepted": 1, + "skipped": 2, + "failed": 1, + "results": [ + { "hostname": "web01", "outcome": "accepted", "id": "9f2c1b0a4d5e6f70", + "command": "/usr/libexec/muc/upd" }, + { "hostname": "db01", "outcome": "skipped", "code": "not_opted_in", + "reason": "This host has not opted into remote updates (set allow_remote_updates: true in its client config)" }, + { "hostname": "ghost01", "outcome": "skipped", "code": "unknown_host", + "reason": "System not found" }, + { "hostname": "web02", "outcome": "unreachable", "code": "not_listening", + "reason": "No response from web02: it is offline, or its client is no longer accepting update commands" } + ] +}
    +
    + 403 Forbidden + Remote updates are disabled on this server +
    +
    + 404 Not Found + No such group +
    +
    + +
    +
    Result Fields
    +
    +
    requested / accepted / skipped / failed
    +
    Counts for the summary line. failed aggregates the + refused, unreachable and failed outcomes; the result + list keeps them apart.
    +
    +
    +
    results[].outcome
    +
    One of accepted, skipped, + refused, unreachable, failed. The distinction between + skipped and refused is worth reading: skipped is what the server knew + before it sent anything, refused is what the host said back.
    +
    +
    +
    results[].code
    +
    The machine-readable reason, so a client need not match on prose: + unknown_host, not_opted_in, no_reboot_pending, + host_refused, not_listening, request_failed.
    +
    +
    +
    results[].reason
    +
    The same sentence the single-host route would have returned as its error + body, so both paths explain themselves identically.
    +
    +
    +
    results[].id
    +
    The host's own request id, present only when it accepted.
    +
    +
    +
    + +
    +
    + POST + /api/groups/{group}/reboot +
    +
    + Reboots every member of the group that reports a pending reboot, all at once. +

    + Gated by remote_reboot on the server and allow_remote_reboot on each host, + and like the single-host route it insists the host's last check-in reported a pending reboot. + Members with nothing pending are skipped with no_reboot_pending, which is also what + makes the control safe to press after a group update: it reboots what needs it and leaves the rest. +

    + The response is the shape documented under /api/groups/{group}/update. +

    + This is the single most consequential thing this server can do, and nothing authenticates the + caller. Every group action is logged at WARN with the group, the action, the requesting + address and the counts. +
    +
    +
    GET diff --git a/server/web/templates/index.html b/server/web/templates/index.html index 41c8038..a33055d 100644 --- a/server/web/templates/index.html +++ b/server/web/templates/index.html @@ -15,9 +15,9 @@ } })(); - + - +