Skip to content

Build guided setup and the LLooM browser experience - #33

Merged
data-angel merged 21 commits into
mainfrom
feat/lloom-overhaul
Sep 23, 2026
Merged

data-angel merged 21 commits into
mainfrom
feat/lloom-overhaul

Conversation

@data-angel

@data-angel data-angel commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

LLooM now opens guided browser setup on the first interactive run and gives installed gateways Live, Models, Machines, Clients, and Settings views. Observed requests animate the topology. The Models view shows proportional blocks for backend memory, system/other apps, and free space, with protected headroom inside the free block. Pointer, keyboard-focus, and selection previews animate a model’s expected memory impact on its own machine. Normal use starts models automatically; optional manual preparation and memory release live under Options & details.

Setup and model installation use reviewed, expiring plans and observable jobs. Backend and download commands remain fixed through apply and retry. Custom Hugging Face imports require immutable commit links and staged acquisition verification. Setup publishes configuration exclusively and verifies a model response through the gateway before reporting readiness. CLI, offline, and scripted paths remain available.

Admission now counts live host memory even when policy specifies only a reserve or absolute budget. Newly launched local backends are monitored through loading and warmup. A threshold breach or unavailable telemetry aborts the load and cleans up only its owned process tree or Docker container. An independent supervisor process checks memory before spawning, and the manager waits for that check before accepting health. Forced starts cannot bypass the guard. Explicit runtimePolicy.memorySafety.mode: "yolo" disables memory admission and the load guard, with a persistent dashboard warning. The guard polls in userspace; it is not a kernel allocation quota. Automatic retries remain blocked until a manual retry or gateway restart; persistent suspension survives restarts.

Memory attribution uses bounded passive process observations with a five-second cache and overlapping process-tree deduplication. macOS reads physical footprint, including charged Metal/compressed memory, with RSS fallback. Known local model-cache observations distinguish a healthy service from a loaded model. Estimates are marked, remote safety policies stay node-specific, and routing/admission status skips local dashboard sampling.

The streaming stall watchdog rechecks elapsed time if its timer wakes early, preventing a premature event from being ignored by stall classification.

Readiness preferences persist without restarting active models. Changes wait for admission, queued restores recheck intent, and pending hard pins protect models from eviction. Browser management rejects foreign origins and rebound loopback hosts; remote writes require an admin credential.

Validation:

  • Full npm test passed locally on macOS / Node 22 with the memory map and Mac footprint sampler. After review fixes, all 36 memory-map/sampler tests and 33 maintenance tests passed, plus cluster, runtime policy, dashboard-status, and the 29-test UX suite. GitHub CI passed on Node 20 and Node 22 at 2609b87, including full tests, dependency audit, interchange, and package installation. The memory safety and supervisor suites passed all 16 cases, including pressure during load and warmup, process-tree cleanup, immutable Docker ownership, independent cutoff, stale health, and the supervisor startup race.
  • Check, lint, formatting, interchange, and packaged-install validation passed. Lint retains one pre-existing unused-variable warning in research-worker tests.
  • The clean package from 2609b87 was installed on the Mac and its installed sources were verified against the running service (artifact SHA-256 ea526b1df64f59c83e18913267cf08b9f131b7386ad77bf8e46f24a82db0f64e). Configuration was preserved; enforcement reports a 12 GiB reserve, 90% ceiling, and 250 ms polling. An isolated test of the installed HTTP API returned 503 runtime_memory_safety_abort before spawning its tiny fixture, despite force: true and disabled predictive policy.
  • Visually checked the real installed dashboard with all 21 visible models: the embedding process tree reports 5.7 GiB from physical-footprint accounting, and focusing FLUX animates its roughly 32 GiB incoming estimate without starting it. Desktop and 390px checks covered memory blocks, keyboard focus, persistent selection, mobile controls, unknown and over-budget models, and node-specific previews. A read-only two-machine fixture exercised remote reserve behavior.
  • Bonsai remains persistently suspended following the reported memory incident. No Spark deployment was performed.

A signed public installer and automatic nearby discovery/pairing remain separate work. Configured federation is available; independent nodes do not pair automatically in this change.

dependabot Bot and others added 21 commits September 7, 2026 21:07
Bumps [undici](https://github.com/nodejs/undici) from 7.29.0 to 8.10.2.
- [Release notes](https://github.com/nodejs/undici/releases)
- [Commits](nodejs/undici@v7.29.0...v8.10.2)

---
updated-dependencies:
- dependency-name: undici
  dependency-version: 8.10.2
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps [eslint](https://github.com/eslint/eslint) from 10.10.0 to 10.11.0.
- [Release notes](https://github.com/eslint/eslint/releases)
- [Commits](eslint/eslint@v10.10.0...v10.11.0)

---
updated-dependencies:
- dependency-name: eslint
  dependency-version: 10.11.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps [prettier](https://github.com/prettier/prettier) from 3.9.6 to 3.9.8.
- [Release notes](https://github.com/prettier/prettier/releases)
- [Changelog](https://github.com/prettier/prettier/blob/main/CHANGELOG.md)
- [Commits](prettier/prettier@3.9.6...3.9.8)

---
updated-dependencies:
- dependency-name: prettier
  dependency-version: 3.9.8
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
- macOS host memory: parse memory_pressure page counts (free+inactive+
  speculative+purgeable, XNU's own availability definition) instead of the
  opaque system-wide free percentage; percentage stays as fallback.
- /gateway/status and /gateway/node sample process memory usage only with
  ?memoryUsage=1; the dashboard asks for it on the live view and cluster
  node polls keep it. Plain status polls no longer spawn ps/lsof/python3.
- Invalid memorySafety config now raises MemorySafetyConfigError (500,
  config-error code) instead of a 503 transient load-abort.
- First-run job reuses the verify stage instead of appending a duplicate.
Profiles live in <configDir>/profiles/<name>.json and describe routes
(alias -> route profile or member id), keep-warm residency per runtime,
and defaults overrides. Applying composes the profile onto the config
source; mutateConfigSource validates the staged candidate with the real
loader before the atomic rename, so a swap is all-or-nothing. After a
swap the registry hot-reloads; local models admit lazily on first
request and unrouted models shed per residency policy.

- src/config-profiles.mjs: format validation, planning, composition,
  apply/save controller; active profile marked at fleet.activeProfile.
- Endpoints: GET /gateway/fleet/profiles, GET/POST
  /gateway/fleet/profiles/:name (?apply=1 applies, else saves).
- CLI: lloom fleet list|show|use|save with the usual plan/apply gate.
- Dashboard: Fleet Profiles band in Operations with one-click apply and
  save-current capture.
- Tests: test/config-profiles.test.mjs (format, plan, compose, atomic
  apply, save capture, traversal rejection); wired into check:js and
  test:unit.
…ion' into feat/lloom-overhaul

# Conflicts:
#	package.json
…-3.9.8' into feat/lloom-overhaul

# Conflicts:
#	package-lock.json
#	package.json
Predictive admission (5cc4f4c) folded host-wide memory usage into the
modeled runtime budget: any machine running a browser or IDE would
report projected usage far above the budget with no evictable path back,
so every failover-to-local admission returned
runtime_capacity_impossible. Admission now decides on the modeled budget
(loaded + requested) and uses the live host signal only for the hard
reserve check: after the admission and planned evictions, available
memory must still cover reserve plus the requested add. Host-wide
projection stays in the plan payload for observability.

Undici 8 speaks a dispatcher contract Node 22's internal undici v6
rejects (invalid onRequestStart). The gateway's upstream calls, the
cluster coordinator, and the CLI gateway client now pair the package
Agent with the package fetch; the long-prefill and openrouter-provider
tests use the v8 MockAgent global-dispatcher pattern with loopback net
allowed for client-side requests, and model maintenance accepts an
upstreamDispatcher for test composition. Smoke accepts the new
macos-memory-pages telemetry source.
@data-angel
data-angel merged commit 528c12e into main Sep 23, 2026
@data-angel
data-angel deleted the feat/lloom-overhaul branch September 23, 2026 04:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant