Build guided setup and the LLooM browser experience - #33
Merged
Merged
Conversation
Bumps [undici](https://github.com/nodejs/undici) from 7.29.0 to 8.10.2. - [Release notes](https://github.com/nodejs/undici/releases) - [Commits](nodejs/undici@v7.29.0...v8.10.2) --- updated-dependencies: - dependency-name: undici dependency-version: 8.10.2 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com>
Bumps [eslint](https://github.com/eslint/eslint) from 10.10.0 to 10.11.0. - [Release notes](https://github.com/eslint/eslint/releases) - [Commits](eslint/eslint@v10.10.0...v10.11.0) --- updated-dependencies: - dependency-name: eslint dependency-version: 10.11.0 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com>
Bumps [prettier](https://github.com/prettier/prettier) from 3.9.6 to 3.9.8. - [Release notes](https://github.com/prettier/prettier/releases) - [Changelog](https://github.com/prettier/prettier/blob/main/CHANGELOG.md) - [Commits](prettier/prettier@3.9.6...3.9.8) --- updated-dependencies: - dependency-name: prettier dependency-version: 3.9.8 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com>
- macOS host memory: parse memory_pressure page counts (free+inactive+ speculative+purgeable, XNU's own availability definition) instead of the opaque system-wide free percentage; percentage stays as fallback. - /gateway/status and /gateway/node sample process memory usage only with ?memoryUsage=1; the dashboard asks for it on the live view and cluster node polls keep it. Plain status polls no longer spawn ps/lsof/python3. - Invalid memorySafety config now raises MemorySafetyConfigError (500, config-error code) instead of a 503 transient load-abort. - First-run job reuses the verify stage instead of appending a duplicate.
Profiles live in <configDir>/profiles/<name>.json and describe routes (alias -> route profile or member id), keep-warm residency per runtime, and defaults overrides. Applying composes the profile onto the config source; mutateConfigSource validates the staged candidate with the real loader before the atomic rename, so a swap is all-or-nothing. After a swap the registry hot-reloads; local models admit lazily on first request and unrouted models shed per residency policy. - src/config-profiles.mjs: format validation, planning, composition, apply/save controller; active profile marked at fleet.activeProfile. - Endpoints: GET /gateway/fleet/profiles, GET/POST /gateway/fleet/profiles/:name (?apply=1 applies, else saves). - CLI: lloom fleet list|show|use|save with the usual plan/apply gate. - Dashboard: Fleet Profiles band in Operations with one-click apply and save-current capture. - Tests: test/config-profiles.test.mjs (format, plan, compose, atomic apply, save capture, traversal rejection); wired into check:js and test:unit.
…ion' into feat/lloom-overhaul # Conflicts: # package.json
…into feat/lloom-overhaul
…0.11.0' into feat/lloom-overhaul
…-3.9.8' into feat/lloom-overhaul # Conflicts: # package-lock.json # package.json
….10.2' into feat/lloom-overhaul
Predictive admission (5cc4f4c) folded host-wide memory usage into the modeled runtime budget: any machine running a browser or IDE would report projected usage far above the budget with no evictable path back, so every failover-to-local admission returned runtime_capacity_impossible. Admission now decides on the modeled budget (loaded + requested) and uses the live host signal only for the hard reserve check: after the admission and planned evictions, available memory must still cover reserve plus the requested add. Host-wide projection stays in the plan payload for observability. Undici 8 speaks a dispatcher contract Node 22's internal undici v6 rejects (invalid onRequestStart). The gateway's upstream calls, the cluster coordinator, and the CLI gateway client now pair the package Agent with the package fetch; the long-prefill and openrouter-provider tests use the v8 MockAgent global-dispatcher pattern with loopback net allowed for client-side requests, and model maintenance accepts an upstreamDispatcher for test composition. Smoke accepts the new macos-memory-pages telemetry source.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
LLooM now opens guided browser setup on the first interactive run and gives installed gateways Live, Models, Machines, Clients, and Settings views. Observed requests animate the topology. The Models view shows proportional blocks for backend memory, system/other apps, and free space, with protected headroom inside the free block. Pointer, keyboard-focus, and selection previews animate a model’s expected memory impact on its own machine. Normal use starts models automatically; optional manual preparation and memory release live under Options & details.
Setup and model installation use reviewed, expiring plans and observable jobs. Backend and download commands remain fixed through apply and retry. Custom Hugging Face imports require immutable commit links and staged acquisition verification. Setup publishes configuration exclusively and verifies a model response through the gateway before reporting readiness. CLI, offline, and scripted paths remain available.
Admission now counts live host memory even when policy specifies only a reserve or absolute budget. Newly launched local backends are monitored through loading and warmup. A threshold breach or unavailable telemetry aborts the load and cleans up only its owned process tree or Docker container. An independent supervisor process checks memory before spawning, and the manager waits for that check before accepting health. Forced starts cannot bypass the guard. Explicit
runtimePolicy.memorySafety.mode: "yolo"disables memory admission and the load guard, with a persistent dashboard warning. The guard polls in userspace; it is not a kernel allocation quota. Automatic retries remain blocked until a manual retry or gateway restart; persistent suspension survives restarts.Memory attribution uses bounded passive process observations with a five-second cache and overlapping process-tree deduplication. macOS reads physical footprint, including charged Metal/compressed memory, with RSS fallback. Known local model-cache observations distinguish a healthy service from a loaded model. Estimates are marked, remote safety policies stay node-specific, and routing/admission status skips local dashboard sampling.
The streaming stall watchdog rechecks elapsed time if its timer wakes early, preventing a premature event from being ignored by stall classification.
Readiness preferences persist without restarting active models. Changes wait for admission, queued restores recheck intent, and pending hard pins protect models from eviction. Browser management rejects foreign origins and rebound loopback hosts; remote writes require an admin credential.
Validation:
npm testpassed locally on macOS / Node 22 with the memory map and Mac footprint sampler. After review fixes, all 36 memory-map/sampler tests and 33 maintenance tests passed, plus cluster, runtime policy, dashboard-status, and the 29-test UX suite. GitHub CI passed on Node 20 and Node 22 at2609b87, including full tests, dependency audit, interchange, and package installation. The memory safety and supervisor suites passed all 16 cases, including pressure during load and warmup, process-tree cleanup, immutable Docker ownership, independent cutoff, stale health, and the supervisor startup race.2609b87was installed on the Mac and its installed sources were verified against the running service (artifact SHA-256ea526b1df64f59c83e18913267cf08b9f131b7386ad77bf8e46f24a82db0f64e). Configuration was preserved; enforcement reports a 12 GiB reserve, 90% ceiling, and 250 ms polling. An isolated test of the installed HTTP API returned503 runtime_memory_safety_abortbefore spawning its tiny fixture, despiteforce: trueand disabled predictive policy.A signed public installer and automatic nearby discovery/pairing remain separate work. Configured federation is available; independent nodes do not pair automatically in this change.