Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
0c256cc
Bump undici from 7.29.0 to 8.10.2
dependabot[bot] Sep 7, 2026
d1e76d9
Bump eslint from 10.10.0 to 10.11.0
dependabot[bot] Sep 21, 2026
ed393a6
Bump prettier from 3.9.6 to 3.9.8
dependabot[bot] Sep 21, 2026
a0387ee
Build guided setup and the LLooM browser experience
data-angel Sep 22, 2026
e35c935
Bind reviewed downloads and verify imported model acquisition
data-angel Sep 22, 2026
4612fe2
Accept compatible CPU recipes in dashboard library smoke check
data-angel Sep 22, 2026
5cc4f4c
Abort unsafe model loads with supervised memory thresholds
data-angel Sep 22, 2026
534c80a
Wait for the measured stall deadline before reporting watchdog failure
data-angel Sep 22, 2026
9781b5d
Show interactive memory blocks and make model use automatic
data-angel Sep 22, 2026
2609b87
Keep memory previews honest across suspended and federated models
data-angel Sep 22, 2026
b67e1a7
Enforce configured OpenRouter provider restrictions
data-angel Sep 22, 2026
5c3eb6e
Clarify provider preferences on protocol bridges
data-angel Sep 22, 2026
b8fbb42
Harden memory telemetry and trim first-run stage noise
data-angel Sep 22, 2026
d02d8d2
Add named fleet profiles: one-file route and residency swaps
data-angel Sep 22, 2026
0ccb029
Discover NVIDIA Sync peers without replacing federation or model plac…
data-angel Sep 23, 2026
dc19d4c
Merge remote-tracking branch 'origin/fix/openrouter-provider-restrict…
data-angel Sep 23, 2026
461b8bc
Merge remote-tracking branch 'origin/fix/nvidia-sync-ring-discovery' …
data-angel Sep 23, 2026
85cee50
Merge remote-tracking branch 'origin/dependabot/npm_and_yarn/eslint-1…
data-angel Sep 23, 2026
620ed52
Merge remote-tracking branch 'origin/dependabot/npm_and_yarn/prettier…
data-angel Sep 23, 2026
73aba29
Merge remote-tracking branch 'origin/dependabot/npm_and_yarn/undici-8…
data-angel Sep 23, 2026
528c12e
Fix predictive admission and undici 8 dispatcher contract after merges
data-angel Sep 23, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .gitignore
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
node_modules/
node_modules
.lloom/
.DS_Store
.env
Expand Down
29 changes: 16 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
[![CI](https://github.com/enntity/lloom/actions/workflows/ci.yml/badge.svg)](https://github.com/enntity/lloom/actions/workflows/ci.yml)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)

LLooM is a local-first LLM gateway for people who run serious open models on their own hardware. It treats NVIDIA systems—including DGX Spark / GB10—and Apple Silicon Macs as first-class platforms. LLooM sits in front of vLLM, SGLang, MLX, MTPLX, llama.cpp, Ollama, image generators, and other local runtimes, then exposes stable OpenAI-compatible and Anthropic-compatible APIs to agent tools.
LLooM installs, manages, and serves AI models on your hardware. It treats NVIDIA systems—including DGX Spark / GB10—and Apple Silicon Macs as first-class platforms. LLooM sits in front of vLLM, SGLang, MLX, MTPLX, llama.cpp, Ollama, image generators, and other local runtimes, then exposes stable OpenAI-compatible and Anthropic-compatible APIs to agent tools.

The goal is simple: install one bridge, let it inspect the machine, choose the best agentic model recipe from the LLooM community library, install the backend needed for that recipe, download and configure the model, keep it warm, and point Codex, Claude Code, OMP, OpenCode, Hermes, Zero, or any OpenAI-compatible client at one base URL.

Expand All @@ -15,11 +15,11 @@ The planned public community host is `https://lloom.enntity.com`; source checkou

## First-Class Platforms

| NVIDIA / DGX Spark | Apple Silicon |
| -------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- |
| CUDA, Blackwell, DGX Spark / GB10, and Linux NVIDIA hosts | M-series Macs with unified memory |
| vLLM and SGLang are the primary high-throughput backends | MLX, MTPLX, OptiQ, and llama.cpp are the primary native backends |
| Managed Docker runtimes, GPU-memory admission, warm/on-demand lanes, and Spark recipes | Native processes, unified-memory-aware recipes, model-root reuse, and Mac recipes |
| NVIDIA / DGX Spark | Apple Silicon |
| --------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- |
| CUDA, Blackwell, DGX Spark / GB10, and Linux NVIDIA hosts | M-series Macs with unified memory |
| vLLM and SGLang are the primary high-throughput backends | MLX, MTPLX, OptiQ, and llama.cpp are the primary native backends |
| Managed Docker runtimes, GPU-memory admission, warm/on-demand lanes, and Spark recipes | Native processes, unified-memory-aware recipes, model-root reuse, and Mac recipes |
| See [`docs/dgx-spark.md`](docs/dgx-spark.md) and [`docs/clusters.md`](docs/clusters.md) | See the bundled `apple-silicon-*` recipes |

Both platforms get the same gateway APIs, runtime policy, per-connection telemetry, live dashboard, client integrations, external-provider passthrough, and community recipe/benchmark workflow. Independent LLooM gateways can also form a heterogeneous lab behind one central endpoint; node profiles and optional GPU telemetry degrade cleanly across CUDA, Metal, ROCm, and CPU-only hosts.
Expand All @@ -34,7 +34,6 @@ cd lloom
npm ci
npm link
lloom
lloom up --go
```

`npm link` installs the same `lloom` and `lloom-host` commands from your checkout. After the first npm release, `npm install -g lloom` will be the supported package install path.
Expand All @@ -48,7 +47,9 @@ curl -sS http://127.0.0.1:8100/v1/models

Dashboard: [http://127.0.0.1:8100/](http://127.0.0.1:8100/)

That is the 1.0 path. A bare `lloom` is a dry run first: it inspects the machine, asks the LLooM community host for the best known recipe pack and backend catalog, shows what will be installed, and refuses writes until you rerun it with `--go`. `up` is the named alias for the same first-run flow. `--go` applies the plan, confirms noninteractive writes, and starts the selected keep-warm runtime after setup. Use `--offline` when you want to ignore the host and select from only the local recipe library.
On a new installation in an interactive terminal, `lloom` opens a local browser setup. It detects your hardware, asks what you want to use AI for, and recommends a compatible vendor recipe from the bundled library. Review the plan and choose **Set up my AI** to install it. The terminal stays open during setup. Chat setup checks a real response through the gateway before showing **Ready**; media installations show when output still needs verification.

Use `lloom up --browser` to request browser setup explicitly, or `lloom ui` to open an installed gateway. The CLI remains available: `lloom --no-browser` previews the community-based plan; `lloom up --go` installs, integrates, and starts it. Scripted, JSON, offline, and explicit recipe commands keep their CLI behavior. See [the browser experience](docs/browser-experience.md) for the flow and current boundaries.

The default gateway endpoint is `127.0.0.1:8100`; managed backend runtimes default to `8201-8299`. This source checkout defaults community lookup to the local development host at `127.0.0.1:8110`, starts it automatically if it is not already running, serves signed seed host data from `community/`, and requires signed recipe packs by default. Local imports still land in `recipes/` and `benchmarks/community/`. A production package should point at the signed public LLooM host. Most users should not need to care.

Expand All @@ -60,12 +61,12 @@ Use this path when validating the repository before a package release:

```bash
npm install -g .
lloom up
lloom up --no-browser
lloom up --go
lloom doctor --no-runtimes
```

`lloom up` should show the detected machine profile, the trusted community recommendation, the selected recipe, benchmark evidence, and the exact apply command.
`lloom up --no-browser` should show the detected machine profile, the trusted community recommendation, the selected recipe, benchmark evidence, and the exact apply command.

- On NVIDIA Linux, LLooM detects CUDA devices, compute capability, Blackwell, and DGX Spark / GB10 markers. Spark recipes use vLLM or SGLang, managed Docker containers, and explicit GPU-memory/runtime policy. The checked-in Spark deployment demonstrates a warm primary chat model, warm embedding model, and an on-demand alternate chat lane.
- On a 96 GB Apple Silicon machine, the bundled development host should recommend `apple-silicon-qwen36-35b-a3b-mtplx` and select `Youssofal/Qwen3.6-35B-A3B-MTPLX-Optimized-Speed-FP16`. Lower-memory Macs should fall back to the 27B MTPLX recipe.
Expand Down Expand Up @@ -99,14 +100,16 @@ Then open OMP normally. The generated OMP config points at `http://127.0.0.1:810
- Backend recipes for vLLM, SGLang, MTPLX, MLX LM, llama.cpp, Ollama, OptiQ, and stable-diffusion.cpp, with dedicated DGX Spark / GB10 and Apple Silicon recipes.
- Community recipe packs and hardware-matched benchmark evidence so machines can select the best known model/backend recipe automatically instead of blindly chasing global tok/s.
- Generated client profiles for OMP, OpenCode, Codex-compatible, Claude-compatible, Hermes, Zero, and any OpenAI-compatible client.
- A small dashboard at `/` for local status and guarded setup actions, with a live topology that can switch between the default columnar racks and an action view whose camera and cards follow live models.
- A browser dashboard with Live, Models, Machines, Clients, and Settings. Inspect real topology and memory blocks, preview a model’s expected footprint by pointing or focusing, and use models directly through chat or connected apps. LLooM prepares models automatically; optional readiness and memory controls live under Options & details. The action camera follows serving models while preserving manual zoom.

## Daily Commands

Primary ladder (see `lloom help`; full catalog under `lloom help advanced`):

```bash
lloom # preview plan
lloom # browser setup on first interactive run
lloom ui # open the installed gateway
lloom --no-browser # preview the CLI plan
lloom up --go # install + integrate + start
lloom down # stop the gateway and all managed model backends
lloom doctor --no-runtimes
Expand All @@ -123,7 +126,7 @@ lloom add-model 'openai:http://127.0.0.1:8000/v1#my-model' --default --apply --y
lloom serve --config ~/.lloom/config.json
```

Bare `lloom`, `up`, and `onboard` all route to the same first-run flow. By default, the community request asks for the best known `agentic-coding` recipe with `tools`, `reasoning`, and `long-context`; use repeated `--workload`, `--capability`, or `--tag` flags to target a different kind of local model. `doctor` is the readiness view for humans and automation. `integrate` repairs or writes client configs from the registry. `add-model` imports an ad hoc Hugging Face, local, or Ollama model outside the community recipe library.
The CLI `onboard` flow remains available independently of browser setup. By default, the community request asks for the best known `agentic-coding` recipe with `tools`, `reasoning`, and `long-context`; use repeated `--workload`, `--capability`, or `--tag` flags to target a different kind of local model. `doctor` is the readiness view for humans and automation. `integrate` repairs or writes client configs from the registry. `add-model` imports an ad hoc Hugging Face, local, or Ollama model outside the community recipe library.

After `~/.lloom/config.json` exists, operational commands such as `doctor`, `models`, `serve`, `integrate`, `add-model`, and runtime controls automatically read that installed config when `--config` is not supplied. Before that file exists, those commands return a `not-installed` report with the exact `lloom up` command to run, instead of silently operating on bundled model defaults. Read-only planning commands such as `lloom`, `lloom up`, `onboard`, and `integrations` can still preview from the packaged gateway shell plus community data. Use `--config` whenever you want to inspect or operate a different config file.

Expand Down
3 changes: 3 additions & 0 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,8 +24,11 @@ The maintainers will acknowledge reports as soon as practical, validate the issu
- A signature proves which key signed a recipe pack. Trusting keys served by the same community host is equivalent to trusting that host and its TLS connection; use explicit local trusted keys for stronger publisher pinning.
- Development keys and the checked-in public seed key are not production trust roots.
- Model weights and external runtimes have their own licenses and security posture. LLooM does not make untrusted model code safe.
- Browser management requests must come from the gateway's own origin. First-run setup additionally requires a local session token and applies only reviewed plans. Keep the setup terminal and tokens private. See [the browser security boundary](docs/browser-experience.md#local-security).
- The public community MVP is read-only by design. Proposals arrive through reviewed pull requests; anonymous recipe and benchmark uploads are disabled in production.
- Remote community feeds must use HTTPS and a locally pinned public signing key. Do not treat a signing key downloaded from the same remote host as an independent trust root.
- The production community deployment is isolated from inference, databases, and Docker control-plane access. See [`deploy/community/README.md`](deploy/community/README.md).

See [docs/architecture.md](docs/architecture.md) for route authorization and network-binding defaults.

Remote admin writes require `security.allowRemoteAdmin=true` and a key in `security.adminApiKeys`. Inference credentials alone do not grant remote process or installation control.
Loading
Loading