Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,8 @@ All notable changes to LLooM will be documented in this file. The format follows

- LLooM Hear CPU audio analysis, with structured estimates, optional dashboard images, and opt-in upstream interpretation. File and URL inputs require operator configuration.

- Four standalone NVIDIA Spark audio recipes (Qwen3-TTS 1.7B CustomVoice, Base voice cloning and VoiceDesign, and Whisper large-v3-turbo). Each runs one pinned, integrity-checked checkpoint in its own container from a locally built `spark-audio` image that replaces a host-local prebuilt image.

- Fourteen per-model NVIDIA ComfyUI media recipes and a public-source backend build for image, video and music generation. Each model runs in its own container with read-only mounts of only its files, so LLooM admits, evicts and restores each one independently. The media launcher refuses to start without `LLOOM_MEDIA_MODEL`, and re-applying a recipe moves a model off the retired shared `comfyui-media` runtime, dropping it once unused.
- OpenRouter Lyria audio generation through `/v1/audio/generations`.

Expand Down
55 changes: 55 additions & 0 deletions backends/catalog.json
Original file line number Diff line number Diff line change
Expand Up @@ -1047,6 +1047,61 @@
"notes": "Per-request cancellation is checked between denoising steps. The backend retains its sole generation slot until the worker thread exits; encoding and decoding are not immediately interruptible."
}
},
{
"id": "spark-audio",
"name": "Spark audio (NVIDIA Docker)",
"kind": "openai-compatible-server",
"description": "Source-built CUDA adapter serving one local Qwen3-TTS or Whisper checkpoint per container through the OpenAI speech and transcription APIs.",
"platforms": [
"linux-arm64",
"linux-x64"
],
"features": [
"audio-speech",
"tts",
"tts-voice-clone",
"audio-transcription",
"stt",
"cuda",
"docker"
],
"commands": [
"docker",
"python3"
],
"setup": [
{
"id": "build-spark-audio",
"title": "Build or reuse the source-pinned Spark audio backend",
"action": "command",
"command": "python3",
"args": [
"${repoRoot}/backends/spark-audio/install.py",
"--install-root",
"${installRoot}",
"--shim-dir",
"${shimDir}"
],
"alwaysRun": true
}
],
"server": {
"protocol": "openai",
"healthPath": "/health",
"speechPath": "/v1/audio/speech",
"transcriptionPath": "/v1/audio/transcriptions"
},
"cancellation": {
"trigger": "client-disconnect",
"computeReclamation": "partial",
"modes": [
"speech",
"transcription",
"streaming"
],
"notes": "Buffered speech and transcription run to completion on the single GPU worker, which keeps the adapter's only slot until it exits. A streamed PCM voice clone on the optional faster engine stops at the next chunk and closes its generator."
}
},
{
"id": "hear",
"name": "LLooM Hear",
Expand Down
2 changes: 2 additions & 0 deletions backends/spark-audio/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,3 +26,5 @@ environment with FastAPI, httpx, numpy, soundfile, python-multipart, and pytest.
These tests use fake models. Hardware acceptance must additionally prove PCM
headers, named-profile routing, streaming onset, interruption/recovery, and
transcription through the selected gateway on the actual host.

Per-model recipes and setup commands are documented in [Spark audio](../../docs/spark-audio.md).
66 changes: 66 additions & 0 deletions backends/spark-audio/install.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
#!/usr/bin/env python3
"""Build/reuse the local image identified by its bundled source digest."""
import argparse
import hashlib
import json
from pathlib import Path
import shutil
import subprocess
import sys

ROOT = Path(__file__).resolve().parent


def source_digest(root=ROOT):
files = [root / name for name in ('Dockerfile', 'lloom_audio_cuda_server.py')]
digest = hashlib.sha256()
for file in sorted(files):
digest.update(file.relative_to(root).as_posix().encode() + b'\0')
digest.update(file.read_bytes() + b'\0')
return digest.hexdigest()


def image_name():
return 'lloom/spark-audio:source-' + source_digest()


def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('--install-root', type=Path)
parser.add_argument('--shim-dir', type=Path)
parser.add_argument('--print-image', action='store_true')
args = parser.parse_args()
image = image_name()
if args.print_image:
print(image)
return
if args.install_root is None or args.shim_dir is None:
parser.error('--install-root and --shim-dir are required')
if sys.platform != 'linux':
parser.error('The CUDA backend requires Linux and NVIDIA Container Toolkit')
found = subprocess.run(['docker', 'image', 'inspect', image], capture_output=True, text=True)
expected = source_digest()
if found.returncode == 0:
labels = json.loads(found.stdout)[0].get('Config', {}).get('Labels', {}) or {}
if labels.get('dev.lloom.source-sha256') != expected:
raise SystemExit('Refusing an existing image with an unexpected source identity: ' + image)
print('Reusing ' + image)
else:
subprocess.run(['docker', 'build', '--label', 'dev.lloom.source-sha256=' + expected,
'--tag', image, str(ROOT)], check=True)
if not shutil.which('hf'):
venv = args.install_root / 'spark-audio' / 'huggingface'
if not (venv / 'bin/python').exists():
subprocess.run([sys.executable, '-m', 'venv', str(venv)], check=True)
subprocess.run([str(venv / 'bin/python'), '-m', 'pip', 'install', 'huggingface-hub==1.7.1'], check=True)
args.shim_dir.mkdir(parents=True, exist_ok=True)
shim = args.shim_dir / 'hf'
if shim.exists() or shim.is_symlink():
if shim.resolve() != (venv / 'bin/hf').resolve():
raise SystemExit('Refusing to replace an unrelated hf shim')
else:
shim.symlink_to(venv / 'bin/hf')


if __name__ == '__main__':
main()
2 changes: 2 additions & 0 deletions docs/backends.md
Original file line number Diff line number Diff line change
Expand Up @@ -139,3 +139,5 @@ Authenticated hosted OpenAI-compatible providers use the same config-only unmana
`lloom-host` publishes a lightweight backend index at `GET /v1/backends` and the full portable backend setup document at `GET /v1/backends/catalog`. Recipe authors and independent installers should use the catalog endpoint when they need setup actions, platform filters, server contracts, and idempotency guards rather than just the stable backend IDs.

The `comfyui-media` backend builds a source-pinned NVIDIA Docker image for the [per-model media recipes](comfyui-media.md). Its idempotent build step uses `alwaysRun: true` so setup verifies the external image even when a previous installation was recorded as complete.

The `spark-audio` backend builds the same kind of source-identified image for the [per-model speech and transcription recipes](spark-audio.md). Each recipe runs one checkpoint in its own container.
2 changes: 2 additions & 0 deletions docs/recipes.md
Original file line number Diff line number Diff line change
Expand Up @@ -341,3 +341,5 @@ Only the stable active file participates in planning and automatic recommendatio
LLooM intentionally does not use stale model fallback aliases to make an index pass. Recipe `model` and `gatewayModel` values must be exact advertised IDs.

Per-model image, video and music recipes, each in its own ComfyUI runtime, are documented in [ComfyUI media](comfyui-media.md).

Per-model NVIDIA speech and transcription recipes are documented in [Spark audio](spark-audio.md).
57 changes: 57 additions & 0 deletions docs/spark-audio.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
# Spark audio speech and transcription recipes

Each `linux-nvidia-spark-audio-*` recipe installs one checkpoint and runs it in
its own LLooM-managed container, port and runtime. All four recipes share only the
image, which is built locally from `backends/spark-audio`.

| Recipe | Gateway model | Kind |
| ----------------------------------------------------- | -------------------------------------- | ------------------------- |
| `linux-nvidia-spark-audio-qwen3-tts-1-7b-customvoice` | `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice` | speech, built-in speakers |
| `linux-nvidia-spark-audio-qwen3-tts-1-7b-base` | `Qwen/Qwen3-TTS-12Hz-1.7B-Base` | speech, voice cloning |
| `linux-nvidia-spark-audio-qwen3-tts-1-7b-voicedesign` | `Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign` | speech, voice design |
| `linux-nvidia-spark-audio-whisper-large-v3-turbo` | `openai/whisper-large-v3-turbo` | transcription |

## Install

Use Linux on NVIDIA hardware with Docker, NVIDIA Container Toolkit and Python 3
with venv support. Preview the complete plan, then apply and start:

```sh
lloom setup --recipe linux-nvidia-spark-audio-qwen3-tts-1-7b-base --additive --json
lloom setup --recipe linux-nvidia-spark-audio-qwen3-tts-1-7b-base --additive --apply --yes --start
```

Repeat with another recipe ID to add a model. `--additive` preserves the existing
catalog and defaults. Speech models reserve 10 GiB and Whisper reserves 6 GiB in
LLooM's admission configuration. Each runtime accepts one request at a time and
is not kept warm.

The backend installer builds `lloom/spark-audio:source-<sha256>` from the digest-pinned
NGC PyTorch base, the Dockerfile and the adapter. The tag is the SHA-256 of those
build inputs; tests and documentation are excluded. Reapplying setup checks the
image's source label and reuses a matching image. Nothing is pulled from a private
registry.

Downloads use immutable Hugging Face revisions with per-file SHA-256 checks. Each
container mounts only its own model directory, read-only, with Hugging Face
offline mode enabled. Backend ports stay bound to loopback; clients use the
authenticated LLooM gateway.

openai-whisper cannot load the Transformers-format `openai/whisper-large-v3-turbo`
repository. The Whisper recipe therefore downloads OpenAI's original
`large-v3-turbo.pt` checkpoint from an unmodified mirror,
[`dataangel/whisper-large-v3-turbo-openai`](https://huggingface.co/dataangel/whisper-large-v3-turbo-openai),
copied from OpenAI's download URL. Its SHA-256 matches the value published in
openai-whisper v20250625.

Voice cloning uses the stock Qwen engine by default. Streamed PCM voice clones
and cancellation between chunks require `LLOOM_TTS_ENGINE=faster` in the
runtime's container environment; see the [adapter README](../backends/spark-audio/README.md).
Buffered synthesis and transcription run to completion after a client disconnects.

Run CPU validation:

```sh
python -m pytest -q backends/spark-audio/test_lloom_audio_cuda_server.py
node test/spark-audio-recipes.test.mjs
```
Loading
Loading