Skip to content

Per-model Spark speech recipes on a source-built audio backend - #41

Merged
data-angel merged 5 commits into
mainfrom
feat/spark-audio-recipes
Sep 27, 2026
Merged

data-angel merged 5 commits into
mainfrom
feat/spark-audio-recipes

Conversation

@data-angel

@data-angel data-angel commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Replaces a host-local recipe that ran four speech models from one 32.8 GB prebuilt image with a source-built backend and one recipe per model. Stacked on #37.

  • backends/spark-audio/install.py builds or reuses lloom/spark-audio:source-<digest>, the same pattern as the other source-built backends. The digest covers only the Dockerfile and the server. The backend is registered in backends/catalog.json.
  • Four recipes, each with its own runtime and container, a pinned Hugging Face revision, per-file sha256, and a read-only mount of only that model's directory:
    • linux-nvidia-spark-audio-qwen3-tts-1-7b-customvoice
    • linux-nvidia-spark-audio-qwen3-tts-1-7b-base (voice clone)
    • linux-nvidia-spark-audio-qwen3-tts-1-7b-voicedesign
    • linux-nvidia-spark-audio-whisper-large-v3-turbo

Whisper weights source

The adapter loads Whisper with openai-whisper, which needs OpenAI's original large-v3-turbo.pt. openai/whisper-large-v3-turbo only publishes Transformers weights. The recipe downloads the file from dataangel/whisper-large-v3-turbo-openai at a pinned revision. That repo is an unmodified mirror copied from OpenAI's download URL, with Whisper's MIT license. The sha256 is enforced and matches the one openai-whisper v20250625 publishes.

Test plan

  • node test/spark-audio-recipes.test.mjs: each recipe alone, and all four together in both orders, get independent runtimes and ports, the pinned image, one read-only mount, and no host paths
  • node --test test/speech-stream.test.mjs (12 pass), node test/tts-catalog.test.mjs
  • pytest backends/spark-audio (53 pass)
  • npm run check, npm run test:unit, node scripts/check-package.mjs
  • Build and run the image on a Spark host. The memory estimates (10 GB TTS, 6 GB STT) come from the old host recipe, not from measurement.

🤖 Generated with Claude Code

data-angel and others added 5 commits September 27, 2026 12:13
Replace the host-local prebuilt media image with a locally built
lloom/spark-audio:source-<sha256> image and one public recipe per model:
Qwen3-TTS 1.7B CustomVoice, Base (voice cloning) and VoiceDesign, and
Whisper large-v3-turbo. Each recipe runs its checkpoint in its own
managed container with a loopback port, a read-only mount of only that
model's directory, and offline Hugging Face mode. Downloads are pinned
to immutable revisions with per-file SHA-256.

openai-whisper cannot load the Transformers-format repository, so the
Whisper recipe fetches OpenAI's original large-v3-turbo.pt from a
byte-identical Hugging Face copy whose hash matches openai-whisper
v20250625.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace the third-party Hugging Face copy with an unmodified mirror of
OpenAI's large-v3-turbo.pt; the enforced SHA-256 is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts:
#	CHANGELOG.md
#	docs/recipes.md
#	package.json
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Base automatically changed from feat/spark-speech-streaming to main September 27, 2026 20:29
@data-angel
data-angel merged commit 441440b into main Sep 27, 2026
2 checks passed
@data-angel
data-angel deleted the feat/spark-audio-recipes branch September 27, 2026 20:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant