Skip to content

Stream Spark speech with bounded audio and canonical voice profiles - #37

Merged
data-angel merged 5 commits into
mainfrom
feat/spark-speech-streaming
Sep 27, 2026
Merged

data-angel merged 5 commits into
mainfrom
feat/spark-speech-streaming

Conversation

@data-angel

@data-angel data-angel commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Adds backends/spark-audio: a one-checkpoint-per-process Qwen3-TTS / Whisper adapter for Spark hosts, with an optional faster-qwen3-tts CUDA-graph engine.
  • POST /v1/audio/speech with stream: true and response_format: pcm now streams raw PCM16 through the gateway (JSON and multipart). WAV stays buffered, and upstream validation errors still come back as buffered JSON with their status.
  • Named voice profiles send their reference audio inline, so the adapter never opens a caller-supplied path. The gateway only expands the profile's own reference, after a realpath containment check against the voice registry, an O_NOFOLLOW open, and a 16 MiB cap.
  • Relays x-audio-sample-rate / x-audio-channels / x-audio-format, and ends a binary response by destroying the socket instead of appending JSON to audio.

Also folds in the ennspark03-only fix 306d84b (upload installed voice references to remote clone backends). Its resolver-side inlining replaces the duplicate server-side copy, and now has the same registry containment, O_NOFOLLOW and size checks, with a symlink-escape test.

Rebased from improve/avatar-calls-20260923. That branch's checkpoint commit (ACE-Step diffusers + Qwen 3.8 NVFP4 recipe) is left out because it lands with the video-workflow work.

Test plan

  • node --test test/speech-stream.test.mjs (12 pass)
  • node test/voice-profiles.test.mjs, node test/tts-catalog.test.mjs, node test/protocol.test.mjs
  • pytest backends/spark-audio/test_lloom_audio_cuda_server.py (53 pass, fake models)
  • eslint + prettier
  • Hardware acceptance on a Spark host (PCM onset, profile routing, interruption)

🤖 Generated with Claude Code

data-angel and others added 3 commits September 27, 2026 11:55
The gateway inlined installed profile references in two places. Keep the
resolver's version, which also covers remote clone backends, and give it
the registry containment, no-follow open and size checks the server copy
had.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@data-angel
data-angel merged commit 9a0f95f into main Sep 27, 2026
2 checks passed
@data-angel
data-angel deleted the feat/spark-speech-streaming branch September 27, 2026 20:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant