From 18e76158501291b59ac92065b1eff14a95bc328d Mon Sep 17 00:00:00 2001 From: Corey Weathers Date: Thu, 1 Oct 2026 09:45:25 -0400 Subject: [PATCH 01/17] docs(api): regenerate references; cover Flux TTS inline controls, Flux STT Warning and numerals, Voice Agent REST --- skills/api/SKILL.md | 37 ++++++++++++++++----------------- skills/api/references/listen.md | 15 ++++++++++++- skills/api/references/speak.md | 22 +++++++++++--------- 3 files changed, 44 insertions(+), 30 deletions(-) diff --git a/skills/api/SKILL.md b/skills/api/SKILL.md index b62748a..c6cf0ab 100644 --- a/skills/api/SKILL.md +++ b/skills/api/SKILL.md @@ -25,11 +25,11 @@ All API requests require authentication via API key or JWT: Base servers: - REST & STT/TTS WebSocket: `https://api.deepgram.com` -- Voice Agent WebSocket **and Voice Agent REST**: `https://agent.deepgram.com` +- Voice Agent WebSocket **and `GET /v1/agent/settings/think/models`**: `https://agent.deepgram.com` -Voice Agent's REST endpoints live on the `agent.` host too, not on `api.`: -`GET /v1/agent/settings/think/models` returns 404 on `api.deepgram.com` and 200 on -`agent.deepgram.com`. Everything else REST stays on `api.deepgram.com`. +`GET /v1/agent/settings/think/models` lives on the `agent.` host too, not on `api.`: it +returns 404 on `api.deepgram.com` and 200 on `agent.deepgram.com`. Everything else REST +stays on `api.deepgram.com`. ### Regional endpoints @@ -54,14 +54,13 @@ unavailable they fail rather than fall back. | `POST /v1/read` | Yes | | `wss://…/v1/agent/converse` | Yes | | `GET /v1/models` | Yes | -| `POST /v1/auth/grant`, `GET /v1/auth/token` | Yes | +| `POST /v1/auth/grant` | Yes | | `/v1/projects/*` (keys, members, usage, billing) | **No — 404** | Two host rules that catch people out: 1. **Voice Agent moves onto the `api.` host regionally.** There is no `agent.eu.deepgram.com` - (the name does not resolve). Use `wss://api.eu.deepgram.com/v1/agent/converse`. The Agent - REST endpoints move with it. Globally it stays on `agent.deepgram.com`. + (the name does not resolve). Use `wss://api.eu.deepgram.com/v1/agent/converse`. `GET /v1/agent/settings/think/models` moves with it. Globally it stays on `agent.deepgram.com`. 2. **Keep management calls on `api.deepgram.com`.** Point a client's management calls at a regional host and `/v1/projects` returns 404, so split the base URL by call type if your app both transcribes and manages keys. @@ -143,7 +142,7 @@ Both model families are actively maintained and industry-leading. They solve dif | Turn detection | Manual (`utterance_end_ms`, VAD events) | Built-in (EOT, eager-EOT, turn_index) | | Transports | REST + WebSocket | WebSocket only | | Intelligence overlays | Yes — `summarize`, `sentiment`, `topics`, `intents`, `diarize_model`, `redact`, etc. | No — smaller focused param set; no `smart_format` / `diarize_model` / `punctuate` | -| Mid-session reconfig | No (reconnect to change) | Yes (`Configure` message updates EOT thresholds + keyterms live) | +| Mid-session reconfig | No (reconnect to change) | Yes (`Configure` message updates EOT thresholds, keyterms, language hints, and `numerals` live) | **Pick Nova (`/v1/listen`, `model=nova-3`) when:** - Generating captions, subtitles, or transcripts for recorded media @@ -155,7 +154,7 @@ Both model families are actively maintained and industry-leading. They solve dif - Building an interactive voice agent or assistant - You want end-of-turn detection handled for you - You need low-latency turn signals and barge-in support -- You want to update EOT thresholds or keyterms mid-session without reconnecting +- You want to update EOT thresholds, keyterms, language hints, or `numerals` mid-session without reconnecting Migrating from Nova 3 to Flux STT? See the official [Nova 3 → Flux migration guide](https://developers.deepgram.com/docs/flux/nova-3-migration). @@ -204,9 +203,9 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// | Listen v2 — STT, Flux STT (conversational) | — | `wss://api.deepgram.com/v2/listen` | [listen.md](references/listen.md) | | Speak v1 — TTS, Aura models | `POST /v1/speak` | `wss://api.deepgram.com/v1/speak` | [speak.md](references/speak.md) | | Speak v2 — TTS, Flux TTS (turn-based) | `POST /v2/speak` | `wss://api.deepgram.com/v2/speak` | [speak.md](references/speak.md) | -| Voice Agent | `GET agent.deepgram.com/v1/agent/settings/think/models` | `wss://agent.deepgram.com/v1/agent/converse` | [agent.md](references/agent.md) | +| Voice Agent | `GET agent.deepgram.com/v1/agent/settings/think/models`; reusable agent configurations at `/v1/projects/{project_id}/agents` (`GET`, `POST`) and `/v1/projects/{project_id}/agents/{agent_id}` (`GET`, `PUT`, `DELETE`); agent variables at `/v1/projects/{project_id}/agent-variables` (`GET`, `POST`) and `/v1/projects/{project_id}/agent-variables/{variable_id}` (`GET`, `PATCH`, `DELETE`) | `wss://agent.deepgram.com/v1/agent/converse` | [agent.md](references/agent.md) | | Read (Intelligence) | `POST /v1/read` | — | [read.md](references/read.md) | -| Models | `GET /v1/models` | — | [models.md](references/models.md) | +| Models | `GET /v1/models`, `GET /v1/models/{model_id}`, `GET /v1/projects/{project_id}/models`, `GET /v1/projects/{project_id}/models/{model_id}`; `include_outdated=true` on either list call also returns non-latest model versions | — | [models.md](references/models.md) | | Projects | `/v1/projects/*` | — | [projects.md](references/projects.md) | | Auth | `POST /v1/auth/grant` | — | [auth.md](references/auth.md) | | Self-Hosted | `/v1/projects/*/self-hosted/*` | — | [self-hosted.md](references/self-hosted.md) | @@ -215,7 +214,7 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// ### All APIs -1. **Feature flags are query params — except for Voice Agent and the v2 mid-session updates.** For `/v1/listen`, `/v2/listen`, `/v1/speak`, and `/v2/speak`, initial options go on the URL. The request body carries only audio data (REST) or audio frames (WebSocket). Exceptions: `/v1/agent/converse` has no URL query params at all (all config goes in the `Settings` message); `/v2/listen` supports a `Configure` message after connection to update EOT thresholds and keyterms mid-session; and `/v2/speak` supports a `Configure` message that updates `speed` only. Also note that `/v2/listen` has a much smaller param set than `/v1/listen` — flags like `smart_format`, `diarize_model`, and `punctuate` are not available. +1. **Feature flags are query params — except for Voice Agent and the v2 mid-session updates.** For `/v1/listen`, `/v2/listen`, `/v1/speak`, and `/v2/speak`, initial options go on the URL. The request body carries only audio data (REST) or audio frames (WebSocket). Exceptions: `/v1/agent/converse` has no URL query params at all (all config goes in the `Settings` message); `/v2/listen` supports a `Configure` message after connection to update EOT thresholds, keyterms, language hints, and `numerals` mid-session; and `/v2/speak` supports a `Configure` message that updates `speed` only. Also note that `/v2/listen` has a much smaller param set than `/v1/listen` — flags like `smart_format`, `diarize_model`, and `punctuate` are not available. 2. **Rate limits are concurrent connections, not total requests.** A 429 means too many simultaneous open connections, not too high a request volume. Diarization and other compute-heavy features reduce your concurrency allowance further. @@ -243,7 +242,7 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// 11. **Streaming is raw audio only, and rejects anything it doesn't recognize.** The WebSocket emits non-containerized audio, so `encoding` is limited to `linear16` (default), `mulaw`, or `alaw`. The compressed and containerized encodings (`mp3`, `opus`, `flac`, `aac`) and the `container`, `bit_rate`, `callback`, `callback_method`, and `priority` params are **batch-only** — sending them to the socket fails the connection, as does any unknown or misspelled param. Use the batch REST transport when you need compressed output. -12. **Insert whitespace between separate generations — the server won't.** Text normalization runs before synthesis, but successive `Speak` messages are concatenated verbatim. Sending `"Hello world."` then `"How are you?"` is processed as `"Hello world.How are you?"`, which causes sentence-boundary artifacts. Add a single space (or the right separator for non-whitespace languages) when you stitch a reply, a tool-call result, and another reply together. Send plain text: SSML and other markup is stripped, with an `INPUT_MARKUP_STRIPPED` warning. +12. **Insert whitespace between separate generations — the server won't.** Text normalization runs before synthesis, but successive `Speak` messages are concatenated verbatim. Sending `"Hello world."` then `"How are you?"` is processed as `"Hello world.How are you?"`, which causes sentence-boundary artifacts. Add a single space (or the right separator for non-whitespace languages) when you stitch a reply, a tool-call result, and another reply together. Send plain text: SSML is not interpreted, and the only markup Flux TTS honors is its own escaped inline controls. A pronunciation override `\{"word":"...","pronounce":""\}` is honored on both transports (Early Access) but only with `speed` 1.0, and a pause marker `\{pause:500ms\}` is batch-only. A pause marker on the socket, or a pronunciation control on a socket whose `speed` is not 1.0, fails the connection with `DATA-0002`. See [Speed, Pause, Pronunciation](https://developers.deepgram.com/docs/tts-voice-controls). ### Voice Agent (`/v1/agent/converse`) @@ -254,19 +253,19 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// { "agent": { "speak": { "provider": { "type": "deepgram", "version": "v2", "model": "flux-alexis-en" } } } } ``` -15. **The Voice Agent REST endpoints live on `agent.deepgram.com`, not `api.deepgram.com`.** `GET /v1/agent/settings/think/models` — the list of LLMs you can name in `agent.think.provider` — returns **404 on `api.deepgram.com`** and 200 on `agent.deepgram.com`. Same key, same path; only the host differs, so a client with one hardcoded base URL silently gets a 404 that looks like a missing feature. The three regional `api.*` hosts serve it as well. +15. **`GET /v1/agent/settings/think/models` lives on `agent.deepgram.com`, not `api.deepgram.com`.** `GET /v1/agent/settings/think/models` — the list of LLMs you can name in `agent.think.provider` — returns **404 on `api.deepgram.com`** and 200 on `agent.deepgram.com`. Same key, same path; only the host differs, so a client with one hardcoded base URL silently gets a 404 that looks like a missing feature. The three regional `api.*` hosts serve it as well. ### Flux STT model (`/v2/listen`) 16. **Use `/v2/listen` and a `flux-general-*` model.** Two are served: `flux-general-en` (English) and `flux-general-multi` (multilingual, and the only model that accepts `language_hint` / `language_hints`). `/v1/listen` does not support Flux STT, and `model=flux` alone is not a valid value. Do not include `language` or `encoding` params for containerized audio. -17. **Use `Configure` to update EOT thresholds and keyterms mid-session.** Unlike `/v1/listen`, Flux STT supports live reconfiguration after connection — no need to reconnect to change turn detection sensitivity or boost new keyterms: +17. **Use `Configure` to update EOT thresholds, keyterms, language hints, and `numerals` mid-session.** Unlike `/v1/listen`, Flux STT supports live reconfiguration after connection — no need to reconnect to change turn detection sensitivity, boost new keyterms, re-bias language detection (`language_hints`, `flux-general-multi` only), or switch `numerals` on for a PIN or order number: ```json - { "type": "Configure", "thresholds": { "eot_threshold": "0.8", "eot_timeout_ms": "3000" }, "keyterms": ["Deepgram"] } + { "type": "Configure", "thresholds": { "eot_threshold": 0.8, "eot_timeout_ms": 3000 }, "keyterms": ["Deepgram"] } ``` - The server responds with `ConfigureSuccess` (echoing back applied values) or `ConfigureFailure`. Omitted threshold fields keep their current values. + The server responds with `ConfigureSuccess` (echoing back applied values) or `ConfigureFailure`, which carries `code` and `description` identifying the rejected configuration. Omitted threshold fields keep their current values. -18. **`ForceEndTurn` outside a turn is a `Warning`, not an error — and the socket stays open.** Sending `{"type":"ForceEndTurn"}` while no turn is in progress returns `{"type":"Warning","code":"FORCE_END_TURN_NO_ACTIVE_TURN","description":"Received ForceEndTurn while no turn was active; the request was ignored."}` and the connection continues. Do not treat it as fatal or reconnect. Neither the `Warning` message nor this code is in the AsyncAPI spec yet, so `references/listen.md` cannot show them. When `ForceEndTurn` *does* land mid-turn, the resulting `TurnInfo` carries `event: "EndOfTurn"` with `trigger: "manual"` — `trigger` is `model` | `manual` | `timeout`, it appears on `EndOfTurn` and nowhere else, and it is an open enum, so tolerate values you do not recognize. +18. **`ForceEndTurn` outside a turn is a `Warning`, not an error — and the socket stays open.** Sending `{"type":"ForceEndTurn"}` while no turn is in progress returns `{"type":"Warning","code":"FORCE_END_TURN_NO_ACTIVE_TURN","description":"Received ForceEndTurn while no turn was active; the request was ignored."}` and the connection continues. Do not treat it as fatal or reconnect. `references/listen.md` shows the message shape (`ListenV2Warning`: `code`, `description`, `request_id`, `sequence_id`); `code` is a free string there, so the individual codes such as `FORCE_END_TURN_NO_ACTIVE_TURN` come from the [Force End Turn](https://developers.deepgram.com/docs/flux/force-end-turn) docs. When `ForceEndTurn` *does* land mid-turn, the resulting `TurnInfo` carries `event: "EndOfTurn"` with `trigger: "manual"` — `trigger` is `model` | `manual` | `timeout`, it appears on `EndOfTurn` and nowhere else, and it is an open enum, so tolerate values you do not recognize. ### Nova diarization (`/v1/listen`) @@ -274,7 +273,7 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// ### Text and Audio Intelligence (`/v1/read`, `/v1/listen`) -20. **`language` is required on `/v1/read`, and it is validated before anything else.** There is no default, despite what `references/read.md` says: omitting it returns `400 INVALID_QUERY_PARAMETER` — "Failed to deserialize query parameters: missing field `language`" — which masks every other problem in the request. English only — `language=multi` is rejected, and `en-US` is accepted but echoed back as `en`. Two more `/v1/read` shapes worth knowing: the JSON body takes **exactly one** of `text` or `url` (both or neither gives `PAYLOAD_ERROR`, and `url` must point at a plain-text document — audio gives `REMOTE_CONTENT_ERROR`), and it is POST-only (`GET` and a WebSocket upgrade both return 405). `summarize` on `/v1/read` accepts `v2` as well as `true`, contrary to the reference. Result paths differ per endpoint: `/v1/read` returns `results.summary.text`, `/v1/listen` returns `results.summary.short`, so code that handles both has to branch. (`sentiment` maps to `results.sentiments` on both.) +20. **`language` is required on `/v1/read`, and it is validated before anything else.** There is no default, despite what `references/read.md` says: omitting it returns `400 INVALID_QUERY_PARAMETER` — "Failed to deserialize query parameters: missing field `language`" — which masks every other problem in the request. English only — `language=multi` is rejected, and `en-US` is accepted but echoed back as `en`. Two more `/v1/read` shapes worth knowing: the JSON body takes **exactly one** of `text` or `url` (both or neither gives `PAYLOAD_ERROR`, and `url` must point at a plain-text document — audio gives `REMOTE_CONTENT_ERROR`), and it is POST-only (`GET` and a WebSocket upgrade both return 405). `summarize` on `/v1/read` accepts `v2` as well as `true`; `references/read.md` types it `v2` | boolean, and only its description still says boolean-only, so trust the type. Result paths differ per endpoint: `/v1/read` returns `results.summary.text`, `/v1/listen` returns `results.summary.short`, so code that handles both has to branch. (`sentiment` maps to `results.sentiments` on both.) 21. **On the Nova streaming socket, only `detect_entities` works — and the other four fail in three different ways.** `detect_entities=true` is supported and puts `entities` at the **top level** of each `Results` message, beside `channel`, not inside `channel.alternatives[0]`. The other four are prerecorded-only: `summarize` fails the handshake with `400 "Summarization is not available for streaming."`; `topics` and `intents` fail it with `403 UNAUTHORIZED_FEATURES_REQUESTED`, which reads like a key-permissions problem even when the same key's prerecorded `topics`/`intents` calls return 200; and `sentiment` is the trap — the handshake succeeds, no error is ever sent, and sentiment simply never appears in the results. diff --git a/skills/api/references/listen.md b/skills/api/references/listen.md index a283621..44285b8 100644 --- a/skills/api/references/listen.md +++ b/skills/api/references/listen.md @@ -252,7 +252,7 @@ for natural voice conversations - `language_hints` string[] — Language hints to constrain and prioritize language detection. Only valid when the model is flux-general-multi. If this field is not supplied, the session will continue to use the currently configured value. - - `numerals` boolean (default: `false`) — Numerals converts numbers from written format to numerical format. Applies to turns transcribed after the update. + - `numerals` boolean (default: `false`) — Numerals converts numbers from written format to numerical format. Applies to transcripts Flux STT sends after it processes the update. #### Server → Client Messages @@ -316,6 +316,7 @@ for natural voice conversations (for example, `keyterm=customer%20service`). Do not separate keyterms with commas, semicolons, or line breaks. - `language_hints` string[] — The currently active language hints. Only applicable to the flux-general-multi model. + - `numerals` boolean (default: `false`) — Whether numeral formatting is enabled for transcripts Flux STT sends after it processes the update. - `sequence_id` integer **(required)** — Starts at `0` and increments for each message the server sends to the client. This includes messages of other types, like `TurnInfo` messages. @@ -327,6 +328,18 @@ for natural voice conversations - `sequence_id` integer **(required)** — Starts at `0` and increments for each message the server sends to the client. This includes messages of other types, like `TurnInfo` messages. + - `code` string — Failure code identifying the rejected configuration + - `description` string — A human-readable description of the configuration failure + +**ListenV2Warning** — Receive a warning; the server keeps the connection open + + - `type` `Warning` **(required)** — Message type identifier + - `request_id` string **(required)** — The unique identifier of the request + - `sequence_id` integer **(required)** — Starts at `0` and increments for each message the server sends + to the client. This includes messages of other types, like + `TurnInfo` messages. + - `code` string **(required)** — Warning code identifying the condition, in `SCREAMING_SNAKE_CASE` + - `description` string **(required)** — A human-readable description of the warning **ListenV2FatalError** — Receive a fatal error message diff --git a/skills/api/references/speak.md b/skills/api/references/speak.md index 7e53952..bb251b9 100644 --- a/skills/api/references/speak.md +++ b/skills/api/references/speak.md @@ -74,19 +74,19 @@ Synthesize a complete block of text into a single audio response using Deepgram' - `expressivity` `-2` | `-1` | `0` | `1` | `2` (default: `0`) — Expressive range of the generated speech, on a calm-to-animated axis. Accepted values: `-2`, `-1`, `0`, `1`, `2`. `0` (the default) is the voice's tuned delivery and the production-validated setting, with `-2` the calm end of the range and `2` the animated end. Supported on all Flux voices; applies to the whole request. Beta: behavior may change in future model versions, and non-default values increase the risk of hallucinations and pronunciation errors; audition before shipping. An invalid value is rejected with a `400` — `EXPRESSIVITY_OUT_OF_RANGE` for a value outside the range, `EXPRESSIVITY_INCREMENT_INVALID` for a fractional value. See [Expressivity](/docs/tts-expressivity). - `model` string **(required)** — Flux TTS model used to synthesize the submitted text, in the form `flux-{voice}-{language}` (for example, `flux-alexis-en`). Required; unlike the v1 (Aura) endpoint there is no default and only flux models are accepted. English-only at launch. - `sample_rate` `8000` | `16000` | `24000` | `32000` | `44100` | `48000` | `8000` | `16000` | `8000` | `16000` | `8000` | `16000` | `22050` | `32000` | `48000` (default: `24000`) — Sample Rate specifies the sample rate for the output audio. Based on the encoding, different sample rates are supported. For some encodings, the sample rate is not configurable -- `speed` number (default: `1`) — Speaking rate multiplier that adjusts the pace of generated speech while preserving natural prosody and voice quality. Accepted values run `0.5` to `1.5` in `0.05` increments. Not yet supported in all languages. +- `speed` number (default: `1`) — Speaking rate multiplier that adjusts the pace of generated speech while preserving natural prosody and voice quality. Accepted values run `0.5` to `1.5` in `0.05` increments. Not yet supported in all languages. When the text contains an inline pause marker, speed is capped at `1.15` (`PAUSE_SPEED_CAP_EXCEEDED` above that). A value other than `1.0` cannot be combined with inline pronunciation controls (`CONTROL_COMBINATION_INVALID`). - `priority` `low` — Processing priority for asynchronous (callback) requests. The only supported value is low. #### Request Body **application/json** -- `text` string **(required)** — The text content to be converted to speech. The server normalizes and preprocesses the text before synthesis. Inline pause and pronunciation controls are not yet applied; they are stripped from the text before synthesis. +- `text` string **(required)** — The text content to be converted to speech. The server normalizes and preprocesses the text before synthesis. May contain inline pause controls (`\{pause:500ms\}`, 500-3000 ms in 100 ms steps, at most 8 per request) and inline pronunciation controls (`\{"word": "...", "pronounce": ""\}`, Early Access). Pronunciation cannot be combined with pause or with a `speed` other than `1.0`, and `speed` is capped at `1.15` when a pause is present. See [Speed, Pause, Pronunciation](/docs/tts-voice-controls). #### Responses **200**: Returns the synthesized audio in the requested encoding as a binary stream. When a `callback` URL is supplied, the request is processed asynchronously and the response body is instead a JSON acknowledgement (Content-Type `application/json`) of the form {"request_id": "..."}, with the audio delivered to the callback URL. Because this endpoint is typed as a binary audio stream, SDK callers that set `callback` receive this JSON acknowledgement through the audio byte iterator as raw bytes and must join the chunks and parse `request_id` themselves. -**400**: Invalid Request. Inline pause and pronunciation controls are not applied and are stripped rather than rejected. +**400**: Invalid Request. Inline control violations return a structured error whose `err_code` names the rule: `CONTROL_COMBINATION_INVALID` (pronunciation combined with speed or pause, or all three together), `PAUSE_SPEED_CAP_EXCEEDED` (a pause marker with `speed` above `1.15`), `BREAK_OUT_OF_RANGE` (a pause outside 500-3000 ms), `BREAK_INCREMENT_INVALID` (a pause off the 100 ms grid), `BREAKS_LIMIT_EXCEEDED` (more than 8 pause markers, or two pauses with no text between them), `BREAK_SYNTAX_INVALID` (a malformed pause marker, such as a simple marker without backslashes or an escaped structured marker), plus the existing pronunciation and speed codes. A `speed` of `1.0` does not count as a speed control for the combination rules. See [Speed, Pause, Pronunciation](/docs/tts-voice-controls). ## WebSocket API @@ -164,7 +164,7 @@ per-turn billing and timing. - `model` string — The Flux TTS model used to synthesize speech. Required on every connection. Model strings follow the format `flux-{voice}-{language}` (e.g. `flux-alexis-en`). An Aura model string is rejected on `/v2/speak`; use `/v1/speak` for Aura voices. - `encoding` `linear16` | `mulaw` | `alaw` (default: `linear16`) — Encoding of the raw output audio. The streaming WebSocket emits raw (non-containerized) audio, so only streaming-compatible encodings are supported. Compressed and containerized encodings (`mp3`, `opus`, `flac`, `aac`) are available on the batch REST transport only. - `sample_rate` `8000` | `16000` | `24000` | `32000` | `44100` | `48000` — Output sample rate in Hz. With `linear16`, valid values are `8000`, `16000`, `24000`, `32000`, `44100`, and `48000`. With `mulaw` or `alaw`, valid values are `8000` and `16000`. Defaults to the model's native sample rate. -- `speed` number (default: `1`) — Speech-rate multiplier. `1.0` is the model's nominal rate; lower is slower. Accepted values run `0.5` to `1.5` in `0.05` increments. A value outside that range is rejected with `SPEED_OUT_OF_RANGE`; a value inside it but off the `0.05` increment with `SPEED_INCREMENT_INVALID`. Models and languages without runtime speed control reject any value with `SPEED_NOT_SUPPORTED`. +- `speed` number (default: `1`) — Speech-rate multiplier. `1.0` is the model's nominal rate; lower is slower. Accepted values run `0.5` to `1.5` in `0.05` increments. A value outside that range is rejected with `SPEED_OUT_OF_RANGE`; a value inside it but off the `0.05` increment with `SPEED_INCREMENT_INVALID`. Models and languages without runtime speed control reject any value with `SPEED_NOT_SUPPORTED`. A speed other than `1.0` cannot be combined with inline pronunciation controls; see [Speed, Pause, Pronunciation](/docs/tts-voice-controls). - `expressivity` `-2` | `-1` | `0` | `1` | `2` (default: `0`) — Expressive range of the generated speech, on a calm-to-animated axis. Accepted values: `-2`, `-1`, `0`, `1`, `2`. `0` (the default) is the voice's tuned delivery and the production-validated setting, with `-2` the calm end of the range and `2` the animated end. Supported on all Flux voices. Fixed for the connection — not settable via `Configure`. Beta: behavior may change in future model versions, and non-default values increase the risk of hallucinations and pronunciation errors; audition before shipping. An invalid value fails the connection with a `400` — `EXPRESSIVITY_OUT_OF_RANGE` for a value outside the range, `EXPRESSIVITY_INCREMENT_INVALID` for a fractional value. See [Expressivity](/docs/tts-expressivity). - `mip_opt_out` boolean (default: `false`) — Opts out requests from the Deepgram Model Improvement Program. Refer to our Docs for pricing impacts before setting this to true. https://dpgr.am/deepgram-mip - `tag` string | string[] — Label your requests for the purpose of identification during usage reporting @@ -174,7 +174,7 @@ per-turn billing and timing. **SpeakV2Speak** — Send text to be synthesized into the active turn - `type` `Speak` **(required)** — Message type identifier - - `text` string **(required)** — The input text to synthesize. Inline pause and pronunciation controls are not yet applied; they are stripped from the text before synthesis. + - `text` string **(required)** — The input text to synthesize. May contain inline pronunciation controls (`\{"word": "...", "pronounce": ""\}`), which are in Early Access. Inline pause controls are supported on the batch (REST) transport only; a pause marker sent over the WebSocket fails the connection with `DATA-0002`. Pronunciation cannot be combined with a `speed` other than `1.0`: text carrying a pronunciation control on a session opened with `speed`, or after a `Configure` that set it, also fails the connection with `DATA-0002`. See [Speed, Pause, Pronunciation](/docs/tts-voice-controls). **SpeakV2Flush** — End the active turn and generate the remaining audio @@ -190,7 +190,7 @@ per-turn billing and timing. **SpeakV2Configure** — Update synthesis configuration mid-session - `type` `Configure` **(required)** — Message type identifier - - `speed` number (default: `1`) — Speech-rate multiplier. `1.0` is the model's nominal rate; lower is slower. Accepted values run `0.5` to `1.5` in `0.05` increments. A value outside that range is rejected with `SPEED_OUT_OF_RANGE`; a value inside it but off the `0.05` increment with `SPEED_INCREMENT_INVALID`. Models and languages without runtime speed control reject any value with `SPEED_NOT_SUPPORTED`. + - `speed` number (default: `1`) — Speech-rate multiplier. `1.0` is the model's nominal rate; lower is slower. Accepted values run `0.5` to `1.5` in `0.05` increments. A value outside that range is rejected with `SPEED_OUT_OF_RANGE`; a value inside it but off the `0.05` increment with `SPEED_INCREMENT_INVALID`. Models and languages without runtime speed control reject any value with `SPEED_NOT_SUPPORTED`. A speed other than `1.0` cannot be combined with inline pronunciation controls; see [Speed, Pause, Pronunciation](/docs/tts-voice-controls). **SpeakV2Close** — Gracefully close the connection, draining all remaining and queued audio @@ -220,7 +220,7 @@ per-turn billing and timing. - `audio_duration_ms` integer **(required)** — Total audio duration produced for this turn, in milliseconds - `input_character_count` integer **(required)** — Raw input character count for this turn, before text normalization - `billable_character_count` integer **(required)** — Billable character count for this turn — the input character count with stripped control characters removed. Always less than or equal to `input_character_count`. - - `controls_applied` { pronunciations_applied: integer, breaks_applied: integer, pronunciation_warnings: integer } **(required)** — Counts of the inline controls the server acted on during the turn. Inline pause and pronunciation controls are not applied at launch — support is coming soon — so every count is currently `0`. + - `controls_applied` { pronunciations_applied: integer, breaks_applied: integer, pronunciation_warnings: integer } **(required)** — Counts of the inline controls the server acted on during the turn. A pronunciation override that triggers an IPA warning is still applied best-effort and counted in `pronunciations_applied`; the warning is reported separately through a `Warning` and `pronunciation_warnings`. **SpeakV2SpeechInterrupted** — Receive what the user heard, and the interrupted turn's billing, after an Interrupt @@ -250,7 +250,7 @@ per-turn billing and timing. **SpeakV2ConfigureFailure** — Receive notice that a Configure was rejected or failed to apply; the prior configuration is retained - `type` `ConfigureFailure` **(required)** — Message type identifier - - `code` `SPEED_OUT_OF_RANGE` | `SPEED_INCREMENT_INVALID` | `SPEED_NOT_SUPPORTED` | `INTERNAL_ERROR` **(required)** — Failure code, in `SCREAMING_SNAKE_CASE`. `SPEED_OUT_OF_RANGE`: outside the range the model publishes. `SPEED_INCREMENT_INVALID`: inside the published range but off the `0.05` increment. `SPEED_NOT_SUPPORTED`: this model or language has no runtime speed control at all. `INTERNAL_ERROR`: the configuration was acceptable but the server could not apply it — unlike the others, a server-side failure rather than a statement about the request. + - `code` `SPEED_OUT_OF_RANGE` | `SPEED_INCREMENT_INVALID` | `SPEED_NOT_SUPPORTED` | `CONTROL_COMBINATION_INVALID` | `INTERNAL_ERROR` **(required)** — Failure code, in `SCREAMING_SNAKE_CASE`. `SPEED_OUT_OF_RANGE`: outside the range the model publishes. `SPEED_INCREMENT_INVALID`: inside the published range but off the `0.05` increment. `SPEED_NOT_SUPPORTED`: this model or language has no runtime speed control at all. `CONTROL_COMBINATION_INVALID`: `speed` was set while a turn buffered behind the active one still carries a pronunciation control; pronunciation and speed cannot be combined, so flush that turn before setting speed (pronunciations in the active turn do not block the change). `INTERNAL_ERROR`: the configuration was acceptable but the server could not apply it — unlike the others, a server-side failure rather than a statement about the request. - `field` `speed` — The configuration field the failure is about. Absent when the failure is not tied to one field. - `value` number — The rejected value for `field`. Absent when there is no offending value to echo — `SPEED_NOT_SUPPORTED` names the field but carries no value, because the rejection is a property of the model. - `description` string **(required)** — A human-readable description of the failure @@ -262,7 +262,9 @@ per-turn billing and timing. Turn-scoped codes: `NO_ACTIVE_SPEECH` (a speech-scoped message arrived with no active turn), `NO_SYNTHESIZABLE_TEXT` (the turn's text was entirely whitespace or punctuation, so it produced no audio and is completed with a zero-duration `SpeechMetadata`), and `SYNTHESIS_RETRYING` (a synthesis request failed and is being retried). - Inline-control codes are reserved and not currently emitted, because inline pause and pronunciation controls are not yet applied: `BREAKS_LIMIT_EXCEEDED` (too many pause controls, or two pauses with no intervening text), `BREAK_TOKENS_OUT_OF_RANGE` (pause durations outside the range the model supports), `BREAK_TOKENS_WITH_INVALID_INCREMENTS` (pause durations off the model's supported increment), `PRONUNCIATION_WARNINGS` (a pronunciation override contained invalid IPA), `PRONUNCIATION_TOO_LONG` (an IPA string exceeded the length limit), `PRONUNCIATIONS_LIMIT_EXCEEDED` (too many pronunciation controls in one turn). + Pronunciation codes: `PRONUNCIATION_WARNINGS` (a pronunciation override contained invalid IPA; it is still applied best-effort and counted in `pronunciations_applied`), `PRONUNCIATION_TOO_LONG` (an IPA string exceeded the length limit), `PRONUNCIATIONS_LIMIT_EXCEEDED` (too many pronunciation controls in one turn). + + Pause codes (`BREAKS_LIMIT_EXCEEDED`, `BREAK_TOKENS_OUT_OF_RANGE`, `BREAK_TOKENS_WITH_INVALID_INCREMENTS`) are reserved and not emitted: inline pause controls are batch-only, and a pause marker on the WebSocket fails the connection instead. Interrupt-scoped codes, each meaning the `Interrupt` was ignored: `NO_AUDIO_GENERATED` (the session has produced no audio yet, so there is nothing to interrupt), `INTERRUPT_IN_PROGRESS` (an earlier `Interrupt` is still being processed — at most one is handled at a time), `INVALID_INTERRUPT_OFFSET` (the `playback_offset` did not advance past the position a prior interrupt established). - `description` string **(required)** — A human-readable description of the warning @@ -270,5 +272,5 @@ per-turn billing and timing. **SpeakV2Error** — Receive a fatal error message followed by a WebSocket close - `type` `Error` **(required)** — Message type identifier - - `code` `MESSAGE-0000` | `DATA-0000` | `DATA-0002` | `BIG-0000` | `NET-0000` | `NET-0001` | `NET-0002` | `NET-0003` | `NET-0004` **(required)** — A code identifying the error, e.g. `MESSAGE-0000` or `NET-0000`. + - `code` `MESSAGE-0000` | `DATA-0000` | `DATA-0002` | `BIG-0000` | `NET-0000` | `NET-0001` | `NET-0002` | `NET-0003` | `NET-0004` **(required)** — A code identifying the error, e.g. `MESSAGE-0000` or `NET-0000`. `DATA-0002` covers invalid inline controls and speed, including an inline pause marker (pause is batch-only) and a pronunciation control combined with a `speed` other than `1.0`; `description` names the specific rule. - `description` string **(required)** — Prose description of the error From 7a955a8016c3dfda69e46fd98f46f47740970c56 Mon Sep 17 00:00:00 2001 From: Corey Weathers Date: Thu, 1 Oct 2026 09:45:25 -0400 Subject: [PATCH 02/17] docs(text-to-speech): Flux TTS inline controls, Interrupt offsets, NET-0003, Go SDK Flux TTS client --- skills/text-to-speech/SKILL.md | 56 ++++++++++++++++++++++++++-------- 1 file changed, 44 insertions(+), 12 deletions(-) diff --git a/skills/text-to-speech/SKILL.md b/skills/text-to-speech/SKILL.md index bb135ef..52c5281 100644 --- a/skills/text-to-speech/SKILL.md +++ b/skills/text-to-speech/SKILL.md @@ -75,26 +75,54 @@ A session is a sequence of turns. Stream tokens in, then end the turn: - Audio starts streaming before you `Flush`. `Flush` ends the turn; the server then sends `Flushed` and `SpeechMetadata` with billing and timing. Treat `SpeechMetadata` as the end of the turn; `Flushed` arrives earlier. - The server assigns `speech_id` per turn in `SpeechStarted` and `SpeechMetadata`. Never send one. +- The first message is `Connected`, with `request_id`, `model_name`, `model_version`, and `model_uuids`. + The last is `SessionMetadata`, with cumulative session totals; an `Interrupt` rebases its + `total_audio_duration_ms` onto the audio the client actually played. - On barge-in, stop local playback first, then send `{"type":"Interrupt","playback_offset":{"type":"time_ms","value":2340}}`. `SpeechInterrupted` returns `text_spoken` and `text_remaining`; feed `text_spoken` back into the LLM context. Without a - `playback_offset` the split is omitted. + `playback_offset` the split is omitted. `value` is milliseconds played since the session started, not + since this turn, and each `Interrupt` must exceed the previous offset or it is ignored with + `INVALID_INTERRUPT_OFFSET`. Read `audio_played_ms` from `SpeechInterrupted` as the baseline for the + next offset. - `{"type":"Configure","speed":1.15}` changes speed mid-session. `speed` runs `0.5` to `1.5` in `0.05` increments, default `1.0`. `0.45` and `1.55` return `'speed' must be between 0.5 and 1.5`, and `1.07` returns `'speed' must be provided in increments of 0.05`. Errors: - `SPEED_OUT_OF_RANGE`, `SPEED_INCREMENT_INVALID`, `SPEED_NOT_SUPPORTED`. + `SPEED_OUT_OF_RANGE`, `SPEED_INCREMENT_INVALID`, `SPEED_NOT_SUPPORTED`, and + `CONTROL_COMBINATION_INVALID`, which means a queued turn still carries a pronunciation control; + `Flush` that turn first. - `expressivity` runs `-2` (calm) to `2` (animated), default `0`. Values must be whole numbers; a fractional value returns `EXPRESSIVITY_INCREMENT_INVALID` and an out-of-range one `EXPRESSIVITY_OUT_OF_RANGE`. It is beta, fixed per connection (`Configure` cannot change it), and only `0` is validated for production. +- Pronunciation control (Early Access) is an escaped JSON object in the text, on the socket and on + batch: `\{"word":"dupilumab","pronounce":"duːˈpɪljuːmæb"\}`, at most 500 per request, IPA at most + 128 characters. It works only with `speed` `1.0`: on the socket, a pronunciation sent on a session + opened with another speed, or after a `Configure` that set one, fails the connection with + `DATA-0002`; on batch the request is a 400 `CONTROL_COMBINATION_INVALID`. Invalid IPA is still + applied best-effort and reported as a `PRONUNCIATION_WARNINGS` warning on the socket and in the + `dg-warnings` header on batch. Each turn's `SpeechMetadata.controls_applied` counts + `pronunciations_applied`, `breaks_applied`, and `pronunciation_warnings`; batch returns + `dg-pronunciations-applied` and `dg-breaks-applied` response headers. Syntax and IPA guidance: + https://developers.deepgram.com/docs/tts-voice-controls +- Pause control `\{pause:500ms\}` is batch only: 500 to 3000 ms in 100 ms steps, at most 8 per request, + text between adjacent pauses, and `speed` capped at `1.15` while a pause is present + (`PAUSE_SPEED_CAP_EXCEEDED`). A pause marker on the socket fails the connection with `DATA-0002`, and + a pause combined with a pronunciation is rejected with `CONTROL_COMBINATION_INVALID`. Other batch 400 + codes: `BREAK_OUT_OF_RANGE`, `BREAK_INCREMENT_INVALID`, `BREAKS_LIMIT_EXCEEDED`, and + `BREAK_SYNTAX_INVALID` (a marker without the backslashes). - The socket emits raw `linear16` (default), `mulaw`, or `alaw`. Batch-only parameters (`container`, `bit_rate`, `callback`, `callback_method`, `priority`) and any unknown parameter fail the connection. -- Idle sessions close after 60 seconds (`NET-0004`). Send a WebSocket Ping between quiet turns. +- Idle sessions close after 60 seconds (`NET-0004`); send a WebSocket Ping between quiet turns. Every + session closes at 1 hour (`NET-0003`). - Batch: `POST https://api.deepgram.com/v2/speak?model=flux-haley-en` with `{"text": "..."}` returns one - audio response, `mp3` by default, and accepts `opus`, `flac`, `aac`, `container`, `bit_rate`. -- SDKs: every Deepgram SDK except Go ships a Flux TTS client. Python, JavaScript, and Java name it - `speak.v2`; .NET ships `FluxSpeakRESTClient` and `FluxSpeakWebSocketClient`; Rust ships - `speak::flux`. In Go, use the WebSocket directly. + audio response, `mp3` by default, accepts `opus`, `flac`, `aac`, `container`, `bit_rate`, and is the + only transport that honors inline pauses. +- SDKs: every Deepgram SDK ships a Flux TTS client. Python, JavaScript, and Java name it `speak.v2`; + .NET ships `FluxSpeakRESTClient` and `FluxSpeakWebSocketClient`; Rust ships `speak::flux`; Go ships + `pkg/client/speak/v2` from v3.8.0. Only the Rust and .NET SDK skills document Flux TTS; the JS, + Python, Java, and Go `text-to-speech` SDK skills cover `/v1/speak` only, so take `/v2/speak` message + shapes from this skill. ## Voices @@ -125,11 +153,14 @@ quote figures from memory. 5. Asking a WebSocket for `mp3`. Streaming is raw audio on both endpoints. Use REST for compressed output. 6. Dropping the space between LLM generations on Flux TTS. `Speak` texts are concatenated verbatim, so `"Hello world."` then `"How are you?"` becomes `"Hello world.How are you?"`. Insert a space when you - stitch a reply, a tool result, and another reply. SSML is stripped with an `INPUT_MARKUP_STRIPPED` - warning; send plain text. + stitch a reply, a tool result, and another reply. SSML is not supported; the only markup Flux TTS + interprets is its own escaped controls, pronunciation on both transports and pause on batch. Aura-2 + pronunciation control is GA on `/v1/speak` with the same syntax, for English and Spanish voices, with + a 2000-character input limit; Aura-2 has no pause control. 7. Pointing a Voice Agent at api.deepgram.com. The Voice Agent API lives at `wss://agent.deepgram.com` and picks the TTS family from `agent.speak.provider.version`: `v2` for Flux TTS, `v1` for Aura. - Omitting `agent.speak` gives Flux TTS with `flux-kit-en`. + Omitting `agent.speak` gives Flux TTS with `flux-kit-en`. Speed and `expressivity` inside an agent go + on `agent.speak.provider`; see https://developers.deepgram.com/docs/voice-agent-tts-controls. 8. Using Aura `speed` on a German, French, Dutch, Italian, or Japanese voice. Aura-2 speed control covers English and Spanish only. @@ -146,7 +177,8 @@ quote figures from memory. - You want the docs inside your coding tool: `setup-mcp` skill. - You want idiomatic code in one language: the `deepgram-{js,python,java,go,rust,dotnet}-text-to-speech` skills from the SDK repositories (`npx skills add deepgram/deepgram-python-sdk`, and so on). Every SDK - but Go carries a Flux TTS client; see the SDK note above for what each one calls it. + carries a Flux TTS client (Go from v3.8.0); only the Rust and .NET SDK skills document it, so pair the + others with this skill for `/v2/speak`. - You want Deepgram to run speech-to-text, the LLM, and TTS in one connection: `voice-agent` skill and `deepgram-{lang}-voice-agent`. - You are transcribing rather than synthesizing: `speech-to-text` skill. Note that "Flux" names both a @@ -159,4 +191,4 @@ All pages fetched September 2026 as Markdown (append `.md` to any URL); index at - Aura: https://developers.deepgram.com/docs/text-to-speech https://developers.deepgram.com/docs/streaming-text-to-speech https://developers.deepgram.com/docs/tts-models https://developers.deepgram.com/docs/tts-voice-controls https://developers.deepgram.com/docs/tts-encoding https://developers.deepgram.com/docs/tts-media-output-settings https://developers.deepgram.com/docs/tts-ws-flush https://developers.deepgram.com/docs/tts-ws-clear - Flux TTS: https://developers.deepgram.com/docs/flux-tts/overview https://developers.deepgram.com/docs/flux-tts/quickstart https://developers.deepgram.com/docs/flux-tts/batch https://developers.deepgram.com/docs/flux-tts/batch-vs-streaming https://developers.deepgram.com/docs/flux-tts/voices https://developers.deepgram.com/docs/flux-tts/client-messages https://developers.deepgram.com/docs/flux-tts/server-messages https://developers.deepgram.com/docs/flux-tts/interrupt-handling https://developers.deepgram.com/docs/flux-tts/migrating https://developers.deepgram.com/docs/flux-tts/voice-agent https://developers.deepgram.com/docs/flux-tts/template-apps https://developers.deepgram.com/docs/tts-expressivity - API reference: https://developers.deepgram.com/reference/text-to-speech/speak-request https://developers.deepgram.com/reference/text-to-speech/speak-streaming https://developers.deepgram.com/reference/text-to-speech/speak-flux https://developers.deepgram.com/reference/speak/v-2/audio/generate https://developers.deepgram.com/reference/manage/models/list -- Auth, errors, agent, pricing: https://developers.deepgram.com/reference/authentication https://developers.deepgram.com/reference/auth/tokens/grant https://developers.deepgram.com/docs/errors https://developers.deepgram.com/docs/voice-agent-tts-models https://developers.deepgram.com/reference/voice-agent/voice-agent https://deepgram.com/pricing +- Auth, errors, agent, pricing: https://developers.deepgram.com/reference/authentication https://developers.deepgram.com/reference/auth/tokens/grant https://developers.deepgram.com/docs/errors https://developers.deepgram.com/docs/voice-agent-tts-models https://developers.deepgram.com/docs/voice-agent-tts-controls https://developers.deepgram.com/reference/voice-agent/voice-agent https://deepgram.com/pricing From 931216f847f8da7783fb4be2fa15174895f1aa2e Mon Sep 17 00:00:00 2001 From: Corey Weathers Date: Thu, 1 Oct 2026 09:45:25 -0400 Subject: [PATCH 03/17] docs(speech-to-text): numerals on Configure, redact values, SDK sendConfigure; audio-intelligence: entities only on is_final --- skills/audio-intelligence/SKILL.md | 6 ++++-- skills/speech-to-text/SKILL.md | 14 ++++++++------ 2 files changed, 12 insertions(+), 8 deletions(-) diff --git a/skills/audio-intelligence/SKILL.md b/skills/audio-intelligence/SKILL.md index d735bc9..fa4e458 100644 --- a/skills/audio-intelligence/SKILL.md +++ b/skills/audio-intelligence/SKILL.md @@ -67,8 +67,10 @@ with `model_uuid`, `input_tokens`, and `output_tokens`. Entity labels come back (`NAME`, `ORGANIZATION`, `LOCATION_CITY`, `MONEY`, `DATE_INTERVAL`); Deepgram documents over 50 types. [6] -On the live socket, `detect_entities=true` adds a **top-level** `entities` array to each `Results` -message, next to `channel` — not inside `channel.alternatives[0]`. Same field shape as above. +On the live socket, `detect_entities=true` adds a **top-level** `entities` array to `Results` +messages whose `is_final` is `true`, next to `channel` and not inside `channel.alternatives[0]`. +Interim results carry no `entities` key, and a final result with nothing detected carries +`"entities": []`. Same field shape as above. ## Narrowing topics and intents diff --git a/skills/speech-to-text/SKILL.md b/skills/speech-to-text/SKILL.md index 0d53aa1..2772771 100644 --- a/skills/speech-to-text/SKILL.md +++ b/skills/speech-to-text/SKILL.md @@ -21,7 +21,7 @@ Deepgram transcribes audio with two model families on two endpoints. Pick the fa | Endpoint | `/v1/listen`, REST and WebSocket | `/v2/listen`, WebSocket only | | Output | A transcript stream | `TurnInfo` events carrying turn state and a transcript per turn | | Turn detection | None built in; you use endpointing and your own logic | Built in: `StartOfTurn`, `EagerEndOfTurn`, `TurnResumed`, `EndOfTurn` | -| Formatting and analysis | `smart_format`, `diarize_model`, `summarize`, `sentiment`, `topics`, `intents`, redaction | Word timestamps, `numerals`, number redaction, `keyterm`; no smart formatting, no diarization | +| Formatting and analysis | `smart_format`, `diarize_model`, `summarize`, `sentiment`, `topics`, `intents`, redaction | Word timestamps, `numerals`, number redaction (`redact=numbers` or `redact=aggressive_numbers`; any other `redact` value fails the handshake with 400), `keyterm`; no smart formatting, no diarization | | Language | `language=`, or `language=multi` for code-switching | The model name selects the language; `language_hint` biases `flux-general-multi` | Decision rule: @@ -69,7 +69,7 @@ The server sends `Connected`, then a stream of `TurnInfo` messages. Each carries - `EndOfTurn`: the speaker finished. Send the transcript to your language model. It carries `trigger`: `model`, `manual`, or `timeout`. - `EagerEndOfTurn` and `TurnResumed`: emitted only when you set `eager_eot_threshold`. Start drafting a reply on the first; cancel it on the second. -Three query parameters tune turn detection, and all three can change mid-stream: +Three query parameters tune turn detection, and all three can change mid-stream through the `Configure` message below: | Parameter | Range | Default | Effect | |---|---|---|---| @@ -79,11 +79,11 @@ Three query parameters tune turn detection, and all three can change mid-stream: Client control messages, each a JSON text frame: -- `{"type":"Configure","thresholds":{"eot_threshold":0.8},"keyterms":["Deepgram"]}` changes thresholds, keyterms, or `language_hints` without reconnecting. Omitted fields keep their values; a `keyterms` array replaces the whole list. The reply is `ConfigureSuccess` or `ConfigureFailure`. -- `{"type":"ForceEndTurn"}` (added August 28, 2026) ends the current turn on your own signal: a push-to-talk release, a DTMF tone, a send button. Flux emits `EndOfTurn` with `"trigger":"manual"`. With no active turn the message is ignored and a `Warning` with code `FORCE_END_TURN_NO_ACTIVE_TURN` comes back. Set `eot_threshold=1.0` to drive every turn yourself. +- `{"type":"Configure","thresholds":{"eot_threshold":0.8},"keyterms":["Deepgram"],"numerals":true}` changes thresholds, keyterms, `language_hints`, or `numerals` without reconnecting. Omitted fields keep their values; a `keyterms` array replaces the whole list. `numerals` starts from the `numerals` query parameter and applies to transcripts Flux STT sends after it processes the update, never to transcripts already sent. It must be a JSON boolean: the string `"true"` fails schema validation, Flux STT returns an `Error` with code `UNPARSABLE_CLIENT_MESSAGE`, and the connection closes. The reply is `ConfigureSuccess`, which echoes the full active configuration including `numerals`, or `ConfigureFailure`, which carries `code` and `description` naming the rejected field and leaves the previous configuration in place. +- `{"type":"ForceEndTurn"}` (added August 28, 2026) ends the current turn on your own signal: a push-to-talk release, a DTMF tone, a send button. Flux STT emits `EndOfTurn` with `"trigger":"manual"`. With no active turn the message is ignored and a `Warning` with code `FORCE_END_TURN_NO_ACTIVE_TURN` comes back. Set `eot_threshold=1.0` to drive every turn yourself. - `{"type":"CloseStream"}` closes the stream. -For non-English or mixed-language calls use `model=flux-general-multi`, optionally with repeated `language_hint=` parameters (for example `language_hint=en&language_hint=es`). Without hints the model detects the language itself. `TurnInfo` then includes `languages` and `languages_hinted`. +For non-English or mixed-language calls use `model=flux-general-multi`, optionally with repeated `language_hint=` parameters (for example `language_hint=en&language_hint=es`). Without hints the model detects the language itself. `TurnInfo` then includes `languages` and `languages_hinted`. `numerals` formats every number on `flux-general-en`; on `flux-general-multi` it formats English, Spanish, French, German, Russian, Portuguese, Italian, and Dutch, and leaves Hindi and Japanese numbers as spoken. ## Common mistakes @@ -107,7 +107,7 @@ Deepgram bills speech-to-text per minute of audio. Figures change, so read them - You want a runnable app with a UI: `starters` skill (the `transcription`, `live-transcription`, and `flux` features). - You want a one-feature snippet under 50 lines: `recipes` skill, https://github.com/deepgram/recipes. - You are wiring Deepgram into Twilio, LiveKit, Pipecat, LangChain, or another platform: `examples` skill. -- You want language-idiomatic SDK code: install `deepgram-{js,python,java,go,rust,dotnet}-speech-to-text` for Nova and `deepgram-{lang}-conversational-stt` for Flux from the matching SDK repository (`npx skills add deepgram/deepgram-python-sdk`, and so on). +- You want language-idiomatic SDK code: install `deepgram-{js,python,java,go,rust,dotnet}-speech-to-text` for Nova and `deepgram-{lang}-conversational-stt` for Flux STT from the matching SDK repository (`npx skills add deepgram/deepgram-python-sdk`, and so on). Mid-stream `numerals` through `Configure` is in the JavaScript SDK from 5.13.0 (`socket.sendConfigure({type:"Configure", numerals:true})`), the Python SDK from 7.11.0 (`connection.send_configure(ListenV2Configure(numerals=True))`), and the Java SDK from 0.10.2 (`sendConfigure(ListenV2Configure.builder().numerals(true).build())`). The `deepgram-{lang}-conversational-stt` skills do not cover it, so take the message shape from the `Configure` bullet in this skill. - You want analysis and not just the transcript (`summarize`, `sentiment`, `topics`, `intents`, `detect_entities` on `/v1/listen`): `audio-intelligence` skill. For text you already have, `/v1/read` and the `text-intelligence` skill. - You want text-to-speech or a full voice agent: the `text-to-speech` or `voice-agent` skill. - You want a shell command rather than application code: `cli` skill. @@ -126,7 +126,9 @@ Deepgram bills speech-to-text per minute of audio. Figures change, so read them - Flux compared with Nova-3: https://developers.deepgram.com/docs/flux/flux-nova-3-comparison - Nova-3 to Flux migration: https://developers.deepgram.com/docs/flux/nova-3-migration - Flux state machine: https://developers.deepgram.com/docs/flux/state +- Flux STT turn-detection parameters (the threshold table): https://developers.deepgram.com/docs/flux/configuration - Flux control messages: https://developers.deepgram.com/docs/flux/configure, https://developers.deepgram.com/docs/flux/force-end-turn, https://developers.deepgram.com/docs/flux/close-stream +- Numerals, including the Flux STT language list and mid-stream toggling: https://developers.deepgram.com/docs/numerals - ForceEndTurn release note: https://developers.deepgram.com/changelog/2026/8/28 - Flux multilingual: https://developers.deepgram.com/docs/flux/language-prompting - Authentication: https://developers.deepgram.com/guides/fundamentals/authenticating and https://developers.deepgram.com/guides/fundamentals/token-based-authentication From 1145da1d44869406100b7c97dd0889a115ba4b40 Mon Sep 17 00:00:00 2001 From: Corey Weathers Date: Thu, 1 Oct 2026 09:45:25 -0400 Subject: [PATCH 04/17] docs(voice-agent): reusable agent configurations, FunctionCallCancelled, defer_until_eot, speak speed and expressivity, ForceEndTurn warnings --- skills/voice-agent/SKILL.md | 31 ++++++++++++++++++++++++------- 1 file changed, 24 insertions(+), 7 deletions(-) diff --git a/skills/voice-agent/SKILL.md b/skills/voice-agent/SKILL.md index 9d952c5..5f1544d 100644 --- a/skills/voice-agent/SKILL.md +++ b/skills/voice-agent/SKILL.md @@ -7,6 +7,7 @@ description: > updates, and function calling that your own client executes. Use when someone says "voice agent", "voice bot", "speech-to-speech", "talk to an AI on the phone", "agent.deepgram.com", "Settings message", "FunctionCallRequest", "barge-in", + "reusable agent configuration", "defer_until_eot", "Twilio voice agent", or asks whether to build on Deepgram directly or through LiveKit Agents, Pipecat, Vapi, or Retell. Routes to the api, docs, starters, recipes, examples, and per-language SDK skills for the full reference. @@ -26,7 +27,7 @@ Both paths are supported; Deepgram publishes guides for LiveKit Agents and Pipec | A Deepgram-managed LLM (OpenAI, Anthropic, Google, NVIDIA) billed through your Deepgram account is fine, or you point `think.endpoint` at your own OpenAI-compatible endpoint. [6] | You need per-stage control the agent does not expose: your own LLM loop, a TTS vendor Deepgram does not proxy, custom voice activity detection, or your own turn logic. | | Your tools can run in your client or behind an HTTP endpoint you own (`FunctionCallRequest` / `FunctionCallResponse`). [9][10] | Your tools live inside the framework's agent runtime. | -For the orchestrator path, load the `examples` skill (LiveKit, Pipecat) and the SDK `conversational-stt` and `text-to-speech` skills. Deepgram's Pipecat guide runs Flux STT and Flux TTS (`flux-alexis-en`) underneath; the LiveKit guide starts on `nova-3` and `aura-2-thalia-en` and shows `flux-general-en` and `flux-alexis-en` as the Flux swap. [14] The rest of this skill covers the Voice Agent API path. +For the orchestrator path, load the `examples` skill (LiveKit, Pipecat) and the SDK `conversational-stt` and `text-to-speech` skills. Deepgram's Pipecat guide runs Flux STT and Flux TTS (`flux-alexis-en`) underneath; the LiveKit guide starts on `nova-3` and `aura-2-thalia-en` and shows `flux-general-en` and `flux-alexis-en` as the Flux STT and Flux TTS swap. [14] The rest of this skill covers the Voice Agent API path. ## First request @@ -69,10 +70,18 @@ Then open the WebSocket to `wss://agent.deepgram.com/v1/agent/converse` with the Field notes, from the configure and model pages [3][6][7][8]: -- `listen`: Flux (`flux-general-en`, or `flux-general-multi` with `language_hints`) requires `"version": "v2"` and gives model-integrated end-of-turn detection. Nova (`nova-3`) uses `v1`, the default, and adds `smart_format` and `language`. Drop `version` with a Flux model and the agent falls back to the v1 endpoint, where `flux-general-en` is not a valid model. [7][20] -- `think`: `provider.type` is `open_ai`, `anthropic`, `google`, or `nvidia` (managed; `endpoint` optional) or `groq` or `aws_bedrock` (`endpoint` required). Bring your own LLM by keeping `type: open_ai` and setting `endpoint.url` to any OpenAI Chat Completions-compatible URL, with `endpoint.headers` for its auth. Pass an array of providers to get an ordered fallback chain. Managed-LLM prompts are limited to 25,000 characters. [6] -- `speak`: `"version": "v2"` selects Flux TTS (`flux-{voice}-{language}`); `v1`, the default when you name a provider, selects Aura (`aura-2-thalia-en`). Omit `agent.speak` entirely and you get Flux TTS with `flux-kit-en`. Flux TTS streams raw audio only: `encoding` must be `linear16`, `mulaw`, or `alaw`, `container` must be `none`, and `mp3` or `wav` returns `INVALID_SETTINGS`. Third-party TTS (`open_ai`, `eleven_labs`, `cartesia`, `aws_polly`) takes an `endpoint`, except Deepgram-managed Cartesia, which needs none. [8] -- `agent.context.messages` replays earlier turns as `{"type":"History","role":"user","content":"..."}` so a new session continues an old one. [3] +- `listen`: Flux STT (`flux-general-en`, or `flux-general-multi` with `language_hints`) requires `"version": "v2"` and gives model-integrated end-of-turn detection. Nova (`nova-3`) uses `v1`, the default, and adds `smart_format` and `language`. Drop `version` with a Flux STT model and the agent falls back to the v1 endpoint, where `flux-general-en` is not a valid model. [7][20] +- `think`: `provider.type` is `open_ai`, `anthropic`, `google`, or `nvidia` (managed; `endpoint` optional) or `groq` or `aws_bedrock` (bring your own; `endpoint` required). `nvidia` is on the LLM models page but not in the `references/agent.md` provider enum, so check its model name against the models endpoint above. `aws_bedrock` authenticates with `provider.credentials` (`type` `iam`, or `sts` plus `session_token`, with `region`, `access_key_id`, and `secret_access_key`) and points `endpoint.url` at `https://bedrock-runtime.{region}.amazonaws.com/`. Bring your own LLM by keeping `type: open_ai` and setting `endpoint.url` to any OpenAI Chat Completions-compatible URL, with `endpoint.headers` for its auth. Pass an array of providers to get an ordered fallback chain. Managed-LLM prompts are limited to 25,000 characters. [6] +- `speak`: `"version": "v2"` selects Flux TTS (`flux-{voice}-{language}`); `v1`, the default when you name a provider, selects Aura (`aura-2-thalia-en`). Omit `agent.speak` entirely and you get Flux TTS with `flux-kit-en`. Flux TTS streams raw audio only: `encoding` must be `linear16`, `mulaw`, or `alaw`, `container` must be `none`, and `mp3` or `wav` returns `INVALID_SETTINGS`. `provider.speed` (default `1.0`) is `0.5` to `1.5` in `0.05` steps on Flux TTS and any value from `0.7` to `1.5` on Aura; a value the family does not accept ends the session with `FAILED_TO_SPEAK`. `provider.expressivity` (whole numbers `-2` to `2`, default `0`) is Flux TTS (`v2`) only and fixed for the session; it is beta and `0` is the only value validated for production. Third-party TTS (`open_ai`, `eleven_labs`, `cartesia`, `aws_polly`) takes an `endpoint` with `url` and `headers`, and `wss` URLs are accepted for Eleven Labs only; `aws_polly` also requires `credentials` (`type` `sts` or `iam`, with `region`, `access_key_id`, `secret_access_key`, and `session_token` for STS). Deepgram-managed Cartesia (`type: cartesia` with no `endpoint`) is the exception. [8][34] +- `agent.context.messages` replays earlier turns as `{"type":"History","role":"user","content":"..."}` or `{"type":"History","function_calls":[{"id","name","client_side","arguments","response"}]}` so a new session continues an old one. While `Settings.flags.history` is `true` (the default) the server sends `History` messages in the same two shapes; set it to `false` to turn them off. [3][35] +- Other knobs, with ranges in `references/agent.md`: `think.context_length` (`max` or a character count; custom `think.endpoint` only), `think.provider.reasoning_mode` (`none` to `high`, on `open_ai` and `groq`), and the top-level `tags`, `experimental`, and `mip_opt_out`. [3][12] +- Conversational Mode (`agent.think_conversational.provider`: backchanneling, presence checks, frustration detection) is invite-only Early Access; a `Settings` message that names it from an unenrolled project is rejected with `UNPARSABLE_CLIENT_MESSAGE`, and its page is not in the documentation index. [36] + +## Reusable agent configurations + +`Settings.agent` is either the full `agent` object above or a Reusable Agent Configuration UUID string, the same `agent: "YOUR_AGENT_ID"` form the Browser Agent SDK takes. Create one with `POST https://api.deepgram.com/v1/projects/{project_id}/agents` and a body whose `config` is the JSON string of the `agent` block (plus optional `metadata`); the response's `agent_id` is the UUID. `GET .../agents` lists them, `GET .../agents/{agent_id}` reads one, `PUT .../agents/{agent_id}` changes `metadata` only (`config` is immutable; delete and recreate to change it), and `DELETE .../agents/{agent_id}` removes it. Deleting a configuration that a running service still references breaks that service, so move its sessions to a new UUID first. A `Settings` message with a UUID that does not resolve ends the session with `INVALID_AGENT_ID`; `AGENT_ID_NOT_SUPPORTED` means the server does not resolve UUIDs at all (a self-hosted build in unauthenticated mode). [32][13] + +Template variables, at `POST/GET /v1/projects/{project_id}/agent-variables` and `GET/PATCH/DELETE .../agent-variables/{variable_id}`, hold values a `config` references by key in the `DG_` form (uppercase letters, digits, `_`, `-`), written unquoted inside the JSON string. A variable can stand in for any JSON value, a whole provider object included, and `is_sensitive` must be `false`. Every project member can read configurations and variables, so keep API keys and passwords out of them. Full request and response schemas: `references/agent.md`. [32] ## Message lifecycle @@ -84,6 +93,7 @@ Field notes, from the configure and model pages [3][6][7][8]: | server | `ConversationText` (`role` is `user` or `assistant`, `content`) | Show the transcript. [11] | | server | `AgentThinking` (`content`) | Optional status. The LLM is working, possibly choosing a function. [11] | | server | `FunctionCallRequest` | See the next section. [9] | +| server | `FunctionCallCancelled` (`functions[]` with `id`, `name`) | The user started speaking again. Stop work on each `id` and do not send a `FunctionCallResponse` for it; a late one is dropped. [33] | | server | `AgentStartedSpeaking` | The reply's audio is starting. [12] | | server | `LatencyReport` | Per-turn latency breakdown, sent automatically after each turn: `stt_latency`, `ttt_token_latency`, `ttt_text_latency`, `ttt_tool_latency`, `ttt_thinking_latency`, `tts_latency`, `total_latency`. All are floats in seconds and each is optional, so read them defensively. [31] | | server | binary frames | Agent audio. Queue it for playback. [5] | @@ -95,7 +105,7 @@ Mid-call updates, each acknowledged by a matching `*Updated` event [16]: - `UpdatePrompt` `{"type":"UpdatePrompt","prompt":"..."}` adds to the current prompt; it does not replace it. Ack: `PromptUpdated`. [16] - `UpdateSpeak` `{"type":"UpdateSpeak","speak":{"provider":{...}}}` changes the voice. With Flux TTS the new voice starts on the next turn. Ack: `SpeakUpdated`. [16] -- `UpdateListen` adjusts Flux end-of-turn thresholds, keyterms, and language hints. `UpdateThink` replaces the whole think block, functions included. Acks: `ListenUpdated`, `ThinkUpdated`. `ForceEndTurn` ends the user's turn now and needs a Flux (`v2`) listen provider. [16][17] +- `UpdateListen` changes the listen `model` and `language` mid-session and, on a Flux STT (`v2`) provider, the end-of-turn thresholds and language hints; keyterms update mid-session on Flux STT models only. `UpdateThink` replaces the whole think block, functions included. Acks: `ListenUpdated`, `ThinkUpdated`. `ForceEndTurn` ends the user's turn now and needs a Flux STT (`v2`) listen provider: with any other listen provider the server sends a `FORCE_END_TURN_UNSUPPORTED` warning and the turn does not end, and with no turn in progress it is ignored silently. [16][17] - `InjectAgentMessage` `{"type":"InjectAgentMessage","message":"...","behavior":"default"}` makes the agent speak. `default` and `queue` are refused with `InjectionRefused` while the user is speaking; `queue` waits behind the agent's own turn; only `interrupt` is never refused. `InjectUserMessage` `{"type":"InjectUserMessage","content":"..."}` sends typed user text. [18][5] ## Function calling: your client runs the call @@ -107,6 +117,8 @@ The server sends one `FunctionCallRequest` with a `functions` array. Each item h - `client_side: true`: run the function, then send `{"type":"FunctionCallResponse","id":"","name":"get_weather","content":""}`. Pass `thought_signature` back unchanged when present. The agent speaks once your response arrives. [9][10] - `client_side: false`: the server ran it (an `endpoint` function). No client action; the server's own `FunctionCallResponse` is informational. [10] +Calls dispatch speculatively. The agent starts thinking as soon as speech-to-text is moderately confident the user has stopped, and a function call goes out the moment the LLM emits it, before the turn is confirmed. If the user keeps talking, the turn resumes and the server sends `FunctionCallCancelled` for every call you already received. Set `defer_until_eot: true` on a function whose side effect cannot be undone (ending a call, booking, charging a card): a deferred call is held until the turn is confirmed and discarded if the turn resumes, and deferring one function does not delay the others. An `endpoint` function that already reached your server is not rolled back, which is the reason to defer rather than rely on cancellation. [33] + During a slow call, send `InjectAgentMessage` with `behavior: "queue"` ("One moment while I look that up"). [18] A call to a name you did not define ends the session with `NON_EXISTENT_FUNCTION_CALLED`. [13] ## Telephony @@ -138,7 +150,7 @@ The Voice Agent API is billed per minute of WebSocket connection time, and a Dee - You want a minimal snippet for one feature (`connect`, `custom-llm`, `custom-tts`, `function-calling`): `recipes` skill. [29] - You are wiring a third-party platform (Twilio, LiveKit, Pipecat, Vonage, SignalWire, CrewAI, OpenAI Agents SDK): `examples` skill. [21] - The agent runs in a browser: `browser-agent` skill, for the four Browser Agent SDK packages on npm (`@deepgram/agents`, `@deepgram/react`, `@deepgram/ui`, `@deepgram/agents-widget`). They wrap the same socket this skill documents, including the `Sec-WebSocket-Protocol` token handshake above. -- You want language-idiomatic code: `deepgram-js-voice-agent`, `deepgram-python-voice-agent`, `deepgram-java-voice-agent`, `deepgram-rust-voice-agent`, `deepgram-dotnet-voice-agent`, or `deepgram-go-voice-agent`. The Go SDK v3 ships an agent WebSocket client under `pkg/client/agent/v1/websocket`. The raw protocol above works in any language. [30] +- You want language-idiomatic code: `deepgram-js-voice-agent`, `deepgram-python-voice-agent`, `deepgram-java-voice-agent`, `deepgram-rust-voice-agent`, `deepgram-dotnet-voice-agent`, or `deepgram-go-voice-agent`. The Go SDK v3 ships an agent WebSocket client under `pkg/client/agent/v1/websocket`. The SDKs carry `FunctionCallCancelled` and `defer_until_eot` from JS 5.12.0, Python 7.10.0, and Java 0.10.1, but their voice-agent skills do not describe them, so take the message shapes from this skill. The raw protocol above works in any language. [30] - You only need transcription with turn detection, or only synthesis: the SDK `conversational-stt`, `speech-to-text`, or `text-to-speech` skills. - You want to find a docs page: `docs` skill. You want the MCP server: `setup-mcp` skill. @@ -175,3 +187,8 @@ The Voice Agent API is billed per minute of WebSocket connection time, and a Dee 29. https://github.com/deepgram/recipes/blob/main/COVERAGE.md 30. https://github.com/deepgram/deepgram-go-sdk (`.agents/skills/deepgram-go-voice-agent`, module `github.com/deepgram/deepgram-go-sdk/v3`, agent client at `pkg/client/agent/v1/websocket`) 31. https://developers.deepgram.com/docs/voice-agent-latency-report +32. https://developers.deepgram.com/docs/reusable-agent-configurations (base URL `https://api.deepgram.com/v1`, `config` as a JSON string, immutable `config`, delete warning, `DG_` variables, no secrets) +33. https://developers.deepgram.com/docs/voice-agent-speculative-replies and https://developers.deepgram.com/docs/voice-agent-function-call-cancelled +34. https://developers.deepgram.com/docs/voice-agent-tts-controls (`speed` ranges per family, `expressivity` on Flux TTS only) +35. https://developers.deepgram.com/docs/voice-agent-history +36. https://developers.deepgram.com/docs/voice-agent-conversational-mode (invite-only Early Access; absent from https://developers.deepgram.com/llms.txt) From c584f12c0b18aac798bfd3eb11870a4532e2b9e2 Mon Sep 17 00:00:00 2001 From: Corey Weathers Date: Thu, 1 Oct 2026 09:45:26 -0400 Subject: [PATCH 05/17] docs(cli,setup-mcp): track deepctl 0.3.1; align upgrade advice; warn about the third-party npm deepgram-mcp --- skills/cli/SKILL.md | 32 ++++++++++++++++---------------- skills/setup-mcp/SKILL.md | 12 +++++++++--- 2 files changed, 25 insertions(+), 19 deletions(-) diff --git a/skills/cli/SKILL.md b/skills/cli/SKILL.md index d6a3e2d..6982eff 100644 --- a/skills/cli/SKILL.md +++ b/skills/cli/SKILL.md @@ -11,7 +11,7 @@ description: > # Deepgram CLI (`deepctl`) -One PyPI package, `deepctl`, installs three interchangeable binaries: `dg`, `deepctl`, and `deepgram`. All three report the same version and take the same arguments. This skill uses `dg`. The current release is 0.3.0, published 2026-08-19. +One PyPI package, `deepctl`, installs three interchangeable binaries: `dg`, `deepctl`, and `deepgram`. All three report the same version and take the same arguments. This skill uses `dg`. The current release is 0.3.1, published 2026-09-29. ## Decision rule @@ -34,7 +34,7 @@ On Windows: `iwr https://deepgram.com/install.ps1 -useb | iex`. Upgrade the way you installed: `brew upgrade deepgram`, re-run `install.sh`, or `pip install --upgrade deepctl`. `dg update --check-only` is the version check, printing `current_version`, `latest_version`, and `update_available`. Bare `dg update` reports `"installation_method": null` on a pip install, so treat it as a reporter and upgrade through your installer. -Homebrew lags PyPI. pip, uv, and pipx install 0.3.0, but `Formula/deepgram.rb` in the tap pins `deepctl-0.2.26`, a release before the 0.3.0 one that added Flux TTS and Flux STT support and enforced exit codes. Install from PyPI if you need Flux STT or Flux TTS from the CLI. +Homebrew lags PyPI. pip, uv, and pipx install 0.3.1, but `Formula/deepgram.rb` in the tap pins `deepctl-0.2.26`, a release before the 0.3.0 one that added Flux TTS and Flux STT support and enforced exit codes. Install from PyPI if you need Flux STT or Flux TTS from the CLI. `dg listen --mic` needs an extra on the PyPI installs: `pip install 'deepctl-cmd-listen[mic]'`. The Homebrew install brings the audio dependencies itself. @@ -48,9 +48,9 @@ Three paths, in order of preference: The config file is `config.yaml` in the platform config directory: `~/Library/Application Support/deepctl/` on macOS and `~/.config/deepctl/` on Linux. -The env-var path is not guaranteed to stay off disk. Any command that persists configuration copies the key it read from the environment into that cleartext file. On 0.3.0 with the env var set and no config file, `dg whoami`, `dg listen`, `dg projects --list`, and `dg models` write nothing, but `dg update --check-only` creates it with `api_key:` in it, and `dg projects --set-default` does the same. Treat the file as a secret, and give CI an ephemeral `HOME`. +The env-var path is not guaranteed to stay off disk. Any command that persists configuration copies the key it read from the environment into that cleartext file. On 0.3.1 with the env var set and no config file, `dg whoami`, `dg listen`, `dg projects --list`, and `dg models` write nothing, but `dg update --check-only` creates it with `api_key:` in it, and `dg projects --set-default` does the same. Treat the file as a secret, and give CI an ephemeral `HOME`. -`dg whoami` shows status; add `-o json` for `authenticated`, `project_id`, and `base_url`. It labels the key source `config file` even when the key came from the environment and no config file exists, so do not read that field as a location. +`dg whoami` shows status; add `-o json` for `authenticated`, `project_id`, `base_url`, and `key_source`, which names where the key came from, for example `DEEPGRAM_API_KEY (env)`. With no key, every API-backed command prints this, with your own config path in the parentheses: @@ -63,7 +63,7 @@ Error: DEEPGRAM_API_KEY is not set in the configuration file ## Command surface -23 top-level commands in 0.3.0. `dg transcribe` also survives as a hidden, deprecated alias of `dg listen`. +23 top-level commands in 0.3.1. `dg transcribe` also survives as a hidden, deprecated alias of `dg listen`. | Group | Commands | |---|---| @@ -75,9 +75,9 @@ Error: DEEPGRAM_API_KEY is not set in the configuration file Account commands are flag-based, so the action is a flag rather than a subcommand: `dg projects --list`, not `dg projects list`. -Global flags go before the subcommand: `dg -o json listen file.wav`, never `dg listen -o json`. `-o` takes `json`, `yaml`, `table`, or `csv`. The CLI also detects agent environments, including Claude Code, Aider, and OpenAI Codex, and switches to JSON with plain-text status lines; force that with `CI=true` or `--non-interactive`. A piped stdout alone is not the trigger, so pass `-o json` rather than relying on redirection. +Global flags go before the subcommand: `dg --base-url listen file.wav`, never `dg listen --base-url `. `--base-url`, `--api-key`, `-p`, `-c`, and `--timing` all fail with `No such option` after it. On 0.3.1, `-o`, `-q`, and `-v` are also accepted after the subcommand, so `dg listen file.wav -o json` and `dg -o json listen file.wav` are equivalent, with one exception: on `dg speak`, `-o` after the subcommand is the output file, so put the format flag before `speak`. `-o` takes `json`, `yaml`, `table`, or `csv`. The CLI also detects agent environments, including Claude Code, Aider, and OpenAI Codex, and switches to JSON with plain-text status lines; force that with `CI=true` or `--non-interactive`. A piped stdout alone is not the trigger, so pass `-o json` rather than relying on redirection. -Separately, every command accepts `--agent-friendly`, which prints a machine-readable spec of that command, covering its description, examples, every parameter, `requires_auth`, and `requires_project`, then exits without calling the API. +Separately, every command except the `skills`, `debug`, and `plugin` groups accepts `--agent-friendly`, which prints a machine-readable spec of that command, covering its description, examples, every parameter, `requires_auth`, and `requires_project`, then exits without calling the API. ## Command reference @@ -118,7 +118,7 @@ dg api /v1/projects `dg skills status` lists eight assistants it can detect: Claude Code, OpenAI Codex, Gemini CLI, Amazon Q Developer, Aider, OpenCode, Cursor, and Cline. `dg skills install --all` and `dg skills setup` write files for the detected ones; `dg skills list`, `update`, and `remove` manage them. State lives in `~/.deepctl/skills/skills.json`. -It is not a substitute for `npx skills add deepgram/skills`. On 0.3.0: +It is not a substitute for `npx skills add deepgram/skills`. On 0.3.1: - It downloads four hardcoded skills from this repository, `api`, `docs`, `setup-mcp`, and `starters`, from `raw.githubusercontent.com/deepgram/skills/main`. Every other skill in the repository, including the capability on-ramps, is never fetched, and the list does not grow when the repository adds one. - For Claude Code it writes them to `~/.claude/commands/deepgram/*.md`, the slash-command directory at user scope, not `~/.claude/skills/`, and keeps the `name:` and `description:` skill frontmatter, which is not the slash-command schema. @@ -138,25 +138,25 @@ dg init node-transcription --dir ./my-app --no-install --no-start It does more than `git clone`: it leaves `.git` intact so you can add your own remote, and it writes your key into the clone's `.env` as `DEEPGRAM_API_KEY=`. That is a live secret on disk, so check `.gitignore` before committing. It also refuses to run until `git`, `node`, `npm`, `make`, and `curl` are all present, even with `--no-install`, failing with `Missing tools: git, node, npm, make, curl`. -Three gallery caveats: the gallery carries no `flux` or `flux-tts` templates, so `dg init --list --search flux` returns `No templates found`; the `nextjs-*` entries live under the `deepgram-devs` org while every other template is under `deepgram-starters`; and `sinatra-transcription` points at an archived repository. Use the `starters` skill for the full starter matrix and for anything on Flux STT or Flux TTS. +Three gallery caveats: the gallery carries no `flux` or `flux-tts` templates, so `dg init --list --search flux` returns `No templates found`; the `nextjs-*` entries live under the `deepgram-devs` org while every other template is under `deepgram-starters`; and `sinatra-transcription` points at an archived, private repository, so its URL returns 404. Use the `starters` skill for the full starter matrix and for anything on Flux STT or Flux TTS. ## `dg mcp` -`dg mcp` runs a stdio MCP proxy, with `--transport sse --port 8000` for SSE. On 0.3.0 it advertises server `deepgram-mcp` 0.1.10 and exposes exactly one tool, `search_deepgram_knowledge_sources`, a semantic search over Deepgram's documentation. It exposes no transcription, synthesis, or project tools, so call `dg` directly for those. Use the `setup-mcp` skill for editor wiring. +`dg mcp` runs a stdio MCP proxy, with `--transport sse --port 8000` for SSE. On 0.3.1 it advertises server `deepgram-mcp` 0.1.10 and exposes exactly one tool, `search_deepgram_knowledge_sources`, a semantic search over Deepgram's documentation. It exposes no transcription, synthesis, or project tools, so call `dg` directly for those. Use the `setup-mcp` skill for editor wiring. ## Regional and custom hosts The global `--base-url` flag and the `DEEPGRAM_BASE_URL` env var both retarget the host, and REST and WebSocket endpoints are derived from whichever you set. `dg whoami` echoes the value back. Use them for self-hosted and staging deployments. -The regional hosts do not work on 0.3.0. Every command first checks credentials against `{base_url}/v1/projects`, and the management API is not served regionally, so `dg --base-url https://api.eu.deepgram.com listen file.wav` fails with `Error: Unexpected error: HTTP 404` while the same transcription request succeeds with curl against `api.eu.deepgram.com/v1/listen`. Use curl or an SDK for `api.eu.deepgram.com`, `api.au.deepgram.com`, and `api.in.deepgram.com`. +The regional hosts do not work on 0.3.1. Every command first checks credentials against `{base_url}/v1/projects`, and the management API is not served regionally, so `dg --base-url https://api.eu.deepgram.com listen file.wav` fails with `Error: Unexpected error: HTTP 404` while the same transcription request succeeds with curl against `api.eu.deepgram.com/v1/listen`. Use curl or an SDK for `api.eu.deepgram.com`, `api.au.deepgram.com`, and `api.in.deepgram.com`. ## Common mistakes -1. Trusting the exit code. 0.3.0 enforces exit codes, where earlier versions always exited 0: now 0 is success, 1 is an error or bad usage, and 2 is an interrupt. The contract has holes. `dg listen --mic` and `dg mcp` swallow Ctrl-C and exit 0, and a missing API key prints `Error: DEEPGRAM_API_KEY is not set …` and still exits 0, on both `dg listen` and `dg projects`. API errors and missing files do exit 1. In CI, check `"status"` in `-o json` as well as `$?`. -2. Expecting caption files from `--srt` or `--webvtt`. The docs label them "SRT subtitles" and "WebVTT captions", but on 0.3.0 `dg listen file.wav --srt --save-to out.srt` writes the plain transcript with no cue numbers and no timestamps, for a local file or a URL, under `-o table` or `-o json`. Use `dg -o json listen` and build cues from the word timestamps, or call the API. +1. Trusting the exit code. 0.3.0 enforces exit codes, where earlier versions always exited 0: now 0 is success, 1 is an error or bad usage, and 2 is an interrupt. The contract has holes. `dg listen --mic` and `dg mcp` swallow Ctrl-C and exit 0. API errors, missing files, and a missing API key do exit 1. In CI, check `"status"` in `-o json` as well as `$?`. +2. Expecting caption files from `--srt` or `--webvtt`. The docs label them "SRT subtitles" and "WebVTT captions", but on 0.3.1 `dg listen file.wav --srt --save-to out.srt` writes the plain transcript with no cue numbers and no timestamps, for a local file or a URL, under `-o table` or `-o json`. Use `dg -o json listen` and build cues from the word timestamps, or call the API. 3. Running `dg config set …`. Every help screen ends with `Disable: dg config set telemetry.enabled false`, but there is no `config` command, and the attempt returns `Error: No such command 'config'`. Set `DEEPCTL_TELEMETRY_DISABLED=1` instead, after which the footer reads `Telemetry is off.` -4. Putting a global flag after the subcommand. `dg listen -v file.wav` returns `Error: No such option '-v'`. Write `dg -v listen file.wav`. -5. Treating `dg models` as the model catalog. It returns 549 rows, strips the family prefix so Aura voices appear as `asteria` rather than `aura-2-asteria-en`, includes deprecated versions, and lists no Flux STT or Flux TTS models at all. Take model names from the `speech-to-text` and `text-to-speech` skills. +4. Putting a global flag after the subcommand. `dg listen --base-url file.wav` returns `Error: No such option '--base-url'` plus a hint that it is a global option. Write `dg --base-url listen file.wav`. The same holds for `--api-key`, `-p`, `-c`, and `--timing`; `-o`, `-q`, and `-v` work in either position. +5. Treating `dg models` as the model catalog. It returns 553 rows, 942 with `--include-outdated`. Each row's `name` is the bare voice, `asteria`, and `canonical_name` carries the full identifier, `aura-2-asteria-en`, so script against `canonical_name`. Rows carry a `deprecated` flag, and the list holds no Flux STT or Flux TTS models at all. Take model names from the `speech-to-text` and `text-to-speech` skills. 6. Pointing Flux STT at a file. `dg listen file.wav -m flux-general-en` returns `Flux STT (flux-general-en) is streaming-only and cannot transcribe a file or URL.` Pipe stdin or use `--mic`, since Flux STT has no prerecorded mode. 7. Acting on `Warning: API key format doesn't match expected pattern`. A valid 40-character key triggers it and then verifies fine. 8. Expecting `dg login --profile ` to persist without a keyring. The profile then appears in neither `dg profiles --list` nor `config.yaml`, which still hold only `default`. @@ -182,7 +182,7 @@ The regional hosts do not work on 0.3.0. Every command first checks credentials - MCP server: https://developers.deepgram.com/developer-tools/cli/mcp-server - Shell completion: https://developers.deepgram.com/developer-tools/cli/shell-completion - Plugins: https://developers.deepgram.com/developer-tools/cli/plugins -- Agentic tools overview: https://developers.deepgram.com/agentic-tools +- Agentic tools overview: https://developers.deepgram.com/developer-tools/agentic-tools - Source and releases: https://github.com/deepgram/cli - Package: https://pypi.org/project/deepctl/ - Homebrew tap: https://github.com/deepgram/homebrew-tap diff --git a/skills/setup-mcp/SKILL.md b/skills/setup-mcp/SKILL.md index e3dffb9..7e0b39b 100644 --- a/skills/setup-mcp/SKILL.md +++ b/skills/setup-mcp/SKILL.md @@ -83,8 +83,10 @@ pipx install deepctl iwr https://deepgram.com/install.ps1 -useb | iex ``` -To upgrade, use the CLI's own updater: `dg update` (add `--check-only` to check without -installing). If it was installed with Homebrew, `brew upgrade deepgram` also works. +To upgrade, use the installer that put it there: `pip install -U deepctl`, +`uv tool upgrade deepctl`, `pipx upgrade deepctl`, `brew upgrade deepgram`, or re-run the install +script. `dg update --check-only` reports whether a newer release exists; on a pip install, bare +`dg update` reports `installation_method: null` instead of upgrading. ### A2. Authenticate — required @@ -165,6 +167,9 @@ pip install deepgram-mcp export DEEPGRAM_API_KEY=your_key_here ``` +`deepgram-mcp` is a PyPI package. The npm package of the same name is unrelated third-party code +that also asks for `DEEPGRAM_API_KEY`, so do not run `npx deepgram-mcp`. + #### Claude Code ```sh @@ -308,7 +313,8 @@ server name if the user wants to keep both. the API serves right now, not what the package version implies. Reconnect to pick up new tools. **Anything else on Path A** -→ Verify `dg --version` works and `dg mcp` runs in a terminal without errors, then `dg update`. +→ Verify `dg --version` works and `dg mcp` runs in a terminal without errors, then +`dg update --check-only` to see whether a newer release exists. ## Sources From 41ed61eff4a39693cb8e277a00810158e325e6d1 Mon Sep 17 00:00:00 2001 From: Corey Weathers Date: Thu, 1 Oct 2026 09:45:26 -0400 Subject: [PATCH 06/17] docs: sweep self-hosted references, examples, recipes, starters, and AGENTS.md against live repositories --- AGENTS.md | 7 +++++++ skills/examples/SKILL.md | 15 +++++++++------ skills/recipes/SKILL.md | 4 ++-- skills/self-hosted/references/docker-podman.md | 1 - skills/self-hosted/references/kubernetes.md | 2 +- skills/self-hosted/references/sagemaker.md | 14 +++++++++++++- skills/starters/SKILL.md | 9 +++++---- 7 files changed, 37 insertions(+), 15 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index be17e91..f6bacfc 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -36,6 +36,12 @@ npx skills add deepgram/skills --agent claude-code -y # one agent, every skill npx skills add deepgram/skills --skill api -y # one skill ``` +`npx skills add` clones the repository with `git`. On an image without it (a +bare `node:22-alpine`, for example) every target fails with `Failed to clone +...: Error: spawn git ENOENT`, yet the command exits 0 and installs nothing. +Install `git` first and check for the `SKILL.md` files rather than trusting +the exit code. + Claude Code plugin route: `/plugin marketplace add deepgram/skills`, then `/plugin install deepgram@deepgram-agent-skills`. @@ -80,6 +86,7 @@ Adding a skill also means adding its path to `plugins[0].skills` in | Symptom | Cause | Fix | |---------|-------|-----| | `npx skills add deepgram/` fails with a not-found or auth error | the target repository is private or does not exist | only the six public SDK repositories listed in README.md carry installable skills | +| every target fails with `spawn git ENOENT`, exit code 0, nothing installed | `git` is absent from the container or CI image; `npx skills add` shells out to it | install `git` (`apk add git` on Alpine) before running the installer | | `bun: command not found` | bun not installed | install from https://bun.sh; the generation scripts are bun-only | | `ENOENT ... specs/openapi.yml` from `generate-skills.ts` | `fetch-specs.ts` was not run first; `specs/` is gitignored, so it is absent in a fresh clone | run both regeneration commands in order | | Regenerated `api` skill shows unexpected churn | the upstream specs moved | inspect the spec diff first; the specs are the source of truth | diff --git a/skills/examples/SKILL.md b/skills/examples/SKILL.md index ba8f19f..92a1ef1 100644 --- a/skills/examples/SKILL.md +++ b/skills/examples/SKILL.md @@ -5,8 +5,9 @@ description: > Use whenever someone wants to integrate Deepgram with Twilio, LiveKit, LangChain, Vercel AI SDK, Discord, Vonage, Pipecat, Expo, FastAPI, Cloudflare Workers, Slack, Telegram, LlamaIndex, Zoom, Next.js, Nuxt, Django, SvelteKit, NestJS, Spring Boot, - CrewAI, Riverside, SignalWire, and more. Examples are full runnable integration - demos, not minimal feature snippets. + CrewAI, Riverside, SignalWire, Telnyx, Plivo, Webex, Jitsi, Microsoft Teams, OBS + Studio, Haystack, Semantic Kernel, Gin, and more. Examples are full runnable + integration demos, not minimal feature snippets. --- # Deepgram Examples @@ -34,16 +35,18 @@ Examples are numbered (010, 020, ...) and each is a self-contained integration. | Category | Integrations | Common STT choice | |---|---|---| -| **Telephony** | Twilio, Vonage, SignalWire, Daily.co, Asterisk/FreeSWITCH | Nova live (`/v1/listen`) for call transcription; Flux STT (`/v2/listen`) for AI-agent calls | +| **Telephony** | Twilio, Vonage, SignalWire, Telnyx, Plivo, Daily.co, Asterisk/FreeSWITCH | Nova live (`/v1/listen`) for call transcription; Flux STT (`/v2/listen`) for AI-agent calls | | **Voice AI frameworks** | LiveKit Agents, Pipecat, OpenAI Agents SDK, CrewAI | Flux STT (`/v2/listen`) — built-in turn detection; or Voice Agent (`/v1/agent/converse`) for full-pipeline | | **Chat platforms** | Discord, Slack, Telegram | Nova prerecorded (`/v1/listen`) for attachments | -| **Web frameworks** | Next.js, Nuxt, Django, SvelteKit, NestJS, Express + React, FastAPI, Spring Boot | Nova live (`/v1/listen`) for captions; Nova prerecorded for batch | +| **Web frameworks** | Next.js, Nuxt, Django, SvelteKit, NestJS, Express + React, FastAPI, Spring Boot, Gin | Nova live (`/v1/listen`) for captions; Nova prerecorded for batch | | **Mobile / desktop** | Expo, Flutter, Swift iOS, Kotlin Android, Tauri, Electron | Nova live (`/v1/listen`); Flux STT (`/v2/listen`) if the app is a voice agent | | **Cloud / serverless** | AWS Lambda, Cloudflare Workers | Nova prerecorded (`/v1/listen`) — best fit for request/response | -| **Recording platforms** | Zoom, Riverside.fm | Nova prerecorded (`/v1/listen`) | +| **Meeting platforms** | Zoom, Riverside.fm, Webex (recordings); Jitsi, Microsoft Teams (live meetings) | Nova prerecorded (`/v1/listen`) for recordings; Nova live (`/v1/listen`) for the Jitsi bridge and the Teams bot | | **Browser / no-bundler** | Vanilla JavaScript | Nova live (`/v1/listen`) via `@deepgram/sdk` in the browser | -| **LLM frameworks** | LangChain, LlamaIndex, Vercel AI SDK | Nova or Flux STT depending on streaming vs batch | +| **LLM frameworks** | LangChain, LlamaIndex, Vercel AI SDK, Haystack, Semantic Kernel | Nova or Flux STT depending on streaming vs batch | | **Low-code / automation** | n8n community nodes | Nova (`/v1/listen`) for event-driven transcription | +| **Broadcast** | OBS Studio (native C plugin) | Nova live (`/v1/listen`) for on-screen captions | +| **Infrastructure** | Deepgram API proxy servers (`520-node-deepgram-proxy`, `521-deepgram-proxy-python-uv`), Silero VAD segmentation (`530-silero-vad-speech-segmentation-python`), multi-provider LLM proxy for Voice Agent (`530-voice-agent-multi-provider-proxy-python`; two directories share the number 530) | The proxies forward Nova prerecorded, Nova live, and Aura TTS; Silero VAD sends each detected speech segment to Nova prerecorded; the Voice Agent proxy is the agent's `think.endpoint.url`, with STT and TTS staying on the agent | The column above is the STT half of each integration. On the TTS side, Aura (`/v1/speak`) is what these examples use today — **no example in the repo covers Flux TTS (`/v2/speak`) yet**, even in the telephony and voice-AI-framework categories where it fits best. If you're wiring Flux TTS into one of these platforms, take the integration's transport and auth handling from the example and the Flux TTS contract from the `api` skill. diff --git a/skills/recipes/SKILL.md b/skills/recipes/SKILL.md index 07d5df3..24a6d28 100644 --- a/skills/recipes/SKILL.md +++ b/skills/recipes/SKILL.md @@ -43,8 +43,8 @@ recipes/{language}/{product}/{version}/{recipe}/ | Product | Recipe examples | |---|---| -| Speech-to-Text — Nova (`/v1/listen`) | transcribe-url, transcribe-file, paragraphs, diarize, smart-format, utterances, summarize, sentiment, topics, intents, detect-entities, detect-language, redact, search, keywords, streaming | -| Speech-to-Text — Flux STT (`/v2/listen`) | streaming conversational transcription, EOT / eager-EOT, mid-session `Configure`, keyterms | +| Speech-to-Text — Nova (`/v1/listen`) | 26 recipes: transcribe-url, transcribe-file, streaming, streaming-file, punctuate, smart-format, paragraphs, utterances, diarize, multichannel, numerals, measurements, dictation, filler-words, profanity-filter, redact, replace, search, keywords, keyterm, detect-language, detect-entities, summarize, sentiment, topics, intents | +| Speech-to-Text — Flux STT (`/v2/listen`) | Two recipes. `streaming` opens the `/v2/listen` WebSocket with `model=flux-general-en`, `encoding=linear16`, `sample_rate=16000` and prints `TurnInfo` events (transcript, turn index, event type) in place of v1 interim/final pairs. `transcribe-url` sets `model=flux-general-en` on the SDK's prerecorded transcribe-URL call with `smart_format`. The CLI has only `transcribe-url`. No recipe is dedicated to EOT or eager EOT thresholds, mid-session `Configure`, or keyterms | | Text-to-Speech — Aura (`/v1/speak`) | generate-audio, stream-audio, websocket-streaming, select-model, select-encoding, bit-rate | | Audio Intelligence (`/v1/listen`) | summarize, sentiment, topics, intents, entities | | Voice Agents | connect, custom-llm, custom-tts, function-calling | diff --git a/skills/self-hosted/references/docker-podman.md b/skills/self-hosted/references/docker-podman.md index 4c627f4..f41ba1d 100644 --- a/skills/self-hosted/references/docker-podman.md +++ b/skills/self-hosted/references/docker-podman.md @@ -201,7 +201,6 @@ On FIPS images, MP3 and FLAC output is a known issue — set `encoding` explicit - Docker/Podman overview: https://developers.deepgram.com/docs/dockerpodman - Deploy STT services: https://developers.deepgram.com/docs/deploy-stt-services - Deploy TTS services: https://developers.deepgram.com/docs/deploy-tts-services -- Deploy Deepgram services: https://developers.deepgram.com/docs/deploy-deepgram-services - Flux STT self-hosted: https://developers.deepgram.com/docs/flux-self-hosted - Flux TTS self-hosted: https://developers.deepgram.com/docs/deploy-flux-tts - Per-cloud and bare metal: https://developers.deepgram.com/docs/aws-docker-podman, https://developers.deepgram.com/docs/gcp-docker-podman, https://developers.deepgram.com/docs/oci-docker-podman, https://developers.deepgram.com/docs/azure-docker-podman, https://developers.deepgram.com/docs/bare-metal diff --git a/skills/self-hosted/references/kubernetes.md b/skills/self-hosted/references/kubernetes.md index 7e26da1..fe51433 100644 --- a/skills/self-hosted/references/kubernetes.md +++ b/skills/self-hosted/references/kubernetes.md @@ -181,7 +181,7 @@ const deepgram = new DeepgramClient({ Enable the **Billing** container, which validates a license locally and journals usage instead of calling `license.deepgram.com`. -- Architecture: `API/Engine → Billing`, or `API/Engine → License Proxy → Billing` for HA. +- Architecture: `API/Engine → Billing`, or `API/Engine → License Proxy → Billing` for HA. The chained form needs a chart newer than `0.46.0`: on `0.46.0` and earlier, with `billing.enabled` and `licenseProxy.enabled` both `true`, the `billing` condition takes precedence, so API and Engine connect to Billing directly and the deployed License Proxy receives no traffic. The fix is the `Unreleased` entry in `charts/deepgram-self-hosted/CHANGELOG.md`. - Obtain from Deepgram: a license key, a license file (`.dg`, a one-line JSON file), and registry access for `quay.io/deepgram/*` including the Billing image. - Configure `billing.enabled`, `billing.licenseFile.secretRef` (key `license.dg` by default), and `global.deepgramLicenseSecretRef`. - Billing listens on `8443` for license verification and `8080` for the `/v1/certificates` endpoint. diff --git a/skills/self-hosted/references/sagemaker.md b/skills/self-hosted/references/sagemaker.md index 36301fc..83c610d 100644 --- a/skills/self-hosted/references/sagemaker.md +++ b/skills/self-hosted/references/sagemaker.md @@ -16,6 +16,8 @@ Choose it when you are AWS-only and want less operational surface than Docker or The tradeoffs versus running containers yourself, and SageMaker pricing, are laid out at [Amazon SageMaker](https://developers.deepgram.com/docs/amazon-sagemaker). +AWS field employees can reach Deepgram models through the [AWS Marketplace Field Demonstration Program](https://docs.aws.amazon.com/marketplace/latest/userguide/field-demonstration-program.html); Deepgram is an eligible provider. + ## Product listings Deepgram publishes to [AWS Marketplace](https://aws.amazon.com/marketplace/search/results?searchTerms=deepgram&CREATOR=6efa21f9-9a33-4cae-ba44-756436fa71dd&FULFILLMENT_OPTION_TYPE=SAGEMAKER_MODEL&filters=CREATOR%2CFULFILLMENT_OPTION_TYPE) (no AWS login needed to browse). @@ -31,6 +33,16 @@ Individual languages are delivered as **versions** of a model package. One monol Every product needs a GPU instance. Request [SageMaker quota](https://developers.deepgram.com/docs/request-sagemaker-quota) before creating an endpoint. +Deploy on an ordered **instance pool** rather than a single instance type. A single type has no fallback: when the Availability Zone is short of that GPU, the endpoint goes `Failed` (`Request to service failed` a few minutes in, or `InsufficientInstanceCapacity`), and that happens routinely for popular GPU types. With [instance pools](https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints-heterogeneous.html), SageMaker tries each type in priority order and falls back to the next when one is capacity-constrained. Order the pool: + +1. The listing's recommended type first (`ml.g6.2xlarge` for STT): the type Deepgram validated the model on, and the best price for the performance. +2. Same-or-newer generations with similar per-instance capacity next (`g6`, then `g6e`, then `g7`). Similar capacity matters if you autoscale, because the predefined scaling metrics are per instance and do not account for a mixed fleet. +3. Older generations last, as insurance (`g5`, and `g4dn` where supported). +4. Never a type the product does not support: `g4dn` for Flux STT, `g5` and `g4dn` for Flux TTS, any single-GPU type for Aura-2. +5. Up to 5 types; three is the sweet spot. + +`VariantInstanceProvisionTimeoutInSeconds` is the per-type wait before SageMaker moves to the next type: `300` is recommended (AWS allows `60` to `3600`), so a three-type pool can sit in `Creating` for about 15 minutes before it fails. Quota does not fall back: SageMaker validates the quota of every type in the pool at `CreateEndpoint`, and a type with a regional quota below `1` fails the call with `ResourceLimitExceeded` regardless of which type would have been used. CLI and Boto3 examples: [Choose instance types](https://developers.deepgram.com/docs/deploy-amazon-sagemaker#choose-instance-types). + | Product | Recommended | Also supported | Not supported | |---|---|---|---| | Nova-3 STT | `ml.g6.2xlarge` | `ml.g7.2xlarge`, `ml.g7e.2xlarge`, `ml.g6e.2xlarge`, `ml.g5.2xlarge`, `ml.g4dn.2xlarge` | — | @@ -158,7 +170,7 @@ Examples: `examples/stt.mjs`, `tts.mjs`, `flux.mjs`, `flux-tts.mjs`, `live-mic.m ### Java -Requires **Java 11+** and Deepgram Java SDK **v0.4.0+** — the `default ReconnectOptions reconnectOptions()` hook on `DeepgramTransportFactory` is what enables storm absorption. The transport's README pins `0.4.0` in its install snippet; Maven Central's latest Java SDK is `0.10.0`, which satisfies the floor. Pin deliberately and test the pairing. +Requires **Java 11+** and Deepgram Java SDK **v0.4.0+**: the `default ReconnectOptions reconnectOptions()` hook on `DeepgramTransportFactory` is what enables storm absorption. The transport's README pins `0.4.0` in its install snippet; Maven Central's latest Java SDK is `0.10.2`, which satisfies the floor. Pin deliberately and test the pairing. ```groovy dependencies { diff --git a/skills/starters/SKILL.md b/skills/starters/SKILL.md index 2c9cef6..8e45f5e 100644 --- a/skills/starters/SKILL.md +++ b/skills/starters/SKILL.md @@ -113,10 +113,11 @@ git -c url."https://github.com/".insteadOf="git@github.com:" \ `dg init` is also marked alpha, and its templates gallery is a separate list from the matrix below rather than a subset of it. It carries 44 templates with no `flux` or `flux-tts` entries; -it still lists `sinatra-transcription`, whose repository is archived; and it lists `nextjs-*` -templates that now redirect out of `deepgram-starters` to `deepgram-devs`, which is why there is -no `nextjs` row below. Treat the matrix as authoritative and fall back to `git clone`. See the -`cli` skill for installing `deepctl` and for the rest of `dg init`. +it still lists `sinatra-transcription`, whose repository is archived and private, so the clone +returns 404 for anyone outside Deepgram; and it lists `nextjs-*` templates that now redirect out +of `deepgram-starters` to `deepgram-devs`, which is why there is no `nextjs` row below. Treat +the matrix as authoritative and fall back to `git clone`. See the `cli` skill for installing +`deepctl` and for the rest of `dg init`. ## The `{feature}-html` repos are not starters From 7a3bdd9071906c3b1c0504de139e4d829b5604c3 Mon Sep 17 00:00:00 2001 From: Corey Weathers Date: Thu, 1 Oct 2026 09:45:26 -0400 Subject: [PATCH 07/17] chore: bump version to 1.7.0 and update changelog --- .claude-plugin/marketplace.json | 2 +- CHANGELOG.md | 66 ++++++++++++++++++++++++++++++++- 2 files changed, 66 insertions(+), 2 deletions(-) diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index 9b8c74a..7dcf739 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -6,7 +6,7 @@ }, "metadata": { "description": "Deepgram skills for AI coding tools", - "version": "1.6.0" + "version": "1.7.0" }, "plugins": [ { diff --git a/CHANGELOG.md b/CHANGELOG.md index c4d72f2..f1de910 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,15 +7,79 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] +[Unreleased]: https://github.com/deepgram/skills/compare/deepgram-skills-v1.7.0...HEAD + +## [1.7.0] - 2026-10-01 + +Catch-up with the September API, spec, SDK, CLI, and documentation changes. Flux TTS inline pause and pronunciation controls, the Flux STT `Warning` message and mid-stream `numerals`, the Voice Agent reusable-configuration and agent-variable REST surface, `FunctionCallCancelled` and `defer_until_eot`, the Go SDK's Flux TTS client, and `deepctl` 0.3.1 are the headline items. No skill is added or removed, so the `deepgram` plugin still lists 14 skills. + ### Added +- API skill: the API Domains table now lists the whole Voice Agent REST surface, not just `GET agent.deepgram.com/v1/agent/settings/think/models`. Reusable agent configurations are `GET` and `POST /v1/projects/{project_id}/agents` plus `GET`, `PUT`, and `DELETE /v1/projects/{project_id}/agents/{agent_id}`; agent variables are `GET` and `POST /v1/projects/{project_id}/agent-variables` plus `GET`, `PATCH`, and `DELETE /v1/projects/{project_id}/agent-variables/{variable_id}`. All ten were already rendered in `references/agent.md`; the table is the router an agent reads first, and it had never pointed at them. The OpenAPI carries one `servers` block for the whole document, so these ten paths are listed without a host +- API skill: the Models row names the four model endpoints rather than one: `GET /v1/models`, `GET /v1/models/{model_id}`, `GET /v1/projects/{project_id}/models`, and `GET /v1/projects/{project_id}/models/{model_id}`, and records that `include_outdated=true` on either list call also returns non-latest model versions +- API skill, Flux TTS mistake 12: the inline-controls rule for `/v2/speak`. A pronunciation override `\{"word":"...","pronounce":""\}` is honored on both transports (Early Access) but only with `speed` 1.0, a pause marker `\{pause:500ms\}` is batch-only, and a pause marker on the socket, or a pronunciation control on a socket whose `speed` is not 1.0, fails the connection with `DATA-0002`. Links the Speed, Pause, Pronunciation page. The sentence that said SSML is stripped with an `INPUT_MARKUP_STRIPPED` warning is gone, for the reason given under the text-to-speech skill below; mistake 12 now says SSML is not interpreted and that the only markup Flux TTS honors is its own escaped inline controls +- API skill, Flux STT mistake 17: `ConfigureFailure` carries `code` and `description` identifying the rejected configuration +- Text-to-speech skill: Flux TTS inline controls, which the skill had not covered. Pronunciation control (Early Access) is an escaped JSON object in the text, `\{"word":"...","pronounce":""\}`, accepted on both the `/v2/speak` WebSocket and batch `POST /v2/speak`, at most 500 per request with IPA of at most 128 characters, and only with `speed` `1.0`: on the socket a pronunciation sent on a session opened with another speed, or after a `Configure` that set one, fails the connection with `DATA-0002`, and on batch the request is a 400 `CONTROL_COMBINATION_INVALID`. Pause control `\{pause:500ms\}` is batch only, 500 to 3000 ms in 100 ms steps, at most 8 per request, with `speed` capped at `1.15` while a pause is present (`PAUSE_SPEED_CAP_EXCEEDED`); a pause marker on the socket fails the connection with `DATA-0002`. The remaining batch 400 codes are listed (`BREAK_OUT_OF_RANGE`, `BREAK_INCREMENT_INVALID`, `BREAKS_LIMIT_EXCEEDED`, `BREAK_SYNTAX_INVALID`), as is how each transport reports what it applied: `SpeechMetadata.controls_applied` on the socket, `dg-pronunciations-applied` and `dg-breaks-applied` response headers on batch, with invalid IPA still applied best-effort and surfaced as a `PRONUNCIATION_WARNINGS` warning or the `dg-warnings` header +- Text-to-speech skill: the `Connected` message (`request_id`, `model_name`, `model_version`, `model_uuids`) and the `SessionMetadata` message (cumulative session totals, rebased by an `Interrupt` onto the audio the client actually played), neither of which the skill had named +- Text-to-speech skill: `Interrupt` offset semantics. `playback_offset.value` is cumulative milliseconds since the session started, not since the current turn; each `Interrupt` must exceed the previous offset or it is ignored with `INVALID_INTERRUPT_OFFSET`; `audio_played_ms` from `SpeechInterrupted` is the baseline for the next offset. The old text produced wrong offsets after the first turn +- Text-to-speech skill: `CONTROL_COMBINATION_INVALID` as the fourth `ConfigureFailure` code, raised when a queued turn still carries a pronunciation control, with the fix (`Flush` that turn first) +- Text-to-speech skill: `NET-0003`, the 1-hour session cap, beside the existing `NET-0004` idle close +- Text-to-speech skill: a pointer from the Voice Agent mistake to https://developers.deepgram.com/docs/voice-agent-tts-controls, where `speed` and `expressivity` are set on `agent.speak.provider` +- Speech-to-text skill: the Flux STT `Configure` message takes `numerals` alongside thresholds, keyterms, and `language_hints`. The skill's example message carries `"numerals":true`, states that the value starts from the `numerals` query parameter and applies only to transcripts Flux STT sends after it processes the update, and that it must be a JSON boolean: the string `"true"` fails schema validation, Flux STT answers with an `Error` of code `UNPARSABLE_CLIENT_MESSAGE`, and the connection closes. `ConfigureSuccess` echoes the full active configuration including `numerals`; `ConfigureFailure` carries `code` and `description` naming the rejected field and leaves the previous configuration in place +- Speech-to-text skill: the `numerals` language scope on Flux STT. `flux-general-en` formats every number; `flux-general-multi` formats English, Spanish, French, German, Russian, Portuguese, Italian, and Dutch, and leaves Hindi and Japanese numbers as spoken +- Speech-to-text skill: SDK routing for mid-stream `numerals`. The JavaScript SDK has it from 5.13.0 (`socket.sendConfigure({type:"Configure", numerals:true})`), the Python SDK from 7.11.0 (`connection.send_configure(ListenV2Configure(numerals=True))`), and the Java SDK from 0.10.2 (`sendConfigure(ListenV2Configure.builder().numerals(true).build())`). The `deepgram-{lang}-conversational-stt` SDK skills do not cover it, so the skill tells the reader to take the message shape from its own `Configure` bullet +- Speech-to-text skill: two Sources entries, https://developers.deepgram.com/docs/flux/configuration for the end-of-turn threshold table and https://developers.deepgram.com/docs/numerals for the Flux STT numerals language list and mid-stream toggling +- Voice-agent skill: a "Reusable agent configurations" section. `Settings.agent` is either the full `agent` object or a Reusable Agent Configuration UUID string, the same `agent: "YOUR_AGENT_ID"` form the browser-agent skill already shows. The section gives the REST surface on `https://api.deepgram.com/v1` (`POST/GET /projects/{project_id}/agents`, `GET/PUT/DELETE .../agents/{agent_id}`), that `config` is the JSON string of the `agent` block and is immutable once created (`PUT` changes `metadata` only), that deleting a configuration a running service references breaks that service, the `INVALID_AGENT_ID` and `AGENT_ID_NOT_SUPPORTED` error codes, and the template-variable endpoints (`/projects/{project_id}/agent-variables`) with the `DG_` key form, unquoted substitution of any JSON value, `is_sensitive: false`, and the rule that configurations and variables are readable by every project member so hold no secrets +- Voice-agent skill: `FunctionCallCancelled` in the message lifecycle table (the user started speaking again; stop work on each `id` and send no `FunctionCallResponse`, a late one is dropped), and a speculative-dispatch paragraph in the function-calling section. Function calls go out before the turn is confirmed by default; `defer_until_eot: true` holds a call until the turn is confirmed and discards it if the turn resumes, for actions that cannot be undone, and an `endpoint` function that already ran is not rolled back +- Voice-agent skill: `agent.speak.provider.speed` (Flux TTS `0.5` to `1.5` in `0.05` steps, Aura `0.7` to `1.5`; an unaccepted value ends the session with `FAILED_TO_SPEAK`) and `expressivity` (whole numbers `-2` to `2`, Flux TTS `v2` only, fixed for the session, beta with `0` the only production-validated value) in the `speak` field notes +- Voice-agent skill: `Settings.flags.history` (default `true`) and the server `History` message in both its conversation-text and `function_calls` shapes, alongside the existing `agent.context.messages` input; a pointer bullet for `think.context_length`, `think.provider.reasoning_mode`, and the top-level `tags`, `experimental`, and `mip_opt_out`, with ranges in `references/agent.md`; a one-line note that Conversational Mode (`agent.think_conversational.provider`) is invite-only Early Access answered with `UNPARSABLE_CLIENT_MESSAGE` from unenrolled projects +- Voice-agent skill: five new sources (reusable agent configurations, speculative replies, function call cancelled, TTS controls, history, conversational mode) and two new description triggers, "reusable agent configuration" and "defer_until_eot" +- Self-hosted skill: the SageMaker reference now carries the ordered instance-pool recommendation above the instance table. A single instance type has no fallback, so when the Availability Zone is short of that GPU the endpoint goes `Failed` (`Request to service failed` or `InsufficientInstanceCapacity`), which happens routinely for popular GPU types. The pool order is the listing's recommended type first, same-or-newer generations with similar per-instance capacity next (`g6`, `g6e`, `g7`), older generations last as insurance, never an unsupported type (`g4dn` for Flux STT, `g5` and `g4dn` for Flux TTS, single-GPU types for Aura-2), up to 5 types with three as the sweet spot. `VariantInstanceProvisionTimeoutInSeconds` is the per-type wait (`300` recommended, `60` to `3600` allowed), and quota does not fall back: every pooled type needs a regional quota of at least `1` or `CreateEndpoint` fails with `ResourceLimitExceeded`. Also one line that AWS field employees can reach Deepgram models through the AWS Marketplace Field Demonstration Program, for which Deepgram is an eligible provider +- Self-hosted skill: the Kubernetes reference states that the `API/Engine -> License Proxy -> Billing` chain for air-gapped HA needs a chart newer than `0.46.0`. On `0.46.0` and earlier, with `billing.enabled` and `licenseProxy.enabled` both `true`, the `billing` condition takes precedence, so API and Engine connect to Billing directly and the deployed License Proxy receives no traffic; the fix is the `Unreleased` entry in `charts/deepgram-self-hosted/CHANGELOG.md` +- Examples skill: the category map and description now cover the integrations the live `deepgram/examples` tree carries that the skill had omitted: Telnyx and Plivo (telephony), Webex, Jitsi, and Microsoft Teams (the "Recording platforms" row is now "Meeting platforms", split into recordings on Nova prerecorded and live meetings on Nova live, since the Jitsi bridge and the Teams bot stream), Haystack and Semantic Kernel (LLM frameworks), Gin (web frameworks), a **Broadcast** row for the OBS Studio native C captioning plugin, and an **Infrastructure** row for the Node and Python Deepgram API proxy servers, Silero VAD speech segmentation, and the multi-provider chat-completions proxy that serves as a Voice Agent `think.endpoint.url`. The two `530-*` directories are named in full because they share a number. The statement that no example covers Flux TTS stands: the Voice Agent proxy keeps TTS on `aura-2` +- `AGENTS.md`: the install section records that `npx skills add` clones with `git`, so an image without it (a bare `node:22-alpine`, for example) fails every target with `Failed to clone ...: Error: spawn git ENOENT` while exiting 0 and installing nothing. The common-failure-modes table gains the matching row with `apk add git` as the fix - API skill: the reference generator now emits JSON-Schema bounds, which it had been discarding for every parameter. `minimum` and `maximum` are the only bounds the specs carry today (20 values in `openapi.yml`, 29 in `asyncapi.yml`); `minLength`, `maxLength`, `minItems`, `maxItems`, `exclusiveMinimum`, and `exclusiveMaximum` are handled so an upstream spec that starts using one needs no further change, and `multipleOf` is handled alongside them. `enum` is left as it was, because `formatType` already renders it as a literal union. A bound whose numbers the description states in prose is suppressed rather than repeated, so `limit`, which ends "Range [1,1000]", does not also render "range: `1` to `1000`". `ttl_seconds` on `POST /v1/auth/grant` gains `range: 1 to 3600`, the `/v1/speak` and Speak v1 WebSocket `speed` parameters gain `range: 0.7 to 1.5`, and `turn_index` on Flux STT `TurnInfo` gains `minimum: 0` +### Changed + +- API skill: `references/listen.md` and `references/speak.md` regenerated from the public specs. `listen.md` gains the Flux STT `Warning` message (`ListenV2Warning`: `type`, `request_id`, `sequence_id`, `code`, `description`, all required), `code` and `description` on `ConfigureFailure`, and `numerals` echoed on `ConfigureSuccess`. `speak.md` replaces every "inline pause and pronunciation controls are not yet applied; they are stripped" sentence with the live rules: `\{pause:500ms\}` markers of 500 to 3000 ms in 100 ms steps, at most 8 per batch request, `speed` capped at `1.15` while a pause is present, pronunciation controls that cannot be combined with a pause or with a `speed` other than `1.0`, the six `400` `err_code` values on `POST /v2/speak` (`CONTROL_COMBINATION_INVALID`, `PAUSE_SPEED_CAP_EXCEEDED`, `BREAK_OUT_OF_RANGE`, `BREAK_INCREMENT_INVALID`, `BREAKS_LIMIT_EXCEEDED`, `BREAK_SYNTAX_INVALID`), the new `CONTROL_COMBINATION_INVALID` value on `SpeakV2ConfigureFailure.code`, the `DATA-0002` meaning on `SpeakV2Error`, and the pronunciation `Warning` codes (`PRONUNCIATION_WARNINGS`, `PRONUNCIATION_TOO_LONG`, `PRONUNCIATIONS_LIMIT_EXCEEDED`) now described as emitted rather than reserved. The other six reference files are unchanged +- API skill: the `/v2/listen` `Configure` scope reads "EOT thresholds, keyterms, language hints, and `numerals`" in all four places it is described (the Nova vs Flux STT comparison table, the "Pick Flux STT" list, all-APIs mistake 1, and Flux STT mistake 17). It had said "EOT thresholds and keyterms", which omitted the two fields `ListenV2Configure` also carries +- API skill, Flux STT mistake 18: the `ForceEndTurn` warning note no longer says the `Warning` message is absent from the AsyncAPI and cannot appear in `references/listen.md`. The reference shows the `ListenV2Warning` shape; its `code` is a free string, so the note now says the individual codes such as `FORCE_END_TURN_NO_ACTIVE_TURN` come from the Force End Turn docs page, and links it +- API skill, mistake 20: the `summarize` sentence on `/v1/read` said the reference contradicted the `v2` value. `references/read.md` types the parameter `v2` | boolean, so the skill now says the type is right and only the description still reads boolean-only +- API skill: the host notes for Voice Agent REST are scoped to the one endpoint they were measured on. "Voice Agent's REST endpoints live on the `agent.` host", "The Agent REST endpoints move with it", and the mistake 15 heading now name `GET /v1/agent/settings/think/models`, so the ten `/v1/projects/{project_id}/agents` and `agent-variables` paths added to the domain table are not read as living on `agent.deepgram.com` +- Text-to-speech skill: the Flux TTS SDK bullets now state that every SDK ships a Flux TTS client, Go from v3.8.0 as `pkg/client/speak/v2`, where they had said every SDK except Go and told Go users to use the WebSocket directly. They also record that only the Rust and .NET SDK skills document Flux TTS and that the JS, Python, Java, and Go `text-to-speech` SDK skills cover `/v1/speak` only, so `/v2/speak` message shapes come from this skill +- Text-to-speech skill: the batch bullet notes that batch is the only transport that honors inline pauses +- Text-to-speech skill: Common mistake 6 now says SSML is not supported and that the only markup Flux TTS interprets is its own escaped controls (pronunciation on both transports, pause on batch), and adds that Aura-2 pronunciation control is GA on `/v1/speak` with the same syntax for English and Spanish voices, a 2000-character input limit, and no pause control. The previous sentence claimed SSML is stripped with an `INPUT_MARKUP_STRIPPED` warning; a live `/v2/speak` session fed `` and `` markup returned audio with no `Warning` of any kind, so the claim is gone +- Speech-to-text skill: the model-family table names the two `redact` values Flux STT accepts, `numbers` and `aggressive_numbers`, and states that any other value fails the WebSocket handshake with 400, where it had said only "number redaction" +- Speech-to-text skill: the threshold-table introduction points at the `Configure` message as the mid-stream mechanism, and the `ForceEndTurn` bullet says "Flux STT" rather than bare "Flux" +- Voice-agent skill: `ForceEndTurn` spells out both failure modes. With a listen provider other than Flux STT (`v2`) the server sends a `FORCE_END_TURN_UNSUPPORTED` warning and the turn does not end; with no turn in progress the message is ignored silently. `UpdateListen` is described as changing `model` and `language` mid-session on any provider, with thresholds and language hints on Flux STT and keyterm updates on Flux STT models only +- Voice-agent skill: the `think` provider note labels `groq` and `aws_bedrock` as bring-your-own, gives `aws_bedrock` its `provider.credentials` block (`type` `iam`, or `sts` plus `session_token`, with `region`, `access_key_id`, `secret_access_key`) and the `https://bedrock-runtime.{region}.amazonaws.com/` endpoint, and notes that `nvidia` is on the LLM models page but absent from the `references/agent.md` provider enum, so its model name should be checked against `/v1/agent/settings/think/models` +- Voice-agent skill: the third-party TTS note states that `endpoint` takes `url` and `headers`, that `wss` URLs are accepted for Eleven Labs only, and that `aws_polly` also requires `credentials`; the Deepgram-managed Cartesia exception (no `endpoint`) stays, as the TTS models page documents it +- Voice-agent skill: the SDK routing line records that `FunctionCallCancelled` and `defer_until_eot` ship in JS 5.12.0, Python 7.10.0, and Java 0.10.1 while the SDK voice-agent skills do not describe them, so message shapes come from this skill +- CLI skill: tracks `deepctl` 0.3.1, published 2026-09-29. The version-scoped statements that still hold on 0.3.1 are relabelled from 0.3.0: the config-file write by `dg update --check-only`, the 23-command surface, the four hardcoded `dg skills` downloads, the single-tool `dg mcp` proxy, the HTTP 404 on regional hosts, and the plain-transcript output of `--srt` and `--webvtt`. pip, uv, and pipx install 0.3.1; the Homebrew tap still pins `deepctl-0.2.26` +- CLI skill: the global-flag rule is now exact. `-o`, `-q`, and `-v` are accepted after the subcommand on 0.3.1, so `dg listen file.wav -o json` and `dg -o json listen file.wav` are equivalent, except on `dg speak`, where `-o` after the subcommand is the output file. `--base-url`, `--api-key`, `-p`, `-c`, and `--timing` still fail with `No such option` after the subcommand, and the "global flag after the subcommand" mistake uses `--base-url` as its example because `dg listen -v file.wav` now succeeds +- CLI skill: `dg models` is described by its 0.3.1 output: 553 rows, 942 with `--include-outdated`, each row carrying both a bare `name` (`asteria`) and a `canonical_name` (`aura-2-asteria-en`) plus a `deprecated` flag. It still lists no Flux STT or Flux TTS models, so the advice to take model names from the `speech-to-text` and `text-to-speech` skills stands +- CLI skill: `dg whoami -o json` now has a `key_source` field that names where the key came from, for example `DEEPGRAM_API_KEY (env)`; the claim that it mislabelled an environment key as `config file` is gone +- CLI skill: `--agent-friendly` is accepted by every command except the `skills`, `debug`, and `plugin` groups, which reject it with `No such option` +- Setup-mcp skill: the upgrade advice for Path A no longer tells the user to run `dg update`, which the CLI skill documents as a reporter that prints `installation_method: null` on a pip install. Both skills now say to upgrade through the installer that put `deepctl` there (`pip install -U deepctl`, `uv tool upgrade deepctl`, `pipx upgrade deepctl`, `brew upgrade deepgram`, or the install script) and to use `dg update --check-only` to see whether a newer release exists. The Path A troubleshooting fallback says the same +- Recipes skill: the Nova row lists all 26 `speech-to-text/v1` recipes rather than 16; `punctuate`, `multichannel`, `streaming-file`, `filler-words`, `replace`, `keyterm`, `profanity-filter`, `dictation`, `numerals`, and `measurements` were missing. The Flux STT row now describes the two recipes that exist under `speech-to-text/v2`: `streaming` (the `/v2/listen` WebSocket with `model=flux-general-en`, `encoding=linear16`, `sample_rate=16000`, printing `TurnInfo` events in place of v1 interim/final pairs) and `transcribe-url` (`model=flux-general-en` on the SDK's prerecorded call with `smart_format`), with the CLI carrying only `transcribe-url`. It no longer claims recipes for EOT, eager EOT, mid-session `Configure`, or keyterms, none of which exist +- Starters skill: `sinatra-transcription`, still listed in the `dg init` gallery, is described as archived and private, so the clone returns 404 for anyone outside Deepgram, rather than only archived + ### Fixed +- API skill, Flux STT mistake 17: the `Configure` example sent `"eot_threshold": "0.8"` and `"eot_timeout_ms": "3000"` as JSON strings. `eot_threshold` is a number and `eot_timeout_ms` an integer, and the Flux STT Configure docs send them unquoted, so the example now reads `0.8` and `3000` +- API skill: removed `GET /v1/auth/token` from the regional-endpoints table. No such path exists; the only auth path is `POST /v1/auth/grant` +- Audio-intelligence skill: streaming `entities` on `/v1/listen` `Results` are present only on messages whose `is_final` is `true`, and a final result with nothing detected carries `"entities": []`. The skill had said the array was on every `Results` message, which sent readers looking for entities on interim results that never carry the key +- Voice-agent skill: three bare "Flux" mentions (the LiveKit swap sentence, the `listen` field note, and the `UpdateListen`/`ForceEndTurn` bullet) now read "Flux STT" or "Flux TTS" +- CLI skill: the exit-code mistake no longer says a missing API key exits 0. On 0.3.1 both `dg listen` and `dg projects` print `Error: DEEPGRAM_API_KEY is not set …` and exit 1; the Ctrl-C hole on `dg listen --mic` and `dg mcp` is still documented +- CLI skill: the `dg init` gallery caveat says `sinatra-transcription` points at an archived, private repository whose URL returns 404, not merely an archived one +- CLI skill: the agentic-tools source link moved from `/agentic-tools` to `/developer-tools/agentic-tools` +- Setup-mcp skill: Path B now says `deepgram-mcp` is a PyPI package and that the npm package of the same name is unrelated third-party code that also asks for `DEEPGRAM_API_KEY`, so `npx deepgram-mcp` must not be used +- Self-hosted skill: the SageMaker reference's Java section said Maven Central's latest Java SDK was `0.10.0`; it is `0.10.2`, which still satisfies the `0.4.0` floor the transport needs +- Self-hosted skill: the Docker/Podman reference dropped its link to `/docs/deploy-deepgram-services`, which redirects to `/docs/deploy-stt-services`, a page the same list already cites - Browser-agent skill: the `ttl` versus `ttl_seconds` mistake no longer claims that browser-agent documentation snippets still show `ttl`. Those snippets were corrected upstream. The mistake itself, the observed `expires_in` values, and the advice to read `expires_in` rather than trust the field name all stand; the reason given for that advice is now the behavior that causes it, which is that `/v1/auth/grant` ignores any field it does not recognize and still answers HTTP 200 -[Unreleased]: https://github.com/deepgram/skills/compare/deepgram-skills-v1.6.0...HEAD +[1.7.0]: https://github.com/deepgram/skills/compare/deepgram-skills-v1.6.0...deepgram-skills-v1.7.0 ## [1.6.0] - 2026-09-18 From 9a4319eeda21384965f76ecc365b29b8d278bc9a Mon Sep 17 00:00:00 2001 From: Corey Weathers Date: Thu, 1 Oct 2026 09:48:25 -0400 Subject: [PATCH 08/17] docs(api,voice-agent): remove em dashes from edited lines; drop source-comparison wording --- skills/api/SKILL.md | 12 ++++++------ skills/voice-agent/SKILL.md | 6 +++--- 2 files changed, 9 insertions(+), 9 deletions(-) diff --git a/skills/api/SKILL.md b/skills/api/SKILL.md index c6cf0ab..a245433 100644 --- a/skills/api/SKILL.md +++ b/skills/api/SKILL.md @@ -214,7 +214,7 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// ### All APIs -1. **Feature flags are query params — except for Voice Agent and the v2 mid-session updates.** For `/v1/listen`, `/v2/listen`, `/v1/speak`, and `/v2/speak`, initial options go on the URL. The request body carries only audio data (REST) or audio frames (WebSocket). Exceptions: `/v1/agent/converse` has no URL query params at all (all config goes in the `Settings` message); `/v2/listen` supports a `Configure` message after connection to update EOT thresholds, keyterms, language hints, and `numerals` mid-session; and `/v2/speak` supports a `Configure` message that updates `speed` only. Also note that `/v2/listen` has a much smaller param set than `/v1/listen` — flags like `smart_format`, `diarize_model`, and `punctuate` are not available. +1. **Feature flags are query params, except for Voice Agent and the v2 mid-session updates.** For `/v1/listen`, `/v2/listen`, `/v1/speak`, and `/v2/speak`, initial options go on the URL. The request body carries only audio data (REST) or audio frames (WebSocket). Exceptions: `/v1/agent/converse` has no URL query params at all (all config goes in the `Settings` message); `/v2/listen` supports a `Configure` message after connection to update EOT thresholds, keyterms, language hints, and `numerals` mid-session; and `/v2/speak` supports a `Configure` message that updates `speed` only. Also note that `/v2/listen` has a much smaller param set than `/v1/listen`: flags like `smart_format`, `diarize_model`, and `punctuate` are not available. 2. **Rate limits are concurrent connections, not total requests.** A 429 means too many simultaneous open connections, not too high a request volume. Diarization and other compute-heavy features reduce your concurrency allowance further. @@ -242,7 +242,7 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// 11. **Streaming is raw audio only, and rejects anything it doesn't recognize.** The WebSocket emits non-containerized audio, so `encoding` is limited to `linear16` (default), `mulaw`, or `alaw`. The compressed and containerized encodings (`mp3`, `opus`, `flac`, `aac`) and the `container`, `bit_rate`, `callback`, `callback_method`, and `priority` params are **batch-only** — sending them to the socket fails the connection, as does any unknown or misspelled param. Use the batch REST transport when you need compressed output. -12. **Insert whitespace between separate generations — the server won't.** Text normalization runs before synthesis, but successive `Speak` messages are concatenated verbatim. Sending `"Hello world."` then `"How are you?"` is processed as `"Hello world.How are you?"`, which causes sentence-boundary artifacts. Add a single space (or the right separator for non-whitespace languages) when you stitch a reply, a tool-call result, and another reply together. Send plain text: SSML is not interpreted, and the only markup Flux TTS honors is its own escaped inline controls. A pronunciation override `\{"word":"...","pronounce":""\}` is honored on both transports (Early Access) but only with `speed` 1.0, and a pause marker `\{pause:500ms\}` is batch-only. A pause marker on the socket, or a pronunciation control on a socket whose `speed` is not 1.0, fails the connection with `DATA-0002`. See [Speed, Pause, Pronunciation](https://developers.deepgram.com/docs/tts-voice-controls). +12. **Insert whitespace between separate generations, because the server won't.** Text normalization runs before synthesis, but successive `Speak` messages are concatenated verbatim. Sending `"Hello world."` then `"How are you?"` is processed as `"Hello world.How are you?"`, which causes sentence-boundary artifacts. Add a single space (or the right separator for non-whitespace languages) when you stitch a reply, a tool-call result, and another reply together. Send plain text: SSML is not interpreted, and the only markup Flux TTS honors is its own escaped inline controls. A pronunciation override `\{"word":"...","pronounce":""\}` is honored on both transports (Early Access) but only with `speed` 1.0, and a pause marker `\{pause:500ms\}` is batch-only. A pause marker on the socket, or a pronunciation control on a socket whose `speed` is not 1.0, fails the connection with `DATA-0002`. See [Speed, Pause, Pronunciation](https://developers.deepgram.com/docs/tts-voice-controls). ### Voice Agent (`/v1/agent/converse`) @@ -253,19 +253,19 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// { "agent": { "speak": { "provider": { "type": "deepgram", "version": "v2", "model": "flux-alexis-en" } } } } ``` -15. **`GET /v1/agent/settings/think/models` lives on `agent.deepgram.com`, not `api.deepgram.com`.** `GET /v1/agent/settings/think/models` — the list of LLMs you can name in `agent.think.provider` — returns **404 on `api.deepgram.com`** and 200 on `agent.deepgram.com`. Same key, same path; only the host differs, so a client with one hardcoded base URL silently gets a 404 that looks like a missing feature. The three regional `api.*` hosts serve it as well. +15. **`GET /v1/agent/settings/think/models` lives on `agent.deepgram.com`, not `api.deepgram.com`.** `GET /v1/agent/settings/think/models`, the list of LLMs you can name in `agent.think.provider`, returns **404 on `api.deepgram.com`** and 200 on `agent.deepgram.com`. Same key, same path; only the host differs, so a client with one hardcoded base URL silently gets a 404 that looks like a missing feature. The three regional `api.*` hosts serve it as well. ### Flux STT model (`/v2/listen`) 16. **Use `/v2/listen` and a `flux-general-*` model.** Two are served: `flux-general-en` (English) and `flux-general-multi` (multilingual, and the only model that accepts `language_hint` / `language_hints`). `/v1/listen` does not support Flux STT, and `model=flux` alone is not a valid value. Do not include `language` or `encoding` params for containerized audio. -17. **Use `Configure` to update EOT thresholds, keyterms, language hints, and `numerals` mid-session.** Unlike `/v1/listen`, Flux STT supports live reconfiguration after connection — no need to reconnect to change turn detection sensitivity, boost new keyterms, re-bias language detection (`language_hints`, `flux-general-multi` only), or switch `numerals` on for a PIN or order number: +17. **Use `Configure` to update EOT thresholds, keyterms, language hints, and `numerals` mid-session.** Unlike `/v1/listen`, Flux STT supports live reconfiguration after connection, so there is no need to reconnect to change turn detection sensitivity, boost new keyterms, re-bias language detection (`language_hints`, `flux-general-multi` only), or switch `numerals` on for a PIN or order number: ```json { "type": "Configure", "thresholds": { "eot_threshold": 0.8, "eot_timeout_ms": 3000 }, "keyterms": ["Deepgram"] } ``` The server responds with `ConfigureSuccess` (echoing back applied values) or `ConfigureFailure`, which carries `code` and `description` identifying the rejected configuration. Omitted threshold fields keep their current values. -18. **`ForceEndTurn` outside a turn is a `Warning`, not an error — and the socket stays open.** Sending `{"type":"ForceEndTurn"}` while no turn is in progress returns `{"type":"Warning","code":"FORCE_END_TURN_NO_ACTIVE_TURN","description":"Received ForceEndTurn while no turn was active; the request was ignored."}` and the connection continues. Do not treat it as fatal or reconnect. `references/listen.md` shows the message shape (`ListenV2Warning`: `code`, `description`, `request_id`, `sequence_id`); `code` is a free string there, so the individual codes such as `FORCE_END_TURN_NO_ACTIVE_TURN` come from the [Force End Turn](https://developers.deepgram.com/docs/flux/force-end-turn) docs. When `ForceEndTurn` *does* land mid-turn, the resulting `TurnInfo` carries `event: "EndOfTurn"` with `trigger: "manual"` — `trigger` is `model` | `manual` | `timeout`, it appears on `EndOfTurn` and nowhere else, and it is an open enum, so tolerate values you do not recognize. +18. **`ForceEndTurn` outside a turn is a `Warning`, not an error, and the socket stays open.** Sending `{"type":"ForceEndTurn"}` while no turn is in progress returns `{"type":"Warning","code":"FORCE_END_TURN_NO_ACTIVE_TURN","description":"Received ForceEndTurn while no turn was active; the request was ignored."}` and the connection continues. Do not treat it as fatal or reconnect. `references/listen.md` shows the message shape (`ListenV2Warning`: `code`, `description`, `request_id`, `sequence_id`); `code` is a free string there, so the individual codes such as `FORCE_END_TURN_NO_ACTIVE_TURN` come from the [Force End Turn](https://developers.deepgram.com/docs/flux/force-end-turn) docs. When `ForceEndTurn` *does* land mid-turn, the resulting `TurnInfo` carries `event: "EndOfTurn"` with `trigger: "manual"`. `trigger` is `model` | `manual` | `timeout`, it appears on `EndOfTurn` and nowhere else, and it is an open enum, so tolerate values you do not recognize. ### Nova diarization (`/v1/listen`) @@ -273,7 +273,7 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// ### Text and Audio Intelligence (`/v1/read`, `/v1/listen`) -20. **`language` is required on `/v1/read`, and it is validated before anything else.** There is no default, despite what `references/read.md` says: omitting it returns `400 INVALID_QUERY_PARAMETER` — "Failed to deserialize query parameters: missing field `language`" — which masks every other problem in the request. English only — `language=multi` is rejected, and `en-US` is accepted but echoed back as `en`. Two more `/v1/read` shapes worth knowing: the JSON body takes **exactly one** of `text` or `url` (both or neither gives `PAYLOAD_ERROR`, and `url` must point at a plain-text document — audio gives `REMOTE_CONTENT_ERROR`), and it is POST-only (`GET` and a WebSocket upgrade both return 405). `summarize` on `/v1/read` accepts `v2` as well as `true`; `references/read.md` types it `v2` | boolean, and only its description still says boolean-only, so trust the type. Result paths differ per endpoint: `/v1/read` returns `results.summary.text`, `/v1/listen` returns `results.summary.short`, so code that handles both has to branch. (`sentiment` maps to `results.sentiments` on both.) +20. **`language` is required on `/v1/read`, and it is validated before anything else.** There is no default, despite what `references/read.md` says: omitting it returns `400 INVALID_QUERY_PARAMETER` with the message "Failed to deserialize query parameters: missing field `language`", which masks every other problem in the request. English only: `language=multi` is rejected, and `en-US` is accepted but echoed back as `en`. Two more `/v1/read` shapes worth knowing: the JSON body takes **exactly one** of `text` or `url` (both or neither gives `PAYLOAD_ERROR`, and `url` must point at a plain-text document, since audio gives `REMOTE_CONTENT_ERROR`), and it is POST-only (`GET` and a WebSocket upgrade both return 405). `summarize` on `/v1/read` accepts `v2` as well as `true`; `references/read.md` types it `v2` | boolean, and only its description still says boolean-only, so trust the type. Result paths differ per endpoint: `/v1/read` returns `results.summary.text`, `/v1/listen` returns `results.summary.short`, so code that handles both has to branch. (`sentiment` maps to `results.sentiments` on both.) 21. **On the Nova streaming socket, only `detect_entities` works — and the other four fail in three different ways.** `detect_entities=true` is supported and puts `entities` at the **top level** of each `Results` message, beside `channel`, not inside `channel.alternatives[0]`. The other four are prerecorded-only: `summarize` fails the handshake with `400 "Summarization is not available for streaming."`; `topics` and `intents` fail it with `403 UNAUTHORIZED_FEATURES_REQUESTED`, which reads like a key-permissions problem even when the same key's prerecorded `topics`/`intents` calls return 200; and `sentiment` is the trap — the handshake succeeds, no error is ever sent, and sentiment simply never appears in the results. diff --git a/skills/voice-agent/SKILL.md b/skills/voice-agent/SKILL.md index 5f1544d..6cc2a85 100644 --- a/skills/voice-agent/SKILL.md +++ b/skills/voice-agent/SKILL.md @@ -71,11 +71,11 @@ Then open the WebSocket to `wss://agent.deepgram.com/v1/agent/converse` with the Field notes, from the configure and model pages [3][6][7][8]: - `listen`: Flux STT (`flux-general-en`, or `flux-general-multi` with `language_hints`) requires `"version": "v2"` and gives model-integrated end-of-turn detection. Nova (`nova-3`) uses `v1`, the default, and adds `smart_format` and `language`. Drop `version` with a Flux STT model and the agent falls back to the v1 endpoint, where `flux-general-en` is not a valid model. [7][20] -- `think`: `provider.type` is `open_ai`, `anthropic`, `google`, or `nvidia` (managed; `endpoint` optional) or `groq` or `aws_bedrock` (bring your own; `endpoint` required). `nvidia` is on the LLM models page but not in the `references/agent.md` provider enum, so check its model name against the models endpoint above. `aws_bedrock` authenticates with `provider.credentials` (`type` `iam`, or `sts` plus `session_token`, with `region`, `access_key_id`, and `secret_access_key`) and points `endpoint.url` at `https://bedrock-runtime.{region}.amazonaws.com/`. Bring your own LLM by keeping `type: open_ai` and setting `endpoint.url` to any OpenAI Chat Completions-compatible URL, with `endpoint.headers` for its auth. Pass an array of providers to get an ordered fallback chain. Managed-LLM prompts are limited to 25,000 characters. [6] +- `think`: `provider.type` is `open_ai`, `anthropic`, `google`, or `nvidia` (managed; `endpoint` optional) or `groq` or `aws_bedrock` (bring your own; `endpoint` required). `nvidia` model names are listed on the LLM models page; confirm the exact string against the models endpoint above before sending it. `aws_bedrock` authenticates with `provider.credentials` (`type` `iam`, or `sts` plus `session_token`, with `region`, `access_key_id`, and `secret_access_key`) and points `endpoint.url` at `https://bedrock-runtime.{region}.amazonaws.com/`. Bring your own LLM by keeping `type: open_ai` and setting `endpoint.url` to any OpenAI Chat Completions-compatible URL, with `endpoint.headers` for its auth. Pass an array of providers to get an ordered fallback chain. Managed-LLM prompts are limited to 25,000 characters. [6] - `speak`: `"version": "v2"` selects Flux TTS (`flux-{voice}-{language}`); `v1`, the default when you name a provider, selects Aura (`aura-2-thalia-en`). Omit `agent.speak` entirely and you get Flux TTS with `flux-kit-en`. Flux TTS streams raw audio only: `encoding` must be `linear16`, `mulaw`, or `alaw`, `container` must be `none`, and `mp3` or `wav` returns `INVALID_SETTINGS`. `provider.speed` (default `1.0`) is `0.5` to `1.5` in `0.05` steps on Flux TTS and any value from `0.7` to `1.5` on Aura; a value the family does not accept ends the session with `FAILED_TO_SPEAK`. `provider.expressivity` (whole numbers `-2` to `2`, default `0`) is Flux TTS (`v2`) only and fixed for the session; it is beta and `0` is the only value validated for production. Third-party TTS (`open_ai`, `eleven_labs`, `cartesia`, `aws_polly`) takes an `endpoint` with `url` and `headers`, and `wss` URLs are accepted for Eleven Labs only; `aws_polly` also requires `credentials` (`type` `sts` or `iam`, with `region`, `access_key_id`, `secret_access_key`, and `session_token` for STS). Deepgram-managed Cartesia (`type: cartesia` with no `endpoint`) is the exception. [8][34] - `agent.context.messages` replays earlier turns as `{"type":"History","role":"user","content":"..."}` or `{"type":"History","function_calls":[{"id","name","client_side","arguments","response"}]}` so a new session continues an old one. While `Settings.flags.history` is `true` (the default) the server sends `History` messages in the same two shapes; set it to `false` to turn them off. [3][35] - Other knobs, with ranges in `references/agent.md`: `think.context_length` (`max` or a character count; custom `think.endpoint` only), `think.provider.reasoning_mode` (`none` to `high`, on `open_ai` and `groq`), and the top-level `tags`, `experimental`, and `mip_opt_out`. [3][12] -- Conversational Mode (`agent.think_conversational.provider`: backchanneling, presence checks, frustration detection) is invite-only Early Access; a `Settings` message that names it from an unenrolled project is rejected with `UNPARSABLE_CLIENT_MESSAGE`, and its page is not in the documentation index. [36] +- Conversational Mode (`agent.think_conversational.provider`: backchanneling, presence checks, frustration detection) is invite-only Early Access; a `Settings` message that names it from an unenrolled project is rejected with `UNPARSABLE_CLIENT_MESSAGE`. [36] ## Reusable agent configurations @@ -191,4 +191,4 @@ The Voice Agent API is billed per minute of WebSocket connection time, and a Dee 33. https://developers.deepgram.com/docs/voice-agent-speculative-replies and https://developers.deepgram.com/docs/voice-agent-function-call-cancelled 34. https://developers.deepgram.com/docs/voice-agent-tts-controls (`speed` ranges per family, `expressivity` on Flux TTS only) 35. https://developers.deepgram.com/docs/voice-agent-history -36. https://developers.deepgram.com/docs/voice-agent-conversational-mode (invite-only Early Access; absent from https://developers.deepgram.com/llms.txt) +36. https://developers.deepgram.com/docs/voice-agent-conversational-mode (invite-only Early Access) From 6e6010bc3eeee0739e18a00a093c7f1cbc166275 Mon Sep 17 00:00:00 2001 From: Corey Weathers Date: Thu, 1 Oct 2026 09:50:45 -0400 Subject: [PATCH 09/17] docs(voice-agent): enumerate reasoning_mode values; changelog: six sources --- CHANGELOG.md | 2 +- skills/voice-agent/SKILL.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f1de910..745cf41 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -33,7 +33,7 @@ Catch-up with the September API, spec, SDK, CLI, and documentation changes. Flux - Voice-agent skill: `FunctionCallCancelled` in the message lifecycle table (the user started speaking again; stop work on each `id` and send no `FunctionCallResponse`, a late one is dropped), and a speculative-dispatch paragraph in the function-calling section. Function calls go out before the turn is confirmed by default; `defer_until_eot: true` holds a call until the turn is confirmed and discards it if the turn resumes, for actions that cannot be undone, and an `endpoint` function that already ran is not rolled back - Voice-agent skill: `agent.speak.provider.speed` (Flux TTS `0.5` to `1.5` in `0.05` steps, Aura `0.7` to `1.5`; an unaccepted value ends the session with `FAILED_TO_SPEAK`) and `expressivity` (whole numbers `-2` to `2`, Flux TTS `v2` only, fixed for the session, beta with `0` the only production-validated value) in the `speak` field notes - Voice-agent skill: `Settings.flags.history` (default `true`) and the server `History` message in both its conversation-text and `function_calls` shapes, alongside the existing `agent.context.messages` input; a pointer bullet for `think.context_length`, `think.provider.reasoning_mode`, and the top-level `tags`, `experimental`, and `mip_opt_out`, with ranges in `references/agent.md`; a one-line note that Conversational Mode (`agent.think_conversational.provider`) is invite-only Early Access answered with `UNPARSABLE_CLIENT_MESSAGE` from unenrolled projects -- Voice-agent skill: five new sources (reusable agent configurations, speculative replies, function call cancelled, TTS controls, history, conversational mode) and two new description triggers, "reusable agent configuration" and "defer_until_eot" +- Voice-agent skill: six new sources (reusable agent configurations, speculative replies, function call cancelled, TTS controls, history, conversational mode) and two new description triggers, "reusable agent configuration" and "defer_until_eot" - Self-hosted skill: the SageMaker reference now carries the ordered instance-pool recommendation above the instance table. A single instance type has no fallback, so when the Availability Zone is short of that GPU the endpoint goes `Failed` (`Request to service failed` or `InsufficientInstanceCapacity`), which happens routinely for popular GPU types. The pool order is the listing's recommended type first, same-or-newer generations with similar per-instance capacity next (`g6`, `g6e`, `g7`), older generations last as insurance, never an unsupported type (`g4dn` for Flux STT, `g5` and `g4dn` for Flux TTS, single-GPU types for Aura-2), up to 5 types with three as the sweet spot. `VariantInstanceProvisionTimeoutInSeconds` is the per-type wait (`300` recommended, `60` to `3600` allowed), and quota does not fall back: every pooled type needs a regional quota of at least `1` or `CreateEndpoint` fails with `ResourceLimitExceeded`. Also one line that AWS field employees can reach Deepgram models through the AWS Marketplace Field Demonstration Program, for which Deepgram is an eligible provider - Self-hosted skill: the Kubernetes reference states that the `API/Engine -> License Proxy -> Billing` chain for air-gapped HA needs a chart newer than `0.46.0`. On `0.46.0` and earlier, with `billing.enabled` and `licenseProxy.enabled` both `true`, the `billing` condition takes precedence, so API and Engine connect to Billing directly and the deployed License Proxy receives no traffic; the fix is the `Unreleased` entry in `charts/deepgram-self-hosted/CHANGELOG.md` - Examples skill: the category map and description now cover the integrations the live `deepgram/examples` tree carries that the skill had omitted: Telnyx and Plivo (telephony), Webex, Jitsi, and Microsoft Teams (the "Recording platforms" row is now "Meeting platforms", split into recordings on Nova prerecorded and live meetings on Nova live, since the Jitsi bridge and the Teams bot stream), Haystack and Semantic Kernel (LLM frameworks), Gin (web frameworks), a **Broadcast** row for the OBS Studio native C captioning plugin, and an **Infrastructure** row for the Node and Python Deepgram API proxy servers, Silero VAD speech segmentation, and the multi-provider chat-completions proxy that serves as a Voice Agent `think.endpoint.url`. The two `530-*` directories are named in full because they share a number. The statement that no example covers Flux TTS stands: the Voice Agent proxy keeps TTS on `aura-2` diff --git a/skills/voice-agent/SKILL.md b/skills/voice-agent/SKILL.md index 6cc2a85..ed3def7 100644 --- a/skills/voice-agent/SKILL.md +++ b/skills/voice-agent/SKILL.md @@ -74,7 +74,7 @@ Field notes, from the configure and model pages [3][6][7][8]: - `think`: `provider.type` is `open_ai`, `anthropic`, `google`, or `nvidia` (managed; `endpoint` optional) or `groq` or `aws_bedrock` (bring your own; `endpoint` required). `nvidia` model names are listed on the LLM models page; confirm the exact string against the models endpoint above before sending it. `aws_bedrock` authenticates with `provider.credentials` (`type` `iam`, or `sts` plus `session_token`, with `region`, `access_key_id`, and `secret_access_key`) and points `endpoint.url` at `https://bedrock-runtime.{region}.amazonaws.com/`. Bring your own LLM by keeping `type: open_ai` and setting `endpoint.url` to any OpenAI Chat Completions-compatible URL, with `endpoint.headers` for its auth. Pass an array of providers to get an ordered fallback chain. Managed-LLM prompts are limited to 25,000 characters. [6] - `speak`: `"version": "v2"` selects Flux TTS (`flux-{voice}-{language}`); `v1`, the default when you name a provider, selects Aura (`aura-2-thalia-en`). Omit `agent.speak` entirely and you get Flux TTS with `flux-kit-en`. Flux TTS streams raw audio only: `encoding` must be `linear16`, `mulaw`, or `alaw`, `container` must be `none`, and `mp3` or `wav` returns `INVALID_SETTINGS`. `provider.speed` (default `1.0`) is `0.5` to `1.5` in `0.05` steps on Flux TTS and any value from `0.7` to `1.5` on Aura; a value the family does not accept ends the session with `FAILED_TO_SPEAK`. `provider.expressivity` (whole numbers `-2` to `2`, default `0`) is Flux TTS (`v2`) only and fixed for the session; it is beta and `0` is the only value validated for production. Third-party TTS (`open_ai`, `eleven_labs`, `cartesia`, `aws_polly`) takes an `endpoint` with `url` and `headers`, and `wss` URLs are accepted for Eleven Labs only; `aws_polly` also requires `credentials` (`type` `sts` or `iam`, with `region`, `access_key_id`, `secret_access_key`, and `session_token` for STS). Deepgram-managed Cartesia (`type: cartesia` with no `endpoint`) is the exception. [8][34] - `agent.context.messages` replays earlier turns as `{"type":"History","role":"user","content":"..."}` or `{"type":"History","function_calls":[{"id","name","client_side","arguments","response"}]}` so a new session continues an old one. While `Settings.flags.history` is `true` (the default) the server sends `History` messages in the same two shapes; set it to `false` to turn them off. [3][35] -- Other knobs, with ranges in `references/agent.md`: `think.context_length` (`max` or a character count; custom `think.endpoint` only), `think.provider.reasoning_mode` (`none` to `high`, on `open_ai` and `groq`), and the top-level `tags`, `experimental`, and `mip_opt_out`. [3][12] +- Other knobs, with ranges in `references/agent.md`: `think.context_length` (`max` or a character count; custom `think.endpoint` only), `think.provider.reasoning_mode` (`none`, `minimal`, `low`, `medium`, or `high`, on `open_ai` and `groq`), and the top-level `tags`, `experimental`, and `mip_opt_out`. [3][12] - Conversational Mode (`agent.think_conversational.provider`: backchanneling, presence checks, frustration detection) is invite-only Early Access; a `Settings` message that names it from an unenrolled project is rejected with `UNPARSABLE_CLIENT_MESSAGE`. [36] ## Reusable agent configurations From 606c55d50bd94873294498602d71eaf90ebb72bc Mon Sep 17 00:00:00 2001 From: Corey Weathers Date: Thu, 1 Oct 2026 11:28:27 -0400 Subject: [PATCH 10/17] docs: second October sweep across the 14 skills after a three-pass review Applies the residual findings from the 2026-10-01 audit of the API, specs, SDKs, CLI, and docs on top of the 1.7.0 catch-up, then the findings of a three-pass devrel review (style sweep, changelog and manifest integrity, six technical-accuracy lanes against the docs, the live API, GitHub tags, and the registries, and two coherence reads). Corrections: ForceEndTurn first ships in JS 5.9.0, Python 7.8.0, Java 0.9.0; only the Rust SDK skill documents Flux TTS; nova-3 sends an `entities` key on interim results too, so read entities from `is_final` only; the Read API's documented limit is 150K tokens (TOKEN_LIMIT_EXCEEDED); Java accessor is `client.voiceAgent()`; Homebrew install is `brew install deepgram/tap/deepgram`; `npx skills add` without git exits 1; `/v1/models` lists no Flux voices. Additions: regional hosts across the product skills, Flux STT CloseStream and Configure semantics, Flux TTS message and error inventory, Voice Agent fallback chains and deprecations, per-SDK availability of numerals, FunctionCallCancelled, and the agent-configuration REST clients, FIPS and SageMaker AMI constraints, deepctl 0.3.1 stdout contracts, text-intelligence limits and concurrency. Structure: long bullets split into sub-bullets, duplicate body statements collapsed, hedges replaced with the documented rule, changelog bullets consolidated and the 1.7.0 intro rewritten. AGENTS.md names all three workflows and validate-skills.ts; the spec-drift workflow comment reflects the generator's prune pass. Co-Authored-By: Claude Fable 5.1 --- .github/workflows/spec-drift.yml | 8 +- AGENTS.md | 20 +- CHANGELOG.md | 102 +++++++-- skills/api/SKILL.md | 16 +- skills/audio-intelligence/SKILL.md | 78 ++++--- skills/browser-agent/SKILL.md | 19 +- skills/cli/SKILL.md | 48 +++-- skills/recipes/SKILL.md | 6 +- skills/self-hosted/SKILL.md | 14 +- .../self-hosted/references/docker-podman.md | 3 +- skills/self-hosted/references/kubernetes.md | 4 +- skills/self-hosted/references/sagemaker.md | 6 +- skills/setup-mcp/SKILL.md | 29 ++- skills/speech-to-text/SKILL.md | 37 +++- skills/starters/SKILL.md | 25 ++- skills/text-intelligence/SKILL.md | 84 +++++--- skills/text-to-speech/SKILL.md | 193 +++++++++++------- skills/voice-agent/SKILL.md | 57 ++++-- 18 files changed, 512 insertions(+), 237 deletions(-) diff --git a/.github/workflows/spec-drift.yml b/.github/workflows/spec-drift.yml index e1a50b5..ff4ac0b 100644 --- a/.github/workflows/spec-drift.yml +++ b/.github/workflows/spec-drift.yml @@ -8,10 +8,10 @@ name: API Spec Drift # (`bun run scripts/fetch-specs.ts … && bun run scripts/generate-skills.ts`), and an # auto-commit would land unreviewed API copy in a skill agents read as reference. # -# Known limitation: `generate-skills.ts` overwrites the reference files it emits but does -# not prune ones it no longer emits, so a reference file for a *removed* endpoint group -# shows up as no drift here. Additions and edits are detected. This job is deliberately -# tolerant of that — it reports what it can see rather than asserting the tree is clean. +# `generate-skills.ts` deletes any reference file it did not emit on the run, so a +# reference file for a *removed* endpoint group shows up here as a deletion. Additions, +# edits, and removals are all detected; the `git add -A` in the report step is what makes +# new and deleted files visible to the diff. on: schedule: diff --git a/AGENTS.md b/AGENTS.md index f6bacfc..fc9e0c9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -14,11 +14,11 @@ voice agent, and audio intelligence APIs correctly. |------|---------| | `skills/` | The 14 shipped skills: `api`, `audio-intelligence`, `browser-agent`, `cli`, `docs`, `examples`, `recipes`, `self-hosted`, `setup-mcp`, `speech-to-text`, `starters`, `text-intelligence`, `text-to-speech`, `voice-agent`. `api` and `self-hosted` are the only two that carry a `references/` folder, and only `api`'s is generated | | `template/` | Starting point for a new skill (`SKILL.md` with YAML frontmatter) | -| `scripts/` | `fetch-specs.ts` and `generate-skills.ts` — regenerate the `api` skill from the public OpenAPI and AsyncAPI specs | +| `scripts/` | `fetch-specs.ts` and `generate-skills.ts` regenerate the `api` skill from the public OpenAPI and AsyncAPI specs; `validate-skills.ts` parses every `SKILL.md` frontmatter with the YAML parser the `skills` installer uses and diffs `.claude-plugin/marketplace.json` against the filesystem (`--remote` also checks the SDK plugins' skill paths through `gh api`) | | `.claude-plugin/` | Claude Code plugin-marketplace manifest — `metadata.version` is the released version, and `plugins[0].skills` is the list the installer reads | | `CHANGELOG.md` | Keep a Changelog / SemVer record; every release has an entry | | `package.json`, `bun.lock` | the single `yaml` dependency the generator needs | -| `.github/workflows/` | `context7.yml` only — refreshes Context7 on a published release | +| `.github/workflows/` | `context7.yml` refreshes Context7 on a published release; `spec-drift.yml` regenerates the `api` references from the live specs every Monday and opens or comments on a `spec-drift` issue when they differ, committing nothing; `validate-skills.yml` runs `validate-skills.ts` on every pull request and push to `main`, plus a report-only remote check of the SDK plugin skill paths and a ci-tools lint that skips until a `CI_TOOLS_READ_TOKEN` secret exists | ## Install (consumer side) @@ -38,9 +38,8 @@ npx skills add deepgram/skills --skill api -y # one skill `npx skills add` clones the repository with `git`. On an image without it (a bare `node:22-alpine`, for example) every target fails with `Failed to clone -...: Error: spawn git ENOENT`, yet the command exits 0 and installs nothing. -Install `git` first and check for the `SKILL.md` files rather than trusting -the exit code. +...: Error: spawn git ENOENT` and exits 1 with nothing installed. Install +`git` first, and check for the `SKILL.md` files after any headless install. Claude Code plugin route: `/plugin marketplace add deepgram/skills`, then `/plugin install deepgram@deepgram-agent-skills`. @@ -48,14 +47,16 @@ Claude Code plugin route: `/plugin marketplace add deepgram/skills`, then ## Regenerate and check headlessly (maintainer side) Requires [bun](https://bun.sh). There is no test suite; regeneration -completing and a clean `git diff` (or an intended one) is the check. +completing, a clean `git diff` (or an intended one), and `validate-skills.ts` +printing every skill as valid are the checks. ```bash bun run scripts/fetch-specs.ts https://dpgr.am/openapi.yml https://dpgr.am/asyncapi.yml bun install && bun run scripts/generate-skills.ts +bun run scripts/validate-skills.ts ``` -## Versions and conventions (as of 2026-09-18) +## Versions and conventions (as of 2026-10-01) - The generated `api` skill tracks the hourly-mirrored public specs at `https://dpgr.am/openapi.yml` and `https://dpgr.am/asyncapi.yml`. @@ -79,14 +80,15 @@ two files together and then ships a tag: triggers `.github/workflows/context7.yml`. Adding a skill also means adding its path to `plugins[0].skills` in -`.claude-plugin/marketplace.json`, or the installer will not offer it. +`.claude-plugin/marketplace.json`, or the installer will not offer it, and a +row to the skills table in `README.md`. ## Common failure modes | Symptom | Cause | Fix | |---------|-------|-----| | `npx skills add deepgram/` fails with a not-found or auth error | the target repository is private or does not exist | only the six public SDK repositories listed in README.md carry installable skills | -| every target fails with `spawn git ENOENT`, exit code 0, nothing installed | `git` is absent from the container or CI image; `npx skills add` shells out to it | install `git` (`apk add git` on Alpine) before running the installer | +| every target fails with `spawn git ENOENT`, exit code 1, nothing installed | `git` is absent from the container or CI image; `npx skills add` shells out to it | install `git` (`apk add git` on Alpine) before running the installer | | `bun: command not found` | bun not installed | install from https://bun.sh; the generation scripts are bun-only | | `ENOENT ... specs/openapi.yml` from `generate-skills.ts` | `fetch-specs.ts` was not run first; `specs/` is gitignored, so it is absent in a fresh clone | run both regeneration commands in order | | Regenerated `api` skill shows unexpected churn | the upstream specs moved | inspect the spec diff first; the specs are the source of truth | diff --git a/CHANGELOG.md b/CHANGELOG.md index 745cf41..2893a90 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -11,65 +11,117 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [1.7.0] - 2026-10-01 -Catch-up with the September API, spec, SDK, CLI, and documentation changes. Flux TTS inline pause and pronunciation controls, the Flux STT `Warning` message and mid-stream `numerals`, the Voice Agent reusable-configuration and agent-variable REST surface, `FunctionCallCancelled` and `defer_until_eot`, the Go SDK's Flux TTS client, and `deepctl` 0.3.1 are the headline items. No skill is added or removed, so the `deepgram` plugin still lists 14 skills. +Catch-up with the September API, spec, SDK, CLI, and documentation changes. The headline items are Flux TTS inline pause and pronunciation controls, the Flux STT `Warning` message and mid-stream `numerals`, the Voice Agent reusable-configuration and agent-variable REST surface, and `FunctionCallCancelled` with `defer_until_eot`. Per-SDK availability is corrected across the product skills: which SDKs ship a Flux TTS client, `ForceEndTurn`, mid-stream `numerals`, and the agent-configuration REST clients. The product skills gain the regional hosts `api.eu`, `api.au`, and `api.in.deepgram.com`, the CLI skill tracks `deepctl` 0.3.1 and the Homebrew 6 install form, and the self-hosted skill carries the FIPS constraints and the SageMaker AMI requirement. No skill is added or removed, so the `deepgram` plugin still lists 14 skills. ### Added - API skill: the API Domains table now lists the whole Voice Agent REST surface, not just `GET agent.deepgram.com/v1/agent/settings/think/models`. Reusable agent configurations are `GET` and `POST /v1/projects/{project_id}/agents` plus `GET`, `PUT`, and `DELETE /v1/projects/{project_id}/agents/{agent_id}`; agent variables are `GET` and `POST /v1/projects/{project_id}/agent-variables` plus `GET`, `PATCH`, and `DELETE /v1/projects/{project_id}/agent-variables/{variable_id}`. All ten were already rendered in `references/agent.md`; the table is the router an agent reads first, and it had never pointed at them. The OpenAPI carries one `servers` block for the whole document, so these ten paths are listed without a host - API skill: the Models row names the four model endpoints rather than one: `GET /v1/models`, `GET /v1/models/{model_id}`, `GET /v1/projects/{project_id}/models`, and `GET /v1/projects/{project_id}/models/{model_id}`, and records that `include_outdated=true` on either list call also returns non-latest model versions -- API skill, Flux TTS mistake 12: the inline-controls rule for `/v2/speak`. A pronunciation override `\{"word":"...","pronounce":""\}` is honored on both transports (Early Access) but only with `speed` 1.0, a pause marker `\{pause:500ms\}` is batch-only, and a pause marker on the socket, or a pronunciation control on a socket whose `speed` is not 1.0, fails the connection with `DATA-0002`. Links the Speed, Pause, Pronunciation page. The sentence that said SSML is stripped with an `INPUT_MARKUP_STRIPPED` warning is gone, for the reason given under the text-to-speech skill below; mistake 12 now says SSML is not interpreted and that the only markup Flux TTS honors is its own escaped inline controls -- API skill, Flux STT mistake 17: `ConfigureFailure` carries `code` and `description` identifying the rejected configuration -- Text-to-speech skill: Flux TTS inline controls, which the skill had not covered. Pronunciation control (Early Access) is an escaped JSON object in the text, `\{"word":"...","pronounce":""\}`, accepted on both the `/v2/speak` WebSocket and batch `POST /v2/speak`, at most 500 per request with IPA of at most 128 characters, and only with `speed` `1.0`: on the socket a pronunciation sent on a session opened with another speed, or after a `Configure` that set one, fails the connection with `DATA-0002`, and on batch the request is a 400 `CONTROL_COMBINATION_INVALID`. Pause control `\{pause:500ms\}` is batch only, 500 to 3000 ms in 100 ms steps, at most 8 per request, with `speed` capped at `1.15` while a pause is present (`PAUSE_SPEED_CAP_EXCEEDED`); a pause marker on the socket fails the connection with `DATA-0002`. The remaining batch 400 codes are listed (`BREAK_OUT_OF_RANGE`, `BREAK_INCREMENT_INVALID`, `BREAKS_LIMIT_EXCEEDED`, `BREAK_SYNTAX_INVALID`), as is how each transport reports what it applied: `SpeechMetadata.controls_applied` on the socket, `dg-pronunciations-applied` and `dg-breaks-applied` response headers on batch, with invalid IPA still applied best-effort and surfaced as a `PRONUNCIATION_WARNINGS` warning or the `dg-warnings` header +- API skill, Flux TTS mistake 12: the inline-controls rule for `/v2/speak`. A pronunciation override `\{"word":"...","pronounce":""\}` is honored on both transports (Early Access) but only with `speed` 1.0, a pause marker `\{pause:500ms\}` is batch-only, and a pause marker on the socket, or a pronunciation control on a socket whose `speed` is not 1.0, fails the connection with `DATA-0002`. On batch the same violations are a 400 whose `err_code` is one of six: `CONTROL_COMBINATION_INVALID`, `PAUSE_SPEED_CAP_EXCEEDED`, `BREAK_OUT_OF_RANGE`, `BREAK_INCREMENT_INVALID`, `BREAKS_LIMIT_EXCEEDED`, and `BREAK_SYNTAX_INVALID`, each given with the rule it names; a `speed` of exactly `1.0` never counts as a speed control. Links the Speed, Pause, Pronunciation page +- Text-to-speech skill: Flux TTS inline controls, which the skill had not covered. Pronunciation control (Early Access) is an escaped JSON object in the text, `\{"word":"...","pronounce":""\}`, accepted on both `/v2/speak` transports, at most 500 per request with IPA of at most 128 characters, and only with `speed` `1.0`. Pause control `\{pause:500ms\}` is batch only, 500 to 3000 ms in 100 ms steps, at most 8 per request, with `speed` capped at `1.15` while a pause is present. A violation on the socket fails the connection with `DATA-0002`; on batch it is a 400 carrying one of the six batch 400 codes. Each transport reports what it applied: `SpeechMetadata.controls_applied` on the socket, `dg-pronunciations-applied` and `dg-breaks-applied` response headers on batch, with invalid IPA still applied best-effort and surfaced as a `PRONUNCIATION_WARNINGS` warning or the `dg-warnings` header - Text-to-speech skill: the `Connected` message (`request_id`, `model_name`, `model_version`, `model_uuids`) and the `SessionMetadata` message (cumulative session totals, rebased by an `Interrupt` onto the audio the client actually played), neither of which the skill had named - Text-to-speech skill: `Interrupt` offset semantics. `playback_offset.value` is cumulative milliseconds since the session started, not since the current turn; each `Interrupt` must exceed the previous offset or it is ignored with `INVALID_INTERRUPT_OFFSET`; `audio_played_ms` from `SpeechInterrupted` is the baseline for the next offset. The old text produced wrong offsets after the first turn -- Text-to-speech skill: `CONTROL_COMBINATION_INVALID` as the fourth `ConfigureFailure` code, raised when a queued turn still carries a pronunciation control, with the fix (`Flush` that turn first) -- Text-to-speech skill: `NET-0003`, the 1-hour session cap, beside the existing `NET-0004` idle close +- Text-to-speech skill: the Flux TTS server message list in full, with `Error` (fatal, every `DOMAIN-NNNN` code is followed by a WebSocket close) distinguished from `Warning`, and `ConfigureSuccess`/`ConfigureFailure` named with their `field`/`value` shape. `ConfigureFailure` gains its fourth code, `CONTROL_COMBINATION_INVALID` (a queued turn still carries a pronunciation control; `Flush` that turn first), and its fifth, `INTERNAL_ERROR`, which the AsyncAPI reference lists for an acceptable configuration the server could not apply +- Text-to-speech skill: Aura `/v1/speak` WebSocket limits (2000 characters per request `BIG-0001`, 2400 characters per minute `DATA-0001`, 60 minutes per connection `NET-0003`), the same `NET-0003` closing a Flux TTS session at 1 hour beside the existing `NET-0004` idle close, and the regional hosts `api.eu`, `api.au`, and `api.in.deepgram.com` for `/v1/speak` and `/v2/speak` - Text-to-speech skill: a pointer from the Voice Agent mistake to https://developers.deepgram.com/docs/voice-agent-tts-controls, where `speed` and `expressivity` are set on `agent.speak.provider` - Speech-to-text skill: the Flux STT `Configure` message takes `numerals` alongside thresholds, keyterms, and `language_hints`. The skill's example message carries `"numerals":true`, states that the value starts from the `numerals` query parameter and applies only to transcripts Flux STT sends after it processes the update, and that it must be a JSON boolean: the string `"true"` fails schema validation, Flux STT answers with an `Error` of code `UNPARSABLE_CLIENT_MESSAGE`, and the connection closes. `ConfigureSuccess` echoes the full active configuration including `numerals`; `ConfigureFailure` carries `code` and `description` naming the rejected field and leaves the previous configuration in place - Speech-to-text skill: the `numerals` language scope on Flux STT. `flux-general-en` formats every number; `flux-general-multi` formats English, Spanish, French, German, Russian, Portuguese, Italian, and Dutch, and leaves Hindi and Japanese numbers as spoken -- Speech-to-text skill: SDK routing for mid-stream `numerals`. The JavaScript SDK has it from 5.13.0 (`socket.sendConfigure({type:"Configure", numerals:true})`), the Python SDK from 7.11.0 (`connection.send_configure(ListenV2Configure(numerals=True))`), and the Java SDK from 0.10.2 (`sendConfigure(ListenV2Configure.builder().numerals(true).build())`). The `deepgram-{lang}-conversational-stt` SDK skills do not cover it, so the skill tells the reader to take the message shape from its own `Configure` bullet +- Speech-to-text skill: SDK routing for mid-stream `numerals`. The JavaScript SDK has it from 5.13.0 (`socket.sendConfigure({type:"Configure", numerals:true})`), the Python SDK from 7.11.0 (`connection.send_configure(ListenV2Configure(numerals=True))`), and the Java SDK from 0.10.2 (`sendConfigure(ListenV2Configure.builder().numerals(true).build())`); the Go 3.8.0, Rust 0.11.0, and .NET 7.1.1 Configure types carry no `numerals` field, so the raw JSON message or the `numerals` query parameter is the path on those SDKs. The `deepgram-{lang}-conversational-stt` SDK skills do not cover it, so the skill tells the reader to take the message shape from its own `Configure` bullet - Speech-to-text skill: two Sources entries, https://developers.deepgram.com/docs/flux/configuration for the end-of-turn threshold table and https://developers.deepgram.com/docs/numerals for the Flux STT numerals language list and mid-stream toggling - Voice-agent skill: a "Reusable agent configurations" section. `Settings.agent` is either the full `agent` object or a Reusable Agent Configuration UUID string, the same `agent: "YOUR_AGENT_ID"` form the browser-agent skill already shows. The section gives the REST surface on `https://api.deepgram.com/v1` (`POST/GET /projects/{project_id}/agents`, `GET/PUT/DELETE .../agents/{agent_id}`), that `config` is the JSON string of the `agent` block and is immutable once created (`PUT` changes `metadata` only), that deleting a configuration a running service references breaks that service, the `INVALID_AGENT_ID` and `AGENT_ID_NOT_SUPPORTED` error codes, and the template-variable endpoints (`/projects/{project_id}/agent-variables`) with the `DG_` key form, unquoted substitution of any JSON value, `is_sensitive: false`, and the rule that configurations and variables are readable by every project member so hold no secrets - Voice-agent skill: `FunctionCallCancelled` in the message lifecycle table (the user started speaking again; stop work on each `id` and send no `FunctionCallResponse`, a late one is dropped), and a speculative-dispatch paragraph in the function-calling section. Function calls go out before the turn is confirmed by default; `defer_until_eot: true` holds a call until the turn is confirmed and discards it if the turn resumes, for actions that cannot be undone, and an `endpoint` function that already ran is not rolled back - Voice-agent skill: `agent.speak.provider.speed` (Flux TTS `0.5` to `1.5` in `0.05` steps, Aura `0.7` to `1.5`; an unaccepted value ends the session with `FAILED_TO_SPEAK`) and `expressivity` (whole numbers `-2` to `2`, Flux TTS `v2` only, fixed for the session, beta with `0` the only production-validated value) in the `speak` field notes - Voice-agent skill: `Settings.flags.history` (default `true`) and the server `History` message in both its conversation-text and `function_calls` shapes, alongside the existing `agent.context.messages` input; a pointer bullet for `think.context_length`, `think.provider.reasoning_mode`, and the top-level `tags`, `experimental`, and `mip_opt_out`, with ranges in `references/agent.md`; a one-line note that Conversational Mode (`agent.think_conversational.provider`) is invite-only Early Access answered with `UNPARSABLE_CLIENT_MESSAGE` from unenrolled projects +- Voice-agent skill: the SDK routing bullets record that `FunctionCallCancelled` and `defer_until_eot` ship in JS 5.12.0, Python 7.10.0, and Java 0.10.1 while the SDK voice-agent skills do not describe them, so message shapes come from this skill; that Go 3.8.0 and .NET 7.1.1 have no typed `FunctionCallCancelled` or `defer_until_eot`, so the raw message is parsed and the field sent as plain JSON; and the reusable-configuration REST clients (JS `client.voiceAgent`, Python `client.voice_agent`, Java `client.voiceAgent()`, .NET `AgentManage`; Go on `main` after 3.8.0 only; Rust none) - Voice-agent skill: six new sources (reusable agent configurations, speculative replies, function call cancelled, TTS controls, history, conversational mode) and two new description triggers, "reusable agent configuration" and "defer_until_eot" - Self-hosted skill: the SageMaker reference now carries the ordered instance-pool recommendation above the instance table. A single instance type has no fallback, so when the Availability Zone is short of that GPU the endpoint goes `Failed` (`Request to service failed` or `InsufficientInstanceCapacity`), which happens routinely for popular GPU types. The pool order is the listing's recommended type first, same-or-newer generations with similar per-instance capacity next (`g6`, `g6e`, `g7`), older generations last as insurance, never an unsupported type (`g4dn` for Flux STT, `g5` and `g4dn` for Flux TTS, single-GPU types for Aura-2), up to 5 types with three as the sweet spot. `VariantInstanceProvisionTimeoutInSeconds` is the per-type wait (`300` recommended, `60` to `3600` allowed), and quota does not fall back: every pooled type needs a regional quota of at least `1` or `CreateEndpoint` fails with `ResourceLimitExceeded`. Also one line that AWS field employees can reach Deepgram models through the AWS Marketplace Field Demonstration Program, for which Deepgram is an eligible provider - Self-hosted skill: the Kubernetes reference states that the `API/Engine -> License Proxy -> Billing` chain for air-gapped HA needs a chart newer than `0.46.0`. On `0.46.0` and earlier, with `billing.enabled` and `licenseProxy.enabled` both `true`, the `billing` condition takes precedence, so API and Engine connect to Billing directly and the deployed License Proxy receives no traffic; the fix is the `Unreleased` entry in `charts/deepgram-self-hosted/CHANGELOG.md` - Examples skill: the category map and description now cover the integrations the live `deepgram/examples` tree carries that the skill had omitted: Telnyx and Plivo (telephony), Webex, Jitsi, and Microsoft Teams (the "Recording platforms" row is now "Meeting platforms", split into recordings on Nova prerecorded and live meetings on Nova live, since the Jitsi bridge and the Teams bot stream), Haystack and Semantic Kernel (LLM frameworks), Gin (web frameworks), a **Broadcast** row for the OBS Studio native C captioning plugin, and an **Infrastructure** row for the Node and Python Deepgram API proxy servers, Silero VAD speech segmentation, and the multi-provider chat-completions proxy that serves as a Voice Agent `think.endpoint.url`. The two `530-*` directories are named in full because they share a number. The statement that no example covers Flux TTS stands: the Voice Agent proxy keeps TTS on `aura-2` -- `AGENTS.md`: the install section records that `npx skills add` clones with `git`, so an image without it (a bare `node:22-alpine`, for example) fails every target with `Failed to clone ...: Error: spawn git ENOENT` while exiting 0 and installing nothing. The common-failure-modes table gains the matching row with `apk add git` as the fix -- API skill: the reference generator now emits JSON-Schema bounds, which it had been discarding for every parameter. `minimum` and `maximum` are the only bounds the specs carry today (20 values in `openapi.yml`, 29 in `asyncapi.yml`); `minLength`, `maxLength`, `minItems`, `maxItems`, `exclusiveMinimum`, and `exclusiveMaximum` are handled so an upstream spec that starts using one needs no further change, and `multipleOf` is handled alongside them. `enum` is left as it was, because `formatType` already renders it as a literal union. A bound whose numbers the description states in prose is suppressed rather than repeated, so `limit`, which ends "Range [1,1000]", does not also render "range: `1` to `1000`". `ttl_seconds` on `POST /v1/auth/grant` gains `range: 1 to 3600`, the `/v1/speak` and Speak v1 WebSocket `speed` parameters gain `range: 0.7 to 1.5`, and `turn_index` on Flux STT `TurnInfo` gains `minimum: 0` +- `AGENTS.md`: the install section records that `npx skills add` clones with `git`, so an image without it (a bare `node:22-alpine`, for example) fails every target with `Failed to clone ...: Error: spawn git ENOENT` and exits 1 with nothing installed; install `git` first and check for the `SKILL.md` files after any headless install. The common-failure-modes table gains the matching row with `apk add git` as the fix +- API skill: the reference generator now emits JSON-Schema bounds, which it had been discarding for every parameter. `minimum` and `maximum` are the only bounds the specs carry (20 values in `openapi.yml`, 29 in `asyncapi.yml`); `minLength`, `maxLength`, `minItems`, `maxItems`, `exclusiveMinimum`, and `exclusiveMaximum` are handled so an upstream spec that starts using one needs no further change, and `multipleOf` is handled alongside them. `enum` is left as it was, because `formatType` already renders it as a literal union. A bound whose numbers the description states in prose is suppressed rather than repeated, so `limit`, which ends "Range [1,1000]", does not also render "range: `1` to `1000`". `ttl_seconds` on `POST /v1/auth/grant` gains `range: 1 to 3600`, the `/v1/speak` and Speak v1 WebSocket `speed` parameters gain `range: 0.7 to 1.5`, and `turn_index` on Flux STT `TurnInfo` gains `minimum: 0` +- API skill: an "Inline controls" row in the Aura vs Flux TTS table. Aura-2 pronunciation is GA for English and Spanish with a 2000-character input limit and no pause control; Flux TTS pronunciation is Early Access on both transports only with `speed` exactly `1.0` (`CONTROL_COMBINATION_INVALID` on batch, `DATA-0002` on the socket), and the `\{pause:500ms\}` marker is batch-only, 500 to 3000 ms in 100 ms steps, at most 8 per request +- API skill: the Documentation list links [API Rate Limits](https://developers.deepgram.com/reference/api-rate-limits), whose tables are per region and apply per project rather than per API key, and [Working with Concurrency Rate Limits](https://developers.deepgram.com/docs/working-with-concurrency-rate-limits); all-APIs mistake 2 points at the same page +- Speech-to-text skill: the Flux STT column of the model-family table now lists `mip_opt_out` and `tag`, which the `/v2/listen` reference accepts, and a new Flux STT bullet states that each redacted span becomes a single `*` rather than `[REDACTED]` or an entity tag, and that `profanity_filter`, accepted by the reference but marked unsupported on the comparison page, is not to be relied on until the two agree +- Speech-to-text skill: `CloseStream` on Flux STT does not finalize the active turn. It emits the remaining `Update` messages and closes without a WebSocket close status code, so no `EndOfTurn` arrives; the skill now says to send `ForceEndTurn` first for a final transcript, and that the reverse order leaves `ForceEndTurn` with no effect +- Speech-to-text skill: `Configure` limits and semantics. `keyterms` holds at most 100 plain terms, `language_hints: []` clears the hints while omitting the field or sending `null` keeps them, and every field in one message applies together or not at all +- Speech-to-text skill: `ForceEndTurn` SDK methods, JS 5.9.0 `sendForceEndTurn`, Python 7.8.0 `send_force_end_turn`, Java 0.9.0 `sendForceEndTurn`, Go 3.8.0 `ForceEndTurn()`, Rust 0.10.1 `force_end_turn()`, .NET 7.1.0 `SendForceEndTurn()` on the concrete `FluxWebSocketClient`, not the interface +- Speech-to-text skill: a common mistake for the `/v1/listen` `{"type":"Configure","features":{"numerals":true}}` form, which Flux STT rejects before closing the connection +- Speech-to-text skill: `EagerEndOfTurn` costs 50 to 70% more LLM calls for roughly 100 to 200 ms less end-to-end latency, and Flux STT has no `KeepAlive`: WebSocket pings replace it with a 60 second timeout +- Speech-to-text skill: regional hosts `wss://api.{eu,au,in}.deepgram.com` for `/v1/listen` and `/v2/listen`, and six Sources entries: feature overview, bring your own turn detection, Flux STT voice agent and eager end of turn, redaction, regional endpoints +- Text-to-speech skill: Aura-2 Spanish `speed` recommended range `0.9` to `1.5`; the five English-Spanish code-switching voices; the Twilio Flux TTS guide's `deepgram-sdk` 7.6.0 floor; Voice Agent notes that an inline pause marker on Flux TTS is silently unspoken and compressed output returns `INVALID_SETTINGS` +- Text-to-speech skill: the IPA length ratio (10x the source word, floor 15), both `BREAK_SYNTAX_INVALID` triggers, the four accepted pause forms, and the about 100 ms pause tolerance +- Text-to-speech skill: nine new sources (models-languages overview, regional endpoints, ws-close, troubleshooting, prompting, Flux TTS feature-overview, state, context, Twilio) +- Voice-agent skill: the `think` note gives the Google `think.provider.version` values (`ai-studio-v1beta`, its alias `v1beta`, and `gemini-enterprise-agent-v1`) with the default per endpoint (`gemini-enterprise-agent-v1` on `api.eu.deepgram.com`, `ai-studio-v1beta` elsewhere), the think and speak provider-array fallback chains (`THINK_REQUEST_FAILED` / `SPEAK_REQUEST_FAILED` per failed provider, `FAILED_TO_THINK` / `FAILED_TO_SPEAK` when all fail), and that a managed-LLM prompt over 25,000 characters is truncated with `PROMPT_TOO_LONG` rather than rejected +- Voice-agent skill: `agent.language` is deprecated in favor of `listen.provider.language` and `speak.provider.language`; a Deepgram speak provider given `language` is rejected with `UNPARSABLE_CLIENT_MESSAGE`; the `audio.output.encoding` enum (`mp3`, `opus`, `flac`, `aac` beside the raw encodings), `container` (`none`, `wav`, `ogg`), and `bitrate`, all Aura-only since Flux TTS answers them with `INVALID_SETTINGS` +- Voice-agent skill: `UpdateListen` clears `language_hints` when omitted; `ConversationText` carries `languages_hinted` and `languages` on `flux-general-multi`; the Error/Warning row names `MAXIMUM_SESSION_LENGTH_REACHED`, `MAXIMUM_SESSION_LENGTH_APPROACHING`, `FAILED_TO_THINK`, and `FAILED_TO_SPEAK`; an observability line (no logging API, persist every non-audio frame keyed to `request_id`) +- Voice-agent skill: six Sources (regional endpoints, media inputs and outputs, conversation context, observability, the Settings/inputs/outputs message indexes, the .NET SDK repository); Source 36 notes the conversational-mode page is absent from llms.txt +- Browser-agent skill: `AgentSession` options `keepAliveInterval` (10,000 ms), `url`, and `reconnect.baseDelay` 500 ms / `maxDelay` 30,000 ms / `jitter` plus or minus 20%; `injectAgentMessage(message, behavior?)`; the `listen-updated`, `latency-report`, and `history` events; `AgentPlayer` decodes `linear16` only; the 4-minute token cache beside the docs' "5-minute keys" figure and the 30-second grant default; `useDeepgramAgent` does not support `useAgentClientTool` +- Browser-agent skill: the two `@deepgram/ui` release notes that matter when pinning. 0.1.5 compiles the standalone `styles.css` export so bundlers receive regular CSS instead of Tailwind source directives, and 0.1.6 preserves the bundled TypeScript declarations after a `vite-plugin-dts` upgrade, so pin 0.1.6 or newer; a Source for the `deepgram/ui` releases and README +- Audio-intelligence skill: the streaming entity finalization rule. The server holds a final `Results` message until the speaker moves on to non-entity speech, 3 seconds of silence pass, or a `Finalize` arrives; `no_delay=true` forces immediate finalization and the docs state it leaves entities missed or incomplete in many cases +- Audio-intelligence skill: `summarize` needs more than 50 words of speech; a shorter input is returned as-is in `summary.short` with no tokens billed as summarization usage +- Audio-intelligence skill: `raw_value` on every entity when a formatting feature is on, per-word `sentiment` and `sentiment_score` on `words[]`, the `sentiment_score` range and 0.333 break point, the `confidence_score` range, and the 100-item cap on `custom_topic` and `custom_intent` +- Text-intelligence skill: a "Hosts and concurrency" section: `/v1/read` on `api.eu.deepgram.com`, `api.au.deepgram.com`, and `api.in.deepgram.com` with no cross-region fallback, and per-project concurrency of 10 per feature in North America, 5 for `sentiment` and `intents` on the regional hosts, Enterprise starting at 10 (20 for `summarize`) +- Text-intelligence skill: the 50-word summarization minimum and zero-token billing for shorter input, the 415 `UNSUPPORTED_MEDIA_TYPE` body, `url` sources serving `text/plain` or `application/json` with a `text` field, callback ports, `dg-token` and Basic Auth, the 10-retry 30-second schedule, tag limits (128 characters, 500 unique per day, immutable), score ranges, the 100-item `custom_topic`/`custom_intent` cap, and the `cli` skill (`dg read`) in the routing list +- CLI skill: the `dg api` output contract. stdout carries two JSON documents, the response body and then a result envelope (`status`, `message`, `method`, `url`, `status_code`, `response_body`, `elapsed_ms`); `-o json` changes nothing and `--raw` only compacts the body, so `json.load` fails with `Extra data` while `jq` filters both documents; `--jq` shells out to a local `jq` and prints a `jq Required` panel without one. Bodies are text only, never binary +- CLI skill: the four soft agent-detection signals (non-TTY stdin, non-TTY stdout, unset or `dumb` `TERM`, `NO_COLOR`), three of which switch output to JSON +- Setup-mcp skill: `/_mcp/server` answers `HEAD` with 404, a `GET` with the MCP `Accept` header with 405, and a bare `GET` with a JSON descriptor; only a `POST` `initialize` reaches the server, with a runnable curl +- Setup-mcp skill: a caveat that the agentic-tools page lists both kapa URLs with no authentication step although both answer an unauthenticated `initialize` with 401, and that its Docs MCP section does not mention `/_mcp/server` +- Self-hosted skill: the FIPS constraints from the FIPS-Compliant Deployment page, in `SKILL.md`, the Kubernetes reference, and the Docker/Podman reference. Flux STT is not supported on FIPS images and runs only on standard images; the FIPS Engine loads `.dgv2` models only, and `.dgv2` and `.dg` files are not interchangeable; the FIPS API image enforces TLS 1.3 exclusively, rejecting TLS 1.2 connections and non-FIPS cipher suites such as ChaCha20 regardless of the `[fips]` flag, with the customer supplying the full-chain PKI certificate for the API's HTTPS endpoint. The MP3 and FLAC known issue now says the request returns `HTTP 200` with an empty body, covers `/v1/speak` as well as batch `/v2/speak`, and names `linear16` and `opus` as unaffected +- Self-hosted skill: the SageMaker reference states that `InferenceAmiVersion=al2023-ami-sagemaker-inference-gpu-4-1` (NVIDIA driver 580, CUDA 13.0) is required on the production variant, because the model packages run a CUDA 13 runtime and without it SageMaker boots the instance family's default AMI and the container fails its CUDA preflight check; the SageMaker AI console cannot set the field, so the endpoint configuration is created with the AWS CLI, Boto3, or Terraform (`inference_ami_version`) +- Starters skill: a note that the `java-flux-tts` README clones with a plain `git clone` while its `.gitmodules` points both submodules at SSH URLs, so following its Maven steps leaves `frontend/` and `contracts/` empty ### Changed - API skill: `references/listen.md` and `references/speak.md` regenerated from the public specs. `listen.md` gains the Flux STT `Warning` message (`ListenV2Warning`: `type`, `request_id`, `sequence_id`, `code`, `description`, all required), `code` and `description` on `ConfigureFailure`, and `numerals` echoed on `ConfigureSuccess`. `speak.md` replaces every "inline pause and pronunciation controls are not yet applied; they are stripped" sentence with the live rules: `\{pause:500ms\}` markers of 500 to 3000 ms in 100 ms steps, at most 8 per batch request, `speed` capped at `1.15` while a pause is present, pronunciation controls that cannot be combined with a pause or with a `speed` other than `1.0`, the six `400` `err_code` values on `POST /v2/speak` (`CONTROL_COMBINATION_INVALID`, `PAUSE_SPEED_CAP_EXCEEDED`, `BREAK_OUT_OF_RANGE`, `BREAK_INCREMENT_INVALID`, `BREAKS_LIMIT_EXCEEDED`, `BREAK_SYNTAX_INVALID`), the new `CONTROL_COMBINATION_INVALID` value on `SpeakV2ConfigureFailure.code`, the `DATA-0002` meaning on `SpeakV2Error`, and the pronunciation `Warning` codes (`PRONUNCIATION_WARNINGS`, `PRONUNCIATION_TOO_LONG`, `PRONUNCIATIONS_LIMIT_EXCEEDED`) now described as emitted rather than reserved. The other six reference files are unchanged - API skill: the `/v2/listen` `Configure` scope reads "EOT thresholds, keyterms, language hints, and `numerals`" in all four places it is described (the Nova vs Flux STT comparison table, the "Pick Flux STT" list, all-APIs mistake 1, and Flux STT mistake 17). It had said "EOT thresholds and keyterms", which omitted the two fields `ListenV2Configure` also carries +- API skill, Flux STT mistake 17: `ConfigureSuccess` echoes the full active configuration, `numerals` included, not only the fields the client sent, and `ConfigureFailure` carries `code` and `description` identifying the rejected configuration - API skill, Flux STT mistake 18: the `ForceEndTurn` warning note no longer says the `Warning` message is absent from the AsyncAPI and cannot appear in `references/listen.md`. The reference shows the `ListenV2Warning` shape; its `code` is a free string, so the note now says the individual codes such as `FORCE_END_TURN_NO_ACTIVE_TURN` come from the Force End Turn docs page, and links it - API skill, mistake 20: the `summarize` sentence on `/v1/read` said the reference contradicted the `v2` value. `references/read.md` types the parameter `v2` | boolean, so the skill now says the type is right and only the description still reads boolean-only -- API skill: the host notes for Voice Agent REST are scoped to the one endpoint they were measured on. "Voice Agent's REST endpoints live on the `agent.` host", "The Agent REST endpoints move with it", and the mistake 15 heading now name `GET /v1/agent/settings/think/models`, so the ten `/v1/projects/{project_id}/agents` and `agent-variables` paths added to the domain table are not read as living on `agent.deepgram.com` -- Text-to-speech skill: the Flux TTS SDK bullets now state that every SDK ships a Flux TTS client, Go from v3.8.0 as `pkg/client/speak/v2`, where they had said every SDK except Go and told Go users to use the WebSocket directly. They also record that only the Rust and .NET SDK skills document Flux TTS and that the JS, Python, Java, and Go `text-to-speech` SDK skills cover `/v1/speak` only, so `/v2/speak` message shapes come from this skill +- API skill: the host notes for Voice Agent REST are scoped to the one endpoint they describe. "Voice Agent's REST endpoints live on the `agent.` host", "The Agent REST endpoints move with it", and the mistake 15 heading now name `GET /v1/agent/settings/think/models`, so the ten `/v1/projects/{project_id}/agents` and `agent-variables` paths added to the domain table are not read as living on `agent.deepgram.com` +- API skill: the `speed` row of the Aura vs Flux TTS table carries the `1.15` cap while a pause marker is present (`PAUSE_SPEED_CAP_EXCEEDED`) and points at the Inline controls row for the pronunciation rule +- API skill: the "Pick Aura" list no longer lists one-shot synthesis or "compressed output from a stream" (both families stream raw audio and both serve `mp3`, `opus`, `flac`, and `aac` on batch REST); it names the one case that remains, compressed output inside a Voice Agent, where Flux TTS returns `INVALID_SETTINGS`. The list now agrees with the text-to-speech skill's Flux TTS first rule, and the Models row's empty host cell reads `none` in place of a dash +- Text-to-speech skill, mistake 6, and API skill, Flux TTS mistake 12: both now say SSML is not interpreted and that the only markup Flux TTS honors is its own escaped inline controls (pronunciation on both transports, pause on batch). Mistake 6 adds that the client-messages page describes known SSML, ElevenLabs, and Cartesia tags stripped with one `INPUT_MARKUP_STRIPPED` warning per `Speak`, which `/v2/speak` does not send for `` and `` markup, so do not wait on it. The Aura REST section states that Aura-2 pronunciation control is GA on `/v1/speak` with the same syntax for English and Spanish voices, a 2000-character input limit, and no pause control, and that the same control is Early Access on Flux TTS `/v2/speak` +- Text-to-speech skill: the Flux TTS SDK bullet now states that every SDK ships a Flux TTS client, Go from v3.8.0 as `pkg/client/speak/v2`, where they had said every SDK except Go and told Go users to use the WebSocket directly. It also records that only the Rust SDK skill documents Flux TTS and that the JS, Python, Java, Go, and .NET `text-to-speech` SDK skills cover `/v1/speak` only, so `/v2/speak` message shapes come from this skill; the .NET SDK itself still ships `FluxSpeakRESTClient` and `FluxSpeakWebSocketClient` - Text-to-speech skill: the batch bullet notes that batch is the only transport that honors inline pauses -- Text-to-speech skill: Common mistake 6 now says SSML is not supported and that the only markup Flux TTS interprets is its own escaped controls (pronunciation on both transports, pause on batch), and adds that Aura-2 pronunciation control is GA on `/v1/speak` with the same syntax for English and Spanish voices, a 2000-character input limit, and no pause control. The previous sentence claimed SSML is stripped with an `INPUT_MARKUP_STRIPPED` warning; a live `/v2/speak` session fed `` and `` markup returned audio with no `Warning` of any kind, so the claim is gone -- Speech-to-text skill: the model-family table names the two `redact` values Flux STT accepts, `numbers` and `aggressive_numbers`, and states that any other value fails the WebSocket handshake with 400, where it had said only "number redaction" +- Speech-to-text skill: the model-family table names the two `redact` values Flux STT accepts, `numbers` and `aggressive_numbers`, where it had said only "number redaction", and the Flux STT section states that any other value fails the WebSocket handshake with 400 - Speech-to-text skill: the threshold-table introduction points at the `Configure` message as the mid-stream mechanism, and the `ForceEndTurn` bullet says "Flux STT" rather than bare "Flux" - Voice-agent skill: `ForceEndTurn` spells out both failure modes. With a listen provider other than Flux STT (`v2`) the server sends a `FORCE_END_TURN_UNSUPPORTED` warning and the turn does not end; with no turn in progress the message is ignored silently. `UpdateListen` is described as changing `model` and `language` mid-session on any provider, with thresholds and language hints on Flux STT and keyterm updates on Flux STT models only -- Voice-agent skill: the `think` provider note labels `groq` and `aws_bedrock` as bring-your-own, gives `aws_bedrock` its `provider.credentials` block (`type` `iam`, or `sts` plus `session_token`, with `region`, `access_key_id`, `secret_access_key`) and the `https://bedrock-runtime.{region}.amazonaws.com/` endpoint, and notes that `nvidia` is on the LLM models page but absent from the `references/agent.md` provider enum, so its model name should be checked against `/v1/agent/settings/think/models` +- Voice-agent skill: the `think` provider note labels `groq` and `aws_bedrock` as bring-your-own, gives `aws_bedrock` its `provider.credentials` block (`type` `iam`, or `sts` plus `session_token`, with `region`, `access_key_id`, `secret_access_key`) and the `https://bedrock-runtime.{region}.amazonaws.com/` endpoint, and tells the reader to confirm the `nvidia` model string against `GET /v1/agent/settings/think/models` - Voice-agent skill: the third-party TTS note states that `endpoint` takes `url` and `headers`, that `wss` URLs are accepted for Eleven Labs only, and that `aws_polly` also requires `credentials`; the Deepgram-managed Cartesia exception (no `endpoint`) stays, as the TTS models page documents it -- Voice-agent skill: the SDK routing line records that `FunctionCallCancelled` and `defer_until_eot` ship in JS 5.12.0, Python 7.10.0, and Java 0.10.1 while the SDK voice-agent skills do not describe them, so message shapes come from this skill - CLI skill: tracks `deepctl` 0.3.1, published 2026-09-29. The version-scoped statements that still hold on 0.3.1 are relabelled from 0.3.0: the config-file write by `dg update --check-only`, the 23-command surface, the four hardcoded `dg skills` downloads, the single-tool `dg mcp` proxy, the HTTP 404 on regional hosts, and the plain-transcript output of `--srt` and `--webvtt`. pip, uv, and pipx install 0.3.1; the Homebrew tap still pins `deepctl-0.2.26` - CLI skill: the global-flag rule is now exact. `-o`, `-q`, and `-v` are accepted after the subcommand on 0.3.1, so `dg listen file.wav -o json` and `dg -o json listen file.wav` are equivalent, except on `dg speak`, where `-o` after the subcommand is the output file. `--base-url`, `--api-key`, `-p`, `-c`, and `--timing` still fail with `No such option` after the subcommand, and the "global flag after the subcommand" mistake uses `--base-url` as its example because `dg listen -v file.wav` now succeeds - CLI skill: `dg models` is described by its 0.3.1 output: 553 rows, 942 with `--include-outdated`, each row carrying both a bare `name` (`asteria`) and a `canonical_name` (`aura-2-asteria-en`) plus a `deprecated` flag. It still lists no Flux STT or Flux TTS models, so the advice to take model names from the `speech-to-text` and `text-to-speech` skills stands - CLI skill: `dg whoami -o json` now has a `key_source` field that names where the key came from, for example `DEEPGRAM_API_KEY (env)`; the claim that it mislabelled an environment key as `config file` is gone - CLI skill: `--agent-friendly` is accepted by every command except the `skills`, `debug`, and `plugin` groups, which reject it with `No such option` -- Setup-mcp skill: the upgrade advice for Path A no longer tells the user to run `dg update`, which the CLI skill documents as a reporter that prints `installation_method: null` on a pip install. Both skills now say to upgrade through the installer that put `deepctl` there (`pip install -U deepctl`, `uv tool upgrade deepctl`, `pipx upgrade deepctl`, `brew upgrade deepgram`, or the install script) and to use `dg update --check-only` to see whether a newer release exists. The Path A troubleshooting fallback says the same +- Setup-mcp skill: the upgrade advice for Path A no longer tells the user to run `dg update`, which the CLI skill documents as a reporter that prints `installation_method: null` on a pip install. Both skills now say to upgrade through the installer that put `deepctl` there (pip, `uv tool upgrade deepctl`, `pipx upgrade deepctl`, `brew upgrade deepgram`, or the install script) and to use `dg update --check-only` to see whether a newer release exists. The Path A troubleshooting fallback says the same - Recipes skill: the Nova row lists all 26 `speech-to-text/v1` recipes rather than 16; `punctuate`, `multichannel`, `streaming-file`, `filler-words`, `replace`, `keyterm`, `profanity-filter`, `dictation`, `numerals`, and `measurements` were missing. The Flux STT row now describes the two recipes that exist under `speech-to-text/v2`: `streaming` (the `/v2/listen` WebSocket with `model=flux-general-en`, `encoding=linear16`, `sample_rate=16000`, printing `TurnInfo` events in place of v1 interim/final pairs) and `transcribe-url` (`model=flux-general-en` on the SDK's prerecorded call with `smart_format`), with the CLI carrying only `transcribe-url`. It no longer claims recipes for EOT, eager EOT, mid-session `Configure`, or keyterms, none of which exist - Starters skill: `sinatra-transcription`, still listed in the `dg init` gallery, is described as archived and private, so the clone returns 404 for anyone outside Deepgram, rather than only archived +- `AGENTS.md`: the layout table describes all three workflows (`context7.yml`, `spec-drift.yml`, `validate-skills.yml`) where it had named `context7.yml` only, names `validate-skills.ts` beside the two regeneration scripts, lists it as a headless check, dates the conventions section 2026-10-01, and says a new skill also needs a row in the `README.md` skills table +- Speech-to-text skill: the Nova `KeepAlive` bullet names the `NET-0001` close after 10 seconds, per the keep-alive page, and notes the comparison page says 12 seconds +- Text-to-speech skill: the family table and rule of thumb follow the models overview page: Flux TTS for all new builds, Aura-2 when the language is not covered, Aura-1 first-generation English only +- Text-to-speech skill: expressivity errors are described as an HTTP 400 at connect time, before the upgrade, with no `Connected` message +- Voice-agent skill: the LiveKit sentence follows the current guide, which starts on LiveKit Inference (`deepgram/aura-2`, voice `thalia`), moves to the Deepgram plugin on `nova-3` and `aura-2-thalia-en`, and swaps to `STTv2 flux-general-en` and `TTSv2 flux-alexis-en` +- Voice-agent skill: the `think/models` curl marks the `Authorization` header optional, since the endpoint is public; the three `references/agent.md` mentions say it is the `api` skill's file; `reasoning_mode` cites the AsyncAPI reference for its five values and notes the configure page lists `low`, `medium`, and `high` +- Voice-agent skill: the `nvidia` note states the two names side by side, `nemotron-3-nano-30B-A3B` on the LLM models page and `nvidia/nemotron-3.5-lightning-30b-a3b` from `GET /v1/agent/settings/think/models`, and that the endpoint's string is what the API accepts +- Voice-agent skill: `UpdatePrompt` keeps "adds to" and notes the conversation-context page says it replaces the system prompt, so confirm on a test session +- Browser-agent skill: mistake 8 says turn detection and barge-in are server-side for any listen model (the default is a `v1` Nova model) and Flux STT `v2` adds model-integrated end-of-turn detection; mistake 6 names the SDK key `audio.output.sampleRate` with the wire field as an aside +- Audio-intelligence skill: "over 50" entity types is now 59 detected and 56 redactable; `cardinal`, `ordinal`, and `percent` are detected but return a 400 when passed to `redact` +- Audio-intelligence skill: the `summarize` non-English mistake distinguishes an explicit `language` (hard 400) from `detect_language=true` (HTTP 200, `metadata.warnings`, and `summary.result: "failure"`) +- Text-intelligence skill: the sample request's text runs to 90 words so `results.summary.text` is a summary rather than the echoed input +- Text-intelligence skill: the `references/read.md` caveat now also covers the `summarize` description reading boolean-only while the API accepts `v2`, and the reference page's double-nested example response +- CLI skill: Homebrew install is `brew install deepgram/tap/deepgram`; Homebrew 6 loads a third-party formula only after it is trusted, and the two-step `brew tap` form fails without `brew trust`. The tap still pins `deepctl-0.2.26` +- CLI skill: 0.2.26 already routes `flux-*` models to `/v2/listen` and `/v2/speak`; what 0.3.0 changed is the `dg speak` default (`flux-alexis-en`), `--speed`, `--expressivity`, `--redact`, `--numerals`, and the exit-code contract +- CLI skill: the `dg skills` markers are `` and ``, used for Codex, Gemini CLI, and OpenCode (`~/.opencode/agents.md`); Cursor's `deepctl.mdc` is overwritten whole with no markers +- CLI skill: `dg mcp` advertises `deepgram-mcp` 0.1.10 while PyPI `deepgram-mcp` is 0.1.1, so the number is not an upgrade signal +- Setup-mcp skill: Homebrew install line and trust note as in the CLI skill +- Self-hosted skill: mistake 13 points at the `0.46.0` section of the Helm chart CHANGELOG for the `release-260915` notes (default tags moved to `release-260915`, `gpu-operator.driver.version` raised from `550.54.15` to `580.173.02`, `gpu-operator.driver.useOpenKernelModules` set to `true`), since the Deepgram changelog's newest self-hosted entry is the August 2026 release (`release-260826`) and carries none for `release-260915` +- Self-hosted skill: the `dg-sagemaker` script count reads 12 scripts plus a shared `_common.py` helper, with model-package ARN lookup and endpoint status added to the coverage list, and `python-tts/` joins the client directories listed under Validating an endpoint +- Starters skill: the submodule section no longer says both `.gitmodules` URLs are SSH in every starter. 80 of the 96 starters use SSH URLs; 16 use HTTPS (12 of the 13 `{framework}-live-transcription` starters, every one except `rust-live-transcription`, plus `csharp-voice-agent`, `django-voice-agent`, `flask-voice-agent`, and `node-voice-agent`). The `insteadOf` rewrite is kept and described as a no-op on the HTTPS starters, and the `make init` and `dg init --install` failure modes are scoped to the 80 SSH starters ### Fixed - API skill, Flux STT mistake 17: the `Configure` example sent `"eot_threshold": "0.8"` and `"eot_timeout_ms": "3000"` as JSON strings. `eot_threshold` is a number and `eot_timeout_ms` an integer, and the Flux STT Configure docs send them unquoted, so the example now reads `0.8` and `3000` - API skill: removed `GET /v1/auth/token` from the regional-endpoints table. No such path exists; the only auth path is `POST /v1/auth/grant` -- Audio-intelligence skill: streaming `entities` on `/v1/listen` `Results` are present only on messages whose `is_final` is `true`, and a final result with nothing detected carries `"entities": []`. The skill had said the array was on every `Results` message, which sent readers looking for entities on interim results that never carry the key - Voice-agent skill: three bare "Flux" mentions (the LiveKit swap sentence, the `listen` field note, and the `UpdateListen`/`ForceEndTurn` bullet) now read "Flux STT" or "Flux TTS" - CLI skill: the exit-code mistake no longer says a missing API key exits 0. On 0.3.1 both `dg listen` and `dg projects` print `Error: DEEPGRAM_API_KEY is not set …` and exit 1; the Ctrl-C hole on `dg listen --mic` and `dg mcp` is still documented - CLI skill: the `dg init` gallery caveat says `sinatra-transcription` points at an archived, private repository whose URL returns 404, not merely an archived one @@ -77,7 +129,21 @@ Catch-up with the September API, spec, SDK, CLI, and documentation changes. Flux - Setup-mcp skill: Path B now says `deepgram-mcp` is a PyPI package and that the npm package of the same name is unrelated third-party code that also asks for `DEEPGRAM_API_KEY`, so `npx deepgram-mcp` must not be used - Self-hosted skill: the SageMaker reference's Java section said Maven Central's latest Java SDK was `0.10.0`; it is `0.10.2`, which still satisfies the `0.4.0` floor the transport needs - Self-hosted skill: the Docker/Podman reference dropped its link to `/docs/deploy-deepgram-services`, which redirects to `/docs/deploy-stt-services`, a page the same list already cites -- Browser-agent skill: the `ttl` versus `ttl_seconds` mistake no longer claims that browser-agent documentation snippets still show `ttl`. Those snippets were corrected upstream. The mistake itself, the observed `expires_in` values, and the advice to read `expires_in` rather than trust the field name all stand; the reason given for that advice is now the behavior that causes it, which is that `/v1/auth/grant` ignores any field it does not recognize and still answers HTTP 200 +- Browser-agent skill: the `ttl` versus `ttl_seconds` mistake no longer claims that browser-agent documentation snippets still show `ttl`. Those snippets were corrected upstream. The mistake itself, the `expires_in` values it quotes, and the advice to read `expires_in` rather than trust the field name all stand; the reason given for that advice is now the behavior that causes it, which is that `/v1/auth/grant` ignores any field it does not recognize and still answers HTTP 200 +- CI: the `spec-drift.yml` header comment said the generator does not prune reference files it no longer emits. It has pruned them since 1.6.0, so a removed endpoint group reports as a deletion; the comment now says so +- Speech-to-text skill: the `keyterm` bullet no longer claims other models return 400 and point at `keywords`, which no page documents; it now says `keyterm` is Nova-3 and Flux STT only and other models such as Nova-2 use `keywords` +- Text-to-speech skill: mistake 2 distinguishes a 400 `No such model/language/tier combination found` (unknown name) from a 403 `INSUFFICIENT_PERMISSIONS` (no access), and no longer claims `GET /v1/models` is scoped to the key; its `tts` list is the public Aura catalog with no `flux-*` entries +- Browser-agent skill: mistake 3 quotes the real error, `useAgentContext must be used inside `, and records that the `deepgram/ui` repository README's install line `npm install @deepgram/ui @deepgram/react @deepgram/agents` produces that tree +- Audio-intelligence skill: streaming `entities` are read only from `is_final: true` messages. The docs say interim results carry no `entities` key; the skill states that `nova-3` sends the key on interim results as well, usually `[]` and sometimes populated, and that those values are not final, so the sentence that said interims never carry the key is gone +- Audio-intelligence skill: three bare "Flux" mentions now read "Flux STT" +- Text-intelligence skill: the claim that a 1 MB / 210,000-token body means there is no size ceiling is replaced by the documented 150K token cap per request, the 400 `TOKEN_LIMIT_EXCEEDED` body, and the advice to chunk a document that approaches 150K tokens rather than send it whole +- CLI skill: on 0.3.1 a URL source prints `Checking URL accessibility...` on stdout, so `dg -o json listen | jq` fails on that line while a local file parses cleanly; download the file first or drop the line with `tail -n +2`. On `main` the line is written to stderr, so the release after 0.3.1 keeps stdout to the JSON body +- CLI skill: `dg -o json init --list` prints the Rich table, a ` template(s)` line, and the JSON array on one stdout; it does not ignore `-o json` +- CLI skill: `dg init` prerequisite failure is a `Missing required tools:` checklist; `Missing tools: git, node, npm, make, curl` is the JSON message +- CLI skill: `dg listen -` on Ctrl-C exits 134 with a buffered-stdin fatal error, alongside the exit-0 holes +- CLI skill: mistake 8 said `dg login --profile` does not persist without a keyring; it writes the profile and the key to `config.yaml` in cleartext and copies the key into `default` +- Self-hosted skill: both `https://deepgram.com/contact-us/` links drop the trailing slash, which returned a 308 +- Recipes and setup-mcp skills: table labels and Sources labels use a colon where they used an em dash [1.7.0]: https://github.com/deepgram/skills/compare/deepgram-skills-v1.6.0...deepgram-skills-v1.7.0 diff --git a/skills/api/SKILL.md b/skills/api/SKILL.md index a245433..487e232 100644 --- a/skills/api/SKILL.md +++ b/skills/api/SKILL.md @@ -176,14 +176,14 @@ Both TTS families are actively maintained. `/v2/speak` is a **new endpoint, not | Batch encodings | `mp3`, `opus`, `flac`, `aac`, `linear16`, `mulaw`, `alaw` + `container` / `bit_rate` | Same — but batch-only; the socket rejects them | | Interruption | `Clear` discards the buffer, no feedback | `Interrupt` → `SpeechInterrupted` with `text_spoken` / `text_remaining` | | Mid-stream reconfig | No (fixed at connection) | Yes — `Configure` updates `speed` only | -| `speed` | `0.7`–`1.5` — Aura-2, English and Spanish only | `0.5`–`1.5` in `0.05` steps | +| `speed` | `0.7` to `1.5`, Aura-2, English and Spanish only | `0.5` to `1.5` in `0.05` steps; capped at `1.15` when the text carries a pause marker (`PAUSE_SPEED_CAP_EXCEEDED` above that); see Inline controls for the pronunciation rule | | `expressivity` | Not supported | `-2`…`2`, default `0` (beta; fixed for the connection) | +| Inline controls | Pronunciation `\{"word":"...","pronounce":""\}` (GA on Aura-2, English and Spanish, input up to 2000 characters, combinable with `speed`); no pause control | Pronunciation (Early Access, both transports) only with `speed` exactly `1.0`: `CONTROL_COMBINATION_INVALID` on batch, `DATA-0002` on the socket; pause `\{pause:500ms\}` on batch only, 500 to 3000 ms in 100 ms steps, at most 8 per request | | Voice Agent `provider.version` | `v1` (the default when a provider is specified) | `v2` (required) | **Pick Aura (`/v1/speak`) when:** - You need a language other than English, or a specific Aura voice -- You want compressed or containerized output (`mp3`, `opus`, `flac`, `aac`) from a stream -- You're doing one-shot synthesis and don't need a turn lifecycle +- You need compressed output (`mp3`, `opus`, `flac`, `aac`) inside a Voice Agent, where Flux TTS returns `INVALID_SETTINGS`; on batch REST both families serve those encodings - You're already on Aura and nothing in Flux TTS is pulling you over — v1 is unchanged **Pick Flux TTS (`/v2/speak`) when:** @@ -205,7 +205,7 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// | Speak v2 — TTS, Flux TTS (turn-based) | `POST /v2/speak` | `wss://api.deepgram.com/v2/speak` | [speak.md](references/speak.md) | | Voice Agent | `GET agent.deepgram.com/v1/agent/settings/think/models`; reusable agent configurations at `/v1/projects/{project_id}/agents` (`GET`, `POST`) and `/v1/projects/{project_id}/agents/{agent_id}` (`GET`, `PUT`, `DELETE`); agent variables at `/v1/projects/{project_id}/agent-variables` (`GET`, `POST`) and `/v1/projects/{project_id}/agent-variables/{variable_id}` (`GET`, `PATCH`, `DELETE`) | `wss://agent.deepgram.com/v1/agent/converse` | [agent.md](references/agent.md) | | Read (Intelligence) | `POST /v1/read` | — | [read.md](references/read.md) | -| Models | `GET /v1/models`, `GET /v1/models/{model_id}`, `GET /v1/projects/{project_id}/models`, `GET /v1/projects/{project_id}/models/{model_id}`; `include_outdated=true` on either list call also returns non-latest model versions | — | [models.md](references/models.md) | +| Models | `GET /v1/models`, `GET /v1/models/{model_id}`, `GET /v1/projects/{project_id}/models`, `GET /v1/projects/{project_id}/models/{model_id}`; `include_outdated=true` on either list call also returns non-latest model versions | none | [models.md](references/models.md) | | Projects | `/v1/projects/*` | — | [projects.md](references/projects.md) | | Auth | `POST /v1/auth/grant` | — | [auth.md](references/auth.md) | | Self-Hosted | `/v1/projects/*/self-hosted/*` | — | [self-hosted.md](references/self-hosted.md) | @@ -216,7 +216,7 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// 1. **Feature flags are query params, except for Voice Agent and the v2 mid-session updates.** For `/v1/listen`, `/v2/listen`, `/v1/speak`, and `/v2/speak`, initial options go on the URL. The request body carries only audio data (REST) or audio frames (WebSocket). Exceptions: `/v1/agent/converse` has no URL query params at all (all config goes in the `Settings` message); `/v2/listen` supports a `Configure` message after connection to update EOT thresholds, keyterms, language hints, and `numerals` mid-session; and `/v2/speak` supports a `Configure` message that updates `speed` only. Also note that `/v2/listen` has a much smaller param set than `/v1/listen`: flags like `smart_format`, `diarize_model`, and `punctuate` are not available. -2. **Rate limits are concurrent connections, not total requests.** A 429 means too many simultaneous open connections, not too high a request volume. Diarization and other compute-heavy features reduce your concurrency allowance further. +2. **Rate limits are concurrent connections, not total requests.** A 429 means too many simultaneous open connections, not too high a request volume. Diarization and other compute-heavy features reduce your concurrency allowance further. Limits apply per project, not per API key, and differ by region; the [API Rate Limits](https://developers.deepgram.com/reference/api-rate-limits) page carries the per-region concurrency tables. ### STT WebSocket (`/v1/listen`) @@ -242,7 +242,7 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// 11. **Streaming is raw audio only, and rejects anything it doesn't recognize.** The WebSocket emits non-containerized audio, so `encoding` is limited to `linear16` (default), `mulaw`, or `alaw`. The compressed and containerized encodings (`mp3`, `opus`, `flac`, `aac`) and the `container`, `bit_rate`, `callback`, `callback_method`, and `priority` params are **batch-only** — sending them to the socket fails the connection, as does any unknown or misspelled param. Use the batch REST transport when you need compressed output. -12. **Insert whitespace between separate generations, because the server won't.** Text normalization runs before synthesis, but successive `Speak` messages are concatenated verbatim. Sending `"Hello world."` then `"How are you?"` is processed as `"Hello world.How are you?"`, which causes sentence-boundary artifacts. Add a single space (or the right separator for non-whitespace languages) when you stitch a reply, a tool-call result, and another reply together. Send plain text: SSML is not interpreted, and the only markup Flux TTS honors is its own escaped inline controls. A pronunciation override `\{"word":"...","pronounce":""\}` is honored on both transports (Early Access) but only with `speed` 1.0, and a pause marker `\{pause:500ms\}` is batch-only. A pause marker on the socket, or a pronunciation control on a socket whose `speed` is not 1.0, fails the connection with `DATA-0002`. See [Speed, Pause, Pronunciation](https://developers.deepgram.com/docs/tts-voice-controls). +12. **Insert whitespace between separate generations, because the server won't.** Text normalization runs before synthesis, but successive `Speak` messages are concatenated verbatim. Sending `"Hello world."` then `"How are you?"` is processed as `"Hello world.How are you?"`, which causes sentence-boundary artifacts. Add a single space (or the right separator for non-whitespace languages) when you stitch a reply, a tool-call result, and another reply together. Send plain text: SSML is not interpreted, and the only markup Flux TTS honors is its own escaped inline controls. A pronunciation override `\{"word":"...","pronounce":""\}` is honored on both transports (Early Access) but only with `speed` 1.0, and a pause marker `\{pause:500ms\}` is batch-only. A pause marker on the socket, or a pronunciation control on a socket whose `speed` is not 1.0, fails the connection with `DATA-0002`. On batch `POST /v2/speak` the same violations are a 400 whose `err_code` names the rule: `CONTROL_COMBINATION_INVALID` (pronunciation with a pause, or with a `speed` other than `1.0`), `PAUSE_SPEED_CAP_EXCEEDED` (a pause marker with `speed` above `1.15`), `BREAK_OUT_OF_RANGE` (a pause outside 500 to 3000 ms), `BREAK_INCREMENT_INVALID` (a pause off the 100 ms grid), `BREAKS_LIMIT_EXCEEDED` (more than 8 pause markers, or two with no text between them), and `BREAK_SYNTAX_INVALID` (a malformed marker, such as a simple marker without backslashes or an escaped structured marker). A `speed` of exactly `1.0` never counts as a speed control, so it triggers none of these. See [Speed, Pause, Pronunciation](https://developers.deepgram.com/docs/tts-voice-controls). ### Voice Agent (`/v1/agent/converse`) @@ -263,7 +263,7 @@ Migrating from Aura? See the official [Migrating from Aura to Flux TTS](https:// ```json { "type": "Configure", "thresholds": { "eot_threshold": 0.8, "eot_timeout_ms": 3000 }, "keyterms": ["Deepgram"] } ``` - The server responds with `ConfigureSuccess` (echoing back applied values) or `ConfigureFailure`, which carries `code` and `description` identifying the rejected configuration. Omitted threshold fields keep their current values. + The server responds with `ConfigureSuccess`, which echoes the full active configuration, `numerals` included, not only the fields you sent, or `ConfigureFailure`, which carries `code` and `description` identifying the rejected configuration. Omitted threshold fields keep their current values. 18. **`ForceEndTurn` outside a turn is a `Warning`, not an error, and the socket stays open.** Sending `{"type":"ForceEndTurn"}` while no turn is in progress returns `{"type":"Warning","code":"FORCE_END_TURN_NO_ACTIVE_TURN","description":"Received ForceEndTurn while no turn was active; the request was ignored."}` and the connection continues. Do not treat it as fatal or reconnect. `references/listen.md` shows the message shape (`ListenV2Warning`: `code`, `description`, `request_id`, `sequence_id`); `code` is a free string there, so the individual codes such as `FORCE_END_TURN_NO_ACTIVE_TURN` come from the [Force End Turn](https://developers.deepgram.com/docs/flux/force-end-turn) docs. When `ForceEndTurn` *does* land mid-turn, the resulting `TurnInfo` carries `event: "EndOfTurn"` with `trigger: "manual"`. `trigger` is `model` | `manual` | `timeout`, it appears on `EndOfTurn` and nowhere else, and it is an open enum, so tolerate values you do not recognize. @@ -328,3 +328,5 @@ Swift and Kotlin SDK skills are not listed because those repositories are not pu - [Self-Hosted Deployments](https://developers.deepgram.com/docs/self-hosted-introduction) - [Regional Endpoints](https://developers.deepgram.com/reference/regional-endpoints) - [Custom Endpoints](https://developers.deepgram.com/reference/custom-endpoints) +- [API Rate Limits](https://developers.deepgram.com/reference/api-rate-limits): per-region concurrency tables for every API; limits apply per project, not per API key +- [Working with Concurrency Rate Limits](https://developers.deepgram.com/docs/working-with-concurrency-rate-limits) diff --git a/skills/audio-intelligence/SKILL.md b/skills/audio-intelligence/SKILL.md index fa4e458..ab823d5 100644 --- a/skills/audio-intelligence/SKILL.md +++ b/skills/audio-intelligence/SKILL.md @@ -25,13 +25,13 @@ from one API call. This skill gets a verified request working and states the lim - **Your input is already text** (a transcript, an email, a chat log, a support ticket): the parameters below do not apply. Use the Read API, `POST /v1/read`. Open the `text-intelligence` skill. - **You only want the transcript**: drop these parameters and open the `speech-to-text` skill. -- **You are streaming live audio**: only `detect_entities` is available. See the matrix. +- **You are streaming live audio**: see the matrix (`detect_entities` only). ## Feature matrix | Parameter | Prerecorded | Streaming (`wss`) | Language | |---|---|---|---| -| `summarize=v2` (or `summarize=true`) | yes | **no** | English only, enforced with a 400 | +| `summarize=v2` (or `summarize=true`) | yes | **no** | English only; an explicit non-English `language` is a 400 (mistake 3 covers `detect_language`) | | `sentiment=true` | yes | **no** | English only | | `topics=true` | yes | **no** | English only | | `intents=true` | yes | **no** | English only | @@ -39,7 +39,7 @@ from one API call. This skill gets a verified request working and states the lim `detect_entities` is the odd one out twice over: it is the only feature that works on the live socket, and it is the only one that does **not** exist on the Read API. Streaming entity detection -runs on Nova, Nova-2, Nova-3, and Enhanced; it is not available on Base models or on Flux. [2] +runs on Nova, Nova-2, Nova-3, and Enhanced; it is not available on Base models or on Flux STT. [2] ## Verified request @@ -57,37 +57,54 @@ Returns 200. The analysis is scattered across the response, not collected in one | Summary | `results.summary.short` (with `results.summary.result` = `"success"`) | | Sentiment per segment | `results.sentiments.segments[]` — `text`, `start_word`, `end_word`, `sentiment`, `sentiment_score` | | Sentiment overall | `results.sentiments.average` — `sentiment`, `sentiment_score` | +| Sentiment per word | `results.channels[0].alternatives[0].words[]` gains `sentiment` and `sentiment_score` | | Topics | `results.topics.segments[].topics[]` — `topic`, `confidence_score` | | Intents | `results.intents.segments[].intents[]` — `intent`, `confidence_score` | -| Entities | `results.channels[0].alternatives[0].entities[]` — `label`, `value`, `confidence`, `start_word`, `end_word` | +| Entities | `results.channels[0].alternatives[0].entities[]`: `label`, `value`, `raw_value`, `confidence`, `start_word`, `end_word` | Note the plural: the parameter is `sentiment`, the result key is `sentiments`. `metadata` gains a `summary_info`, `sentiment_info`, `topics_info`, and `intents_info` block per enabled feature, each -with `model_uuid`, `input_tokens`, and `output_tokens`. Entity labels come back uppercased -(`NAME`, `ORGANIZATION`, `LOCATION_CITY`, `MONEY`, `DATE_INTERVAL`); Deepgram documents over 50 -types. [6] +with `model_uuid`, `input_tokens`, and `output_tokens`. `sentiment_score` runs from `-1` to `1`, +and the break point between `neutral` and `positive` or `negative` is `0.333...` either side of +zero; `confidence_score` on topics and intents runs from `0` to `1`. [3] Entity labels come back +uppercased (`NAME`, `ORGANIZATION`, `LOCATION_CITY`, `MONEY`, `DATE_INTERVAL`, `CARDINAL`). +`value` is the formatted text and `raw_value` is the text as spoken; `raw_value` is present when +a formatting feature such as `smart_format` is on. [4] Deepgram documents 59 entity types, and +`redact` accepts 56 of them: `cardinal`, `ordinal`, and `percent` are detected but are not valid +`redact` values, and passing one to `redact` returns a 400. [6] + +`summarize` needs more than 50 words of speech. For a shorter input, `summary.short` is the +original input returned as-is, and no tokens in or out are billed as summarization usage, so a +`summary_info` with zero tokens on a short clip is not a failure. [3] On the live socket, `detect_entities=true` adds a **top-level** `entities` array to `Results` -messages whose `is_final` is `true`, next to `channel` and not inside `channel.alternatives[0]`. -Interim results carry no `entities` key, and a final result with nothing detected carries -`"entities": []`. Same field shape as above. +messages, next to `channel` and not inside `channel.alternatives[0]`. Read entities only from +messages whose `is_final` is `true`. The docs say interim results carry no `entities` key; `nova-3` +sends the key on interim results as well, usually `[]` and sometimes populated, and +those values are not final. A final result with nothing detected carries `"entities": []`. Same +field shape as above, `raw_value` included when formatting is on. [4] + +To return complete entities, the server holds a final result until the speaker moves on to +non-entity speech, 3 seconds of silence pass, or a `Finalize` message arrives. `no_delay=true` +forces immediate finalization without that wait, and the docs state it will leave entities missed +or incomplete in many cases. Send `no_delay=true` only when latency matters more than entity +accuracy. [4] ## Narrowing topics and intents -`custom_topic` and `custom_intent` (repeatable) add your own labels; `custom_topic_mode` and -`custom_intent_mode` take `extended` (default, your labels plus the model's) or `strict` (your -labels only). `strict` returns `"segments": []` whenever nothing matches your list, which looks -like a broken request but is not. Start with `extended`. +`custom_topic` and `custom_intent` (repeatable, up to 100 of each) add your own labels; +`custom_topic_mode` and `custom_intent_mode` take `extended` (default, your labels plus the +model's) or `strict` (your labels only). `strict` returns `"segments": []` whenever nothing +matches your list, which looks like a broken request but is not. Start with `extended`. [3] ## Common mistakes -1. **Expecting these to work on streaming.** Each failure mode is different, which makes this - confusing. `summarize` fails the WebSocket handshake with 400 `"Summarization is not available +1. **Expecting these to work on streaming.** `summarize` fails the WebSocket handshake with 400 `"Summarization is not available for streaming."`. `topics` and `intents` fail it with 403 `{"err_code":"UNAUTHORIZED_FEATURES_REQUESTED","err_msg":"Project does not have access to the - requested feature/s [\"topics\"]."}` — which reads as a permissions problem but is not one: the + requested feature/s [\"topics\"]."}`, which reads as a permissions problem but is not one: the same key's prerecorded `topics` requests succeed, and the docs matrix lists streaming as - unsupported for both. [2] Do not go asking for an entitlement. `sentiment` is worse still: the + unsupported for both. [2] `sentiment` is worse still: the handshake **succeeds** and no sentiment is ever returned. Only `detect_entities` works. 2. **Expecting a non-English request to fail loudly.** It does not. With `language=es` on real Spanish audio, `sentiment`, `topics`, `intents`, and `detect_entities` return **HTTP 200** with @@ -96,9 +113,14 @@ like a broken request but is not. Start with `extended`. for English."}]`, plus `"Topics are only supported for English."`, `"Intents are only supported for English."`, and `"Entity detection is only supported for English."`. Read `metadata.warnings` before you conclude the model found nothing. -3. **Assuming `summarize` behaves the same way.** It is the exception: non-English is a hard 400, - `{"err_code":"Bad Request","err_msg":"Summarization v2 not supported for non-English languages"}`. - `language=multi` gets the same 400. +3. **Assuming `summarize` behaves the same way.** It is the exception when `language` is explicit: + `language=es` is a hard 400, + `{"err_code":"Bad Request","err_msg":"Summarization v2 not supported for non-English languages"}`, + and `language=multi` gets the same 400. With `detect_language=true` and non-English speech it + behaves like the others: HTTP 200, a `metadata.warnings` entry with `"parameter":"summarize"` + and `"type":"unsupported_language"`, and `results.summary` present with `"result":"failure"` + and a `short` string that says the feature is English only. Check `summary.result` before you + use `summary.short`. [3] 4. **`language=multi` as a workaround.** It is not one. `multi` returns the analysis when the detected speech is English and drops it with the same `metadata.warnings` when it is not, so the same request succeeds or silently degrades depending on what the caller said. @@ -107,12 +129,13 @@ like a broken request but is not. Start with `extended`. returns the same `summary.short` shape. 6. **Looking for `results.summary.text`.** That is the Read API's shape. On `/v1/listen` the summary is at `results.summary.short`. -7. **Putting any of these on Flux.** `/v2/listen` rejects all five at the handshake with 400 +7. **Putting any of these on Flux STT.** `/v2/listen` rejects all five at the handshake with 400 `{"err_code":"INVALID_QUERY_PARAMETER","err_msg":"Unknown query parameters: detect_entities"}`, and the same message naming `summarize`, `sentiment`, `topics`, or `intents`. Transcribe with - Flux, then send the transcript to `/v1/read`. + Flux STT, then send the transcript to `/v1/read`. 8. **Reaching for these to mask PII.** Detection returns entities, it does not remove them. Use - `redact` for that, which is a speech-to-text parameter. [6] + `redact` for that, which is a speech-to-text parameter and rejects `cardinal`, `ordinal`, and + `percent` (see the entity-type counts above). [6] ## Pricing @@ -121,11 +144,12 @@ Enabling these features changes what a request costs. Rates and the billing mode ## Use a different skill when -- Your input is text rather than audio: `text-intelligence` skill (`/v1/read`). +- Your input is text rather than audio, or you need the Read API's input limits: `text-intelligence` + skill (`/v1/read`). - You want every parameter and the full response schema: `api` skill, `references/listen.md`. - You only need transcription, diarization, redaction, or captions: `speech-to-text` skill. - You want a runnable demo app: `starters` skill. Note there is no `audio-intelligence` starter; the - `text-intelligence` feature (13 languages) is the Read API app. + `text-intelligence` feature (13 frameworks) is the Read API app. - You want a snippet under 50 lines: `recipes` skill, "Audio Intelligence `v1`" — `summarize`, `sentiment`, `topics`, `intents`, `entities`, in Python, JavaScript, Go, .NET, Java, Rust, and the CLI. [7] @@ -140,7 +164,7 @@ Enabling these features changes what a request costs. Rates and the billing mode 3. https://developers.deepgram.com/docs/summarization, https://developers.deepgram.com/docs/sentiment-analysis, https://developers.deepgram.com/docs/topic-detection, https://developers.deepgram.com/docs/intent-recognition 4. https://developers.deepgram.com/docs/detect-entities 5. https://developers.deepgram.com/docs/language and https://developers.deepgram.com/docs/models-languages-overview -6. https://developers.deepgram.com/docs/supported-entity-types (over 50 types; `redact` for removal) +6. https://developers.deepgram.com/docs/supported-entity-types (59 entity types, 56 of them valid `redact` values) and https://developers.deepgram.com/docs/redaction 7. https://github.com/deepgram/recipes/blob/main/COVERAGE.md ("Audio Intelligence `v1`" section) 8. https://developers.deepgram.com/reference/speech-to-text/listen-pre-recorded and https://developers.deepgram.com/reference/speech-to-text/listen-streaming 9. https://developers.deepgram.com/docs/errors and https://deepgram.com/pricing diff --git a/skills/browser-agent/SKILL.md b/skills/browser-agent/SKILL.md index 857a893..5704d7d 100644 --- a/skills/browser-agent/SKILL.md +++ b/skills/browser-agent/SKILL.md @@ -49,7 +49,7 @@ Declared runtime dependencies, as published. `@deepgram/agents` depends on `@dee ## Browser auth: never the API key -A browser `WebSocket` cannot set request headers, so the SDK sends a short-lived bearer token as the `Sec-WebSocket-Protocol` handshake value. You supply that token through `tokenFactory`, which the SDK calls before every connect and every reconnect, so a few seconds of TTL is enough. The token only has to be valid at the handshake: once the socket is open, the token expiring does not close it, and a 30-second token is fine for an hour-long call. [1][2] +A browser `WebSocket` cannot set request headers, so the SDK sends a short-lived bearer token as the `Sec-WebSocket-Protocol` handshake value. You supply that token through `tokenFactory`, which the SDK calls before every connect and every reconnect, so a few seconds of TTL is enough. The token only has to be valid at the handshake: once the socket is open, the token expiring does not close it, and a 30-second token is fine for an hour-long call. `AgentSession` wraps `tokenFactory` in a cache that holds a token for 4 minutes by default and is invalidated before each reconnection attempt. The docs call that cache "safe for Deepgram's 5-minute short-lived keys", but `/v1/auth/grant` returns `expires_in: 30` unless you set `ttl_seconds`; with the default TTL a reused cached token is already expired, so either keep the default and rely on the pre-reconnect invalidation, or mint with `ttl_seconds` of 300 or more. [1][2][3] Mint them on your own server. `POST https://api.deepgram.com/v1/auth/grant` needs an API key with Member or higher authorization and returns `{"access_token":"...","expires_in":30}`. Its tokens work on the voice APIs but not on the Manage APIs: @@ -112,7 +112,7 @@ function Agent() { } ``` -The other hooks: `useAgentMode` (`idle`/`listening`/`thinking`/`speaking`), `useAgentMicrophone`, `useAgentPlayer`, `useAgentControls` (grouped lifecycle, messaging, runtime settings, mute), `useAgentClientTool` (register a function-call handler scoped to the component), `useAgentContext`, `useAgentSession` (the raw `AgentSession`), and `useDeepgramAgent` (no provider needed). `config`, `playerSampleRate`, and the initial `autoStart` are read once and pinned for the provider's lifetime. Change a connected agent with the runtime controls, not by mutating `config`. [4] +The other hooks: `useAgentMode` (`idle`/`listening`/`thinking`/`speaking`), `useAgentMicrophone`, `useAgentPlayer`, `useAgentControls` (grouped lifecycle, messaging, runtime settings, mute), `useAgentClientTool` (register a function-call handler scoped to the component), `useAgentContext`, `useAgentSession` (the raw `AgentSession`), and `useDeepgramAgent` (no provider needed). `config`, `playerSampleRate`, and the initial `autoStart` are read once and pinned for the provider's lifetime. Change a connected agent with the runtime controls, not by mutating `config`. `useDeepgramAgent` does not support `useAgentClientTool`; give it `onFunctionCall` in its options, or use the provider pattern for per-component tool registration. [4] ## `@deepgram/ui`: components @@ -130,7 +130,7 @@ import "@deepgram/ui/styles.css"; ``` -Components: `AgentStatus`, `AgentConversation`, `AgentMessage`, `AgentTextInput`, `AgentMicrophoneButton`, `AgentSpeakerButton`, `AgentStartButton`, plus `VoiceButton`, `Orb`, `LiveWaveform`, `BarVisualizer`, `MicSelector`, and `Response`. Styling is Tailwind v4 compiled into `@deepgram/ui/styles.css` and scoped to `[data-dg-agent]`; tokens are shadcn `--color-*` names you override on any `[data-dg-agent]` ancestor. `data-dg-scheme="dark"` on the same element forces dark; without it, components follow `prefers-color-scheme`. Install from npm rather than through a shadcn registry: `deepgram/ui` builds one, but `@deepgram/ui-registry` is a private package and `ui.deepgram.com` serves no registry JSON, so `npx shadcn add` against it returns 404. [5] +Components: `AgentStatus`, `AgentConversation`, `AgentMessage`, `AgentTextInput`, `AgentMicrophoneButton`, `AgentSpeakerButton`, `AgentStartButton`, plus `VoiceButton`, `Orb`, `LiveWaveform`, `BarVisualizer`, `MicSelector`, and `Response`. Styling is Tailwind v4 compiled into `@deepgram/ui/styles.css` and scoped to `[data-dg-agent]`; tokens are shadcn `--color-*` names you override on any `[data-dg-agent]` ancestor. `data-dg-scheme="dark"` on the same element forces dark; without it, components follow `prefers-color-scheme`. Install from npm rather than through a shadcn registry: `deepgram/ui` builds one, but `@deepgram/ui-registry` is a private package and `ui.deepgram.com` serves no registry JSON, so `npx shadcn add` against it returns 404. Two release notes matter when pinning: `ui-v0.1.5` compiles the standalone `styles.css` export so bundlers receive regular CSS instead of Tailwind source directives, and `ui-v0.1.6` preserves the bundled TypeScript declarations after a `vite-plugin-dts` upgrade, so pin 0.1.6 or newer. [11] ## `@deepgram/agents`: any framework @@ -152,7 +152,7 @@ await session.connect(); await mic.start(); ``` -`AgentSession` handles the `Welcome`/`Settings`/`SettingsApplied` handshake, buffers mic frames until `SettingsApplied`, sends `KeepAlive`, and reconnects with jittered exponential backoff (`reconnect.maxAttempts` default 8). Runtime methods mirror the protocol: `updateListen`, `updateSpeak`, `updateThink`, `updatePrompt`, `injectUserMessage`, `injectAgentMessage`, `sendFunctionCallResponse`. Events are the protocol messages in kebab-case plus `audio`, `connecting`, `connected`, `reconnecting`, `disconnected`, `sdk-error`. `AgentMicrophone` and `AgentPlayer` expose `getInputVolume` and `getOutputVolume`, plus `getInputByteFrequencyData` and `getOutputByteFrequencyData`, for visualizers. [3] +`AgentSession` handles the `Welcome`/`Settings`/`SettingsApplied` handshake, buffers mic frames until `SettingsApplied`, sends `KeepAlive` every `keepAliveInterval` ms (default 10,000), and reconnects with jittered exponential backoff: `reconnect.maxAttempts` default 8, `baseDelay` 500 ms, `maxDelay` 30,000 ms, `jitter` true for plus or minus 20%. Its other options are `audio.input` and `audio.output` (`encoding`, `sampleRate`) and `url`, which overrides the socket URL. Runtime methods mirror the protocol: `updateListen`, `updateSpeak`, `updateThink`, `updatePrompt`, `injectUserMessage`, `injectAgentMessage(message, behavior?)` with `behavior` one of `default`, `queue`, or `interrupt` (the docs page shows the one-argument form; the README and the `AgentMessageBehavior` type carry the second), and `sendFunctionCallResponse`. Events are the protocol messages in kebab-case, among them `listen-updated`, `latency-report`, and `history`, plus `audio`, `connecting`, `connected`, `reconnecting`, `disconnected`, `sdk-error`. `AgentPlayer` decodes raw PCM Int16 (`linear16`) only and does not decode compressed output, so keep `audio.output.encoding` at `linear16` or supply your own decoder. `AgentMicrophone` and `AgentPlayer` expose `getInputVolume` and `getOutputVolume`, plus `getInputByteFrequencyData` and `getOutputByteFrequencyData`, for visualizers. [3] ## Upgrade `@deepgram/react` 0.1 to 0.2 @@ -168,12 +168,12 @@ await mic.start(); 1. Shipping the API key to the browser. `{ auth: { apiKey } }` in client-side code publishes a credential anyone can bill against. Use `tokenFactory` against a route you gate. [1] 2. Setting the token lifetime with `ttl` in the `/v1/auth/grant` body. The field is `ttl_seconds`. A body of `{"ttl":300}` is accepted and ignored, and the response comes back `expires_in: 30`; `{"ttl_seconds":300}` returns `expires_in: 300`. The endpoint ignores any field it does not recognize and still answers HTTP 200, so read `expires_in` in the response rather than trusting the field name you sent. [2] -3. Installing `@deepgram/react@0.2.0` next to `@deepgram/ui@0.1.6` and importing the provider from one and the hooks from the other. `@deepgram/ui` declares `@deepgram/react ^0.1.0`, so npm nests a second copy at 0.1.0 and the two packages build separate React contexts: `AgentProvider` from `@deepgram/ui` is not the same function as `AgentProvider` from `@deepgram/react`, and a hook that reads the other context throws "used outside AgentProvider". Either import everything from `@deepgram/ui` alone, or pin one copy with `"overrides": { "@deepgram/react": "0.2.0" }` in npm, `overrides` in pnpm, or `resolutions` in yarn, which dedupes the tree and makes both imports resolve to one module. [6] +3. Installing `@deepgram/react@0.2.0` next to `@deepgram/ui@0.1.6` and importing the provider from one and the hooks from the other. `@deepgram/ui` declares `@deepgram/react ^0.1.0`, so npm nests a second copy at 0.1.0 and the two packages build separate React contexts: `AgentProvider` from `@deepgram/ui` is not the same function as `AgentProvider` from `@deepgram/react`, and a hook that reads the other context throws `useAgentContext must be used inside `. The install line in the `deepgram/ui` repository README, `npm install @deepgram/ui @deepgram/react @deepgram/agents`, produces exactly this tree (the npm README's `npm install @deepgram/ui react react-dom` does not). Either import everything from `@deepgram/ui` alone, or pin one copy with `"overrides": { "@deepgram/react": "0.2.0" }` in npm, `overrides` in pnpm, or `resolutions` in yarn, which dedupes the tree and makes both imports resolve to one module. [6] 4. Forgetting `import "@deepgram/ui/styles.css"` or the `data-dg-agent` attribute on a wrapper. Every `@deepgram/ui` token is scoped to `[data-dg-agent]`, so without it the components render unstyled. [5] 5. Never calling `player.interrupt()` on `user-started-speaking` in a raw-SDK build. Deepgram stops generating, but your queued audio keeps talking over the caller. Only the raw SDK leaves this to you: `@deepgram/react`'s provider already interrupts the player on that event, and the UI and widget layers inherit it. [3][4] -6. Mismatching sample rates. `AgentPlayer`'s `sampleRate` (default 24000) must equal `audio.output.sample_rate` in your agent settings, and `AgentMicrophone`'s (default 16000) must equal `audio.input.sample_rate`. [3] +6. Mismatching sample rates. `AgentPlayer`'s `sampleRate` (default 24000) must equal the session's `audio.output.sampleRate` (the SDK key; on the wire it is `audio.output.sample_rate`), and `AgentMicrophone`'s (default 16000) must equal `audio.input.sampleRate`. [3] 7. Leaving `latest` in a production CDN `