Transcribe interviews, meetings and podcasts with speaker labels and timestamps, entirely on your own machine, using faster-whisper and pyannote.audio. Optionally, summarize the conversation with the LLM of your choice.
[00:00:00 --> 00:00:12] Speaker 1: Thanks for joining us. Could you introduce yourself?
[00:00:12 --> 00:00:31] Speaker 2: Sure. I've been working on sustainable building materials for about ten years, mostly with local authorities.
- Local processing: audio never leaves your computer. Uses an NVIDIA GPU when available, CPU otherwise.
- Accurate transcription with Whisper (
large-v3-turboby default) and word-level timestamps. - Speaker diarization with pyannote's
speaker-diarization-community-1pipeline. - Sentence-level speaker attribution: speaker changes inside a Whisper segment are caught, and sentences are not cut in half by imprecise turn boundaries.
- One timestamped folder per run: nothing gets overwritten, and each run is documented (options, versions, timings) in
run.json. - Several outputs: readable text, SRT subtitles and JSON.
- Optional summary with Gemini, OpenAI, Anthropic, Mistral, any OpenAI-compatible service, or a local model through Ollama. Interviews are summarized question by question; open discussions are summarized by topic.
- Reuse a transcription to re-run only the diarization, for instance with a known number of speakers.
- Hotwords to help Whisper with names, acronyms and jargon.
- Works on older GPUs: falls back to int8 when the card doesn't support float16 (e.g. Pascal generation).
- Python 3.10 or newer
- An NVIDIA GPU with CUDA 12 is strongly recommended. CPU works, but is much slower.
- A free Hugging Face account, to download the pyannote model
- For summaries only: an API key from an LLM provider (the Gemini free tier is enough), or Ollama installed locally
- No FFmpeg installation needed: audio is decoded with PyAV, which ships with faster-whisper.
git clone https://github.com/kheinzz/echoes.git
cd echoes
# with conda (or use python -m venv)
conda create -n echoes python=3.11
conda activate echoes
# 1. PyTorch: pick the command matching your system on https://pytorch.org/get-started/locally/
pip install torch --index-url https://download.pytorch.org/whl/cu126
# 2. echoes and its dependencies
pip install -e .echoes reads its secrets from a .env file, so that they never appear in your commands or in the code:
-
Copy
.env.exampleto a file named.env, in the folder from which you runechoes:cp .env.example .env # Windows: copy .env.example .env -
Fill in the values you need:
# Hugging Face read token, for the diarization model (required) HF_TOKEN=hf_xxxxxxxxxxxxxxxx # Key of the LLM provider used for summaries (optional) GEMINI_API_KEY=xxxxxxxxxxxxxxxx
- To keep the file elsewhere, pass its path with
--env-file, e.g.--env-file ~/.config/echoes.env. - Variables already set in the environment take precedence over the file, so
HF_TOKEN=... echoes ...or a system-wide variable also works. .envis ignored by git. Never commit it, and never write a key in the code. If a key leaks, revoke it and create a new one.
The pyannote models are free but gated:
- Accept the conditions on pyannote/speaker-diarization-community-1.
- Create a read token on huggingface.co/settings/tokens. For a fine-grained token, tick "Read access to contents of all public gated repos you can access".
- Put the token in your
.envfile asHF_TOKEN=.... Alternatively, runhf auth loginonce, or pass--hf-token.
# simplest form: language is detected automatically
echoes interview.mp3
# French interview, French speaker labels, names Whisper should know
echoes interview.mp3 --language fr --speaker-label Locuteur --hotwords "Dupont, CNRS, RGPD"
# the number of speakers is known
echoes interview.mp3 --num-speakers 3
# transcribe, then summarize with Gemini
echoes interview.mp3 --language fr --summary gemini
# re-run only the diarization, reusing the transcription of a previous run
echoes interview.mp3 --num-speakers 3 \
--transcription echoes_runs/interview_2026-09-17_15-07-00/transcription.json
# summarize an existing run (see "Summary" below)
echoes summarize echoes_runs/interview_2026-09-17_15-07-00 --provider geminipython -m echoes works too. Run echoes --help and echoes summarize --help for all options.
Video files (mp4, mkv, mov, webm...) are accepted everywhere an audio file is: their sound track is read directly, no conversion needed.
| Option | Default | Description |
|---|---|---|
-o, --output-dir |
echoes_runs/ next to the audio |
Where run folders are created |
-m, --model |
large-v3-turbo |
Whisper model: tiny, base, small, medium, large-v3, large-v3-turbo... |
-l, --language |
auto-detected | Language code (fr, en...) |
--hotwords |
Names and jargon to help recognition | |
-n, --num-speakers |
auto | Exact number of speakers |
--min-speakers, --max-speakers |
auto | Bounds on the number of speakers |
--speaker-label |
Speaker |
Prefix of speaker names (Speaker 1, Speaker 2...) |
--formats |
txt,srt,json |
Transcript formats to write |
--transcription |
Reuse a transcription.json and skip Whisper |
|
--batch-size |
1 |
Transcribe N chunks in parallel (faster, uses more GPU memory) |
--device |
auto |
cuda or cpu |
--compute-type |
auto |
float16 if the GPU supports it, int8 otherwise |
--no-vad |
Disable the filter that skips silences | |
--summary |
no summary | Summarize with gemini, openai, anthropic, mistral or ollama |
--summary-model |
see below | LLM model |
--summary-base-url |
provider's API | Address of an OpenAI-compatible service or of a remote Ollama server |
--env-file |
.env |
File with the keys and tokens |
from pathlib import Path
from echoes.runner import Options, run, summarize_run
# the keys are read from the environment (the .env file is only loaded by the CLI),
# or passed with hf_token=... and summary_api_key=...
run_dir = run(Options(audio=Path("interview.mp3"), language="fr", num_speakers=2, summary="gemini"))
# summarize an existing run
summary_file = summarize_run(run_dir, "ollama", model="qwen3")With --summary PROVIDER, a fourth step sends the transcript to an LLM and writes summary.md in the run folder. The model first decides what kind of conversation it is:
- interview: one section per question, with its timestamp and a summary of the answer, followed by the key points;
- open discussion: a summary by topic, followed by the decisions and next steps, if any.
The summary is written in the language of the transcript.
A free API key is enough: the Gemini free tier summarizes a long interview at no cost. Free tiers limit how many requests each model accepts per minute and per day, and a long interview may need a model with a large context window (all the default models above have one).
--summary |
Default model | Key in .env |
Get a key |
|---|---|---|---|
gemini |
gemini-flash-latest |
GEMINI_API_KEY |
Google AI Studio |
openai |
gpt-5-mini |
OPENAI_API_KEY |
OpenAI platform |
anthropic |
claude-sonnet-5 |
ANTHROPIC_API_KEY |
Anthropic console |
mistral |
mistral-medium-latest |
MISTRAL_API_KEY |
Mistral console |
ollama |
none: set --summary-model |
no key | runs locally, see below |
-
Another model:
--summary-model, e.g.--summary gemini --summary-model gemini-pro-latest. -
OpenAI-compatible services (OpenRouter, Groq, LM Studio, vLLM...): use
--summary openaiwith--summary-base-url, and put that service's key inOPENAI_API_KEY. A local server that needs no key works too.echoes interview.mp3 --summary openai --summary-base-url https://openrouter.ai/api/v1 --summary-model MODEL_NAME
-
Ollama (fully local): install Ollama, download a model (
ollama pull qwen3), then:echoes interview.mp3 --summary ollama --summary-model qwen3
echoes asks Ollama for a context window large enough for the whole transcript, since the default one is often too small. Large models need a lot of memory. Use
--summary-base-urlfor an Ollama server running on another machine.
echoes summarize RUN_DIR summarizes a run that is already transcribed, without running Whisper or pyannote again:
echoes summarize echoes_runs/interview_2026-09-17_15-07-00 --provider gemini
echoes summarize echoes_runs/interview_2026-09-17_15-07-00 --provider ollama --model qwen3- It uses
transcript.txt, including your edits. For instance, replaceSpeaker 1with the person's name before summarizing. - Existing summaries are kept: new ones are written to
summary_2.md,summary_3.md, and so on. - Each summary file starts with a small header that records the provider, the exact model version and the date.
- If the summary step of a run fails (quota, network...), the transcript is kept and echoes prints the
echoes summarizecommand to retry.
Each run creates its own folder, by default in an echoes_runs folder next to the audio file:
echoes_runs/
└── interview_2026-09-17_15-07-00/
├── transcript.txt readable transcript, one paragraph per speaker turn
├── transcript.srt subtitles, one cue per sentence
├── transcript.json speaker turns, for further processing
├── summary.md summary, with --summary only
├── transcription.json raw Whisper output with word timestamps (reusable)
├── diarization.json raw speaker turns from pyannote
└── run.json options, package versions, timings and status
If a run fails or is interrupted, run.json records the error. Files from the steps that completed are kept, so a finished transcription can be reused with --transcription.
- The audio is decoded once to 16 kHz mono.
- Transcription: Whisper produces text segments with word-level timestamps. Silences are skipped by a voice activity filter, which also limits hallucinations.
- Diarization: pyannote finds who speaks when. The exclusive variant is used, with at most one speaker at a time.
- Alignment: words are grouped into sentences, split on punctuation or on pauses longer than one second. Each sentence goes to the speaker who talks the most during its words, or to the nearest speaker when nothing overlaps. Consecutive sentences from the same speaker are then merged into turns.
- Summary (optional): the readable transcript is sent to the chosen LLM with instructions to find the questions, or the topics in an open discussion. Missing keys are detected before the transcription starts, and temporary errors are retried a few times, waiting the delay the provider asks for.
- GPU memory:
large-v3needs noticeably more memory thanlarge-v3-turbo. On a card with little memory, preferlarge-v3-turbo, and close other applications that use the GPU. - Too many speakers? A short noise or a laugh can create a spurious speaker. Re-run with
--num-speakersor--max-speakers, reusing the transcription to save time. - Wrong language? Detection only listens to the first 30 seconds. Set
--languageif the recording starts with silence or music. - Misspelled names? Add them to
--hotwords. - "HTTP 429" or "HTTP 503" from the summary provider? The model is busy, or you reached a rate limit. echoes retries a few times, but a daily quota only comes back the next day. Run
echoes summarizeon the run folder later, or choose another model with--summary-model: limits apply per model, so a lighter one is often still available. Gemini shows your usage and limits on ai.dev/rate-limit.
- Transcription and diarization happen locally. Models are downloaded once from Hugging Face and cached.
- pyannote.audio 4 sends anonymous usage metrics (such as file duration and number of speakers) by default. echoes turns them off, unless you explicitly set
PYANNOTE_METRICS_ENABLED=1. - Summaries are the exception: with any provider other than
ollama, the transcript text (not the audio) is sent to that provider and handled under its terms. Free tiers may use your data to improve their products; for instance, see the Gemini API terms about unpaid services. For confidential interviews, useollama, or a paid plan whose terms suit you. - Recordings and transcripts often contain personal data. The
.gitignoreexcludes audio files and run folders, but handle them according to your local regulations (e.g. GDPR).
pip install -e ".[dev]"
pytestThe unit tests cover the alignment, export and summary logic. They need neither a GPU, nor any model, nor network access: calls to LLM providers are simulated.
MIT. The models have their own terms: see the model cards of Whisper and pyannote speaker-diarization-community-1, and the terms of the LLM provider you use for summaries.
Built on faster-whisper, OpenAI Whisper and pyannote.audio.