Skip to content

Repository files navigation

AI-Voice-Assistant — Lightweight Local Voice Assistant

A 100% local voice assistant, with no LLM required, designed to run comfortably on a modest CPU.

Features

  • Local wake word detection (openWakeWord)
  • Speech-to-Text (faster-whisper) and Text-to-Speech (Piper)
  • Control of your local music library (play, pause, etc.)
  • Command understanding via rules, or an optional Zero-Shot intent engine
  • Optional AI agent (Ollama, Gemini, OpenAI, Anthropic)

Requirements

winget install VideoLAN.VLC
winget install Gyan.FFmpeg

Installation

uv sync

Usage

1. Index your music

Run this once, then again whenever you add new songs. The program scans your music folder and generates a JSON file in data/.

python main.py --parse

2. Start the assistant

python main.py

Wait until 'Listen...' is displayed, then say "Hey Jarvis" (the pre-trained wake word shipped by default), followed by your command. For example:

"Hey Jarvis, play Billie Jean by Michael Jackson"

3. Quit the assistant

Say "Hey Jarvis" followed by "au revoir", "bye" or ciao to exit the program.

"Hey Jarvis, goodbye"

The assistant replies with a farewell message and shuts down cleanly.

You can also stop it at any time from the terminal with Ctrl+C.

The list of exit phrases is handled by the exit_command/ component.

Architecture

Each folder is a component that can be replaced independently (see the docstring at the top of each file).

Folder Role
activation_sound/ Plays a melody when the wake word is triggered
ai_agent/ AI agent that answers your questions
audio/ Microphone capture + wake word detection (openWakeWord)
config/ Configuration files to customize the assistant's behavior
data/ JSON file containing all your songs
exit_command/ Command to quit the program
handle/ Executes the recognized command
log/ Log files
music/ Functions to play and control music
nlp/ Understanding: rules.py (default) or sentiment.py (optional)
parse/ Scans your music folder to generate the file stored in data/
reply/ Predefined sentences the assistant uses to reply
speech/ Speech-to-Text (faster-whisper) and Text-to-Speech (Piper)
voice_assistant Entry point that assembles the voice assistant

The shared contract is the Command class (nlp/command.py): any understanding engine (rules, a future LLM, a future trained model) must produce this same object.

Configuration

Behavior is configured in config/config.yaml.

Music

Before starting the application, make sure to update the music_dir value in the configuration file with the path to your music folder.

player:
  music_dir: "The path for your music folder"

Replace The path for your music folder with the actual path to your music directory.
For example:

player:
  music_dir: "/home/user/Music"

NLP engine

The zeroshot engine is optional and recommended for more powerful machines.

nlp:
  engine: "rules"      # "rules" (default) or "zeroshot", "laya"

The laya engine is optional and recommended for more powerful machines.

nlp:
  engine: "laya"      # "rules" (default) or "zeroshot", "laya"

AI agent

After the wake word, say "demande a Ollama" followed by your question.
For example:

"Hey Jarvis, demande a ollama qelle est la météo à Paris aujourd'hui"

Available backends: ollama, gemini, openai, anthropic.

ai_agent:
  backend: "ollama"
  model: "qwen3:1.7b"

Environment variables (.env)

API keys for cloud backends are read from a .env file at the project root. Never commit this file: it contains secrets.

  1. Copy the example file:
   copy .env.example .env
  1. Fill in only the keys for the backends you use:
   # Only needed if ai_agent.backend = "gemini"
   GEMINI_API_KEY=your_gemini_key

   # Only needed if ai_agent.backend = "openai"
   OPENAI_API_KEY=your_openai_key

   # Only needed if ai_agent.backend = "anthropic"
   ANTHROPIC_API_KEY=your_anthropic_key
Backend Variable Required
gemini GEMINI_API_KEY Yes
openai OPENAI_API_KEY Yes
anthropic ANTHROPIC_API_KEY Yes

Where to get a key: Google AI Studio · OpenAI · Anthropic Console

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages