diff --git a/README.de.md b/README.de.md
index d9b9095..101a467 100644
--- a/README.de.md
+++ b/README.de.md
@@ -1,4 +1,6 @@
+

+
Blitztext Linux
Dein lokaler KI-Sprachassistent für KDE Plasma & Wayland
@@ -18,9 +20,11 @@
## Features
-- **NEU: Mehrsprachige Oberfläche (EN/DE):** Schalte die App-Oberfläche zwischen Deutsch und Englisch um – unter **Einstellungen → Allgemein → „Sprache"** (die Änderung greift nach einem Neustart der App).
+- **Mehrsprachige Oberfläche (EN/DE):** Schalte die App-Oberfläche zwischen Deutsch und Englisch um – unter **Einstellungen → Allgemein → „Sprache der Oberfläche"** (die Änderung greift nach einem Neustart der App).
+- **Compose-Fenster:** Text eintippen oder einfügen, einen Workflow und Schreibstil wählen und von der KI umschreiben lassen — ganz ohne Mikrofon. Mit Tonfall-Auswahl, eigenem Preset, Varianten-Verlauf und Signatur-Unterstützung.
+- **OpenRouter & eigene LLM-Endpunkte:** Nutze OpenRouter oder eine beliebige OpenAI-kompatible API als Alternative zu OpenAI für alle KI-Workflows.
+- **Audio-Export:** Speichere die Ausgabe der Vorlesefunktion direkt als Audiodatei.
- **Eigennamen / Begriffe:** Erweitere das Vokabular der KI um eigene Begriffe, Namen oder Fachwörter für perfekte Transkriptionen.
-
- **Globale Hotkeys:** Jederzeit von überall im System aufnehmen.
- **Auto-Paste:** Erkennt Sprache und fügt sie direkt dort ein, wo der Cursor ist.
- **LLM-gestützte Workflows:** Lass die KI deine Sätze professionell umformulieren, emotional filtern oder mit passenden Emojis anreichern.
@@ -144,15 +148,32 @@ Blitztext registriert globale Hotkeys via `evdev`. Mit diesen Kombinationen hast
| **Blitztext :)** |
Meta +
Shift +
E | ✅ | Ergänzt deine Nachricht passend mit Emojis. |
> [!NOTE]
-> **LLM-Workflows** (`Blitztext+`, `Blitztext $%&!`, `Blitztext :)`) setzen einen gültigen **OpenAI API-Key** voraus. Lege ihn am einfachsten in `~/.config/blitztext-linux/secrets.env` ab, indem du dort die Variable `OPENAI_API_KEY` mit deinem Key als Wert setzt (Zeilenformat `NAME=WERT`). `./run.sh` und der systemd-Service laden diese Datei automatisch. Ohne diesen Key sind diese Funktionen im Menü und über die Hotkeys deaktiviert bzw. führen zu einer Fehlermeldung.
+> **LLM-Workflows** (`Blitztext+`, `Blitztext $%&!`, `Blitztext :)`) setzen einen gültigen **API-Key** voraus. Lege ihn am einfachsten in `~/.config/blitztext-linux/secrets.env` ab, indem du dort die Variable mit deinem Key als Wert setzt (Zeilenformat `NAME=WERT`, z. B. `OPENAI_API_KEY=sk-…`). `./run.sh` und der systemd-Service laden diese Datei automatisch. Ohne diesen Key sind diese Funktionen im Menü und über die Hotkeys deaktiviert bzw. führen zu einer Fehlermeldung.
## KI-Workflows
-Die KI-Workflows helfen bei Formulierung, Ton und Emojis. Die passenden Einstellungen findest du direkt in der App.
+Die KI-Workflows helfen bei Formulierung, Ton und Emojis. Die passenden Einstellungen findest du unter **Einstellungen → KI-Workflows**:
+
+
+

+
+
+
+### LLM-Anbieter
+
+Blitztext unterstützt drei Anbieter-Modi, wählbar unter **Einstellungen → KI-Workflows → „LLM-Anbieter"**:
+
+| Anbieter | Wann verwenden |
+| :--- | :--- |
+| **OpenAI** (Standard) | Standard-OpenAI-API mit `gpt-4o-mini` oder einem anderen Modell. |
+| **OpenRouter** | Zugriff auf hunderte Modelle über einen einzigen API-Key (`OPENROUTER_API_KEY`). Base-URL: `https://openrouter.ai/api/v1`. |
+| **Eigener Endpunkt** | Jede OpenAI-kompatible API — „Base-URL" und „LLM-Modell" auf den Anbieter anpassen. |
+
+Für OpenRouter `base_url` auf `https://openrouter.ai/api/v1` setzen und Modell wählen (z. B. `openai/gpt-4o`). Der Name der API-Key-Umgebungsvariable wird unter „API-Key-Umgebung" eingestellt.
### Schreibstil-Vorlagen
-Für den Workflow **Blitztext+** (Text-Verbesserer) gibt es vorgefertigte Schreibstil-Vorlagen, die du unter **Einstellungen → KI-Workflows → „Schreibstil-Vorlage"** auswählst:
+Für den Workflow **Blitztext+** (Text-Verbesserer) gibt es vorgefertigte Schreibstil-Vorlagen, die du unter **Einstellungen → KI-Workflows → „Schreibstil-Vorlage"** oder direkt im **Compose-Fenster** auswählst:
| Vorlage | Wirkung |
| --- | --- |
@@ -164,8 +185,34 @@ Für den Workflow **Blitztext+** (Text-Verbesserer) gibt es vorgefertigte Schrei
| **Persönlich (Du-Form)** | Klarer Text in der persönlichen Du-Form. |
| **Höflich (Sie-Form)** | Klarer Text in der höflichen Sie-Form. |
| **Kurz & präzise** | Maximal knapp, ohne Füllwörter und Wiederholungen. |
+| **Eigenes Preset…** | Ein freier System-Prompt, den du selbst unter **Einstellungen → Allgemein → „Eigenes Preset (Compose)"** festlegst. |
+
+> Bei **Standard** wird zusätzlich der eingestellte **Tonfall** angewendet. Jede andere Vorlage bringt ihren eigenen Schreibstil mit und überschreibt den Tonfall. Eigennamen/Begriffe bleiben in allen Vorlagen erhalten.
+
+---
+
+## Compose-Fenster
+
+Das **Compose-Fenster** (`✍ Compose…` im Tray-Kontextmenü) ermöglicht das Umschreiben beliebiger Texte mit der KI — ganz ohne Sprachaufnahme. Es eignet sich ideal zum Überarbeiten fertiger Entwürfe, E-Mails oder Notizen.
+
+**Öffnen:** Klick auf das Tray-Icon → **✍ Compose…**
+
+**Was du im Compose-Fenster tun kannst:**
-> Bei **Standard** wird zusätzlich der eingestellte **Tonfall** angewendet. Jede andere Vorlage bringt ihren eigenen Schreibstil mit und ersetzt den Tonfall. Eigennamen/Begriffe bleiben in allen Vorlagen erhalten.
+| Element | Beschreibung |
+| :--- | :--- |
+| **Entwurf (linkes Feld)** | Text eintippen oder einfügen, der umgeschrieben werden soll. |
+| **Workflow** | Wähle zwischen Blitztext+ (Text-Verbesserer), Blitztext $%&! (Dampfablassen) oder Blitztext :) (Emojis). |
+| **Schreibstil-Vorlage** | Vorlage auswählen oder **Eigenes Preset…** für einen vollständig freien System-Prompt. |
+| **Tonfall** | Locker, neutral oder professionell. Aktiv nur bei **Standard**-Preset + **Blitztext+**; bei allen anderen Vorlagen ausgegraut (Tooltip erklärt warum). |
+| **Verbessern** | Sendet den Entwurf an die KI und zeigt das Ergebnis im rechten Feld. |
+| **Varianten-Verlauf** | Die letzten 10 generierten Ergebnisse der aktuellen Sitzung werden als scrollbare Liste gespeichert — Klick auf einen Eintrag stellt ihn wieder her. |
+| **Signatur** | Hängt deine gespeicherte Signatur an (konfiguriert unter **Einstellungen → Allgemein**). Ersetzt automatisch gängige KI-generierte Platzhalter wie `[Your Name]`, `[Ihr Name]`, `[Vorname Nachname]`, `[Signature]` u. Ä. — kein verlorener Platzhalter bleibt zurück. |
+| **Kopieren** | Kopiert das Ergebnis in die Zwischenablage. |
+| **Einfügen & Schließen** | Fügt das Ergebnis direkt in die aktive Anwendung ein und schließt das Fenster. |
+
+> [!NOTE]
+> Signatur und eigener Preset-Text werden unter **Einstellungen → Allgemein** konfiguriert. Setze dort „Signatur für das Compose-Fenster" und aktiviere „Nach jeder Generierung automatisch anhängen", wenn die Signatur bei jedem Ergebnis ergänzt werden soll.
---
@@ -200,19 +247,34 @@ Das Mikrofon im System-Tray ist dein Indikator für den aktuellen Zustand:
+Das Tray-Kontextmenü gibt dir schnellen Zugriff auf alle Workflows, das Compose-Fenster, Schreibstil-Vorlagen, Diktat-Modus, Verlauf und Einstellungen:
+
+
+
+

+
+
+
> [!NOTE]
> Steht im Desktop-Environment kein Tray-Bereich zur Verfügung, fällt das Icon auf das System-Theme `audio-input-microphone` zurück; die Farbkodierung greift dann ggf. nicht.
---
-## Hauptfenster (grafischer Fallback)
+## Hauptfenster
-Falls du keine Tastatur parat hast oder Hotkeys blockiert sind:
+Das Hauptfenster ist dein grafisches Kontrollzentrum — nützlich, wenn Hotkeys blockiert sind oder du lieber mit der Maus arbeitest:
-- **Maus-Steuerung:** Start/Stopp-Button für die Aufnahme.
-- **Workflow-Menü:** Dropdown für alle 5 Modi.
-- **Abbruch:** Verwirft eine Aufnahme sofort ohne Transkription.
-- **Schnellzugriffe:** Diktat, Verlauf, Vorlesen und Einstellungen.
+
+
+

+
+
+
+- **Workflow-Dropdown:** Alle 5 Aufnahmemodi zur Auswahl.
+- **Start/Stopp-Button:** Klick zum Starten oder Beenden einer Aufnahme.
+- **Abbruch:** Bricht die aktuelle Aufnahme ohne Transkription ab.
+- **Diktat / Verlauf:** Schnellzugriff auf den Diktat-Modus und den Transkript-Verlauf.
+- **Vorlesen / Einstellungen:** Öffnet das Vorlese-Fenster oder den Einstellungs-Dialog.
*Das Fenster öffnet sich beim Start sowie über den Tray-Eintrag **Fenster anzeigen** oder einen Klick auf das Tray-Icon. Schließen versteckt das Fenster nur — die App läuft im Tray weiter.*
@@ -222,11 +284,18 @@ Falls du keine Tastatur parat hast oder Hotkeys blockiert sind:
Zusätzlich zu den Workflows bietet das Tool drei Komfort-Funktionen:
+
+
+

+

+
+
+
| Menüpunkt | Beschreibung |
| :--- | :--- |
| **Diktat-Modus** | Umschalter. Ist er aktiv, werden alle Transkripte als Diktat-Einträge gesammelt und einzeln als Markdown-Datei gespeichert. Im Verlauf erscheint dann eine Schaltfläche **Zusammenführen**, die alle Einträge kombiniert und in die Zwischenablage kopiert. |
| **Verlauf…** | Öffnet ein Fenster mit den letzten Transkripten. Pro Eintrag: In Zwischenablage kopieren oder löschen. |
-| **Vorlesen…** | Lässt dir beliebigen Text vorlesen — lokal per **Piper TTS** (Standard) oder optional über **OpenAI Cloud-TTS** (inklusive Anbieter-, Stimmen- und Modellauswahl)! |
+| **Vorlesen…** | Lässt dir beliebigen Text vorlesen — lokal per **Piper TTS** (Standard) oder optional über **OpenAI Cloud-TTS** (inklusive Anbieter-, Stimmen- und Modellauswahl). Nutze die Schaltfläche **Exportieren**, um die Audioausgabe als Datei zu speichern. |
> [!NOTE]
> **Diktat-Notizen** werden ausschließlich in einen Ordner **innerhalb des Home-Verzeichnisses** geschrieben (Schutz gegen Pfad-Ausbruch), mit Berechtigungen `0o600`.
@@ -248,6 +317,17 @@ Zusätzlich zu den Workflows bietet das Tool drei Komfort-Funktionen:
Alles wird lokal und sicher unter `~/.config/blitztext-linux/config.json` gespeichert. Der OpenAI-Schlüssel wird nicht mehr in dieser Datei abgelegt, sondern aus einer Umgebungsvariable gelesen. Die Konfigurationsdatei lässt sich für erweiterte Prompt- und Workflow-Anpassungen direkt aus den Einstellungen öffnen: **Einstellungen → Allgemein → „Konfigurationsdatei öffnen"**.
+Der Einstellungs-Dialog hat drei Tabs:
+
+
+

+
Spracherkennung — Whisper-Modell, Backend, Sprache, Hotkey-Modus und Aufnahmetaste.
+

+
KI-Workflows — LLM-Anbieter, API-Key, Base-URL, Modell, Tonfall und Schreibstil-Vorlage.
+

+
Allgemein — Auto-Paste, Diktat-Ordner, Verlaufsgröße, Sprache der Oberfläche und Signatur.
+
+
> [!IMPORTANT]
> Die Konfigurationsdatei wird automatisch mit restriktiven Dateiberechtigungen (**`0o600` / `chmod 600`**) gespeichert. Der echte OpenAI-Key liegt stattdessen in `~/.config/blitztext-linux/secrets.env` oder wird als Umgebungsvariable bereitgestellt.
@@ -264,9 +344,16 @@ Alles wird lokal und sicher unter `~/.config/blitztext-linux/config.json` gespei
"openai_api_key_env": "OPENAI_API_KEY",
"autopaste": true,
"audio_device": "@DEFAULT_SOURCE@",
+ "llm_provider": "openai",
+ "base_url": "",
+ "llm_model": "gpt-4o-mini",
+ "compose_signature": "",
+ "compose_signature_auto_append": false,
+ "compose_custom_preset_text": "",
"workflows": {
"text_improver_tone": "neutral",
- "emoji_density": "mittel",
+ "writing_preset": "standard",
+ "emoji_density": "medium",
"dampf_system_prompt": ""
}
}
@@ -279,12 +366,17 @@ Alles wird lokal und sicher unter `~/.config/blitztext-linux/config.json` gespei
- **hotkey_mode**:
- `toggle`: Einmal drücken startet, erneutes Drücken beendet.
- `hold`: Aufnahme läuft solange der Hotkey gedrückt wird.
-- **openai_api_key_env**: Name der Umgebungsvariable für den OpenAI API-Key. Standard: `OPENAI_API_KEY`.
-- Der eigentliche Key liegt nicht in `config.json`, sondern in `~/.config/blitztext-linux/secrets.env` oder einer bereits gesetzten Umgebungsvariable.
+- **openai_api_key_env**: Name der Umgebungsvariable für den API-Key. Standard: `OPENAI_API_KEY`. Für OpenRouter: `OPENROUTER_API_KEY`.
+- **llm_provider**: `openai` (Standard), `openrouter` oder `custom`.
+- **base_url**: Eigene API-Base-URL. Leer = OpenAI-Standard. Für OpenRouter: `https://openrouter.ai/api/v1`.
+- **llm_model**: Modellname beim Anbieter, z. B. `gpt-4o-mini` (OpenAI) oder `openai/gpt-4o` (OpenRouter).
- **autopaste**: Fügt per `ydotool` ein.
- **audio_device**: Name der Audioquelle.
+- **compose_signature**: Signaturtext, der im Compose-Fenster angehängt wird.
+- **compose_signature_auto_append**: Signatur nach jeder Generierung im Compose-Fenster automatisch anhängen (`true`/`false`).
+- **compose_custom_preset_text**: Freier System-Prompt für die Option „Eigenes Preset…" im Compose-Fenster.
- **tts_provider**: TTS-Anbieter für „Vorlesen" — `piper` (lokal, Standard) oder `openai` (Cloud).
-- **tts_openai_model** / **tts_openai_voice**: Modell und Stimme für OpenAI Cloud-TTS (Standard: `gpt-4o-mini-tts`, `marin`).
+- **tts_openai_model** / **tts_openai_voice**: Modell und Stimme für OpenAI Cloud-TTS (Standard: `gpt-4o-mini-tts`, `nova`).
- **tts_openai_consent**: `true`, sobald die einmalige Datenschutz-Bestätigung für Cloud-TTS erteilt wurde. Standard: `false`.
- **workflows**: Feintuning von Tonalität (`text_improver_tone`), Schreibstil-Vorlage (`writing_preset`), Emojis (`emoji_density`) und dem Dampf-Prompt (`dampf_system_prompt`).
@@ -299,7 +391,7 @@ Wir lieben Stabilität! Führe die Tests lokal aus:
pytest
```
-Mit `WHISPER_GUI_TESTS=1 QT_QPA_PLATFORM=offscreen pytest` laufen zusätzlich die GUI-Tests des Hauptfensters.
+Mit `WHISPER_GUI_TESTS=1 QT_QPA_PLATFORM=offscreen pytest` laufen zusätzlich die GUI-Tests (Hauptfenster, Compose-Fenster).
Verzeichnisüberblick
@@ -310,15 +402,20 @@ Mit `WHISPER_GUI_TESTS=1 QT_QPA_PLATFORM=offscreen pytest` laufen zusätzlich di
│ ├── __init__.py
│ ├── audio_recorder.py # PulseAudio/PipeWire-Aufnahme via parec
│ ├── blitztext_linux.py # PyQt6-Hauptanwendung (System-Tray)
+│ ├── compose_window.py # Compose-Fenster für textbasiertes KI-Umschreiben
│ ├── config.py # Konfigurations-Manager
+│ ├── history_panel.py # Transkript-Verlauf-Panel
│ ├── hotkey_service.py # evdev-basierter Hotkey-Daemon
│ ├── i18n.py # Übersetzungen (DE/EN) für die Oberfläche
-│ ├── llm_service.py # OpenAI API Schnittstelle
+│ ├── llm_service.py # OpenAI / OpenRouter / eigene Endpunkte
+│ ├── main_window.py # Hauptanwendungsfenster
│ ├── paste_service.py # Wayland-Clipboard-Integration
│ ├── transcribe.py # Whisper-Transkription
-│ └── workflows.py # Workflows Definition
+│ ├── tts_window.py # Vorlese-Fenster mit Audio-Export
+│ ├── workflows.py # Workflow-Definitionen
+│ └── writing_presets.py # Schreibstil-Vorlagen-Definitionen
├── tests/ # Test-Suite
-└── README.md # Dieses Dokument (englische Fassung)
+└── README.md # Englische Fassung (diese Datei: README.de.md)
```
@@ -328,7 +425,7 @@ Mit `WHISPER_GUI_TESTS=1 QT_QPA_PLATFORM=offscreen pytest` laufen zusätzlich di
- **Linux Exclusive:** Nur für Linux-Systeme.
- **Wayland Fokus:** Entwickelt für Wayland (`wl-clipboard`, `ydotool`).
-- **Datenschutz:** Lokale Workflows bleiben zu 100% auf deinem Rechner. OpenAI wird nur bei Bedarf für LLM-Aufgaben kontaktiert.
+- **Datenschutz:** Lokale Workflows bleiben zu 100% auf deinem Rechner. OpenAI oder OpenRouter wird nur bei Bedarf für LLM- oder Cloud-TTS-Aufgaben kontaktiert.
- **Sicherheit (`evdev` & `input` Gruppe):** Das Tool liest Input global über `/dev/input/event*`. Auf System-Ebene bedeutet dies, dass alle Prozesse des Benutzers Eingaben mitlesen könnten (Trade-off unter Wayland ohne XDG GlobalShortcuts). Nutzen Sie Blitztext nur in Umgebungen, denen Sie vertrauen!
- **Entwickler-Hinweis:** Dieses Projekt wurde mit Unterstützung künstlicher Intelligenz (AI-assisted) entworfen. Architektur, Code und Tests wurden manuell gesichtet und auf Funktion/Sicherheit lokal verifiziert.
diff --git a/README.md b/README.md
index 35a1110..d99c032 100644
--- a/README.md
+++ b/README.md
@@ -1,4 +1,6 @@
+

+
Blitztext Linux
Your local AI voice assistant for KDE Plasma & Wayland
@@ -18,9 +20,11 @@
## Features
-- **NEW: Multilingual interface (EN/DE):** Switch the app interface between German and English under **Settings → General → "Language"** (the change takes effect after restarting the app).
+- **Multilingual interface (EN/DE):** Switch the app interface between German and English under **Settings → General → "Interface language"** (takes effect after restarting the app).
+- **Compose window:** Type or paste any text, select a workflow and writing style, and let the AI rewrite it — no microphone needed. Includes tone selector, custom preset, variant history, and signature support.
+- **OpenRouter & custom LLM endpoints:** Use OpenRouter or any OpenAI-compatible API as an alternative to OpenAI for all AI workflows.
+- **Audio export:** Save read-aloud output as an audio file directly from the Read Aloud window.
- **Custom names / terms:** Extend the AI's vocabulary with your own terms, names, or technical words for perfect transcriptions.
-
- **Global hotkeys:** Record from anywhere in the system at any time.
- **Auto-paste:** Detects speech and pastes it right where your cursor is.
- **LLM-powered workflows:** Let the AI rephrase your sentences professionally, filter them emotionally, or enrich them with fitting emojis.
@@ -144,15 +148,34 @@ Blitztext registers global hotkeys via `evdev`. With these combinations you have
| **Blitztext :)** |
Meta +
Shift +
E | ✅ | Enriches your message with fitting emojis. |
> [!NOTE]
-> **LLM workflows** (`Blitztext+`, `Blitztext $%&!`, `Blitztext :)`) require a valid **OpenAI API key**. The easiest way is to place it in `~/.config/blitztext-linux/secrets.env` by setting the variable `OPENAI_API_KEY` there with your key as the value (line format `NAME=VALUE`). `./run.sh` and the systemd service load this file automatically. Without this key, these functions are disabled in the menu and via the hotkeys, or result in an error message.
+> **LLM workflows** (`Blitztext+`, `Blitztext $%&!`, `Blitztext :)`) require a valid **API key**. The easiest way is to place it in `~/.config/blitztext-linux/secrets.env` using the format `NAME=VALUE` (e.g. `OPENAI_API_KEY=sk-…`). `./run.sh` and the systemd service load this file automatically. Without a key, these functions are disabled in the menu and via hotkeys, or result in an error message.
+
+---
## AI workflows
-The AI workflows help with phrasing, tone, and emojis. You'll find the relevant settings directly in the app.
+The AI workflows help with phrasing, tone, and emojis. You'll find the relevant settings under **Settings → AI Workflows**:
+
+
+

+
+
+
+### LLM providers
+
+Blitztext supports three provider modes, selectable under **Settings → AI Workflows → "LLM provider"**:
+
+| Provider | When to use |
+| :--- | :--- |
+| **OpenAI** (default) | Standard OpenAI API with `gpt-4o-mini` or any other model. |
+| **OpenRouter** | Access hundreds of models via a single API key (`OPENROUTER_API_KEY`). Base URL: `https://openrouter.ai/api/v1`. |
+| **Custom endpoint** | Any OpenAI-compatible API — set "Base URL" and "LLM model" to match your provider. |
+
+For OpenRouter, set `base_url` to `https://openrouter.ai/api/v1` and choose your model (e.g. `openai/gpt-4o`). The API key environment variable name is configured under "API key environment".
### Writing-style presets
-For the **Blitztext+** workflow (text improver) there are ready-made writing-style presets that you select under **Settings → AI Workflows → "Writing-style preset"**:
+For the **Blitztext+** workflow (text improver) there are ready-made writing-style presets that you select under **Settings → AI Workflows → "Writing-style preset"** or directly in the **Compose window**:
| Preset | Effect |
| --- | --- |
@@ -164,12 +187,38 @@ For the **Blitztext+** workflow (text improver) there are ready-made writing-sty
| **Personal (informal)** | Clear text in a personal, informal tone. |
| **Polite (formal)** | Clear text in a polite, formal tone. |
| **Short & precise** | As concise as possible, without filler words and repetitions. |
+| **Custom preset…** | A free-form system prompt you define yourself under **Settings → General → "Custom preset (Compose)"**. |
+
+> With **Standard**, the configured **tone** (casual / neutral / professional) is additionally applied. Every other preset brings its own writing style and overrides the tone setting. Custom names/terms are preserved in all presets.
+
+---
+
+## Compose window
+
+The **Compose window** (`✍ Compose…` in the tray menu) lets you rewrite any text using the AI — without recording your voice. It is ideal for editing existing drafts, emails, or notes.
+
+**How to open:** Click the tray icon → **✍ Compose…**
+
+**What you can do in the Compose window:**
-> With **Standard**, the configured **tone** is additionally applied. Every other preset brings its own writing style and replaces the tone. Custom names/terms are preserved in all presets.
+| Element | Description |
+| :--- | :--- |
+| **Draft (left pane)** | Type or paste the text you want to rewrite. |
+| **Workflow** | Choose between Blitztext+ (text improver), Blitztext $%&! (steam release), or Blitztext :) (emojis). |
+| **Writing-style preset** | Select a preset or **Custom preset…** for a fully custom system prompt. |
+| **Tone** | Choose casual, neutral, or professional. Active only when **Standard** preset + **Blitztext+** is selected; grayed out for all other presets (a tooltip explains why). |
+| **Improve** | Sends your draft to the AI and shows the result in the right pane. |
+| **Variant history** | The last 10 generated results within the current session are kept as a scrollable list — click any entry to restore it. |
+| **Signature** | Appends your saved signature (configured under **Settings → General**). Automatically replaces common AI-generated placeholders such as `[Your Name]`, `[Ihr Name]`, `[Vorname Nachname]`, `[Signature]`, and similar — so no stray placeholder is ever left behind. |
+| **Copy** | Copies the result to the clipboard. |
+| **Insert & Close** | Pastes the result directly into the active application and closes the window. |
+
+> [!NOTE]
+> The signature and custom preset text are configured under **Settings → General**. Set "Signature for Compose window" and toggle "Automatically append after generation" if you want the signature added to every result.
---
-## Tray icon: status colors
+## Tray icon and context menu
The microphone in the system tray is your indicator of the current state:
@@ -200,21 +249,36 @@ The microphone in the system tray is your indicator of the current state:
+The tray context menu gives you quick access to all workflows, the compose window, writing-style presets, dictation mode, history, and settings:
+
+
+
+

+
+
+
> [!NOTE]
> If no tray area is available in the desktop environment, the icon falls back to the system theme `audio-input-microphone`; the color coding may then not apply.
---
-## Main window (graphical fallback)
+## Main window
+
+The main window is your graphical control center — useful when hotkeys are blocked or you prefer mouse control:
-In case you don't have a keyboard handy or hotkeys are blocked:
+
+
+

+
+
-- **Mouse control:** Start/stop button for recording.
-- **Workflow menu:** Dropdown for all 5 modes.
-- **Cancel:** Discards a recording immediately without transcription.
-- **Quick access:** Dictation, history, read-aloud, and settings.
+- **Workflow dropdown:** Select from all 5 recording modes.
+- **Start/Stop button:** Click to begin or end a recording.
+- **Discard:** Cancels the current recording without transcription.
+- **Dictation / History:** Quick access to dictation mode and the transcript history.
+- **Read aloud / Settings:** Open the read-aloud window or the settings dialog.
-*The window opens at startup as well as via the tray entry **Show window** or a click on the tray icon. Closing only hides the window — the app keeps running in the tray.*
+*The window opens at startup and via the tray entry **Show window** or a click on the tray icon. Closing only hides the window — the app keeps running in the tray.*
---
@@ -222,11 +286,19 @@ In case you don't have a keyboard handy or hotkeys are blocked:
In addition to the workflows, the tool offers three convenience functions:
+
+
+

+

+
+
+
+
| Menu item | Description |
| :--- | :--- |
| **Dictation mode** | Toggle. When active, all transcripts are collected as dictation entries and each saved as a Markdown file. The history then shows a **Merge** button that combines all entries and copies them to the clipboard. |
| **History…** | Opens a window with the most recent transcripts. Per entry: copy to clipboard or delete. |
-| **Read aloud…** | Reads any text aloud to you — locally via **Piper TTS** (default) or optionally via **OpenAI Cloud TTS** (including provider, voice, and model selection)! |
+| **Read aloud…** | Reads any text aloud to you — locally via **Piper TTS** (default) or optionally via **OpenAI Cloud TTS** (including provider, voice, and model selection). Use the **Export** button to save the audio as a file. |
> [!NOTE]
> **Dictation notes** are written exclusively into a folder **inside the home directory** (protection against path traversal), with permissions `0o600`.
@@ -246,7 +318,19 @@ In addition to the workflows, the tool offers three convenience functions:
## Configuration
-Everything is stored locally and securely under `~/.config/blitztext-linux/config.json`. The OpenAI key is no longer stored in this file but read from an environment variable. The configuration file can be opened directly from the settings for advanced prompt and workflow adjustments: **Settings → General → "Open configuration file"**.
+Everything is stored locally and securely under `~/.config/blitztext-linux/config.json`. The OpenAI key is not stored in this file but read from an environment variable. The configuration file can be opened directly from the settings: **Settings → General → "Open configuration file"**.
+
+The settings dialog has three tabs:
+
+
+

+
Speech Recognition — Whisper model, backend, language, hotkey mode, and recording key.
+

+
AI Workflows — LLM provider, API key, base URL, model, tone, and writing-style preset.
+

+
General — Auto-Paste, dictation folder, history size, interface language, and signature.
+
+
> [!IMPORTANT]
> The configuration file is automatically saved with restrictive file permissions (**`0o600` / `chmod 600`**). The real OpenAI key instead lives in `~/.config/blitztext-linux/secrets.env` or is provided as an environment variable.
@@ -258,15 +342,22 @@ Everything is stored locally and securely under `~/.config/blitztext-linux/confi
{
"model": "base",
"language": "de",
- "ui_language": "de",
+ "ui_language": "en",
"backend": "openai-whisper",
"hotkey_mode": "toggle",
"openai_api_key_env": "OPENAI_API_KEY",
"autopaste": true,
"audio_device": "@DEFAULT_SOURCE@",
+ "llm_provider": "openai",
+ "base_url": "",
+ "llm_model": "gpt-4o-mini",
+ "compose_signature": "",
+ "compose_signature_auto_append": false,
+ "compose_custom_preset_text": "",
"workflows": {
"text_improver_tone": "neutral",
- "emoji_density": "mittel",
+ "writing_preset": "standard",
+ "emoji_density": "medium",
"dampf_system_prompt": ""
}
}
@@ -279,13 +370,18 @@ Everything is stored locally and securely under `~/.config/blitztext-linux/confi
- **hotkey_mode**:
- `toggle`: press once to start, press again to stop.
- `hold`: recording runs as long as the hotkey is held.
-- **openai_api_key_env**: Name of the environment variable for the OpenAI API key. Default: `OPENAI_API_KEY`.
-- The actual key does not live in `config.json` but in `~/.config/blitztext-linux/secrets.env` or an already-set environment variable.
+- **openai_api_key_env**: Name of the environment variable for the API key. Default: `OPENAI_API_KEY`. For OpenRouter use `OPENROUTER_API_KEY`.
+- **llm_provider**: `openai` (default), `openrouter`, or `custom`.
+- **base_url**: Custom API base URL. Empty = OpenAI default. For OpenRouter: `https://openrouter.ai/api/v1`.
+- **llm_model**: Model name at the provider, e.g. `gpt-4o-mini` (OpenAI) or `openai/gpt-4o` (OpenRouter).
- **autopaste**: Pastes via `ydotool`.
- **audio_device**: Name of the audio source.
+- **compose_signature**: Signature text appended in the Compose window.
+- **compose_signature_auto_append**: Auto-append signature after every generation in Compose (`true`/`false`).
+- **compose_custom_preset_text**: Free-form system prompt for the "Custom preset…" option in the Compose window.
- **tts_provider**: TTS provider for "Read aloud" — `piper` (local, default) or `openai` (cloud).
-- **tts_openai_model** / **tts_openai_voice**: Model and voice for OpenAI Cloud TTS (default: `gpt-4o-mini-tts`, `marin`).
-- **tts_openai_consent**: `true` once the one-time privacy confirmation for Cloud TTS has been granted. Default: `false`.
+- **tts_openai_model** / **tts_openai_voice**: Model and voice for OpenAI Cloud TTS (default: `gpt-4o-mini-tts`, `nova`).
+- **tts_openai_consent**: `true` once the one-time privacy confirmation for Cloud TTS has been granted.
- **workflows**: Fine-tuning of tonality (`text_improver_tone`), writing-style preset (`writing_preset`), emojis (`emoji_density`), and the steam-release prompt (`dampf_system_prompt`).
@@ -299,7 +395,7 @@ We love stability! Run the tests locally:
pytest
```
-With `WHISPER_GUI_TESTS=1 QT_QPA_PLATFORM=offscreen pytest`, the GUI tests of the main window run additionally.
+With `WHISPER_GUI_TESTS=1 QT_QPA_PLATFORM=offscreen pytest`, the GUI tests (main window, compose window) run additionally.
Directory overview
@@ -310,13 +406,18 @@ With `WHISPER_GUI_TESTS=1 QT_QPA_PLATFORM=offscreen pytest`, the GUI tests of th
│ ├── __init__.py
│ ├── audio_recorder.py # PulseAudio/PipeWire recording via parec
│ ├── blitztext_linux.py # PyQt6 main application (system tray)
+│ ├── compose_window.py # Compose window for text-only AI rewriting
│ ├── config.py # Configuration manager
+│ ├── history_panel.py # Transcript history panel
│ ├── hotkey_service.py # evdev-based hotkey daemon
│ ├── i18n.py # Interface translations (DE/EN)
-│ ├── llm_service.py # OpenAI API interface
+│ ├── llm_service.py # OpenAI / OpenRouter / custom endpoint interface
+│ ├── main_window.py # Main application window
│ ├── paste_service.py # Wayland clipboard integration
│ ├── transcribe.py # Whisper transcription
-│ └── workflows.py # Workflow definitions
+│ ├── tts_window.py # Read Aloud window with audio export
+│ ├── workflows.py # Workflow definitions
+│ └── writing_presets.py # Writing-style preset definitions
├── tests/ # Test suite
└── README.md # This document (German version: README.de.md)
```
@@ -328,7 +429,7 @@ With `WHISPER_GUI_TESTS=1 QT_QPA_PLATFORM=offscreen pytest`, the GUI tests of th
- **Linux exclusive:** For Linux systems only.
- **Wayland focus:** Developed for Wayland (`wl-clipboard`, `ydotool`).
-- **Privacy:** Local workflows stay 100% on your machine. OpenAI is only contacted when needed for LLM tasks.
+- **Privacy:** Local workflows stay 100% on your machine. OpenAI or OpenRouter is only contacted when needed for LLM or Cloud TTS tasks.
- **Security (`evdev` & `input` group):** The tool reads input globally via `/dev/input/event*`. At the system level, this means all of the user's processes could read along with input (a trade-off under Wayland without XDG GlobalShortcuts). Only use Blitztext in environments you trust!
- **Developer note:** This project was designed with the support of artificial intelligence (AI-assisted). Architecture, code, and tests were reviewed manually and verified locally for function/security.
diff --git a/docs/screenshots/linux/en/blitztext-ready.png b/docs/screenshots/linux/en/blitztext-ready.png
new file mode 100644
index 0000000..403ae42
Binary files /dev/null and b/docs/screenshots/linux/en/blitztext-ready.png differ
diff --git a/docs/screenshots/linux/en/read-aloud.png b/docs/screenshots/linux/en/read-aloud.png
new file mode 100644
index 0000000..dbd15fb
Binary files /dev/null and b/docs/screenshots/linux/en/read-aloud.png differ
diff --git a/docs/screenshots/linux/en/settings-ai-workflows.png b/docs/screenshots/linux/en/settings-ai-workflows.png
new file mode 100644
index 0000000..518d4f7
Binary files /dev/null and b/docs/screenshots/linux/en/settings-ai-workflows.png differ
diff --git a/docs/screenshots/linux/en/settings-general.png b/docs/screenshots/linux/en/settings-general.png
new file mode 100644
index 0000000..6d18f86
Binary files /dev/null and b/docs/screenshots/linux/en/settings-general.png differ
diff --git a/docs/screenshots/linux/en/settings-speech-recognition.png b/docs/screenshots/linux/en/settings-speech-recognition.png
new file mode 100644
index 0000000..feef9a2
Binary files /dev/null and b/docs/screenshots/linux/en/settings-speech-recognition.png differ
diff --git a/docs/screenshots/linux/en/tray-menu.png b/docs/screenshots/linux/en/tray-menu.png
new file mode 100644
index 0000000..0a26a7f
Binary files /dev/null and b/docs/screenshots/linux/en/tray-menu.png differ