feat: better configuration
This commit is contained in:
@@ -94,17 +94,21 @@ sudo usermod -aG input $USER
|
||||
After installing via AUR:
|
||||
|
||||
1. **Configure hyprvoice interactively:**
|
||||
|
||||
```bash
|
||||
hyprvoice configure
|
||||
```
|
||||
|
||||
This wizard will guide you through setting up your transcription provider, API key, audio preferences, and other settings.
|
||||
|
||||
2. **Enable and start the service:**
|
||||
|
||||
```bash
|
||||
systemctl --user enable --now hyprvoice.service
|
||||
```
|
||||
|
||||
3. **Add keybinding to your window manager:**
|
||||
|
||||
```bash
|
||||
# For Hyprland, add to ~/.config/hypr/hyprland.conf
|
||||
bind = SUPER, R, exec, hyprvoice toggle
|
||||
@@ -203,415 +207,25 @@ hyprvoice status
|
||||
|
||||
## Configuration
|
||||
|
||||
Use the interactive configuration wizard:
|
||||
The recommended way to configure hyprvoice is through the interactive wizard:
|
||||
|
||||
```bash
|
||||
hyprvoice configure
|
||||
```
|
||||
|
||||
This will guide you through setting up:
|
||||
The wizard guides you through all settings with a user-friendly interface:
|
||||
|
||||
- Provider API keys (OpenAI, Groq, Mistral, ElevenLabs)
|
||||
- Transcription provider and model
|
||||
- LLM post-processing options (enabled by default)
|
||||
- Keywords for domain-specific terms
|
||||
- Text injection method (clipboard/typing/fallback)
|
||||
- Notification settings
|
||||
- **Providers** - API keys for OpenAI, Groq, Mistral, ElevenLabs
|
||||
- **Transcription** - Speech-to-text provider and model selection
|
||||
- **LLM** - Post-processing to clean up transcriptions (enabled by default)
|
||||
- **Keywords** - Domain-specific terms for better accuracy
|
||||
- **Injection** - How text is typed (ydotool, wtype, clipboard)
|
||||
- **Notifications** - Desktop notification preferences
|
||||
- **Advanced Settings** - Recording parameters, timeouts
|
||||
|
||||
Configuration is stored in `~/.config/hyprvoice/config.toml` and can also be edited manually. Changes are applied immediately without restarting the daemon.
|
||||
Configuration is stored in `~/.config/hyprvoice/config.toml`. Changes are applied immediately without restarting the daemon.
|
||||
|
||||
### Unified Provider System
|
||||
|
||||
Hyprvoice uses a unified provider system where API keys are configured once and shared between transcription and LLM features:
|
||||
|
||||
```toml
|
||||
# Configure API keys for providers you want to use
|
||||
[providers.openai]
|
||||
api_key = "sk-..." # Or set OPENAI_API_KEY env var
|
||||
|
||||
[providers.groq]
|
||||
api_key = "gsk_..." # Or set GROQ_API_KEY env var
|
||||
|
||||
[providers.mistral]
|
||||
api_key = "..." # Or set MISTRAL_API_KEY env var
|
||||
|
||||
[providers.elevenlabs]
|
||||
api_key = "..." # Or set ELEVENLABS_API_KEY env var
|
||||
```
|
||||
|
||||
**API key resolution order:**
|
||||
1. `[providers.X]` section in config
|
||||
2. Legacy `transcription.api_key` (backward compatible)
|
||||
3. Environment variable (`OPENAI_API_KEY`, `GROQ_API_KEY`, etc.)
|
||||
|
||||
### Transcription Providers
|
||||
|
||||
Hyprvoice supports multiple transcription backends:
|
||||
|
||||
#### OpenAI Whisper API
|
||||
|
||||
Cloud-based transcription using OpenAI's Whisper API:
|
||||
|
||||
```toml
|
||||
[transcription]
|
||||
provider = "openai"
|
||||
language = "" # Empty for auto-detect, or "en", "es", "fr", etc.
|
||||
model = "whisper-1"
|
||||
```
|
||||
|
||||
**Features:**
|
||||
- High-quality transcription
|
||||
- Supports 50+ languages
|
||||
- Auto-detection or specify language for better accuracy
|
||||
|
||||
#### Groq Whisper API (Transcription)
|
||||
|
||||
Fast cloud-based transcription using Groq's Whisper API:
|
||||
|
||||
```toml
|
||||
[transcription]
|
||||
provider = "groq-transcription"
|
||||
language = "" # Empty for auto-detect, or "en", "es", "fr", etc.
|
||||
model = "whisper-large-v3" # Or "whisper-large-v3-turbo" for faster processing
|
||||
```
|
||||
|
||||
**Features:**
|
||||
- Ultra-fast transcription (significantly faster than OpenAI)
|
||||
- Same Whisper model quality
|
||||
- Supports 50+ languages
|
||||
- Free tier available with generous limits
|
||||
|
||||
#### Groq Translation API
|
||||
|
||||
Fast translation of audio to English using Groq's Whisper API:
|
||||
|
||||
```toml
|
||||
[transcription]
|
||||
provider = "groq-translation"
|
||||
language = "es" # Optional: hint source language for better accuracy
|
||||
model = "whisper-large-v3-turbo"
|
||||
```
|
||||
|
||||
**Features:**
|
||||
- Translates any language audio → English text
|
||||
- Ultra-fast processing
|
||||
- Language field hints at source language (improves accuracy)
|
||||
- Always outputs English regardless of input language
|
||||
|
||||
### LLM Post-Processing
|
||||
|
||||
LLM post-processing is **enabled by default** and significantly improves transcription quality. After transcription, the text is processed by an LLM to:
|
||||
|
||||
- Remove stutters and repeated words ("I I I want" → "I want")
|
||||
- Add proper punctuation
|
||||
- Fix grammar errors
|
||||
- Remove filler words ("um", "uh", "like", "you know", etc.)
|
||||
|
||||
#### Basic Configuration
|
||||
|
||||
```toml
|
||||
[llm]
|
||||
enabled = true # Disable with false if you want raw transcriptions
|
||||
provider = "openai" # "openai" or "groq"
|
||||
model = "gpt-4o-mini" # OpenAI: "gpt-4o-mini", Groq: "llama-3.3-70b-versatile"
|
||||
```
|
||||
|
||||
#### Post-Processing Options
|
||||
|
||||
All options are enabled by default. Disable specific ones as needed:
|
||||
|
||||
```toml
|
||||
[llm.post_processing]
|
||||
remove_stutters = true # "I I I want" → "I want"
|
||||
add_punctuation = true # Adds periods, commas, etc.
|
||||
fix_grammar = true # Fixes grammatical errors
|
||||
remove_filler_words = true # Removes "um", "uh", "like", "you know"
|
||||
```
|
||||
|
||||
#### Custom Prompts
|
||||
|
||||
Add custom instructions for specific use cases:
|
||||
|
||||
```toml
|
||||
[llm.custom_prompt]
|
||||
enabled = true
|
||||
prompt = "Format as bullet points"
|
||||
```
|
||||
|
||||
**Use cases for custom prompts:**
|
||||
- "Format as bullet points" - for note-taking
|
||||
- "Keep technical terms exactly as spoken" - for programming dictation
|
||||
- "Use formal language" - for professional documents
|
||||
- "Translate to Spanish" - for translation workflows
|
||||
|
||||
#### LLM Provider Recommendations
|
||||
|
||||
| Provider | Model | Best For |
|
||||
| -------- | ----- | -------- |
|
||||
| OpenAI | gpt-4o-mini | Best quality/cost balance (default) |
|
||||
| Groq | llama-3.3-70b-versatile | Fastest processing, free tier |
|
||||
|
||||
Both providers use the same API key as transcription if you're using OpenAI or Groq for transcription.
|
||||
|
||||
### Keywords
|
||||
|
||||
Keywords help both transcription and LLM understand domain-specific terms, names, and technical vocabulary:
|
||||
|
||||
```toml
|
||||
keywords = ["Hyprland", "Wayland", "PipeWire", "Claude", "TypeScript"]
|
||||
```
|
||||
|
||||
**How keywords work:**
|
||||
- **Transcription**: Passed as initial_prompt to Whisper, improving recognition of these terms
|
||||
- **LLM**: Included in the system prompt to ensure correct spelling
|
||||
|
||||
**When to use keywords:**
|
||||
- Names of people, companies, or products
|
||||
- Technical terminology specific to your field
|
||||
- Acronyms or abbreviations
|
||||
- Words commonly misheard by speech-to-text
|
||||
|
||||
### Example Configurations
|
||||
|
||||
#### Fast Transcription Only (No LLM)
|
||||
|
||||
```toml
|
||||
[providers.groq]
|
||||
api_key = "gsk_..."
|
||||
|
||||
[transcription]
|
||||
provider = "groq-transcription"
|
||||
model = "whisper-large-v3-turbo"
|
||||
|
||||
[llm]
|
||||
enabled = false
|
||||
```
|
||||
|
||||
#### High Quality with OpenAI (Default)
|
||||
|
||||
```toml
|
||||
[providers.openai]
|
||||
api_key = "sk-..."
|
||||
|
||||
[transcription]
|
||||
provider = "openai"
|
||||
model = "whisper-1"
|
||||
|
||||
[llm]
|
||||
enabled = true
|
||||
provider = "openai"
|
||||
model = "gpt-4o-mini"
|
||||
```
|
||||
|
||||
#### Budget-Friendly with Groq
|
||||
|
||||
```toml
|
||||
[providers.groq]
|
||||
api_key = "gsk_..."
|
||||
|
||||
[transcription]
|
||||
provider = "groq-transcription"
|
||||
model = "whisper-large-v3-turbo"
|
||||
|
||||
[llm]
|
||||
enabled = true
|
||||
provider = "groq"
|
||||
model = "llama-3.3-70b-versatile"
|
||||
```
|
||||
|
||||
#### Mixed Providers (Groq Transcription + OpenAI LLM)
|
||||
|
||||
```toml
|
||||
[providers.openai]
|
||||
api_key = "sk-..."
|
||||
|
||||
[providers.groq]
|
||||
api_key = "gsk_..."
|
||||
|
||||
[transcription]
|
||||
provider = "groq-transcription"
|
||||
model = "whisper-large-v3-turbo"
|
||||
|
||||
[llm]
|
||||
enabled = true
|
||||
provider = "openai"
|
||||
model = "gpt-4o-mini"
|
||||
```
|
||||
|
||||
### Migration from Old Config Format
|
||||
|
||||
If you're upgrading from an older version with `transcription.api_key`:
|
||||
|
||||
**Old format (still works):**
|
||||
```toml
|
||||
[transcription]
|
||||
provider = "openai"
|
||||
api_key = "sk-..." # Legacy location
|
||||
model = "whisper-1"
|
||||
```
|
||||
|
||||
**New format (recommended):**
|
||||
```toml
|
||||
[providers.openai]
|
||||
api_key = "sk-..." # Unified location
|
||||
|
||||
[transcription]
|
||||
provider = "openai"
|
||||
model = "whisper-1"
|
||||
|
||||
[llm]
|
||||
enabled = true
|
||||
provider = "openai"
|
||||
model = "gpt-4o-mini"
|
||||
```
|
||||
|
||||
Run `hyprvoice configure` to interactively update your config to the new format.
|
||||
|
||||
#### whisper.cpp Local (Planned) -> Not yet implemented
|
||||
|
||||
Private, offline transcription using local models:
|
||||
|
||||
```toml
|
||||
[transcription]
|
||||
provider = "whisper_cpp"
|
||||
model_path = "~/models/ggml-base.en.bin"
|
||||
threads = 4
|
||||
```
|
||||
|
||||
#### Recording Configuration
|
||||
|
||||
Audio capture settings:
|
||||
|
||||
```toml
|
||||
[recording]
|
||||
sample_rate = 16000 # Audio sample rate in Hz
|
||||
channels = 1 # Number of audio channels (1 for mono)
|
||||
format = "s16" # Audio format (s16 recommended)
|
||||
buffer_size = 8192 # Internal buffer size in bytes
|
||||
device = "" # PipeWire device (empty for default)
|
||||
channel_buffer_size = 30 # Audio frame buffer size
|
||||
timeout = "5m" # Maximum recording duration (prevents runaway recordings)
|
||||
```
|
||||
|
||||
**Recording Timeout:**
|
||||
|
||||
- Prevents accidental long recordings that could consume resources
|
||||
- Default: 5 minutes (`"5m"`)
|
||||
- Format: Go duration strings like `"30s"`, `"2m"`, `"10m"`
|
||||
- Recording automatically stops when timeout is reached
|
||||
|
||||
#### Text Injection
|
||||
|
||||
Configurable text injection with multiple backends:
|
||||
|
||||
```toml
|
||||
[injection]
|
||||
backends = ["ydotool", "wtype", "clipboard"] # Ordered fallback chain
|
||||
ydotool_timeout = "5s"
|
||||
wtype_timeout = "5s"
|
||||
clipboard_timeout = "3s"
|
||||
```
|
||||
|
||||
**Injection Backends:**
|
||||
|
||||
- **`ydotool`**: Uses ydotool (requires `ydotoold` daemon for ydotool v1.0.0+). Most compatible with Chromium/Electron apps.
|
||||
- **`wtype`**: Uses wtype for Wayland. May have issues with some Chromium-based apps (known upstream bug).
|
||||
- **`clipboard`**: Copies text to clipboard only. Most reliable, but requires manual paste.
|
||||
|
||||
**Fallback Chain:**
|
||||
|
||||
Backends are tried in order. The first successful one wins. Example configurations:
|
||||
|
||||
```toml
|
||||
# Clipboard only (safest, always works)
|
||||
backends = ["clipboard"]
|
||||
|
||||
# wtype with clipboard fallback
|
||||
backends = ["wtype", "clipboard"]
|
||||
|
||||
# Full fallback chain (default) - best compatibility
|
||||
backends = ["ydotool", "wtype", "clipboard"]
|
||||
|
||||
# ydotool only (if you have it set up)
|
||||
backends = ["ydotool"]
|
||||
```
|
||||
|
||||
**ydotool Setup:**
|
||||
|
||||
ydotool requires the `ydotoold` daemon running (for ydotool v1.0.0+) and access to `/dev/uinput`:
|
||||
|
||||
```bash
|
||||
# Start ydotool daemon (systemd)
|
||||
systemctl --user enable --now ydotool
|
||||
|
||||
# Or add user to input group
|
||||
sudo usermod -aG input $USER
|
||||
# Then logout/login
|
||||
|
||||
# For Hyprland, add to config to set correct keyboard layout:
|
||||
# device:ydotoold-virtual-device {
|
||||
# kb_layout = us
|
||||
# }
|
||||
```
|
||||
|
||||
**Behavior:**
|
||||
|
||||
- Backends are tried in order until one succeeds
|
||||
- Include `clipboard` in the chain if you want text copied to clipboard as fallback
|
||||
|
||||
#### Notifications
|
||||
|
||||
Desktop notification settings:
|
||||
|
||||
```toml
|
||||
[notifications]
|
||||
enabled = true # Enable/disable notifications
|
||||
type = "desktop" # "desktop", "log", or "none"
|
||||
```
|
||||
|
||||
**Notification Types:**
|
||||
|
||||
- **`desktop`**: Use notify-send for desktop notifications
|
||||
- **`log`**: Log messages to console only
|
||||
- **`none`**: Disable all notifications
|
||||
|
||||
Always keep `type = "desktop"` unless debugging.
|
||||
|
||||
##### Custom Notification Messages
|
||||
|
||||
You can customize notification text via the `[notifications.messages]` section.
|
||||
|
||||
```toml
|
||||
[notifications.messages]
|
||||
[notifications.messages.recording_started]
|
||||
title = "Hyprvoice"
|
||||
body = "Recording Started"
|
||||
[notifications.messages.transcribing]
|
||||
title = "Hyprvoice"
|
||||
body = "Recording Ended... Transcribing"
|
||||
[notifications.messages.llm_processing]
|
||||
title = "Hyprvoice"
|
||||
body = "Processing..."
|
||||
[notifications.messages.config_reloaded]
|
||||
title = "Hyprvoice"
|
||||
body = "Config Reloaded"
|
||||
[notifications.messages.operation_cancelled]
|
||||
title = "Hyprvoice"
|
||||
body = "Operation Cancelled"
|
||||
[notifications.messages.recording_aborted]
|
||||
body = "Recording Aborted"
|
||||
[notifications.messages.injection_aborted]
|
||||
body = "Injection Aborted"
|
||||
```
|
||||
|
||||
### Configuration Hot-Reloading
|
||||
|
||||
The daemon automatically watches the config file for changes and applies them immediately:
|
||||
|
||||
- **Notification settings**: Applied instantly
|
||||
- **Injection settings**: Applied to current and future operations
|
||||
- **Recording/Transcription/LLM settings**: Applied to new recording sessions
|
||||
- **Invalid configs**: Rejected with error notification, daemon continues with previous config
|
||||
For manual configuration and detailed options, see [docs/config.md](docs/config.md).
|
||||
|
||||
### Service Management
|
||||
|
||||
@@ -641,25 +255,25 @@ journalctl --user -u hyprvoice.service -f
|
||||
|
||||
## Development Status
|
||||
|
||||
| Component | Status | Notes |
|
||||
| ---------------------- | ------ | ----------------------------------------------------- |
|
||||
| Core daemon & IPC | ✅ | Unix socket control plane |
|
||||
| Recording workflow | ✅ | Toggle recording via PipeWire |
|
||||
| Audio capture | ✅ | Efficient PipeWire integration |
|
||||
| Desktop notifications | ✅ | Status feedback via notify-send |
|
||||
| OpenAI transcription | ✅ | HTTP API integration |
|
||||
| Groq transcription | ✅ | Fast Whisper API with transcription and translation |
|
||||
| Mistral transcription | ✅ | Voxtral API for European languages |
|
||||
| ElevenLabs transcription| ✅ | Scribe API with 99 language support |
|
||||
| LLM post-processing | ✅ | OpenAI/Groq text cleanup (enabled by default) |
|
||||
| Text injection | ✅ | Clipboard + wtype/ydotool with fallback |
|
||||
| Configuration system | ✅ | TOML-based user settings with hot-reload |
|
||||
| Interactive TUI setup | ✅ | `hyprvoice configure` wizard with section editing |
|
||||
| Unit test coverage | ✅ | Comprehensive test suite (100% pass) |
|
||||
| CI/CD Pipeline | ✅ | Automated builds and releases via GitHub Actions |
|
||||
| Installation (AUR etc) | ✅ | AUR package with automated dependency installation |
|
||||
| Light dictation models | ⏳ | Alternatives to whispers for light and fast dictation |
|
||||
| whisper.cpp support | ⏳ | Local model inference |
|
||||
| Component | Status | Notes |
|
||||
| ------------------------ | ------ | ----------------------------------------------------- |
|
||||
| Core daemon & IPC | ✅ | Unix socket control plane |
|
||||
| Recording workflow | ✅ | Toggle recording via PipeWire |
|
||||
| Audio capture | ✅ | Efficient PipeWire integration |
|
||||
| Desktop notifications | ✅ | Status feedback via notify-send |
|
||||
| OpenAI transcription | ✅ | HTTP API integration |
|
||||
| Groq transcription | ✅ | Fast Whisper API with transcription and translation |
|
||||
| Mistral transcription | ✅ | Voxtral API for European languages |
|
||||
| ElevenLabs transcription | ✅ | Scribe API with 99 language support |
|
||||
| LLM post-processing | ✅ | OpenAI/Groq text cleanup (enabled by default) |
|
||||
| Text injection | ✅ | Clipboard + wtype/ydotool with fallback |
|
||||
| Configuration system | ✅ | TOML-based user settings with hot-reload |
|
||||
| Interactive TUI setup | ✅ | `hyprvoice configure` wizard with section editing |
|
||||
| Unit test coverage | ✅ | Comprehensive test suite (100% pass) |
|
||||
| CI/CD Pipeline | ✅ | Automated builds and releases via GitHub Actions |
|
||||
| Installation (AUR etc) | ✅ | AUR package with automated dependency installation |
|
||||
| Light dictation models | ⏳ | Alternatives to whispers for light and fast dictation |
|
||||
| whisper.cpp support | ⏳ | Local model inference |
|
||||
|
||||
**Legend**: ✅ Complete · ⏳ Planned
|
||||
|
||||
@@ -868,6 +482,7 @@ export PATH="$HOME/.local/bin:$PATH"
|
||||
See [`packaging/RELEASE.md`](packaging/RELEASE.md) for complete release process including AUR deployment.
|
||||
|
||||
Quick start for AUR:
|
||||
|
||||
```bash
|
||||
# After creating your first GitHub release
|
||||
cd packaging/
|
||||
|
||||
Reference in New Issue
Block a user