Files
hyprvoice/docs/config.md
T

11 KiB

Configuration Reference

This document covers manual configuration of hyprvoice via the config.toml file. For most users, the interactive wizard is recommended:

hyprvoice configure

Configuration is stored in ~/.config/hyprvoice/config.toml and changes are applied immediately without restarting the daemon.

Table of Contents

Unified Provider System

Hyprvoice uses a unified provider system where API keys are configured once and shared between transcription and LLM features:

# Configure API keys for providers you want to use
[providers.openai]
  api_key = "sk-..."           # Or set OPENAI_API_KEY env var

[providers.groq]
  api_key = "gsk_..."          # Or set GROQ_API_KEY env var

[providers.mistral]
  api_key = "..."              # Or set MISTRAL_API_KEY env var

[providers.elevenlabs]
  api_key = "..."              # Or set ELEVENLABS_API_KEY env var

API key resolution order:

  1. [providers.X] section in config
  2. Environment variable (OPENAI_API_KEY, GROQ_API_KEY, etc.)

Transcription Providers

Hyprvoice supports multiple transcription backends:

OpenAI Whisper API

Cloud-based transcription using OpenAI's Whisper API:

[transcription]
provider = "openai"
language = ""                   # Empty for auto-detect, or "en", "es", "fr", etc.
model = "whisper-1"

Features:

  • High-quality transcription
  • Supports 50+ languages
  • Auto-detection or specify language for better accuracy

Groq Whisper API (Transcription)

Fast cloud-based transcription using Groq's Whisper API:

[transcription]
provider = "groq-transcription"
language = ""                   # Empty for auto-detect, or "en", "es", "fr", etc.
model = "whisper-large-v3"      # Or "whisper-large-v3-turbo" for faster processing

Features:

  • Ultra-fast transcription (significantly faster than OpenAI)
  • Same Whisper model quality
  • Supports 50+ languages
  • Free tier available with generous limits

Groq Translation API

Fast translation of audio to English using Groq's Whisper API:

[transcription]
provider = "groq-translation"
language = "es"                 # Optional: hint source language for better accuracy
model = "whisper-large-v3"

Features:

  • Translates any language audio → English text
  • Ultra-fast processing
  • Language field hints at source language (improves accuracy)
  • Always outputs English regardless of input language

Mistral Voxtral

Transcription using Mistral's Voxtral API, excellent for European languages:

[transcription]
provider = "mistral-transcription"
language = ""
model = "voxtral-mini-latest"   # Or "voxtral-mini-2507"

ElevenLabs Scribe

Transcription using ElevenLabs' Scribe API with 99 language support:

[transcription]
provider = "elevenlabs"
language = ""
model = "scribe_v1"             # Or "scribe_v2" for real-time, lower latency

LLM Post-Processing

LLM post-processing is enabled by default and significantly improves transcription quality. After transcription, the text is processed by an LLM to:

  • Remove stutters and repeated words ("I I I want" → "I want")
  • Add proper punctuation
  • Fix grammar errors
  • Remove filler words ("um", "uh", "like", "you know", etc.)

Basic Configuration

[llm]
  enabled = true               # Disable with false if you want raw transcriptions
  provider = "openai"          # "openai" or "groq"
  model = "gpt-4o-mini"        # OpenAI: "gpt-4o-mini", Groq: "llama-3.3-70b-versatile"

Post-Processing Options

All options are enabled by default. Disable specific ones as needed:

[llm.post_processing]
  remove_stutters = true       # "I I I want" → "I want"
  add_punctuation = true       # Adds periods, commas, etc.
  fix_grammar = true           # Fixes grammatical errors
  remove_filler_words = true   # Removes "um", "uh", "like", "you know"

Custom Prompts

Add custom instructions for specific use cases:

[llm.custom_prompt]
  enabled = true
  prompt = "Format as bullet points"

Use cases for custom prompts:

  • "Format as bullet points" - for note-taking
  • "Keep technical terms exactly as spoken" - for programming dictation
  • "Use formal language" - for professional documents
  • "Translate to Spanish" - for translation workflows

LLM Provider Recommendations

Provider Model Best For
OpenAI gpt-4o-mini Best quality/cost balance (default)
Groq llama-3.3-70b-versatile Fastest processing, free tier

Keywords

Keywords help both transcription and LLM understand domain-specific terms, names, and technical vocabulary:

keywords = ["Hyprland", "Wayland", "PipeWire", "Claude", "TypeScript"]

How keywords work:

  • Transcription: Passed as initial_prompt to Whisper, improving recognition of these terms
  • LLM: Included in the system prompt to ensure correct spelling

When to use keywords:

  • Names of people, companies, or products
  • Technical terminology specific to your field
  • Acronyms or abbreviations
  • Words commonly misheard by speech-to-text

Recording Configuration

Audio capture settings:

[recording]
sample_rate = 16000        # Audio sample rate in Hz (16000 recommended for speech)
channels = 1               # Number of audio channels (1 = mono, 2 = stereo)
format = "s16"             # Audio format (s16 = 16-bit signed integers)
buffer_size = 8192         # Internal buffer size in bytes (larger = less CPU, more latency)
device = ""                # PipeWire device name (empty = default microphone)
channel_buffer_size = 30   # Audio frame buffer size (frames to buffer)
timeout = "5m"             # Maximum recording duration (e.g., "30s", "2m", "5m")

Recording Timeout

  • Prevents accidental long recordings that could consume resources
  • Default: 5 minutes ("5m")
  • Format: Go duration strings like "30s", "2m", "10m"
  • Recording automatically stops when timeout is reached

Text Injection

Configurable text injection with multiple backends:

[injection]
backends = ["ydotool", "wtype", "clipboard"]  # Ordered fallback chain
ydotool_timeout = "5s"
wtype_timeout = "5s"
clipboard_timeout = "3s"

Injection Backends

  • ydotool: Uses ydotool (requires ydotoold daemon for ydotool v1.0.0+). Most compatible with Chromium/Electron apps.
  • wtype: Uses wtype for Wayland. May have issues with some Chromium-based apps (known upstream bug).
  • clipboard: Copies text to clipboard only. Most reliable, but requires manual paste.

Fallback Chain

Backends are tried in order. The first successful one wins. Example configurations:

# Clipboard only (safest, always works)
backends = ["clipboard"]

# wtype with clipboard fallback
backends = ["wtype", "clipboard"]

# Full fallback chain (default) - best compatibility
backends = ["ydotool", "wtype", "clipboard"]

# ydotool only (if you have it set up)
backends = ["ydotool"]

ydotool Setup

ydotool requires the ydotoold daemon running (for ydotool v1.0.0+) and access to /dev/uinput:

# Start ydotool daemon (systemd)
systemctl --user enable --now ydotool

# Or add user to input group
sudo usermod -aG input $USER
# Then logout/login

# For Hyprland, add to config to set correct keyboard layout:
# device:ydotoold-virtual-device {
#     kb_layout = us
# }

Notifications

Desktop notification settings:

[notifications]
enabled = true             # Enable/disable notifications
type = "desktop"           # "desktop", "log", or "none"

Notification Types

  • desktop: Use notify-send for desktop notifications
  • log: Log messages to console only
  • none: Disable all notifications

Custom Notification Messages

You can customize notification text via the [notifications.messages] section:

[notifications.messages]
  [notifications.messages.recording_started]
    title = "Hyprvoice"
    body = "Recording Started"
  [notifications.messages.transcribing]
    title = "Hyprvoice"
    body = "Recording Ended... Transcribing"
  [notifications.messages.llm_processing]
    title = "Hyprvoice"
    body = "Processing..."
  [notifications.messages.config_reloaded]
    title = "Hyprvoice"
    body = "Config Reloaded"
  [notifications.messages.operation_cancelled]
    title = "Hyprvoice"
    body = "Operation Cancelled"
  [notifications.messages.recording_aborted]
    body = "Recording Aborted"
  [notifications.messages.injection_aborted]
    body = "Injection Aborted"

Emoji-only example (for minimal pill-style notifications):

[notifications.messages.recording_started]
  title = ""
  body = "🎙️"

Example Configurations

Fast Transcription Only (No LLM)

[providers.groq]
  api_key = "gsk_..."

[transcription]
  provider = "groq-transcription"
  model = "whisper-large-v3-turbo"

[llm]
  enabled = false

High Quality with OpenAI (Default)

[providers.openai]
  api_key = "sk-..."

[transcription]
  provider = "openai"
  model = "whisper-1"

[llm]
  enabled = true
  provider = "openai"
  model = "gpt-4o-mini"

Budget-Friendly with Groq

[providers.groq]
  api_key = "gsk_..."

[transcription]
  provider = "groq-transcription"
  model = "whisper-large-v3-turbo"

[llm]
  enabled = true
  provider = "groq"
  model = "llama-3.3-70b-versatile"

Mixed Providers (Groq Transcription + OpenAI LLM)

[providers.openai]
  api_key = "sk-..."

[providers.groq]
  api_key = "gsk_..."

[transcription]
  provider = "groq-transcription"
  model = "whisper-large-v3-turbo"

[llm]
  enabled = true
  provider = "openai"
  model = "gpt-4o-mini"

Migration from Old Config Format

If you're upgrading from an older version with transcription.api_key:

Old format (still works):

[transcription]
  provider = "openai"
  api_key = "sk-..."  # Legacy location
  model = "whisper-1"

New format (recommended):

[providers.openai]
  api_key = "sk-..."  # Unified location

[transcription]
  provider = "openai"
  model = "whisper-1"

[llm]
  enabled = true
  provider = "openai"
  model = "gpt-4o-mini"

Run hyprvoice configure to interactively update your config to the new format.

Configuration Hot-Reloading

The daemon automatically watches the config file for changes and applies them immediately:

  • Notification settings: Applied instantly
  • Injection settings: Applied to current and future operations
  • Recording/Transcription/LLM settings: Applied to new recording sessions
  • Invalid configs: Rejected with error notification, daemon continues with previous config