# Provider Comparison Guide This guide helps you choose the right transcription provider for your use case. ## Transcription Providers | Provider | Type | Models | Languages | Streaming | Speed | Quality | Cost | |----------|------|--------|-----------|-----------|-------|---------|------| | **OpenAI** | Cloud | 4 | 57 | Yes | Fast | Excellent | $0.006/min | | **Groq** | Cloud | 3 | 57 (1 EN-only) | No | Very Fast | Excellent | Free tier | | **Mistral** | Cloud | 2 | 57 | No | Fast | Good | Pay per use | | **ElevenLabs** | Cloud | 4 | 57+ | Yes | Fast | Excellent | Pay per use | | **Deepgram** | Cloud | 4 | 33-42 | Yes | Very Fast | Excellent | Pay per use | | **whisper-cpp** | Local | 12 | 57 (4 EN-only) | No | Varies | Excellent | Free | ### OpenAI The original Whisper provider. Reliable and well-documented. **Models:** - `whisper-1` - Production speech-to-text (batch) - `gpt-4o-transcribe` - High quality with GPT-4o (batch) - `gpt-4o-mini-transcribe` - Faster with GPT-4o Mini (batch) - `gpt-4o-realtime-preview` - Real-time streaming **Best for:** General use, high accuracy requirements, streaming needs ### Groq Extremely fast inference using specialized hardware. OpenAI-compatible API. **Models:** - `whisper-large-v3` - Full Whisper v3, best accuracy - `whisper-large-v3-turbo` - Faster with slightly lower accuracy **Best for:** Speed-critical applications, English-only use cases, budget-conscious users ### Mistral European provider with Voxtral transcription models. **Models:** - `voxtral-mini-latest` - Latest Voxtral, recommended **Notes:** Mistral's streaming responses are not real-time audio streaming; hyprvoice treats Voxtral as batch-only. **Best for:** European data residency requirements, Mistral ecosystem users ### ElevenLabs Known for voice synthesis, also offers excellent transcription via Scribe. **Models:** - `scribe_v1` - 90+ languages, best accuracy (batch) - `scribe_v2` - Lower latency (batch) - `scribe_v2_realtime` - Streaming-only realtime endpoint **Best for:** Applications needing both TTS and STT, ultra-low latency streaming ### Deepgram Streaming-first provider with Nova models. Excellent for real-time applications. **Models:** - `flux-general-en` - Streaming with turn detection (English) - `nova-3` - Best accuracy, 42 languages - `nova-2` - Fast, 33 languages, filler word detection **Notes:** Flux is English-only. **Language Support:** Nova-3 supports 42 languages, Nova-2 supports 33 languages. Not all 57 languages from the master list are available. **Best for:** Real-time transcription, live captions, meeting transcription ### whisper-cpp (Local) Run Whisper models locally on your machine. No API keys, no network latency, complete privacy. **Requires:** `whisper-cli` binary installed on your system. **English-only models (faster):** | Model | Size | Speed | Quality | |-------|------|-------|---------| | `tiny.en` | 75MB | Fastest | Basic | | `base.en` | 142MB | Fast | Good | | `small.en` | 466MB | Medium | Better | | `medium.en` | 1.5GB | Slow | Best EN | **Multilingual models:** | Model | Size | Speed | Quality | |-------|------|-------|---------| | `tiny` | 75MB | Fastest | Basic | | `base` | 142MB | Fast | Good | | `small` | 466MB | Medium | Better | | `medium` | 1.5GB | Slow | Great | | `large-v1` | 2.9GB | Slowest | Best | | `large-v2` | 2.9GB | Slowest | Best | | `large-v3` | 3GB | Slowest | Best | | `large-v3-turbo` | 1.6GB | Slower | Great | **Best for:** Privacy-sensitive applications, offline use, avoiding API costs --- ## LLM Providers Used for post-processing transcriptions (formatting, summarization, etc.) | Provider | Models | Quality | Cost | |----------|--------|---------|------| | **OpenAI** | gpt-4o, gpt-4o-mini | Excellent | Pay per token | | **Groq** | llama-3.3-70b, llama-3.1-8b, mixtral-8x7b | Good-Excellent | Free tier | --- ## Choosing a Provider ### Decision Flowchart ``` Need complete privacy? ├─ Yes → whisper-cpp (local) └─ No └─ Need real-time streaming? ├─ Yes │ └─ Latency critical (<150ms)? │ ├─ Yes → ElevenLabs scribe_v2_realtime (streaming) │ └─ No → Deepgram nova-3 or OpenAI realtime └─ No (batch) └─ Need fastest response? ├─ Yes → Groq whisper-large-v3-turbo └─ No └─ Need highest accuracy? ├─ Yes → OpenAI gpt-4o-transcribe or Groq whisper-large-v3 └─ No → OpenAI whisper-1 (reliable default) ``` ### Quick Recommendations | Use Case | Recommended Provider | Model | |----------|---------------------|-------| | General dictation | OpenAI | whisper-1 | | Fast multilingual | Groq | whisper-large-v3-turbo | | Live captions | Deepgram | nova-3 | | Ultra-low latency | ElevenLabs | scribe_v2_realtime (streaming) | | Offline/privacy | whisper-cpp | base.en or base | | High accuracy | OpenAI | gpt-4o-transcribe | --- ## Language Support All providers support **auto-detect mode** (recommended for most users) which automatically identifies the spoken language. ### Full Language Support (57 languages) OpenAI, Groq, Mistral, ElevenLabs, and whisper-cpp multilingual models support all 57 languages: Afrikaans, Arabic, Armenian, Azerbaijani, Belarusian, Bosnian, Bulgarian, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, Galician, German, Greek, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Italian, Japanese, Kannada, Kazakh, Korean, Latvian, Lithuanian, Macedonian, Malay, Marathi, Maori, Nepali, Norwegian, Persian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swahili, Swedish, Tagalog, Tamil, Thai, Turkish, Ukrainian, Urdu, Vietnamese, Welsh ### English-Only Models These models only support English but are faster: | Provider | Model | |----------|-------| | whisper-cpp | `tiny.en`, `base.en`, `small.en`, `medium.en` | If you select an English-only model with a non-English language, hyprvoice will: 1. **At config time:** Show an error and prevent saving 2. **At runtime:** Fall back to auto-detect with a warning notification ### Deepgram Language Support Deepgram Nova models support a subset of languages: **Nova-3 (42 languages):** Arabic, Belarusian, Bosnian, Bulgarian, Catalan, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hindi, Hungarian, Indonesian, Italian, Japanese, Kannada, Korean, Latvian, Lithuanian, Macedonian, Malay, Marathi, Norwegian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swedish, Tagalog, Tamil, Turkish, Ukrainian, Vietnamese **Nova-2 (33 languages):** Bulgarian, Catalan, Chinese, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Latvian, Lithuanian, Malay, Norwegian, Polish, Portuguese, Romanian, Russian, Slovak, Spanish, Swedish, Thai, Turkish, Ukrainian, Vietnamese --- ## Streaming vs Batch ### Batch Transcription - Send complete audio file - Wait for full transcription - Higher accuracy - Better for: recordings, file processing, dictation ### Streaming Transcription - Send audio chunks in real-time - Get partial results immediately - Lower latency - Better for: live captions, voice commands, interactive apps **Streaming providers:** OpenAI (realtime model), ElevenLabs, Deepgram --- ## Local vs Cloud ### Cloud Providers **Pros:** - No setup required - Always up-to-date models - Scales automatically - Professional support **Cons:** - Requires internet connection - API costs - Data leaves your machine - Potential latency ### Local (whisper-cpp) **Pros:** - Complete privacy - No API costs - Works offline - No network latency - Your data stays on your machine **Cons:** - Requires setup (install whisper-cli) - Need to download models (75MB-3GB) - Uses local CPU/GPU resources - Slower on modest hardware ### When to Choose Local - Sensitive data (medical, legal, personal) - Offline environments - High-volume use (avoiding API costs) - Privacy-first applications - Air-gapped systems ### When to Choose Cloud - Quick setup needed - Best accuracy required - Real-time streaming - Light/occasional use - Mobile or low-power devices