feat: models test and fixes
This commit is contained in:
+21
-15
@@ -108,10 +108,12 @@ language = "" # Empty for auto-detect, or "en", "es", "fr", et
|
||||
|
||||
Transcription using Mistral's Voxtral API, excellent for European languages:
|
||||
|
||||
Note: Mistral's API supports streaming responses, but it is not real-time audio streaming. Hyprvoice treats Voxtral as batch-only.
|
||||
|
||||
```toml
|
||||
[transcription]
|
||||
provider = "mistral-transcription"
|
||||
model = "voxtral-mini-latest" # Or "voxtral-mini-2507"
|
||||
model = "voxtral-mini-latest"
|
||||
language = "" # Empty for auto-detect
|
||||
```
|
||||
|
||||
@@ -122,7 +124,7 @@ Transcription using ElevenLabs' Scribe API with 57+ language support:
|
||||
```toml
|
||||
[transcription]
|
||||
provider = "elevenlabs"
|
||||
model = "scribe_v1" # Or "scribe_v2" for lower latency
|
||||
model = "scribe_v1" # Or "scribe_v2" for lower latency (batch)
|
||||
language = "" # Empty for auto-detect
|
||||
```
|
||||
|
||||
@@ -148,9 +150,9 @@ language = "" # Empty for auto-detect
|
||||
|
||||
**Features:**
|
||||
|
||||
- All models are streaming-only
|
||||
- Nova-3: 42 languages, best accuracy
|
||||
- Nova-2: 33 languages, faster with filler word detection
|
||||
- Flux: streaming-only, English with turn detection
|
||||
- Nova-3: 42 languages, best accuracy (batch+streaming)
|
||||
- Nova-2: 33 languages, faster with filler word detection (batch+streaming)
|
||||
- Excellent for real-time transcription and live captions
|
||||
|
||||
### Local Transcription (whisper-cpp)
|
||||
@@ -182,7 +184,10 @@ threads = 0 # 0 = auto (uses NumCPU - 1)
|
||||
| `base` | 142MB | 57 languages | Daily multilingual use |
|
||||
| `small` | 466MB | 57 languages | Better multilingual |
|
||||
| `medium` | 1.5GB | 57 languages | Great accuracy |
|
||||
| `large-v1` | 2.9GB | 57 languages | Best accuracy |
|
||||
| `large-v2` | 2.9GB | 57 languages | Best accuracy |
|
||||
| `large-v3` | 3GB | 57 languages | Best accuracy |
|
||||
| `large-v3-turbo` | 1.6GB | 57 languages | Faster large-v3 |
|
||||
|
||||
**Threads configuration:**
|
||||
|
||||
@@ -195,12 +200,13 @@ threads = 0 # 0 = auto (uses NumCPU - 1)
|
||||
For real-time transcription, use streaming models:
|
||||
|
||||
```toml
|
||||
# ElevenLabs streaming
|
||||
# ElevenLabs streaming (realtime only)
|
||||
[transcription]
|
||||
provider = "elevenlabs"
|
||||
model = "scribe_v1-streaming" # Or "scribe_v2-streaming" for <150ms latency
|
||||
model = "scribe_v2_realtime"
|
||||
streaming = true
|
||||
|
||||
# Deepgram streaming (all models are streaming)
|
||||
# Deepgram streaming (all models support streaming)
|
||||
[transcription]
|
||||
provider = "deepgram"
|
||||
model = "nova-3"
|
||||
@@ -215,8 +221,8 @@ model = "gpt-4o-realtime-preview"
|
||||
|
||||
| Provider | Model | Latency | Languages |
|
||||
|----------|-------|---------|-----------|
|
||||
| ElevenLabs | `scribe_v1-streaming` | Low | 57+ |
|
||||
| ElevenLabs | `scribe_v2-streaming` | <150ms | 57+ |
|
||||
| ElevenLabs | `scribe_v2_realtime` | <150ms | 57+ |
|
||||
| Deepgram | `flux-general-en` | Very Low | en |
|
||||
| Deepgram | `nova-3` | Low | 42 |
|
||||
| Deepgram | `nova-2` | Very Low | 33 |
|
||||
| OpenAI | `gpt-4o-realtime-preview` | Low | 57 |
|
||||
@@ -257,7 +263,6 @@ Some models only support English. When configuring via `hyprvoice configure`, on
|
||||
|
||||
| Provider | Model |
|
||||
|----------|-------|
|
||||
| Groq | `distil-whisper-large-v3-en` |
|
||||
| whisper-cpp | `tiny.en`, `base.en`, `small.en`, `medium.en` |
|
||||
|
||||
**Deepgram models** support fewer languages than the full 57 - see [providers.md](./providers.md#deepgram-language-support).
|
||||
@@ -270,8 +275,8 @@ Some models only support English. When configuring via `hyprvoice configure`, on
|
||||
```toml
|
||||
# This combination will be rejected at validation:
|
||||
[transcription]
|
||||
provider = "groq-transcription"
|
||||
model = "distil-whisper-large-v3-en" # English only!
|
||||
provider = "whisper-cpp"
|
||||
model = "base.en" # English only!
|
||||
language = "es" # Error: model does not support Spanish
|
||||
```
|
||||
|
||||
@@ -622,8 +627,9 @@ You can customize notification text via the `[notifications.messages]` section:
|
||||
api_key = "..."
|
||||
|
||||
[transcription]
|
||||
provider = "elevenlabs"
|
||||
model = "scribe_v2-streaming" # <150ms latency
|
||||
provider = "elevenlabs"
|
||||
model = "scribe_v2_realtime" # <150ms latency
|
||||
streaming = true
|
||||
language = "" # Auto-detect
|
||||
|
||||
[llm]
|
||||
|
||||
+15
-17
@@ -11,7 +11,7 @@ This guide helps you choose the right transcription provider for your use case.
|
||||
| **Mistral** | Cloud | 2 | 57 | No | Fast | Good | Pay per use |
|
||||
| **ElevenLabs** | Cloud | 4 | 57+ | Yes | Fast | Excellent | Pay per use |
|
||||
| **Deepgram** | Cloud | 4 | 33-42 | Yes | Very Fast | Excellent | Pay per use |
|
||||
| **whisper-cpp** | Local | 9 | 57 (4 EN-only) | No | Varies | Excellent | Free |
|
||||
| **whisper-cpp** | Local | 12 | 57 (4 EN-only) | No | Varies | Excellent | Free |
|
||||
|
||||
### OpenAI
|
||||
|
||||
@@ -32,7 +32,6 @@ Extremely fast inference using specialized hardware. OpenAI-compatible API.
|
||||
**Models:**
|
||||
- `whisper-large-v3` - Full Whisper v3, best accuracy
|
||||
- `whisper-large-v3-turbo` - Faster with slightly lower accuracy
|
||||
- `distil-whisper-large-v3-en` - **English only**, fastest option
|
||||
|
||||
**Best for:** Speed-critical applications, English-only use cases, budget-conscious users
|
||||
|
||||
@@ -42,7 +41,8 @@ European provider with Voxtral transcription models.
|
||||
|
||||
**Models:**
|
||||
- `voxtral-mini-latest` - Latest Voxtral, recommended
|
||||
- `voxtral-mini-2507` - Stable version from July 2025
|
||||
|
||||
**Notes:** Mistral's streaming responses are not real-time audio streaming; hyprvoice treats Voxtral as batch-only.
|
||||
|
||||
**Best for:** European data residency requirements, Mistral ecosystem users
|
||||
|
||||
@@ -52,9 +52,8 @@ Known for voice synthesis, also offers excellent transcription via Scribe.
|
||||
|
||||
**Models:**
|
||||
- `scribe_v1` - 90+ languages, best accuracy (batch)
|
||||
- `scribe_v2` - Lower latency, real-time optimized (batch)
|
||||
- `scribe_v1-streaming` - Real-time transcription
|
||||
- `scribe_v2-streaming` - Real-time with <150ms latency
|
||||
- `scribe_v2` - Lower latency (batch)
|
||||
- `scribe_v2_realtime` - Streaming-only realtime endpoint
|
||||
|
||||
**Best for:** Applications needing both TTS and STT, ultra-low latency streaming
|
||||
|
||||
@@ -63,10 +62,11 @@ Known for voice synthesis, also offers excellent transcription via Scribe.
|
||||
Streaming-first provider with Nova models. Excellent for real-time applications.
|
||||
|
||||
**Models:**
|
||||
- `flux-general-en` - Streaming with turn detection (English)
|
||||
- `nova-3` - Best accuracy, 42 languages
|
||||
- `nova-3-general` - Same as nova-3
|
||||
- `nova-2` - Fast, 33 languages, filler word detection
|
||||
- `nova-2-general` - Same as nova-2
|
||||
|
||||
**Notes:** Flux is English-only.
|
||||
|
||||
**Language Support:** Nova-3 supports 42 languages, Nova-2 supports 33 languages. Not all 57 languages from the master list are available.
|
||||
|
||||
@@ -93,7 +93,10 @@ Run Whisper models locally on your machine. No API keys, no network latency, com
|
||||
| `base` | 142MB | Fast | Good |
|
||||
| `small` | 466MB | Medium | Better |
|
||||
| `medium` | 1.5GB | Slow | Great |
|
||||
| `large-v1` | 2.9GB | Slowest | Best |
|
||||
| `large-v2` | 2.9GB | Slowest | Best |
|
||||
| `large-v3` | 3GB | Slowest | Best |
|
||||
| `large-v3-turbo` | 1.6GB | Slower | Great |
|
||||
|
||||
**Best for:** Privacy-sensitive applications, offline use, avoiding API costs
|
||||
|
||||
@@ -121,14 +124,11 @@ Need complete privacy?
|
||||
└─ Need real-time streaming?
|
||||
├─ Yes
|
||||
│ └─ Latency critical (<150ms)?
|
||||
│ ├─ Yes → ElevenLabs scribe_v2-streaming
|
||||
│ ├─ Yes → ElevenLabs scribe_v2_realtime (streaming)
|
||||
│ └─ No → Deepgram nova-3 or OpenAI realtime
|
||||
└─ No (batch)
|
||||
└─ Need fastest response?
|
||||
├─ Yes
|
||||
│ └─ English only?
|
||||
│ ├─ Yes → Groq distil-whisper-large-v3-en
|
||||
│ └─ No → Groq whisper-large-v3-turbo
|
||||
├─ Yes → Groq whisper-large-v3-turbo
|
||||
└─ No
|
||||
└─ Need highest accuracy?
|
||||
├─ Yes → OpenAI gpt-4o-transcribe or Groq whisper-large-v3
|
||||
@@ -140,10 +140,9 @@ Need complete privacy?
|
||||
| Use Case | Recommended Provider | Model |
|
||||
|----------|---------------------|-------|
|
||||
| General dictation | OpenAI | whisper-1 |
|
||||
| Fast English | Groq | distil-whisper-large-v3-en |
|
||||
| Fast multilingual | Groq | whisper-large-v3-turbo |
|
||||
| Live captions | Deepgram | nova-3 |
|
||||
| Ultra-low latency | ElevenLabs | scribe_v2-streaming |
|
||||
| Ultra-low latency | ElevenLabs | scribe_v2_realtime (streaming) |
|
||||
| Offline/privacy | whisper-cpp | base.en or base |
|
||||
| High accuracy | OpenAI | gpt-4o-transcribe |
|
||||
|
||||
@@ -155,7 +154,7 @@ All providers support **auto-detect mode** (recommended for most users) which au
|
||||
|
||||
### Full Language Support (57 languages)
|
||||
|
||||
OpenAI, Groq (except distil model), Mistral, ElevenLabs, and whisper-cpp multilingual models support all 57 languages:
|
||||
OpenAI, Groq, Mistral, ElevenLabs, and whisper-cpp multilingual models support all 57 languages:
|
||||
|
||||
Afrikaans, Arabic, Armenian, Azerbaijani, Belarusian, Bosnian, Bulgarian, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, Galician, German, Greek, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Italian, Japanese, Kannada, Kazakh, Korean, Latvian, Lithuanian, Macedonian, Malay, Marathi, Maori, Nepali, Norwegian, Persian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swahili, Swedish, Tagalog, Tamil, Thai, Turkish, Ukrainian, Urdu, Vietnamese, Welsh
|
||||
|
||||
@@ -165,7 +164,6 @@ These models only support English but are faster:
|
||||
|
||||
| Provider | Model |
|
||||
|----------|-------|
|
||||
| Groq | `distil-whisper-large-v3-en` |
|
||||
| whisper-cpp | `tiny.en`, `base.en`, `small.en`, `medium.en` |
|
||||
|
||||
If you select an English-only model with a non-English language, hyprvoice will:
|
||||
|
||||
Reference in New Issue
Block a user