5.8 KiB
Hyprvoice
Voice‑powered typing for Wayland/Hyprland desktops — press a global shortcut, speak, and watch words appear in whatever window you're focused on.
🚧 Early Development Status: This project is currently in early development. Not ready for production use yet.
Why does this hyprvoice exist?
Typing is slow, repetitive strain injuries are real. This project aims to give Wayland users a fast, privacy‑respecting alternative that works entirely on their own hardware or any cloud ASR they trust.
Current Implementation Status
| Component | Status | Description |
|---|---|---|
| Hot‑key daemon | ✅ Done | Background service with Unix socket IPC for command handling |
| Desktop notifications | ✅ Done | Recording state changes via notify-send |
| Service management | ✅ Done | Install/remove systemd user service with embedded unit file |
| Live audio capture | 🔄 TODO | PipeWire input with VAD/noise gate |
| ASR backends | 🔄 TODO | Local Whisper or cloud APIs (OpenAI, Deepgram, AssemblyAI) |
| Text injection | 🔄 TODO | wtype/ydotool keystrokes to focused window |
| Configuration | 🔄 TODO | TOML config file with runtime reload |
Quick Start (Arch / Hyprland)
# 1. Install from AUR (source build)
yay -S hyprvoice
# or pre‑built binary
yay -S hyprvoice-bin
# 2. Run the interactive setup helper
hyprvoice-install run --backend whispercpp
# 3. Enable the user service
systemctl --user enable --now hyprvoice.service
# 4. Add a key binding in Hyprland conf
bind = SUPER, R, exec, hyprvoice toggle
Configuration file (~/.config/hyprvoice/config.toml)
Note: Configuration file support is not yet implemented. This shows the planned format:
[asr]
backend = "whispercpp" # whispercpp | openai
model = "medium.en.bin" # used if backend = whispercpp
api_key = "" # used if backend = openai
[vad]
threshold = -45 # dBFS
hang_ms = 300
[inject]
method = "wtype" # wtype | ydotool | clipboard
[keybind]
# If you don't use Hyprland, set an XDG desktop accelerator here
shortcut = "CTRL+ALT+SPACE"
All options are documented in the sample config generated by --init.
Architecture
┌──────────────┐ Unix Socket ┌──────────────┐ PCM ┌──────────┐
│ CLI Client │◄────────────────│ Hot‑key │────────▶│ Audio │
│ (toggle/stop)│ │ Daemon │ │ Capture │
└──────────────┘ │ (hotkeydaemon│ │ (pipewire│
│ + bus) │ │ + VAD) │
└──────────────┘ └──────────┘
│ │
│ notify-send │ chunks
▼ ▼
┌──────────────┐ ┌──────────┐
│ Desktop │ │ ASR │
│ Notification │ │ Adapter │
└──────────────┘ │(whisper/ │
│ openai) │
└──────────┘
│
│ text
▼
┌──────────┐
│ Inject │
│ (wtype/ │
│ ydotool) │
└──────────┘
Planned Components:
cmd/hyprvoice: CLI with Cobra commands (existing:serve,toggle,stop)internal/bus: Unix socket IPC for daemon communication (✅ implemented)internal/hotkeydaemon: Background service managing recording state (✅ implemented)internal/notify: Desktop notifications vianotify-send(✅ implemented)internal/audiocapture: PipeWire input with VAD/noise gate (🔄 TODO)internal/asr: ASR backends (Whisper, OpenAI, etc.) (🔄 TODO)internal/inject: Text injection via wtype/ydotool (🔄 TODO)internal/config: TOML configuration with runtime reload (🔄 TODO)
Each box will be a Go package inside internal/ so you can swap implementations without touching public APIs.
Building from Source
git clone https://github.com/leonardotrapani/hyprvoice.git
cd hyprvoice
# Compile
CGO_ENABLED=1 go build -o hyprvoice ./cmd/hyprvoice
# Run tests
go test ./...
Requires Go 1.22+, a C compiler, and pkg-config with PipeWire headers.
Contributions welcome! Check the issues for good first tasks.
License
This project is licensed under the MIT License. See the LICENSE.md file for details.