Hyprvoice
Voice‑powered typing for Wayland/Hyprland — press to toggle, speak, instant paste. Streams audio while you talk and pastes the final text the moment you toggle off → aims to be the fastest feel on Wayland.
Status: Early development (expect rough edges)
TL;DR
- Toggle workflow (Hyprland‑friendly): press to start, press to stop.
- Pipeline owns state; daemon is a thin control plane (IPC + lifecycle).
- Notifications for key events (recording started/ended, aborted).
- Audio capture via PipeWire (
pw-record) with backpressure. - ASR + clipboard injection are planned; injection is currently stubbed.
Requirements
- Go 1.24.5+ (for building from source)
- Wayland + Hyprland
- PipeWire tools:
pw-recordandpw-cli - systemd --user (service)
- Optional: libnotify/
notify-send(desktop notifications) - Planned/optional:
wl-clipboard(clipboard save/restore),wtype/ydotool(text injection)
Other distros may work, but Arch/Hyprland is the primary target for now.
Install (Arch / Hyprland)
# AUR
yay -S hyprvoice # or: yay -S hyprvoice-bin
# Enable user service
systemctl --user enable --now hyprvoice.service
# Hyprland keybind (toggle)
bind = SUPER, R, exec, hyprvoice toggle
Usage
Basic Usage
- Press your toggle key to start; press again to stop.
- Audio is captured via PipeWire; the pipeline enters
transcribingafter the first frame. - On toggle‑off during
transcribing, aninjectaction is sent. Injection is currently simulated (no clipboard paste yet).
CLI Commands
# Start the daemon
hyprvoice serve
# Toggle recording on/off
hyprvoice toggle
# Check current status
hyprvoice status
# Get protocol version
hyprvoice version
# Stop the daemon
hyprvoice stop
Status
| Component | State | Notes |
|---|---|---|
| Daemon (control plane) | ✅ | IPC server, lifecycle; forwards status from the pipeline |
| Recording control | ✅ | hyprvoice toggle |
| Desktop notifications | ✅ | notify-send (logs fallback) |
| Audio capture | ✅ | PipeWire (pw-record) frames + bounded channels |
| Simple transcriber | ✅ | Collect audio and transcribe when complete |
| OpenAI adapter | ✅ | HTTP API calls with clean audio buffering |
| whisper.cpp adapter | ⏳ | Local inference ready for implementation |
| Text injection | ⏳ | Not implemented; will use clipboard + wtype/ydotool |
| Service management | 🔄 | systemd --user unit example provided |
Legend: ✅ done · 🔄 in progress · ⏳ planned
How it works
- Model: The pipeline owns all runtime state; the daemon is a control plane (IPC + lifecycle) that starts/stops a pipeline instance and forwards status.
- State machine (pipeline):
idle → recording → transcribing → injecting → idle. - Rule: switch to
transcribingas soon as the first audio frame arrives.
Diagrams
flowchart LR
subgraph Client
CLI["CLI/Tool"]
end
subgraph Daemon
D["Control Daemon (lifecycle + IPC)"]
end
subgraph Pipeline
A["Audio Capture"]
T["Transcribing (ASR TBD)"]
I["Injecting (stub)"]
end
N["notify-send/log"]
CLI -- unix socket --> D
D -- start/stop --> A
A -- frames --> T
T -- status --> D
D -- events --> N
D -- inject action --> T
T --> I
I -->|done| D
stateDiagram-v2
[*] --> idle
idle --> recording: toggle
recording --> transcribing: first_frame
transcribing --> injecting: inject_action
injecting --> idle: done
recording --> idle: abort
injecting --> idle: abort
Transcription Strategy
Hyprvoice uses a simple collect-and-transcribe approach for reliable transcription:
- Collect all audio during recording session
- Single transcription when recording stops
- Clean, predictable results with full context
- Provider-agnostic adapter pattern for different backends
Architecture:
Audio Frames → Audio Buffer → Backend Adapter → Transcription
↓
[OpenAI API, whisper.cpp, etc.]
Data flow
toggle(daemon) → create pipeline → recording- First frame arrives → transcribing (daemon may notify
Transcribinglater) - Audio frames → audio buffer (collect all audio during session)
- Second
toggleduring transcribing → sendinjectaction → transcribe collected audio → injecting (simulated) - Complete → idle; pipeline stops; daemon clears reference
- Notifications at key transitions
Build from source
git clone https://github.com/leonardotrapani/hyprvoice.git
cd hyprvoice
# Build the binary
CGO_ENABLED=1 go build -o hyprvoice ./cmd/hyprvoice
# Run tests (when available)
go test ./...
# Install locally
sudo cp hyprvoice /usr/local/bin/
Dependencies
- Cobra CLI - Command-line interface framework
- Go 1.24.5+ - Programming language runtime
Configuration
File Locations
- Socket:
~/.cache/hyprvoice/control.sock- IPC communication - PID file:
~/.cache/hyprvoice/hyprvoice.pid- Process tracking
Systemd Service
In the future, this will be implemented with the command hyprvoice install
The daemon runs as a user service. To create a systemd service file:
# Create service file at ~/.config/systemd/user/hyprvoice.service
mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/hyprvoice.service << 'EOF'
[Unit]
Description=Hyprvoice daemon
After=pipewire.service
[Service]
Type=simple
ExecStart=/usr/local/bin/hyprvoice serve
Restart=on-failure
RestartSec=5
[Install]
WantedBy=default.target
EOF
# Enable and start
systemctl --user daemon-reload
systemctl --user enable --now hyprvoice.service
Development
Project Structure
hyprvoice/
├── cmd/hyprvoice/ # Main CLI application
├── internal/
│ ├── bus/ # IPC (Unix socket) + PID management
│ ├── daemon/ # Control plane (IPC server, lifecycle; no state)
│ ├── notify/ # Desktop notifications
│ ├── pipeline/ # Pipeline + state machine (record/transcribe/inject)
│ ├── recording/ # Audio capture via PipeWire
│ └── transcriber/ # Simple transcriber + adapters (OpenAI, whisper.cpp)
├── go.mod # Go module definition
└── README.md
State Machine
The pipeline operates with these states:
- idle → recording → transcribing → injecting → idle
IPC Protocol
Single-character commands over Unix socket:
t- Toggle recordings- Get statusv- Get protocol versionq- Quit daemon
Running in Development
# Terminal 1: Start daemon with logs
go run ./cmd/hyprvoice serve
# Terminal 2: Test commands
go run ./cmd/hyprvoice toggle
go run ./cmd/hyprvoice status
Direction / Roadmap
- ASR integration: OpenAI adapter complete; whisper.cpp adapter ready for implementation.
- Proper injection: clipboard save/restore + Ctrl+V, with
wtype/ydotoolfallbacks. - Configuration options: devices, sample rate, transcription providers.
- Enhanced features: VAD for auto-stop, improved chunking strategies if needed.
- Tests: comprehensive testing for pipeline state transitions and transcription.
- Direction is flexible; we can adjust based on UX feedback and performance needs.
Troubleshooting
Common Issues
Daemon won't start
# Check if already running
hyprvoice status
# Check PID file
ls -la ~/.cache/hyprvoice/
# Remove stale files
rm ~/.cache/hyprvoice/hyprvoice.pid
rm ~/.cache/hyprvoice/control.sock
No notifications
# Test notify-send
notify-send "Test notification"
# Check if libnotify is installed
which notify-send
Permission errors
# Check socket permissions
ls -la ~/.cache/hyprvoice/control.sock
# Recreate cache directory
rm -rf ~/.cache/hyprvoice
mkdir -p ~/.cache/hyprvoice
Debug Mode
# Run with verbose logging
hyprvoice serve 2>&1 | tee hyprvoice.log
Contributing
- All PRs and issues welcome.
- Follow existing code conventions
- Add tests for new functionality
- Update documentation for user-facing changes
License
MIT — see LICENSE.md