@@ -1,139 +1,135 @@
# Hyprvoice
> **Voice‑ powered typing for Wayland/Hyprland desktops** — press a global shortcut, speak, and watch words appear in whatever window you're focused on.
> **Voice‑ powered typing for Wayland/Hyprland — press to toggle, speak, instant paste.**
> Streams audio while you talk and **pastes the final text the moment you toggle off** → aims to be the **fastest feel** on Wayland.
* * 🚧 Early Development Status:** This project is currently in early development. Not ready for production use yet.
**Status: ** Early development (expect rough edges)
---
## Why does this hyprvoice exist?
## TL;DR
Typing is slow, repetitive strain injuries are real. This project aims to give Wayland users a fast, privacy‑ respecting alternative that works entirely on their own hardware _ or _ any cloud ASR they trust .
- **Toggle workflow** (Hyprland‑ friendly): press to start, press to stop .
- **Cloud streaming ASR** (MVP) → **single final paste ** into the focused window.
- **Daemon with clear states & events**; desktop notifications.
- **Clipboard‑ based injection** (save/restore) with * * `wtype` **
- **Unixy pipeline** (small pieces, bounded channels).
---
## Current Implementation Statu s
## Requirement s
| Component | Status | Description |
| ------------------------- | ------- | ------------------------------------------------------------ |
| **Hot‑ key daemon ** | ✅ Done | Background service with Unix socket IPC for command handling |
| **Desktop notifications ** | ✅ Done | Recording state changes via `notify-send` |
| **Service management ** | ✅ Done | Install/remove systemd user service with embedded unit file |
| **Live audio capture ** | 🔄 TODO | PipeWire input with VAD/noise gate |
| **ASR backends ** | 🔄 TODO | Local Whisper or cloud APIs (OpenAI, Deepgram, AssemblyAI) |
| **Text injection ** | 🔄 TODO | `wtype` /`ydotool` keystrokes to focused window |
| **Configuration ** | 🔄 TODO | TOML config file with runtime reload |
- Wayland + **Hyprland **
- **PipeWire** (audio capture)
- **systemd --user** (service)
- **wl-clipboard** (clipboard save/restore)
- **libnotify**/`notify-send` (optional notifications)
- `wtype` or `ydotool` (optional text injection fallback)
> Other distros may work, but Arch/Hyprland is the primary target for now.
---
## Quick Start (Arch / Hyprland)
## Install (Arch / Hyprland)
``` bash
# 1. Install from AUR (source build)
yay -S hyprvoice
# or pre‑ built binary
yay -S hyprvoice-bin
# AUR
yay -S hyprvoice # or: yay -S hyprvoice-bin
# 2. Run the interactiv e setup helper
hyprvoice-install run --backend whispercpp
# 3. Enable the user service
# Enabl e u ser service
systemctl --user enable --now hyprvoice.service
# 4. Add a key binding in Hyprland conf
# Hyprland keybind (toggle)
bind = SUPER, R, exec, hyprvoice toggle
```
---
## Configuration file (`~/.config/hyprvoice/config.toml`)
## Usage
**Note: ** Configuration file support is not yet implemented. This shows the planned format:
``` toml
[ asr ]
backend = "whispercpp" # whispercpp | openai
model = "medium.en.bin" # used if backend = whispercpp
api_key = "" # used if backend = openai
[ vad ]
threshold = -45 # dBFS
hang_ms = 300
[ inject ]
method = "wtype" # wtype | ydotool | clipboard
[ keybind ]
# If you don't use Hyprland, set an XDG desktop accelerator here
shortcut = "CTRL+ALT+SPACE"
```
All options are documented in the sample config generated by `--init` .
- Press your **toggle ** key to start; press again to stop.
- Audio streams to the cloud ASR while you speak.
- On stop (or VAD endpoint), Hyprvoice **pastes once ** into the focused window.
- Injection flow: **save clipboard → copy final text → send Ctrl+V → restore clipboard ** .
---
## Architecture
## Status
```
┌──────────────┐ Unix Socket ┌──────────────┐ PCM ┌──────────┐
│ CLI Client │◄────────────────│ Hot‑key │────────▶│ Audio │
│ (toggle/stop)│ │ Daemon │ │ Capture │
└──────────────┘ │ (hotkeydaemon│ │ (pipewire│
│ + bus) │ │ + VAD) │
└──────────────┘ └──────────┘
│ │
│ notify-send │ chunks
▼ ▼
┌──────────────┐ ┌──────────┐
│ Desktop │ │ ASR │
│ Notification │ │ Adapter │
└──────────────┘ │(whisper/ │
│ openai) │
└──────────┘
│
│ text
▼
┌──────────┐
│ Inject │
│ (wtype/ │
│ ydotool) │
└──────────┘
```
| Component | State | Notes |
| -------------------------- | ----- | ------------------------------------------- |
| **Daemon (control plane) ** | ✅ | State, IPC, worker orchestration |
| **Recording control ** | ✅ | `hyprvoice toggle` |
| **Desktop notifications ** | ✅ | `notify-send` (logs fallback) |
| **Audio capture ** | 🔄 | PipeWire + VAD |
| **ASR backends ** | 🔄 | Cloud **streaming ** now; local Whisper next |
| **Text injection ** | 🔄 | Clipboard paste → `wtype` → `ydotool` |
| **Service management ** | 🔄 | `systemd --user` |
**Planned Components: **
- `cmd/hyprvoice` : CLI with Cobra commands (existing: `serve` , `toggle` , `stop` )
- `internal/bus` : Unix socket IPC for daemon communication (✅ implemented)
- `internal/hotkeydaemon` : Background service managing recording state (✅ implemented)
- `internal/notify` : Desktop notifications via `notify-send` (✅ implemented)
- `internal/audiocapture` : PipeWire input with VAD/noise gate (🔄 TODO)
- `internal/asr` : ASR backends (Whisper, OpenAI, etc.) (🔄 TODO)
- `internal/inject` : Text injection via wtype/ydotool (🔄 TODO)
- `internal/config` : TOML configuration with runtime reload (🔄 TODO)
Each box will be a Go package inside `internal/` so you can swap implementations without touching public APIs.
Legend: ✅ done · 🔄 in progress · ⏳ planned
---
## Building from Source
## How it works
- **Model:** pipeline + central state (daemon = control plane).
- **State machine:** `idle → recording → transcribing → injecting → idle` .
- **Rule:** switch to * * `transcribing` **\*\* as soon as the first audio frame is sent\*\* to the ASR.
### ASCII diagram
```
+-------------------+ Unix socket IPC +-----------+
CLI cmd → | Control Daemon | <---------------------------- | CLI/Tool |
|-------------------| +-----------+
| State: idle/rec/ |
| transcribing/... | events events
| Event bus (chan) | -----> [Notifications] -----> notify-send/log
| |
| frames finals |
+--+-----------+----+
| |
Audio | | Final Text
Frames v v
+--------+ +--------+ text +-----------+
| Audio |-->| ASR | -------------->| Injection |
| Capture| | Stream | | Worker |
+--------+ +--------+ +-----------+
| ^
+--------------+
backpressure via bounded channels
State (daemon):
idle --toggle--> recording --first frame--> transcribing --final--> injecting --done--> idle
```
### Data flow
1. `toggle` → **recording **
2. First frame sent → **transcribing **
3. Cloud ASR returns **final ** → **injecting **
4. Paste once → **idle **
5. Notifications at each transition
---
## Build from source
``` bash
git clone https://github.com/leonardotrapani/hyprvoice.git
cd hyprvoice
# Compile
CGO_ENABLED = 1 go build -o hyprvoice ./cmd/hyprvoice
# Run tests
go test ./...
```
**Requires ** Go 1.22+, a C compiler, and `pkg-config` with PipeWire headers.
---
Contributions welcome! Check the [issues ]( https://github.com/leonardotrapani/hyprvoice/issues ) for good first tasks.
## Contributing
- All PRs and issues welcome.
---
## License
This project is licensed under the MIT License. See the [LICENSE.md ](LICENSE.md ) file for details.
MIT — see [LICENSE.md ](LICENSE.md )