Files
hope-voice-api/AGENTS.md
T

122 lines
5.0 KiB
Markdown

# Voice — Real-time Voice Assistant
A single-container voice assistant built on LiveKit Agents. Speaks and listens in real time using Azure Speech (STT + TTS) with a local Gemma LLM for reasoning.
## Architecture
One Docker container runs three processes via supervisord:
1. **LiveKit server** — open-source WebRTC media transport (port 7880 TCP, 50000-60000 UDP)
2. **Voice agent** — Python LiveKit Agents pipeline: Azure STT → Gemma LLM → Azure TTS
3. **Web frontend** — static HTML served by a tiny HTTP server (port 8090)
```
Browser ──WebRTC──► LiveKit Server ──audio──► Agent
▲ │
└────────── audio ◄────────────────────────────┘
Azure STT → Gemma LLM → Azure TTS
```
## Quick Start
```bash
# 1. Set your Azure Speech key (already in ~/.hermes/.env as AZURE_SPEECH)
export AZURE_SPEECH_KEY=$(grep '^AZURE_SPEECH=' ~/.hermes/.env | cut -d= -f2)
# 2. Build and start
docker compose up --build -d
# 3. Open the web UI
# http://<host-ip>:8090
```
## Ports
| Port | Protocol | Service | Access |
|-------|----------|----------------------|--------------|
| 7880 | TCP | LiveKit HTTP/WS | LAN |
| 7881 | TCP | LiveKit internal | container |
| 50000-60000 | UDP | LiveKit RTC media | LAN |
| 8090 | TCP | Web frontend | LAN |
## Configuration
All config lives in `.env` (git-ignored). See `.env.example` for the full list.
Key variables:
- `AZURE_SPEECH_KEY` — Azure Speech resource key
- `AZURE_SPEECH_REGION` — default `eastus`
- `AZURE_TTS_VOICE` — default voice (e.g. `en-US-AvaNeural`)
- `GEMMA_BASE_URL` — LLM endpoint (default `http://192.168.86.2:8023/v1`)
- `GEMMA_MODEL` — model name (default `gemma-4-e4b`)
## Voice Selection
The web UI includes a voice picker that sends the selected voice to the agent via LiveKit data channel. The agent updates its TTS voice in real time without restarting.
Supported voices: any Azure Neural voice. See https://learn.microsoft.com/en-us/azure/ai-services/speech-service/language-support?tabs=en-us#neural-voices
## SSML / Expressive Speech
The TTS layer uses full SSML with `<mstts:express-as>` for style and `<prosody>` for rate/volume/pitch. The agent's system prompt instructs the LLM to write in a conversational, spoken style (short sentences, natural phrasing) so the output sounds like speech, not text.
## Agent Prompt
The LLM is instructed to:
- Speak as if talking to someone, not writing for them to read
- Keep responses to 1-3 sentences
- Use contractions, natural fillers sparingly
- Never use markdown, lists, or formatting
- Spell out numbers and abbreviations
## Testing
```bash
# Verify the container is running
docker compose ps
# Check agent logs
docker compose logs -f agent
# Test TTS directly (outside the agent)
curl -s "https://eastus.tts.speech.microsoft.com/cognitiveservices/v1" \
-H "Ocp-Apim-Subscription-Key: $AZURE_SPEECH_KEY" \
-H "Content-Type: application/ssml+xml" \
-H "X-Microsoft-OutputFormat: audio-24khz-48kbitrate-mono-mp3" \
--data-binary '<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis" xml:lang="en-US"><voice name="en-US-AvaNeural">Test</voice></speak>' > /tmp/test.mp3
```
## File Layout
```
~/dev/voice/
├── AGENTS.md ← you are here
├── .env.example ← config template (copy to .env)
├── .gitignore
├── docker-compose.yml ← single container, 3 processes
├── Dockerfile ← multi-stage build
├── livekit.yaml ← LiveKit server config
├── supervisord.conf ← process manager for the 3 services
├── agent/
│ ├── agent.py ← LiveKit Agents voice pipeline
│ └── pyproject.toml ← Python deps (uv)
├── web/
│ ├── index.html ← single-page voice UI
│ ├── app.js ← LiveKit client logic
│ └── style.css ← minimal dark theme
└── docs/
└── ARCHITECTURE.md ← deeper architecture notes
```
## Conventions
- **Single container.** All services (LiveKit, agent, web) run in one Docker container via supervisord. No multi-service compose.
- **No published UDP ports in compose.** LiveKit binds its UDP range directly on the host network (`network_mode: host`). This avoids the docker-proxy process explosion that hit hope-webui.
- **Gemma is a reasoning model.** It sometimes spends tokens on hidden reasoning before producing content. The agent handles this by using `max_tokens=1000` and falling back to `reasoning_content` if `content` is empty.
- **Azure TTS uses SSML, not JSON.** The REST endpoint requires `Content-Type: application/ssml+xml`. The LiveKit Azure plugin handles this internally.
- **Voice changes are live.** The web UI sends a data message to the agent; the agent calls `session.update_options(voice=...)` without restarting.
## Git
Commit frequently. Conventional commits (`feat:`, `fix:`, `chore:`).