diff --git a/FEATURES.md b/FEATURES.md index 8f0ad96b..b8761d0a 100644 --- a/FEATURES.md +++ b/FEATURES.md @@ -51,6 +51,23 @@ It combines: - Safer speech rules for agent-driven bubble content - Bubble behavior that avoids showing code, logs, URLs, paths, or secrets in normal integration speech +## Pet Text-to-Speech (Phase 2) + +- Settings > **Text-to-Speech** panel controls voice output for the pet and floating chat. +- **System voice** uses the OS speech engine through the renderer (`window.speechSynthesis`) with optional voice-name matching and a 0.5×–2.0× speed multiplier. +- **Cloud TTS providers**: OpenAI TTS, ElevenLabs, and an **OpenAI-compatible** preset for OpenRouter, LiteLLM, WaveSpeedAI, or a custom endpoint. +- **Local TTS provider**: Piper (spawned as a local process) for fully offline speech. +- Provider credentials are stored with Electron `safeStorage` and fall back to plain local storage when encryption is unavailable (`apps/desktop/src/tts-credentials.ts`). +- Per-provider model, voice, speed, endpoint preset, and custom endpoint controls. +- Dynamic voice list fetching for ElevenLabs; static voice lists for OpenAI / OpenAI-compatible. +- In-app **Test voice** and **Stop** buttons to preview the configured voice. +- Assistant replies and other speech are spoken through the pet window renderer via a shared TTS service (`apps/desktop/src/tts-service.ts`). +- Respects quiet-hours: speech is skipped while quiet hours are active. +- Audio returned by cloud providers is validated (MP3/WAV magic bytes) and played through a renderer `