fix: stop voice call hanging in "speaking" when TTS fails (#30372)
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run

When the configured TTS provider fails during a voice call, the text answer arrives but the overlay stays in "speaking" with no audio and no error until the user taps to interrupt. The sentence that failed never reaches the audio cache, so the playback loop re-queues it every 200 ms forever.

A failed sentence now marks its message as failed. The playback loop drops that message's unplayed sentences, the rest of the turn requests no more TTS, and the overlay returns to listening once the text finishes. The OpenAI-compatible path now shows the provider error once per turn, the same way Read Aloud and the Kokoro path already do. The next turn tries TTS again.

Failure is tracked per message so an outage (the report shows 16 parallel requests all failing) costs one toast and no further requests. The trade-off is that a one-off failure mutes the rest of that reply.

Verified with the real fetch/playback code in a harness: base loops forever with no toast; with the fix, one toast, the loop ends, later sentences are not requested, a new turn plays normally, and a late failure from a previous turn does not affect the next one.

Fixes #30052
This commit is contained in:
Classic298 2026-09-22 22:57:10 +02:00 • committed by GitHub
parent 2f92635409
commit 24c01bd76d
No known key found for this signature in database
GPG key ID: B5690EEEBB952194

View file

@ -379,6 +379,7 @@
};
let finishedMessages = {};
let failedMessages = {};
let currentMessageId = null;
let currentUtterance: SpeechSynthesisUtterance | null = null;
@ -502,7 +503,9 @@
const emojiCache = new Map();
const fetchAudio = async (content) => {
if (!audioCache.has(content)) {
const id = currentMessageId;
if (!audioCache.has(content) && !failedMessages[id]) {
try {
// Set the emoji for the content if needed
if ($settings?.showEmojiInCall ?? false) {
@ -535,6 +538,10 @@
const res = await synthesizeOpenAISpeech(localStorage.token, getVoiceId(), content).catch(
(error) => {
console.error(error);
if (!failedMessages[id]) {
failedMessages[id] = true;
toast.error(`${error}`);
}
return null;
}
);
@ -550,6 +557,10 @@
} catch (error) {
console.error('Error synthesizing speech:', error);
}
if (!audioCache.has(content)) {
failedMessages[id] = true;
}
}
return audioCache.get(content);
@ -591,7 +602,7 @@
} else {
await speakSpeechSynthesisHandler(content);
}
} else {
} else if (!failedMessages[id]) {
// If not available in the cache, push it back to the queue and delay
messages[id].unshift(content); // Re-queue the content at the start
console.log(`Audio for "${content}" not yet available in the cache, re-queued...`);