From f1e7d9eef7b988199353e4999716c0bd32f50354 Mon Sep 17 00:00:00 2001 From: Yuneng Jiang Date: Wed, 26 Aug 2026 20:58:38 -0700 Subject: [PATCH] fix(e2e): move the vertex realtime suite off the retired Live preview model Google withdrew gemini-live-2.5-flash-preview-native-audio-09-2025 from the Vertex Live API. Every session dies at setup: received 1007 (invalid frame payload data) gemini-live-2.5-flash-preview-native-audio-09-2025 is not supported in the live api. The client sees session.created (the proxy synthesizes it on connect) and then nothing, so both vertex_ai realtime tests time out waiting for session.updated. Confirmed by probing the Vertex Live endpoint directly with the e2e stack's own credentials: gemini-live-2.5-flash-preview-native-audio-09-2025 -> 1007, not supported gemini-live-2.5-flash-native-audio -> setupComplete so this swaps to the non-preview sibling, which is the same native-audio class and is what the cost map already carries for vertex_ai. Not a litellm regression. The suspicion fell on #38395 because it removed the native-audio speechConfig strip, but the setup payload this suite sends is byte-identical either side of that change: the strip only fires when a client sends a voice, and the e2e SessionConfig has no voice field. Google's rejection names the model, not a field. The gemini (Google AI Studio) provider keeps the -09-2025 id, which still works there; only the Vertex endpoint dropped it. (cherry picked from commit a215ecaf3d64ecf50a9f9868b1b8bdcfc91fc955) --- tests/e2e/llm_translation/realtime/REALTIME_COVERAGE_MATRIX.md | 2 +- tests/e2e/llm_translation/realtime/realtime_client.py | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/tests/e2e/llm_translation/realtime/REALTIME_COVERAGE_MATRIX.md b/tests/e2e/llm_translation/realtime/REALTIME_COVERAGE_MATRIX.md index bae858d50af..a6e32b88479 100644 --- a/tests/e2e/llm_translation/realtime/REALTIME_COVERAGE_MATRIX.md +++ b/tests/e2e/llm_translation/realtime/REALTIME_COVERAGE_MATRIX.md @@ -42,7 +42,7 @@ at call time. The provider table below is the source of truth; edit `PROVIDERS` | openai | `openai-realtime` | `openai/gpt-realtime-2` | | azure | `azure-realtime` | `azure/gpt-realtime-2` (GA protocol) | | gemini | `gemini-realtime` | `gemini/gemini-3.1-flash-live-preview` | -| vertex_ai | `vertex-realtime` | `vertex_ai/gemini-live-2.5-flash-preview-native-audio-09-2025` | +| vertex_ai | `vertex-realtime` | `vertex_ai/gemini-live-2.5-flash-native-audio` | Bedrock and xai (`xai/grok-4-1-fast-non-reasoning`) are supported by the proxy but kept commented out in `PROVIDERS` until they pass end-to-end here; re-enable them by diff --git a/tests/e2e/llm_translation/realtime/realtime_client.py b/tests/e2e/llm_translation/realtime/realtime_client.py index 632a9cf7e57..3ffca7e8b88 100644 --- a/tests/e2e/llm_translation/realtime/realtime_client.py +++ b/tests/e2e/llm_translation/realtime/realtime_client.py @@ -78,7 +78,7 @@ PROVIDERS = ( "vertex_ai", "vertex-realtime", LiteLLMParamsBody( - model="vertex_ai/gemini-live-2.5-flash-preview-native-audio-09-2025", + model="vertex_ai/gemini-live-2.5-flash-native-audio", vertex_location="us-central1", vertex_credentials="os.environ/VERTEXAI_CREDENTIALS", ),