Commit graph

39431 commits

Author SHA1 Message Date
mateo-berri
0e13cfabef
fix(gemini realtime): preserve sibling keys on empty toolCall no-op
Replace the early return on `functionCalls` empty/absent with a
`continue` plus a `tool_call_handled` flag that mirrors the existing
`server_content_handled` pattern. The post-loop guard already
distinguishes intentionally-consumed known keys from genuinely-unknown
messages, so adding `toolCall` to that exclusion list lets the loop
continue iterating over any sibling top-level keys in the same Gemini
frame instead of short-circuiting on the first empty toolCall.

In practice Gemini's protobuf places `toolCall`/`serverContent`/
`setupComplete` in a `oneof` so the only realistic sibling is
`usageMetadata` (already filtered as unknown-top-level), but the
uniform handling avoids silently discarding any future sibling key
should the wire format grow.
2026-05-23 01:58:30 +00:00
mateo-berri
70e1169989
fix(realtime): forward sanitized function_call_output on guardrail block
Providers that pair every toolCall with a toolResponse (e.g. Gemini and
Vertex Live) stay in the awaiting-tool-call state until a toolResponse
arrives. Dropping a blocked function_call_output outright left those
providers stalled — the subsequent guardrail clientContent and
response.create were ignored because the prior toolCall had no matching
toolResponse.

When the client-supplied tool output fails the realtime guardrail check,
forward a sanitized placeholder function_call_output (same call_id,
generic policy marker as output) instead of dropping the message
entirely. The placeholder carries no blocked content, so the model never
sees it, while still completing the provider's tool-call cycle so the
session can recover and the violation message reaches the user.
2026-05-23 01:55:32 +00:00
mateo-berri
d3490859a4
fix(ci): restore guardrail injection on duplicate session.created and cast realtime delta event
- Re-enable the one-time guardrail turn_detection update on duplicate
  session.created. `_maybe_send_guardrail_turn_detection_update` is
  already idempotent via `_guardrail_turn_detection_update_sent`, so
  the previous guard was unnecessary and broke the deferred-setup path
  where the synthetic session.created is emitted by llm_http_handler
  outside this loop (no prior chance to inject).

- Cast the response.function_call_arguments.delta dict appended to
  `returned_message: List[OpenAIRealtimeEvents]` so mypy is satisfied.
2026-05-23 01:52:45 +00:00
Cursor Agent
a27e6bcae9
fix(realtime): avoid stale session.created flag triggering guardrail re-injection
Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-23 01:38:17 +00:00
mateo-berri
4ca26e88a5
fix(gemini realtime): emit function_call_arguments.delta before .done
Gemini delivers the full function-call arguments in a single toolCall
frame. The OpenAI Realtime spec orders the streaming events as
output_item.added -> function_call_arguments.delta(+) ->
function_call_arguments.done -> output_item.done. Emit a single delta
carrying the complete arguments string before the matching .done so
spec-compliant SDK clients that accumulate deltas and gate finalisation
on at least one delta arriving do not stall on Gemini tool calls.
2026-05-23 01:29:30 +00:00
mateo-berri
16e44cbee8
fix(realtime): run guardrails on function_call_output content
Tool result outputs are client-controlled and fed to the model, so
they must pass the same content checks as user text messages.
Otherwise an attacker can smuggle blocked content into a
function_call_output and have the model process it.
2026-05-23 01:11:17 +00:00
Cursor Agent
295a9e6e13
fix(gemini realtime): promote nested turn_detection when flat value is not a dict
When the session payload had `turn_detection: None` (or any non-dict value), the
normalizer skipped promoting the GA nested `audio.input.turn_detection` because
it only checked key presence. The stale None then flowed into
`map_automatic_turn_detection` and raised TypeError on `'create_response' in value`.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-23 00:45:57 +00:00
Cursor Agent
ef98c98a12
refactor(gemini realtime): drop unused json_message arg from map_openai_event
Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-23 00:31:39 +00:00
Cursor Agent
f1b76d99e8
fix(gemini realtime): preserve sibling toolCall when serverContent has only transcription
Previously, when a Gemini frame contained both a transcription-only
serverContent and a sibling toolCall, the transcription handler would
early-return and silently drop the toolCall. Instead, mark serverContent
as handled and fall through so the main loop still processes siblings
like toolCall, while preserving the prior no-op behavior for empty/
transcription-only frames.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-23 00:15:54 +00:00
mateo-berri
84791fc3f5
fix(gemini realtime): tolerate sibling-only frames (e.g. standalone usageMetadata)
A Gemini Live frame that contains only metadata keys outside
_KNOWN_GEMINI_TOP_LEVEL_KEYS (e.g. a bare {"usageMetadata": {...}}
emitted between turns) leaves returned_message empty after the
transform loop and was tripping the 'Unknown message type' guard,
which raised ValueError and terminated the WebSocket session.

Treat such frames as no-ops and return the unchanged state instead.
2026-05-22 23:55:00 +00:00
Cursor Agent
1e4f86e19a
fix(realtime): inject guardrail turn_detection on subsequent session.update without one
Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 23:52:35 +00:00
mateo-berri
1a838a1517
fix(gemini realtime): cast maxOutputTokens to int for typeddict assignment 2026-05-22 23:34:44 +00:00
Cursor Agent
3a1a5ae392
fix(gemini realtime): use camelCase maxOutputTokens in response.done
Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 23:25:08 +00:00
mateo-berri
6299633138
fix(gemini realtime): cast maxOutputTokens to int for typeddict assignment 2026-05-22 23:12:13 +00:00
Cursor Agent
61928b9704
fix(gemini realtime): scope dotted-key event lookup and propagate session metadata to tool-call response.done
- map_openai_event: only check the current key/value pair when resolving
  dotted map entries (e.g. serverContent.turnComplete) so a sibling key in
  the same frame can't misclassify the event being processed
  (e.g. toolCall returning RESPONSE_DONE).
- tool-call path: extract generationConfig once and include modalities,
  temperature, and max_output_tokens on response.done so its shape matches
  response.created and the non-tool-call response.done.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 23:02:48 +00:00
Cursor Agent
2c41400bdf
fix(gemini realtime): skip unknown sibling keys in transform loop
Gemini realtime messages can include sibling metadata keys like
usageMetadata alongside primary payload keys (toolCall, serverContent).
Previously, the transform loop called map_openai_event for every
top-level key, raising ValueError for unknown ones and terminating
the WebSocket session.

Skip top-level keys not present in MAP_GEMINI_FIELD_TO_OPENAI_EVENT
to keep the session alive when Gemini emits usage metadata with a
toolCall response.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 22:49:23 +00:00
mateo-berri
42169c8578
test(gemini realtime): wrap test_gemini_tool_call_resets_ids fixture in setup envelope
The cached session_configuration_request the proxy stores is always
serialized as {"setup": ...}; this test passed a bare config dict, so
transform_session_created_event's .get('setup', {}) returned an empty
dict and the responseModalities lookup ran against the default rather
than the fixture. Wrap the fixture in the same shape the production
cache uses.
2026-05-22 22:34:53 +00:00
mateo-berri
d1a9da3514
fix(gemini realtime): cast merged realtimeInputConfig for typeddict assignment
mypy flagged the assignment of the merged dict into
BidiGenerateContentSetup.realtimeInputConfig with [typeddict-item]: the
intermediate variable widens to dict[Any, Any], losing the TypedDict
narrowing the previous dict-literal form had.
2026-05-22 22:29:15 +00:00
mateo-berri
bafc187248
test(gemini realtime): wrap remaining cached session configs in setup envelope
The session_configuration_request the proxy caches is always serialized
as {"setup": ...}; three modality-related tests dumped a bare config
dict instead, so transform_session_created_event's
`.get('setup', {})` quietly returned an empty dict and the
responseModalities lookup ran against the default rather than the
fixture. Wrap the remaining tests in the same shape the production
cache uses so any regression in modality forwarding actually trips.
2026-05-22 22:27:14 +00:00
mateo-berri
e5ffd021f7
fix(gemini realtime): bound _tool_call_id_to_name with an LRU; exercise modality forwarding test
Two minor follow-ups from review:

* Switch _tool_call_id_to_name to a 256-entry LRU OrderedDict so a long
  session with many tool calls doesn't grow the dict without bound,
  while retried function_call_output lookups still hit for recently-seen
  call_ids.
* Fix test_gemini_realtime_transformation_session_created to wrap the
  cached session config in {"setup": ...} so the modality lookup in
  transform_session_created_event actually exercises responseModalities
  forwarding (the prior payload was silently treated as empty).
2026-05-22 22:23:37 +00:00
mateo-berri
20764dd342
fix(realtime): record synthetic session.created in deferred-setup mode
The deferred-setup path emits a synthetic session.created directly to
the client websocket but did not run it through RealTimeStreaming's
store_message, so the event was missing from the session log used by
success_handler / async_success_handler. Call store_message before
forwarding so the synthetic event lands in the same log stream as
provider-driven events.
2026-05-22 22:14:15 +00:00
mateo-berri
459c1973b4
fix(gemini realtime): deep-merge automaticActivityDetection on follow-up session.update
The follow-up setup merge already deep-merged generationConfig and
realtimeInputConfig, but realtimeInputConfig.automaticActivityDetection
itself is a nested dict. A partial VAD update (e.g. the
guardrail-injected disabled=True from create_response=False) silently
dropped unrelated knobs such as silenceDurationMs and prefixPaddingMs
from the original setup. Deep-merge that block too so partial overrides
only touch the fields they specify.
2026-05-22 22:10:57 +00:00
mateo-berri
d56d875ea6
fix(vertex realtime): warn when dropping guardrail turn-detection update
In non-deferred mode the auto-setup is sent on connect, so the audio-transcription
guardrail's subsequent session.update carrying turn_detection.create_response=False
cannot be forwarded as a second setup (Vertex Live closes the WebSocket with 1007).
Surface a warning when this specific drop happens so operators know the model
will auto-respond before the guardrail can gate it, instead of failing silently
at debug level.
2026-05-22 22:04:19 +00:00
mateo-berri
0780e5f69f
fix(gemini realtime): empty toolCall must not terminate the WebSocket
If Gemini sends a toolCall whose functionCalls list is empty (or absent),
the previous `continue` left returned_message empty and the
"Unknown message type" guard fired, killing the WebSocket session.
Return a normal (empty) result instead so the session keeps going.
2026-05-22 21:54:16 +00:00
mateo-berri
c59260fc70
fix(gemini realtime): keep call_id→name mapping across function_call_output retries
A client SDK that retries function_call_output (or sends the same result
twice) would previously hit a missing-name lookup on the second send
because _handle_function_call_output popped the call_id → name entry.
Without name, Gemini may silently reject the response. Use dict.get so
the mapping persists for the lifetime of the session.
2026-05-22 21:50:22 +00:00
mateo-berri
b60cc950f3
fix(gemini realtime): mirror modalities/temperature/max_output_tokens on tool-call response.created
The audio/text response.created preamble includes modalities, temperature,
and max_output_tokens on the response object so spec-compliant clients can
initialise per-response state. The tool-call response.created was missing
these fields, leaving clients without consistent response metadata when a
response starts with a tool call instead of content. Read them from the
cached session_configuration_request the same way the audio/text path
does.
2026-05-22 21:44:43 +00:00
mateo-berri
615a7da9ba
fix(realtime): deep-merge generationConfig and refresh cache on follow-up setup
A subsequent Gemini session.update that touches any generationConfig sub-field
(e.g. just temperature) was clobbering the original generationConfig — silently
dropping responseModalities and switching the session to text-only. Deep-merge
generationConfig so existing keys (responseModalities, maxOutputTokens, ...) are
preserved when the client updates only a subset.

Also drop the early-return in _cache_session_configuration_request so the
cached payload tracks the latest setup sent to the backend. Without this,
downstream readers (transform_session_created_event, modality lookup in
return_new_content_delta_events) keep reading stale modalities/system
instruction after a follow-up setup.
2026-05-22 21:37:04 +00:00
mateo-berri
3efd803cd1
fix(gemini realtime): default conversation_id before tool-call response.done
mypy flagged that response.done's conversation_id (str on the TypedDict)
could be None when current_response_id was already set on entry. Ensure
the fallback runs unconditionally before the response is constructed.
2026-05-22 21:34:34 +00:00
Cursor Agent
a0043494a0
fix(gemini realtime): deep-merge nested config in follow-up session update
Previously, the follow-up setup performed a shallow merge between the
original setup and new overrides. If a session.update touched any field
inside generationConfig (e.g. modalities), the entire generationConfig
would be replaced, silently dropping unrelated sub-keys like temperature
or maxOutputTokens. Apply the same deep-merge to realtimeInputConfig so
partial automatic-activity-detection updates don't drop other realtime
input config fields either.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 21:33:07 +00:00
Cursor Agent
76a82447e8
fix(gemini realtime): include usage on tool-call response.done; coerce non-dict tool output to struct
- Tool-call response.done now includes an empty usage object, matching the
  non-tool-call path so OpenAI-compatible clients always see usage.
- _handle_function_call_output wraps non-dict JSON parses under a 'result'
  key so Gemini's functionResponses[].response (a Struct) always receives a
  mapping.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 21:20:42 +00:00
mateo-berri
ba75de7848
fix(realtime): preserve client session.update fields on follow-up Gemini setup
In non-deferred mode the auto-setup pre-populates session_configuration_request,
so a later client session.update carrying tools or instructions used to fall
into the subsequent path and only forward turn_detection. Rebuild a merged
follow-up setup that overlays the new client fields on top of the original
setup so tools/instructions/etc. are no longer silently dropped.
2026-05-22 21:19:13 +00:00
Cursor Agent
9efcc02776
fix(realtime): correct conversation_id, VAD disable, modality state, empty toolCall
- Gemini tool-call response.done now includes conversation_id so clients
  can match it against the preceding response.created.
- Vertex AI setup no longer overrides an explicit guardrail-injected
  create_response: False back to disabled: False; the guardrail's intent
  to disable VAD auto-response is now respected.
- Modality handler is now passed the locally-updated response/item IDs
  rather than the original input snapshot, preventing stale IDs after a
  prior tool-call/response.done in the same JSON message resets them.
- Skip emitting orphaned response.created/response.done events when
  Gemini sends an empty functionCalls array.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 21:06:06 +00:00
mateo-berri
24b8e17a4b
test(gemini realtime): exercise toolCall → function_call_output name round-trip
Update test_gemini_realtime_function_call_output_transformation to pre-load
the call_id → name mapping by transforming a Gemini toolCall first, then
assert that the resulting Gemini toolResponse functionResponses entry
carries the function name. This pins the production round-trip rather
than the degenerate 'name missing' branch.
2026-05-22 20:48:11 +00:00
Cursor Agent
36f6d84c22
fix(realtime): avoid double-serialization and normalize non-dict turn_detection in guardrail override
- Skip the force-override block when the injection block already ran for
  the same session.update to avoid redundant JSON re-serialization.
- Normalize non-dict client-provided turn_detection values (flat and
  nested audio.input.turn_detection) to a dict before enforcing
  create_response=False, matching the injection block's behavior and
  preventing potential bypass on backends that accept non-dict values.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 20:44:58 +00:00
Cursor Agent
50ef9c8c21
fix(gemini realtime): preserve original setup config on follow-up session.update
Gemini Live treats a second BidiGenerateContentSetup as a full session
replacement, not a partial merge. The guardrail-driven turn_detection-only
session.update was emitting a setup containing only model + realtimeInputConfig,
which would silently drop tools, generationConfig, inputAudioTranscription, and
systemInstruction from the original setup. Carry forward the cached original
setup and only override realtimeInputConfig.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 20:31:15 +00:00
mateo-berri
00cde9da7c
fix(lint): silence PLR0915 on client_ack_messages
The function exceeded the 50-statement limit (64 > 50) after recent
realtime guardrail additions. Matches the existing project pattern for
inherently complex event/message-mapping methods (see _process_event,
translate_messages_to_responses_input, transform_realtime_response,
_arealtime, etc.).
2026-05-22 20:20:11 +00:00
mateo-berri
f2d8a2a8fe
fix(realtime): force create_response=False in all client session.update turn_detection when audio guardrails active
Prevents a client from re-enabling Gemini/GA VAD auto-response (and thereby
bypassing the audio transcription guardrail) by sending a later
session.update with turn_detection.create_response: true.
2026-05-22 20:11:59 +00:00
mateo-berri
fb935426eb
fix(vertex_ai/realtime): normalize all GA-remapped session fields before mapping
Previously _build_vertex_ai_setup_config only lifted nested turn_detection
back to the top level. GA clients' output_modalities and
audio.input.transcription were silently dropped because map_openai_params
only recognises the flat OpenAI-beta keys. Use the parent's
_normalize_session_payload_for_mapping so modalities, transcription, and
turn_detection are all surfaced before mapping.
2026-05-22 19:43:11 +00:00
Cursor Agent
2325041913
fix(gemini realtime): normalize GA-remapped session fields before mapping
map_openai_params only recognises the flat OpenAI-beta keys (modalities,
input_audio_transcription, turn_detection). For GA clients the upstream
shim renames these into the nested GA schema (output_modalities,
audio.input.transcription, audio.input.turn_detection), causing them to
be silently dropped in _handle_session_update. Add a normalization helper
that surfaces the GA-remapped values back at the top level so the
existing mapping logic picks them up. Without this, a GA client
explicitly requesting modalities=['text'] would still default to audio
output.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 19:29:36 +00:00
Cursor Agent
8cf3f4eda1
fix(vertex_ai realtime): keep VAD enabled when guardrails inject create_response: False
map_automatic_turn_detection sets disabled=True whenever create_response is
absent OR False. Transcription guardrails inject create_response: False to
suppress auto-responses while expecting VAD to stay active, but the previous
override in _build_vertex_ai_setup_config only fired when create_response was
absent, leaving disabled=True and silently breaking speech detection and
transcription events. Vertex Live has no 'VAD on, no auto-response' mode, so
always keep VAD active in the setup config.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 19:06:38 +00:00
Cursor Agent
6f45b5df06
fix(realtime): set guardrail turn_detection flag only after successful send
Previously the _guardrail_turn_detection_update_sent flag was set inline
during message rewriting in client_ack_messages, before the modified
session.update was forwarded to the backend. If _send_to_backend raised
(e.g. backend WebSocket disconnect), the exception was caught and the
loop continued, but the flag remained True — permanently disabling the
guardrail create_response=False injection for the rest of the session.
Neither the client_ack_messages path nor the
_maybe_send_guardrail_turn_detection_update backup path would retry.

Track the injection locally and only set the flag after _send_to_backend
returns a truthy sent result, matching the pattern used by
_maybe_send_guardrail_turn_detection_update.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 18:49:47 +00:00
mateo-berri
b721714783 fix(gemini realtime): default response.done modalities to AUDIO and correct audio-done test 2026-05-22 18:34:43 +00:00
Cursor Agent
2f0d8080d1
fix(gemini realtime): default responseModalities to AUDIO in delta events
Align return_new_content_delta_events with the AUDIO defaults used in
_handle_session_update and transform_session_created_event so deferred
session config does not produce TEXT-typed delta events for audio data.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 18:23:05 +00:00
mateo-berri
bcba880c8b
fix(vertex_ai/realtime): drop follow-up session.update to avoid 1007 close
Vertex AI Live treats setup as a first-and-only client message; emitting a
second setup with realtimeInputConfig only closes the websocket with a 1007
policy error. Reverting the follow-up-setup branch restores the pre-existing
no-op behavior for subsequent session.update messages.
2026-05-22 18:11:21 +00:00
Cursor Agent
159982a892
fix(gemini realtime): default synthetic session.created modalities to AUDIO
The synthetic session.created event emitted in deferred setup mode used
TEXT as the default for responseModalities, while _handle_session_update
defaults to AUDIO. Align the default so clients reading modalities from
the initial session.created see the correct value for live sessions.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 18:09:03 +00:00
mateo-berri
38e822e41b
fix(realtime): tolerate non-dict turn_detection in guardrail injection
When a client sends a session.update whose turn_detection field is None or
a non-dict value (e.g. "auto"), the guardrail injection used setdefault
followed by item assignment on the returned value, raising TypeError. The
inner except only caught JSONDecodeError/AttributeError, so the TypeError
escaped to the outer Exception handler that wraps the entire client_ack
loop, killing the connection. Replace non-dict turn_detection with a
fresh dict carrying create_response=False so the guardrail still applies
without crashing the loop.
2026-05-22 17:49:18 +00:00
Cursor Agent
76225f35b1
fix(realtime): normalize Vertex AI nested turn_detection and unify session.created guardrail ordering
- Vertex AI _build_vertex_ai_setup_config now lifts nested
  audio.input.turn_detection to the top level before calling
  map_openai_params, mirroring the parent GeminiRealtimeConfig
  behavior. Without this, guardrail-injected create_response: False
  was silently dropped for GA-protocol Vertex AI clients.
- realtime_streaming session.created handling now sends the
  (possibly re-typed) event first and then triggers the guardrail
  turn-detection update for both first and duplicate cases, removing
  the inconsistent guardrail-then-event ordering for duplicates.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 17:16:34 +00:00
Cursor Agent
8e6fe6c8ad
fix(gemini-realtime): preserve nested turn_detection through map_openai_params
After the GA remap moves session.turn_detection into session.audio.input.turn_detection,
Gemini's map_openai_params only looks at top-level keys and silently drops it. Normalize
the extracted turn_detection back to the top level on first session.update so the guardrail
create_response:False (and any client-provided VAD settings) reach the Gemini setup.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 17:02:12 +00:00
Cursor Agent
24af92663d
fix(realtime): correct deferred-setup session.created modalities and reset IDs after response.done
- Convert provider's real session.created to session.updated when a synthetic
  one was already forwarded so clients receive the authoritative modalities
  derived from their session.update instead of the synthetic placeholder.
- Reset current_response_id / current_output_item_id after Gemini RESPONSE_DONE
  so a toolCall arriving in a later frame starts a fresh response instead of
  reusing the completed response's ID and emitting a duplicate response.done.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-22 16:45:13 +00:00
Sameer Kankute
86ace3ef1a
Merge branch 'litellm_internal_staging' into litellm_live_api_tool_calling_support 2026-05-22 21:56:28 +05:30