* fix(azure): add model router flat fee to azure provider cost tracking
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(azure): price the azure router fee from one canonical entry
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(azure): reuse the azure_ai model router fee path for azure
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure): drop model router fee unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Move the shared test helpers out of litellm-inference's test-support
feature into a publish = false litellm-inference-testing crate used only
as a dev dependency by the format crates.
Also drop the dead src/constants.rs (OPENAI_DEFAULT_API_BASE had no
users) and declare the litellm-http/litellm-llms test-support features
on the crates that actually use them instead of relying on feature
unification through litellm-inference's dev-dependencies.
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): restore key activity search and add model activity search
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): drop rebuilt dashboard bundle from the usage search change
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(ui): format usage search tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): only show model no-match when the range has models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): budget s3 dedupe retries on the deployment so the seeded router num_retries cannot zero them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): prevent nested router retries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): fix provider retry and pytest collection setup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): keep sdk retries on the sync router responses path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(utils): count upstream calls instead of doubling the retry helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): replace Any with proven types in 299 files
Clears 871 basedpyright Any errors (reportAny 6,703 to 6,240, reportExplicitAny 1,651 to 1,243) without adding a cast, an ignore or a suppression, and without touching any budget file
Most edits are annotation-only: a parameter, return or local goes from Any to object, Mapping[str, object] or the concrete type the value always held. Fourteen files validate untyped JSON once where it enters, through a module-level pydantic TypeAdapter or model_validate, and then use real types
No HTTP status, error type or response shape changes. Mistral speech and fal.ai Bria image generation now report a pydantic ValidationError instead of an AttributeError when the provider answers 2xx with a body that is not a JSON object
* refactor(types): make the config locals fix effective and trim no-op edits
The Mapping[str, object] annotation on `locals().copy()` removed no error,
because the checker narrows the variable back to the dict[str, Any] the call
returns. 21 provider config constructors now build the same copy with
dict(locals()), which the checker infers as object values under that
annotation, so each file loses one reportAny.
The same annotation is reverted in 31 other config files where it stayed a
no-op, together with the tests that were added only to cover those lines, and
the one MCP server manager line that no CI coverage shard executes is
reverted too. The pull request drops from 345 to 297 changed files.
* refactor(types): accept only int in the proxy state setter
get_proxy_state_variable is annotated to return int, but
set_proxy_state_variable still took Any, so the checker could not hold
callers to the type the getter promises. The setter now takes int, which is
what its only caller already passes.
* refactor(types): index the proxy state key so the getter returns int
* refactor(types): keep public annotations and provider error text unchanged
Restore every public return, public method parameter, public attribute and exported
alias to its annotation on main so code that type-checks against the package keeps
type-checking, and take the Mistral speech and fal.ai Bria changes back out so no
provider error message differs from main
* refactor(types): leave the Vertex RAG chunking read as it is on main
Take the chunking format validation back out of the Vertex RAG ingestion path. It needs the vertexai SDK, a storage bucket and a RAG corpus to execute, so nothing here could run it end to end, and it cleared only two errors
* test(integration): pin the validated provider boundaries on a live proxy
* test(integration): give the held burst a client that outlasts the gate
The fault cell holds a burst at the upstream for up to 60 seconds while it kills a worker, but sent the burst through the shared 15 second client, so a slow box could time the survivors out before the gate opened. The burst now goes through its own client whose timeout is twice the gate, and the gate length is one named constant.
* refactor(types): keep the license reply handling and experimental MCP signatures as they were
The license check validated the whole reply as a mapping, which changed the error text logged for a reply that is not an object. It now validates only the verify value, so every reply is handled and logged exactly as before while the value is still typed.
Three files under the experimental MCP server changed annotations on public functions and methods (three returns and three parameters). They go back to their previous content so no public signature in the diff is narrowed.
* test(integration): answer the proxy's model-list call in the OpenAI stand-ins
Every 300 seconds each proxy worker asks an OpenAI deployment for GET /v1/models. Four new cells own an OpenAI stand-in that accepted only the call under test, so a refresh landing inside a cell failed it. The stand-ins now answer that call through the suite's own helper and the cells count only the provider calls they drive.
* chore(rust): prune unused inference crate dependencies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): rewrite inference layering docs for the split format crates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): note transcription and RouteError alias exceptions in inference AGENTS.md
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): rename litellm-core to litellm-inference
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): expose inference base API for format crates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): move shared inference test helpers behind test-support
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): wait for the classifier model popup before picking its option
Base UI exposes its select option asynchronously, and the helper waits for the option and its positioner to become clickable before selection
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: isolate two router tests from work leaked by earlier tests
Collect earlier tests' garbage before warning capture so their unawaited coroutines cannot be attributed to the target's warning assertion
Filter success-event callbacks by the request's litellm_call_id so queued logging work cannot replace the current request's captured messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): replace Any with proven types in 14 files
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): keep base parsing in sso userinfo, copilot auth and hf config lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): keep base parsing at unproven provider seams
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): keep base delete in jwt orphan cleanup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor: remove fresh tech debt from the 2026-10-05 window
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(lens): drop the review models left unused by the dead review helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): price a model group from the deployments that serve it
A model group's info read composed the group's own model_group_alias
entry into its deployments, a hop the router never takes: an alias is
resolved exactly once at request time, so a group reached as an alias
target is served by its own deployments. For the chain X -> T -> U the
price read for X included U's deployments too, and the free-model
budget waiver refused a free request to X on an over-budget key, while
GET /model_group/info reported U's providers and price for X.
The group info read now prices a group from the deployments routing
serves it with: the ones named after it, the routing group of that
name, or the wildcard route matching it when neither exists. The
budget waiver, GET /model_group/info, the rate limiters, and the
response headers all read the same set as routing. get_model_list
keeps its behavior for every other caller.
* test(router): give the paid fixtures explicit per-token prices
* test(router): call the routed-group read by name so the router coverage gate sees it
* test(integration): audit the alias chain budget waiver on every route, shape, and outage
Thirty-four cells under the management group prove an over-budget key is served through an alias chain entry at the price of the deployment that serves it, on chat, responses, and messages, sync and streamed, through the OpenAI and Anthropic SDKs and raw httpx on both replicas, with the chain middle, the reverse chain, a cost-map priced middle, a ghost middle, wildcard and routing-group targets, malformed and hostile inputs, a cached reply, a repointed alias, a provider failure, and two chaos bursts (a killed worker, a scripted outage)
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* perf(types): defer pydantic schema builds via shared LiteLLMBaseModel
Add LiteLLMBaseModel with defer_build driven by DEFER_PYDANTIC_BUILD (default true) and move litellm and enterprise pydantic models onto it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(types): build deferred models created by a parent validator; keep lens worker models litellm-free
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(types): link pydantic issue on deferred-build rebuild hook
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): polish Lens runs loading, reload, and time range menu
Port the dashboard-only parts of a0a275e486, 1325389624, 5e7b0afd5c, 03bec959bf, 93e6ccff8d, 3d35d9f920, dc1a2b5b60, 1d0397f1c0, 773bfef066, f3387221ea and 3d049cd4a8 from litellm_lens_server_search onto main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep the Lens timeline on the shown runs' window during a reload
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): drop a redundant fixture comment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): lead the Lens investigation detail with a run report
Port of the dashboard changes from a43f111abe and 29a1f3175d onto main. The detail view now opens with a run report for the selected run: status, a headline, progress or the failure, then cost, duration, coverage and issues, plus a collapsed activity log. An ordered situation table picks the report's one next action (Run now, Stop run, Retry, Raise budget, Connect worker, Review issues, Monitor this), and Run now moves into the investigation actions menu
Main's live review stays as is below the report, the queue reason still shows under queued progress, and partial results keep their own warning state with the run details folded away
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep Stop run on older Lens runs while another run is active
Also drop doc comments that restate the code
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vertex_ai): add regional endpoint uplift to gemini-3.1-flash-image
Google prices Gemini 3.1 Flash Image at 1.1x on non-global endpoints for
input, text output and image output, but the cost map row had no
regional_endpoint_uplift_multiplier, so regional calls billed at the
global rate.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(integration): cover regional uplift spend for gemini-3.1-flash-image
* test(integration): read the proxy salt from the environment in the uplift spend cells
---------
Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
The native response-cache runtime attached through Cache._native_cache is
unreachable since the V2 cache replaced it. Drop the Python branches and
wrapper, the _ResponseCacheRuntime pyclass and its backend/activation/
semantic modules, the python-bridge deps only they used, and the tests and
fixtures dedicated to that path.
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): honor caller stream flag when provider forces SSE
The Responses handlers decided whether to hand back a streaming iterator
from the provider payload's `stream` field, which chatgpt sets
unconditionally because the Codex backend only serves SSE. A caller that
sent `stream: false` therefore received a raw SSE stream on /v1/responses,
and the chat-completions bridge failed with "Unknown items in responses
API response: []" once its recovery path lost the raw SSE it reads from
Transport streaming still follows the provider payload; only the caller's
own `stream` value now decides the response shape. When the provider
forces SSE for a non-streaming caller the body is read and aggregated
through the existing path
* test(chatgpt): inject the authenticator into the responses config so handler tests never log in
* fix(responses): treat an extra_body stream flag as the caller's own and drop a redundant comment
* test(integration): audit the chatgpt caller stream flag across responses, chat and messages
---------
Co-authored-by: SeongWoon Cho <coffee@soylatte.kr>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(chatgpt,github_copilot): refuse device-code login when an event loop is running
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(github_copilot): drop stray whitespace change
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(chatgpt): bound token refresh timeout and drop placeholder assignment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(chatgpt,github_copilot): keep the token file path out of the event-loop 401 message
* fix(auth): refuse device-code login from worker threads too
/v1/messages runs its handler in an executor thread, where the
running-loop check never fires, so a chatgpt or github_copilot model
still started the interactive device-code login there and the request
hung for up to 15 minutes. The guard now also requires the main thread,
so the login only runs where a human can actually answer it.
* test(chatgpt): keep authenticator tests out of the real token directory
* fix(chatgpt): keep the 5 second connect timeout and the operator's request_timeout on the token refresh call
* test(integration): cover the device-code login guard on the proxy and the SDK
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(integration): bedrock_mantle-route basic translation cases on messages, chat completions and responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): rename translation runner run to assert_translation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): use assert_translation and drop the LIT-9196 skips in the bedrock_mantle basic cases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): add traceShareUrl helper for shareable trace links
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens): add copy link button to trace header
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens): cover copy link on trace header
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(proxy): cap batch file records, daily batch uploads, and per-file downloads
Adds three opt-in limits for batch jobs, each settable in general_settings as a
per-key default and overridable in key or team metadata by a proxy admin:
max_batch_file_records rejects a purpose=batch upload with more request lines
than allowed with a 413 before it reaches the provider.
max_batch_file_uploads_per_day counts accepted batch uploads per key and per
team in a UTC day and returns 429 with Retry-After once the count is used.
max_file_downloads_per_minute counts GET /v1/files/{id}/content per key and per
team for each file in a one-minute window and returns 429 with Retry-After.
The existing admin-only guard for batch_enqueued_token_limit now covers all four
metadata keys.
* fix(proxy): keep file usage counters in their own store and gate batch limits on user creation
* fix(proxy): count keyless JWT callers per user, require positive file caps, and take the upload slot after request validation
* refactor(proxy): end file usage cap describers with an explicit return after the match
* fix(proxy): keep file usage counters when more than 200 are live without Redis
The file usage counter store used a default in-memory cache, which holds 200
entries and evicts the one that expires soonest. Without Redis, a caller got a
fresh per-file download allowance after touching about 200 other file ids in
the same minute, and a key got a fresh daily upload allowance once about 200
other keys had uploaded that day. The store now tracks up to 20,000 live
counters per worker, the same bound the login throttle uses
* test(proxy): move the file usage cap tests into the directory the proxy shard runs
Main's shard coverage check found tests/unit/proxy/openai_files_endpoints
claimed by no shard, so its tests would not run in CI. The file moves next to
the other files endpoint tests in tests/unit/proxy/openai_files_endpoint, which
the proxy-endpoints shard already runs
* fix(proxy): declare the file usage counters as rate limit calls
Main's redis producer gate requires every module that writes a shared cache to name its key family, and the file usage counters wrote theirs without one.
* test(files): audit batch file usage caps across processes, Redis outages, and config reloads
* test(files): guard the chaos cells against minute boundaries and open Redis breakers
Two chaos cells each failed once in the audit run. The restart check ran
three sequential downloads with no guard against straddling a UTC minute,
and the exact-cap probe after a Redis outage ran while both workers' Redis
circuit breakers were still open (60 s default recovery), so it counted in
per-process memory and the two workers split the cap
Every burst now carries a window guard, the chaos fixture lowers the
breaker recovery to 2 s, and the post-outage check drives a fresh key to its
cap through a one-worker sibling proxy and then expects the two-worker
candidate to refuse the whole burst, which only the shared Redis count can
produce, polled until the breakers close
* test(files): release the held uploads when the killed-worker cell fails early
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(vertex_ai): apply regional endpoint uplift on image generation cost path
completion_cost had vertex_location but never passed it to the image
generation cost router, so a regional_endpoint_uplift_multiplier on a
Vertex image row would be ignored. No image row carries the multiplier
yet, so nothing is misbilled today. Pass the location through to the
Vertex image calculator for both the token-based price and the
per-image fallback.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(vertex_ai): cover regional image cost through the proxy logging path
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(images): hand image_edit's vertex_location to its cost resolver
* test(integration): audit Vertex image regional uplift billing
* test(integration): require every spend row after a proxy restart
* test(integration): reject extra spend rows after a proxy restart
---------
Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(integration): vertex_ai-route basic translation cases on messages, chat completions and responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): rename translation runner run to assert_translation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): use assert_translation in vertex_ai-route basic translation cases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): judge the free-model budget waiver by the group an alias routes to
A hidden model_group_alias can reuse the name of a real model_name. The
router serves that name from the alias target, but two of the three checks
behind the free-model budget waiver read the deployments of both the alias
target and the real model that shares the name.
So an over-budget key was served through an alias whose name belongs to a
model with an explicit $0 price when the target is $0 only by the cost map,
which the same key is refused on by its own name. The mirror case refused a
free target because the shadowed name belongs to a PTU-priced deployment.
Resolve the alias once and have the explicit-cost and PTU checks read the
routed group, the same group the price check already reads.
* test(proxy): cover the plain alias form of a shadowed free model name
The same wrong verdict exists for a plain string alias, so the unpriced-target case now runs for both alias shapes.
* fix(proxy): judge an alias chain's budget waiver by the deployments it is served from
The explicit-price and PTU checks read the alias target through
Router.get_model_list(), which follows a second alias hop when the target is
itself an alias key. The router never takes that hop, so an alias chain was
judged by a deployment the request never reaches. The checks now take the
deployments named after the routed group, or the wildcard deployment serving
it when none carries its name.
* fix(proxy): refuse the budget waiver when an alias chain is served by a priced wildcard route
* test(integration): cover the shadowing alias budget gate end to end
Adds the audit cells for a hidden alias whose name shadows an explicitly
free group: streamed SDK refusals on chat, responses and messages,
embeddings, the free wildcard and mixed-group paths, per-model budgets,
JWT and custom auth callers, cache hits, alias removal under traffic,
provider failure and fallback, a concurrent outage burst, and a worker
kill on an owned two-worker proxy
* test(integration): cite the cost-map rows the shadowing alias cells rely on
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(ci): namespace claude session ids in tracing seeds and allowlist /v1/logs on backend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): allowlist /v1/logs on the gateway alongside /v1/traces
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(gemini): stop replaying thinking block signatures to Gemini
A thinking block's signature has no provenance, and LiteLLM never fills it
from a Gemini response (Google signs text and functionCall parts, which ride
provider_specific_fields and the tool call id), so a Claude signature replayed
through a mixed model group reached Gemini as a thoughtSignature and Google
answered 400 Invalid thought signature on every later Gemini-served turn. The
same replay also sent the thinking text a second time as a plain text part.
The thinking text now goes out once, as the thought part built from
reasoning_content, and no part is built from thinking_blocks
* test(gemini): type the parts helper and split its comprehension
* test(gemini): cover thinking signature replay on the integration rig
Two integration files from the audit of the foreign thought signature fix: 62 wire cells asserting the model turn Google receives on chat, messages and responses across gemini and vertex_ai, streaming and not, SDK and httpx clients, the sad shapes of thinking_blocks, context caching through cachedContents, and 3 chaos cells (a concurrent burst across endpoints, upstream stream drops, a worker SIGKILL mid burst)
* test(gemini): read the integration salt from the environment
The wire test decrypted Responses ids with a literal salt; tests/integration/_support/process.py boots the proxy with LITELLM_SALT_KEY when it is set, so the test now reads the same variable with the same default
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(integration): bedrock_invoke-route basic translation cases on messages, chat completions and responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): rename translation runner run to assert_translation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): call assert_translation in the bedrock_invoke basic cases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>