A variable declared with both global and user scope is covered by the
global value (globals win in the merge), so the tool-call path must
resolve it from the global rather than raising a 412 when the user has
not filled in the per-user value. Restrict the missing-var check to
referenced user vars that lack a global fallback.
test_tokenizers downloads Xenova/llama-3-tokenizer from the HuggingFace
Hub via create_pretrained_tokenizer. On the CI runners the Hub keeps
returning 429 Too Many Requests, which propagated into the blanket
except and turned a third-party rate-limit into a hard pytest.fail. The
same test already skips its llama2 differentiation assertion when the
Hub is unreachable; this extends that exact handling to the custom
tokenizer download so a HuggingFace outage/rate-limit no longer fails
the suite while still failing on real assertion or logic errors.
The server-row delete is the commit point; a transient failure cleaning the
FK-less per-user env var rows now logs a warning instead of propagating, so a
successful delete is no longer turned into a caller error that triggers a retry
and a 404 for an already-gone server. Also mark the health-check env-var
round-trip test as asyncio so it runs explicitly like its siblings.
The frontend-lint baseline this branch had grown carried three
react-hooks/set-state-in-effect entries and four raw-fetch entries that were
added rather than fixed. Load the per-user env-var data in UserEnvVarsModal and
mcp_servers through React Query (useQuery/useMutation) so the setState-in-effect
findings go away instead of being baselined, and capture the ?fill_env_vars deep
link in lazy initial state so the modal target is derived during render rather
than set from an effect
Delete the unused clearMCPUserEnvVars wrapper, the one new networking fetch with
no caller. The remaining three wrappers (GET status, GET server vars, POST save)
still need a raw fetch because networking.tsx is the only HTTP layer and React
Query consumes it, so they stay grandfathered; the baseline only ratchets down:
UserEnvVarsModal 1->0, mcp_servers set-state 4->2, networking fetch 274->273
When prisma_client is None, _load_user_env_vars returned an empty dict,
which on the tool-call path was indistinguishable from "user has no
stored values" and produced a misleading 412 directing the user to set
up credentials they can never store without a database. Raise instead so
the tool-call path fails with a clear error and the listing path stays
best-effort via its existing catch.
The frontend-lint job (added on the base branch after this branch diverged)
runs prettier and eslint on the UI files a PR touches, measuring eslint errors
against the committed eslint-suppressions.json baseline. Pulling the base in
brings that gate, its config, and the baseline.
Format the touched MCP env var components and networking.tsx so they are
prettier-clean, and extend the suppressions baseline to cover the findings this
branch adds in files that already carry grandfathered entries: the four extra
raw fetch wrappers in networking.tsx (the API layer, where 270 raw fetches are
already grandfathered and there is no React Query alternative) and the
setState-in-effect findings in mcp_servers.tsx and UserEnvVarsModal.tsx, matching
the same rule already baselined in the sibling MCP components.
* fix(gemini-realtime): use GA event names for Pipecat 1.3.x compatibility
Pipecat v1.3.0 adopted the OpenAI Realtime API GA event naming:
response.audio.delta -> response.output_audio.delta
response.text.delta -> response.output_text.delta
response.audio.done -> response.output_audio.done
response.text.done -> response.output_text.done
The proxy was still emitting the old beta names; Pipecat's
`parse_server_event` raises "Unimplemented server event type" for any
unknown type, which killed the receive task handler and broke audio
playback and tool-call delivery.
Also:
- conversation.item.created -> conversation.item.added (already handled)
- client audio is buffered until backend setupComplete in deferred mode
- call_id fallback UUID when Gemini returns empty id
- status_details / token detail fields added to Pydantic-strict events
The _GA_TO_BETA_EVENT_TYPES map in RealTimeStreaming already translates
GA names back to beta for clients that opt in with the openai-beta
header, so legacy clients are unaffected.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(gemini-realtime): address greptile review comments
- emit outputTranscription as response.output_audio_transcript.delta
instead of suppressing it; GA_TO_BETA map handles translation for
legacy clients
- cap pre-setup audio buffer at 200 frames to prevent memory exhaustion;
log a warning when the limit is hit and additional frames are dropped
- log remaining dropped message count on flush error
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(gemini-realtime): address veria review comments
- remove unused OpenAIRealtimeConversationItemCreated import
- fix guardrail bypass: semantic_vad early-return now preserves
create_response when set so a guardrail-injected create_response:false
is not silently dropped
- add per-connection 10 MB byte cap alongside the 200-frame count cap
for the pre-setup audio buffer to prevent memory exhaustion
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(gemini-realtime): fix mypy arg-type on _finalize_gemini_live_setup
setup parameter typed as BidiGenerateContentSetup to match the TypedDict
passed at both call sites; was dict which mypy rejected.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(gemini-realtime): widen _finalize_gemini_live_setup to Dict[str, Any]
BidiGenerateContentSetup (TypedDict) is a subtype of Dict[str,Any] so
both call sites (one passing a plain dict, one passing the TypedDict)
satisfy mypy.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(gemini-realtime): cast BidiGenerateContentSetup to Dict at _finalize call site
mypy rejects TypedDict as dict[str, Any] argument; cast at the call site
where follow_up_setup is BidiGenerateContentSetup to satisfy the checker.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix Gemini realtime beta compatibility
* Fix deferred Gemini setup audio ordering
* fix: preserve Gemini audio transcript ids
* fix(realtime): cap pre-setup client buffer on all append paths
Route every append to the deferred-setup pending buffer through the
per-connection message/byte caps. Previously only the audio-buffer
fast path enforced the caps; once one frame was buffered, a client
that withheld session.update could stream arbitrary frames into
_pending_messages_until_setup unbounded and exhaust proxy memory.
* style(gemini-realtime): apply black formatting to transformation.py
* fix(gemini-realtime): log beta-translation fallback and name native-audio marker
Surface the previously swallowed exception in _send_event_to_client so a
failed GA->beta translation is observable instead of silently forwarding the
untranslated event. Extract the native-audio model substring used by
_finalize_gemini_live_setup into a named constant documenting why speechConfig
is dropped on those setups.
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(proxy): match passthrough registry routes bare-to-bare with SERVER_ROOT_PATH
After #28547, get_request_route strips the deployment prefix while registry
lookup still re-inflated stored paths via SERVER_ROOT_PATH, causing 404s
under paths like /llmproxy/ml. Compare normalized bare routes in both
is_registered_pass_through_route and get_registered_pass_through_route.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(proxy): patch utils.get_server_root_path in passthrough auth tests
After removing get_server_root_path from pass_through_endpoints, route
and JWT tests must mock litellm.proxy.utils where normalization reads it.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(gemini): keep googleSearch with server-side tools and googleMaps JSON schema
Wire include_server_side_tool_invocations through completion() so mixed
google_search and function tools are not dropped on Gemini 3+. Rewrite
generationConfig to responseFormat when googleMaps is used with JSON schema.
Fixes#27479Fixes#29451
Co-authored-by: Cursor <cursoragent@cursor.com>
* address greptile review feedback (greploop iteration 1)
* style: fix black formatting in main.py for py312 compat
* Fix Gemini Google Maps extra_body JSON rewrite
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* ci(ui): add frontend-lint job enforcing prettier and eslint on changed files
Lints only the files a PR adds or modifies under ui/litellm-dashboard,
so new and touched code must be prettier-clean and eslint-clean while the
existing tree is grandfathered. Skips cleanly when a PR touches no
lintable UI files. This lets us adopt the formatters incrementally
without a repo-wide reformat
* ci(ui): write frontend-lint file lists to $RUNNER_TEMP
Keep the prettier/eslint changed-file lists out of the checkout dir so
they cannot collide with a future source file of the same name
* lint(ui): baseline existing eslint findings so only new ones block
Capture the current error-level eslint findings (318 across 183 files)
in a committed suppressions baseline via eslint --suppress-all. Every
rule stays at its error severity, so any newly introduced violation
fails the frontend-lint gate, while the existing tree is grandfathered;
touching a legacy file never forces fixing its pre-existing issues. CI
runs eslint with --pass-on-unpruned-suppressions so that fixing a
baselined issue does not fail on a now-stale suppression, and the
generated baseline is prettier-ignored since eslint owns its format.
Burn the baseline down over time with eslint --prune-suppressions
* lint(ui): enforce a count budget for explicit any
Make @typescript-eslint/no-explicit-any a warning and cap the total
instead of hard-blocking each new one. A frontend-lint step counts the
repo-wide explicit any and fails only when it exceeds the committed
budget in eslint-any-budget.json. max starts at 2031, ten above the
current 2021, so the next ten land as warnings and the build fails once
that headroom is gone. Lower max over time toward target to ratchet the
count down. New anys still surface as warnings on changed files via the
normal eslint step
* lint(ui): enable zero-cost rules no-var, no-self-assign, react/no-danger
These have no existing violations, so they need no baseline; turning them
on purely blocks new instances. react/no-danger guards against new
dangerouslySetInnerHTML (XSS), no-var enforces let/const, and
no-self-assign catches self-assignment typos. no-debugger is already
enforced by the recommended preset
* lint(ui): add baselined complexity rules
Enable complexity:20, max-depth:4, max-params:4, max-nested-callbacks:4,
with thresholds set near the codebase p99 so only genuine outliers are
flagged. The 272 existing over-threshold functions are grandfathered in
the suppressions baseline; new over-threshold functions block. Lower the
thresholds over time to ratchet complexity down. max-lines-per-function
is intentionally left off since React components are legitimately long
* lint(ui): ban new raw fetch, standardize on React Query
Add a no-restricted-syntax rule flagging bare fetch() calls, pointing
contributors at React Query (@tanstack/react-query). The rule is not
exempted anywhere, including the already-bloated networking.tsx, so all
331 existing fetch calls are grandfathered but no new ones can be added
there or elsewhere. New data access goes through React Query, and the
networking layer can be migrated out and pruned from the baseline over
time
* lint(ui): ban new @tremor/react imports
Add a no-restricted-imports rule flagging imports from @tremor/react so
tremor is phased out rather than spread further. The 232 existing tremor
imports are grandfathered in the baseline; new ones block and point at
antd. Migrate components off tremor and prune the baseline over time
* lint(ui): widen explicit-any budget headroom to 2040
Raise max from 2031 to 2040, giving ~19 of slack over the current 2021
instead of 10
* style(ui): prettier-format eslint.config.mjs
The frontend-lint gate flagged its own config file. Format it so the
prettier check on this PR's changed files passes
* lint(ui): soften complexity and max-depth to warnings
These two are smell metrics with arbitrary thresholds where a legit new
function can trip them, so make them advisory rather than hard-blocking.
They drop out of the baseline (now 963). max-params, max-nested-callbacks,
and the react-hooks rules stay strict since those are clear-cut
* lint(ui): move complexity and max-depth to the count-budget pattern
Generalize the explicit-any budget into a shared lint-budget mechanism:
eslint-budgets.json maps a rule to {max, target} and check-lint-budgets.mjs
counts each across the repo and fails when a count exceeds its max.
complexity (129, max 140) and max-depth (61, max 70) now use the same
slack-plus-counter model as explicit-any (2021, max 2040): they warn
per-file and the build only fails if the repo-wide total crosses the
ceiling. Lower each max toward its target over time
* docs(ui): note pruning the eslint suppressions baseline when fixing lint debt
Management endpoints that create or update an MCP server return the
server with decrypted scope=global env var values so the admin edit
form can be pre-filled. management_endpoint_wrapper serializes that
response into an OTEL span and only filtered top-level credential
fields, so the nested env_vars list reached the trace verbatim and an
observability user could read upstream API keys.
Blank env_vars[].value in the telemetry response while keeping names
and scopes; the endpoint's own return value is untouched so the admin
still receives the decrypted values.
Master-key rotation re-encrypted only the credentials column and the
litellm_mcpusercredentials table, leaving the new global env_vars values
and the litellm_mcpuserenvvars values_b64 column encrypted under the old
key. After a rotation those values fail to decrypt, so global ${VAR}
headers are forwarded as empty substitutions and every per-user value
reads back as missing (412). Re-encrypt both new columns alongside the
existing ones, skipping undecryptable entries so a corrupt row is
preserved rather than overwritten.
Read env_vars from the exclude_unset-filtered data_dict (like every other
JSON column) so a partial update that omits env_vars can never overwrite the
stored values. Make the reload path's env-var encryption state explicit by
passing env_vars_are_encrypted=True, since raw DB rows are still encrypted
there unlike the already-decrypted records add_server/update_server receive.
The db.py read/write helpers (get_mcp_server, create_mcp_server,
update_mcp_server) decrypt global env var values in place before returning,
while leaving credentials encrypted. add_server/update_server then passed
those records to build_mcp_server_from_table with the default
credentials_are_encrypted=True, which decrypted the global env var values a
second time. Decrypting an already-plaintext value (e.g. "postgresql")
fails and zeroes it, so the registry entry forwarded the raw ${NAME}
placeholder upstream instead of the interpolated secret, and reload's
updated_at-equality reuse kept the broken entry.
Add an env_vars_are_encrypted flag to build_mcp_server_from_table (defaulting
to credentials_are_encrypted) and have add_server/update_server pass
env_vars_are_encrypted=False so global env var values are decrypted exactly
once.
Streaming responses from the proxy (/chat/completions, /v1/messages,
/v1/responses, assistants) all return through create_response() but never
sent the headers that tell an intermediary reverse proxy not to buffer the
SSE stream. nginx with the default proxy_buffering, k8s ingress-nginx, and
Envoy/Istio sidecars therefore hold the whole stream and release it in one
batch, which looks like a broken/buffered stream to the client even though
litellm is yielding chunks incrementally.
Add Cache-Control: no-cache and X-Accel-Buffering: no to every
StreamingResponse create_response() returns, matching what the proxy already
does for its own usage/policy SSE endpoints. Fixes#28384.
Global env var values are always stored encrypted, so a value that no longer
decrypts (typically a rotated LITELLM_SALT_KEY) was being forwarded into
upstream ${NAME} headers as ciphertext with only a debug log. Drop the value
and log a warning so the failure surfaces instead of silently sending ciphertext.
Per-user env var stores now merge over the existing values instead of replacing
them. Per-user credentials are write-only and never shown back, so requiring the
full set on every save forced users to re-enter credentials they could not see
just to change one field. Omitting (or sending empty) a field now keeps its
stored value; DELETE still clears everything. The modal no longer marks
already-set fields as required.
The (user_id, server_id) unique index cannot serve the
delete_many(where={server_id}) orphan cleanup in delete_mcp_server,
since server_id is not the leading column, so it falls back to a full
table scan. Add a server_id index to cover that delete path
The per-user env var cache is process-local, so in multi-worker
deployments a user who stored values on another worker could hit a stale
cached negative and receive a misleading 412 'missing credentials'
response until the entry expired. Before raising MCPMissingUserEnvVarsError
the resolver now re-reads straight from the DB with force_refresh, so a
process-local stale entry can never mask values stored elsewhere; the hot
success path still serves from cache.
Also drops the inaccurate claim that env vars are interpolated into the
server URL from the schema and type docs; only static_headers are
interpolated.
The health check and initialize-instructions prefetch copied
server.static_headers verbatim, so any auth header backed by a
${NAME} global env var was sent upstream as the literal placeholder.
Servers that migrated to the new ${NAME} convention authenticated fine
on real tool calls but flipped to 'unhealthy' in the dashboard. Both
probes now resolve static headers through
_resolve_static_headers_with_env_vars (no user context, best-effort)
so global values are substituted before the connection opens.
Global-scope MCP env vars hold admin-supplied secrets (API keys, passwords) that interpolate into static headers, but their raw value was serialized into the env_vars JSON column in plaintext, so anyone with read access to the database could recover those upstream credentials. Credentials and the per-user values_b64 column are already encrypted; global env var values now match that, encrypted on write in _prepare_mcp_server_data and decrypted when the server is built into the runtime registry and when records are read back for admin views. Per-user placeholder values are not secrets and stay verbatim.
A stored per-user value for a var that the admin later switched to global
scope was still able to override the admin's global value, because the
header merge applied the full user blob over the globals. Filter the user
blob to vars that are currently user-scoped and let admin globals win.
MCPServer.env_vars is stored as List[Dict[str, Any]] (deserialized from
the JSON column), but LiteLLM_MCPServerTable.env_vars is typed as
List[MCPEnvVar]. Passing the raw dicts relied on pydantic coercion at
runtime and tripped mypy's dataclass_transform __init__ check, failing
the lint job. Normalize the dicts into MCPEnvVar models at both
construction sites.
These values interpolate into Static Headers and Auth via ${NAME};
they are not exported into the MCP server's process environment, so
'Environment Variables' overpromised. The backend env_vars field is
unchanged.
#29612 exempts UI/CLI session tokens from the key budget ceiling when they
create a team key, keyed on data.team_id. That value is read after the
default_key_generate_params loop can populate team_id, so on deployments that
set default_key_generate_params.team_id a request the caller did not scope to a
team is treated as a team key and skips the ceiling. Capture _requested_team_id
before defaults run and key the exemption off it, mirroring how
_requested_max_budget is already captured. Requests the caller did not scope to a
team keep the ceiling.
The per-user value column rendered as a plain input, so it read like a static
value shared by every user, defeating the purpose of per-user variables. Add a
persistent "Hint" addon with an explanatory tooltip, lighten the typed text,
and align the row to the top so the addon group no longer sits lower than its
neighbours and the layout stays put when the name field shows a validation
error.
The GET /v1/mcp/server list and health endpoints build LiteLLM_MCPServerTable
from the in-memory registry via _build_mcp_server_table and health_check_server.
Both copied static_headers but dropped env_vars, so the list always returned
env_vars: null. The admin edit form is populated from that list data, so it
loaded an empty env-var list; saving any edit then persisted env_vars: [],
silently wiping the stored variables. With nothing left to interpolate, the
${VAR} static headers were forwarded upstream verbatim as literal text.
Carry env_vars through both conversions, mirroring static_headers. Add
regression tests asserting both paths round-trip name, scope, value, and
description.
ESLint 9 defaults to flat config and eslint-config-next was pinned at 15
while Next is on 16, so eslint only ran with ESLINT_USE_FLAT_CONFIG=false
and next lint is gone on Next 16. Replace .eslintrc.json with a native
flat eslint.config.mjs (config-next 16 ships flat configs, so no
FlatCompat shim is needed), bump eslint-config-next to 16.2.6, add
@eslint/js and typescript-eslint as explicit devDeps for the recommended
rule sets, and point the lint script at eslint directly.
This only makes eslint runnable on modern tooling; it does not wire it
into CI. The same rules carry over (next/core-web-vitals, eslint and
typescript-eslint recommended, prettier, unused-imports)
Allow realtime event transcript fields to be nullable so GA conversation.item payloads with transcript=null don't fail logging normalization and suppress success callbacks.
Co-authored-by: Cursor <cursoragent@cursor.com>
add_mcp_server wrote the new row and then reloaded the entire registry from the
database inside one try block. A single pre-existing malformed row made the
reload raise, so the endpoint returned 500 even though the new server was already
persisted; callers assumed failure and retried, creating duplicate servers.
Split the flow so the database write is the commit point and still 500s on
failure, while the in-memory registry refresh is best-effort and only logged on
error. Add regression tests for both the refresh-fails-after-commit path and the
db-write-fails path
Non-admin users creating a team key through the UI were rejected with
"max_budget cannot exceed the caller's own max_budget (0.25)". The request is
authenticated by a UI/CLI session token whose max_budget is the per-session chat
spend cap (max_ui_session_budget, default $0.25), and the delegated-authority
budget ceiling (GHSA-q775-qw9r-2r4g) treated that cap as a delegation limit.
Skip the ceiling only when a session token creates a team key (data.team_id set);
that key's spend is bounded by the team budget at request time. Personal keys and
every other non-admin caller keep the ceiling, so a session token cannot mint an
arbitrary-budget personal key.
A stale admin secret left in form state after switching an env var's
scope from instance to per-user was forwarded to the backend as the
user-scope value, which is returned unredacted to authorized non-admin
users. Per-user entries carry no admin value, so drop it on submit.
* Fix remaining VCR live-call leaks
* test(vcr): dedupe live-test helpers and drop spurious kwargs
Extract the duplicated isVertexQuotaError/runVertexRequestOrSkip Vertex
quota-skip helpers into tests/pass_through_tests/vertex_test_helpers.js and the
duplicated _skip_live_prompt_caching_test guard into tests/_live_test_helpers.py
so each lives in one place. In test_aarun_thread_litellm, build a separate
message_data carrying role/content for add_message and a thread_data without
them for run_thread/run_thread_stream/get_messages, which no longer receive the
spurious message fields.
* test(overhead): assert mock transport is exercised in non-streaming and stream tests
The single-server and bulk per-user env var status endpoints echoed the
decrypted credential value back to any holder of the user's LiteLLM token,
so a leaked token could exfiltrate the raw upstream secret (e.g. a personal
access token) for use outside the proxy. Drop the value field from
MCPUserEnvVarSpec and the include_values plumbing so the status reports only
whether each credential is_set; users overwrite a field to rotate it. The
fill-in modal no longer pre-populates from the secret and flags already-set
fields instead.
* fix(ci): keep coverage rename green when a parallel node runs no tests
local_testing_part1 and local_testing_part2 run with parallelism 4. When
CircleCI reruns only the failed tests, the failed test lands on a single
node and the other nodes receive an empty bucket, so pytest never writes
coverage.xml or .coverage. The unguarded "mv coverage.xml ..." then exits
1 and turns the whole job red even though the rerun passed; the next
persist_to_workspace step would fail the same way on the missing paths.
Guard the rename so a node with no coverage emits empty placeholders
instead. coverage combine tolerates the empty files, so the downstream
upload-coverage job keeps the real nodes' data intact.
* fix(ci): pre-create test-results in litellm_router_testing for empty-bucket reruns
litellm_router_testing also runs with parallelism 4. On a rerun of only the
failed tests, a node can receive no tests, so the test command never creates
test-results and the final store_test_results step can fail on the missing
path. Pre-create the directory up front, matching what local_testing_part1
and part2 already do and CircleCI's own guidance for parallel reruns.
* test(openai): retry wildcard chat completion on transient OpenAI 500
build_and_test reddened on test_openai_wildcard_chat_completion when the
real gpt-3.5-turbo-0125 call returned an OpenAI 500 ("The server had an
error while processing your request"). The base branch passed the same
call concurrently, so the 500 is an intermittent OpenAI server error, not
a regression. Add the same pytest-retry marker the sibling real-call tests
in this file already use so a transient upstream 500 no longer fails CI.
getMCPUserEnvVars now throws on non-2xx so UserEnvVarsModal reports the
error instead of silently rendering the empty 'no per-user fields' state.
The per-user env-var endpoints now run the access check before the server
lookup so a non-admin cannot tell a missing server (404) apart from one
they lack access to (403), closing a server-id enumeration leak.
PROXY_ADMIN_VIEW_ONLY callers are treated as admin_view by GET /v1/mcp/server
and GET /v1/mcp/server/{id}, so they received unredacted scope=global env var
values, which can carry upstream API keys used in Authorization headers. Only
a full PROXY_ADMIN needs those values (to pre-fill the edit form); read-only
admins now get the same global-secret redaction already applied to non-admin
and restricted virtual-key views.
The single-server DB fetch returned the raw Prisma model, whose JSONB
env_vars deserialize to plain dicts. The non-admin and virtual-key
sanitizers read env_var.scope as an attribute, so GET /v1/mcp/server/{id}
for a server with env_vars raised AttributeError and returned a 500.
Wrap the result like every bulk fetch helper so env_vars become MCPEnvVar
objects before redaction.
* fix duplicate cost callbacks for anthropic streaming pass-through
Two bugs caused _PROXY_track_cost_callback to see stream=True +
complete_streaming_response=None on every streaming pass-through request,
making the dedup guard in dispatch_success_handlers permanently inactive:
1. pass_through_endpoints.py created the Logging object with stream=False
for all requests. _is_assembled_stream_success short-circuits on
self.stream is not True, so has_dispatched_final_stream_success was
never set and any second dispatch went through unchecked.
Fix: set logging_obj.stream = True after stream detection.
2. _create_anthropic_response_logging_payload set complete_streaming_response
inside the try block after litellm.completion_cost(), so a pricing error
caused an early return without setting it on model_call_details.
Fix: set complete_streaming_response before the try block.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix stream
* add stream to logging obj
* test(pass_through): give mock logging object a real model_call_details dict
The anthropic passthrough logging payload now records the assembled
response on model_call_details before cost calculation, which requires
model_call_details to support item assignment. In production it is always
a dict; the existing unit test stubbed the logging object with a bare Mock
whose attribute is not subscriptable, so the new assignment raised
TypeError. Use a real dict to match the production logging object.
* test(pass_through): cover streaming logging-obj stream flag
The streaming branch of pass_through_request that marks the logging object
as streaming (logging_obj.stream and model_call_details["stream"]) had no
unit coverage, so the patch coverage gate flagged it. Add a regression test
that drives a streaming pass-through request through pass_through_request and
asserts the logging object is flagged as a stream before dispatch.
* test(pass_through): cover SSE-response stream flag fallback branch
The auto-detected streaming branch of pass_through_request (when a request
that was not flagged as streaming returns a text/event-stream response) sets
logging_obj.stream and model_call_details["stream"] but had no unit coverage,
so the codecov patch gate failed at 60%. Drive a non-streaming pass-through
request whose upstream response is SSE through pass_through_request and assert
the logging object is flagged as a stream before dispatch.
* fix(pass_through): gate complete_streaming_response on stream flag
perform_redaction only scrubs complete_streaming_response when
model_call_details["stream"] is True. Setting it unconditionally for
non-streaming Anthropic pass-through responses left the assembled
response unredacted in model_call_details, which is handed to logging
callbacks as kwargs when message logging is disabled. Only record it for
actual streaming responses so redaction always applies.
---------
Co-authored-by: mubashir1osmani <mubashir.osmani777@gmail.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>