Commit graph

49384 commits

Author SHA1 Message Date
mateo-berri
836c20a4b6 test(ai-gateway): drop provider fingerprint and cache identity assertions 2026-09-04 22:31:06 -07:00
mateo-berri
cf275cf442 chore: merge origin/litellm_internal_staging into litellm_lit_7022_azure_ai_passthrough_config 2026-09-04 22:26:05 -07:00
mateo-berri
d98cc2993c chore: merge litellm_internal_staging into litellm_lit_6348_fireworks_responses_api 2026-09-04 22:23:21 -07:00
Mateo Wang
aea5358c48
Merge pull request #39843 from BerriAI/litellm_lit_7015_migration_job_node_selector
feat(helm): render nodeSelector, tolerations, and affinity on the componentized chart migrations Job
2026-09-04 22:19:50 -07:00
mateo-berri
e9a40ad4d2 fix(fireworks_ai): map developer items after pydantic input items are dumped 2026-09-04 22:16:14 -07:00
mateo-berri
4c9537febe chore: merge origin/litellm_internal_staging into litellm_lit_7022_azure_ai_passthrough_config 2026-09-04 22:14:53 -07:00
tin-berri
78ad88f52c
fix(responses): decode JSON-string tool schemas before sending to the provider (#39844)
* fix(responses): decode JSON-string tool schemas before sending to the provider

A caller that hands a tool schema over already JSON-encoded reached the
Responses API with a string `parameters`, and the provider rejected the
request with a 400 naming the routed model instead of the offending tool.
Decode it at the one place every Responses request converges, and refuse
anything that is neither an object nor a string encoding one.

Collapses the duplicated input/tool sanitization block shared by the
request and compact-request builders into a single owner, so the decode
cannot be wired into one path and not the other.

* test(responses): pin null tool schemas as accepted, and type the parametrized cases

The Responses API serves `parameters: null` and an omitted schema alike, so
neither may raise. Pin both against a future tightening, annotate the
parametrized inputs, and trim the docstrings back to what the code does not
already say.
2026-09-05 05:14:44 +00:00
mateo-berri
d3a179f988 fix(azure_ai): route only cohere parse deployment names to Cohere Parse 2026-09-04 22:07:26 -07:00
mateo-berri
221d08d643 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_batch_ui_logs
# Conflicts:
#	litellm/batches/batch_utils.py
2026-09-04 21:51:52 -07:00
Mateo Wang
75736323e6
Merge pull request #39461 from BerriAI/litellm_decrease_anys_opus5_r4
refactor(typing): cut 1,397 Any errors across 183 backend files
2026-09-04 21:44:04 -07:00
mateo-berri
512e730f2f fix(azure_ai): add passthrough config so router-model relays reach the deployment's own endpoint
Every /azure_ai/<router model>/<native path> relay failed with HTTP 500 because
azure_ai had no passthrough config. The new AzureAIPassthroughConfig strips the
router-model prefix from the relayed path, forwards to the deployment's api_base
with its own credential (api-key on Foundry and Azure OpenAI hosts, Bearer
elsewhere, Entra as the fallback), and delegates chat/completions cost logging
to the Azure passthrough config.

The router's provider inference now receives the deployment's api_base so an
OpenAI-family model on a Foundry resource stays azure_ai instead of flipping to
azure through the AZURE_AI_API_BASE env var.
2026-09-04 21:30:48 -07:00
mateo-berri
525d9fb14c chore: merge litellm_internal_staging into litellm_lit_6348_fireworks_responses_api 2026-09-04 21:30:31 -07:00
mateo-berri
8426235290 feat(ocr): add Cohere Parse support for cohere and azure_ai 2026-09-04 21:25:46 -07:00
mateo-berri
412c36bb8e fix(realtime): detect an upstream refusal from received frames, not the session log
The refusal predicate also required the session log to be empty, but that
log is not limited to upstream frames. With gemini_live_defer_setup the
handler stores a synthetic session.created before the relay starts, and
the transcription usage flush appends a usage event before the check
runs, so an upstream policy close with no received frames was still
logged as a $0 success. Key the check off the received-frames flag only
2026-09-04 21:24:36 -07:00
mateo-berri
7bddb656c1 refactor(files): build the next listing page's headers in one handler helper
Staging sits exactly at the LIT002 ceiling, so the duplicated validate_environment call for the next page is shared to keep the merged tree under it
2026-09-04 21:21:52 -07:00
moe-berri
f03f82381e merge origin/litellm_internal_staging, keep the reportPrivateUsage suppression 2026-09-04 21:03:28 -07:00
mateo-berri
29dcd0cc2e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5_r4
# Conflicts:
#	basedpyright-code-budget.json
#	type-discipline-budget.json
2026-09-04 21:02:19 -07:00
mateo-berri
14f8677bfc fix(realtime): mark realtime sessions async so failure hooks fire once
The relay's failure dispatch runs the async handler and then the legacy sync
failure_handler for the proxy's callable callbacks. The realtime logging object
carried no async marker, so failure_handler treated the session as a sync SDK
call and fired every CustomLogger's sync failure hook on top of the async one:
Langfuse recorded two ERROR observations per refused session, and OpenTelemetry,
MLflow, Braintrust, Literal AI, DeepEval and New Relic implement the same sync
hook. Plant the _arealtime marker in litellm_params the way aanthropic_messages
and agenerate_content already do, so both dispatchers classify the session async.
2026-09-04 20:57:51 -07:00
Mateo Wang
9c05c158cb
Merge pull request #39518 from BerriAI/litellm_techdebt_20260903
refactor: clear fresh tech debt from the last 24 hours (2026-09-03, 2026-09-04)
2026-09-04 20:56:42 -07:00
mateo-berri
5c8016e99e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_async_remote_image_fetch 2026-09-04 20:54:41 -07:00
mateo-berri
f09992b730 fix(image_handling): keep DNS, signing, and Vertex Gemini fetches off the event loop
The SSRF check in async_safe_get resolved DNS on the event loop and a blocked
address was retried three times; validate_url now runs in a thread and an
SSRFError fails the fetch on the first attempt in both fetchers. The shared
HTTP handler signed the request and ran pre_call logging on the loop after an
async transform; both now run in a thread. Vertex AI Gemini still fetched
http:// images and https images without an inferrable mime type with the sync
converter inside its async body builder; the walker takes a should_inline
predicate and Vertex AI inlines exactly those URLs, leaving https images with a
known mime type and Files API refs to Google. When one download fails the
other in-flight downloads for that request are now cancelled instead of
finishing in the background
2026-09-04 20:54:40 -07:00
Mateo Wang
3d00ad3f29
Merge pull request #39847 from BerriAI/litellm_lit_5730_bedrock_batch_cancel_e2e
test(e2e/batches): assert Bedrock batch cancel and list in the lifecycle
2026-09-04 20:54:27 -07:00
moe-berri
955baf8a5c Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_shadow_eval_judge_output_cap 2026-09-04 20:54:26 -07:00
mateo-berri
787f2dee0c fix(health): match team public model names when targeting /health by model 2026-09-04 20:52:59 -07:00
tin-berri
8b6ea72845
feat(shadow_eval): scope a job to model groups, ANDed with its key, team, and user targets (#39828)
A shadow eval job could only be scoped by identity, so "this user's traffic on model X
across every key they own" was not expressible and a models field on the start body was
silently dropped. The job now carries a models list that every target is narrowed to,
matched on the requested model group with model_group_alias resolved on both sides. An
unresolvable name is a 400 at start. Empty means every model, which is what every existing
row reads as. The dashboard start form gains an "Only on models" picker and the job
headline shows the scope.
2026-09-04 20:50:46 -07:00
mateo-berri
6294367920 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_realtime_reasoning_double_bill 2026-09-04 20:49:12 -07:00
mateo-berri
93fa989238 fix(router): fall back to the deployment voice without a conditional dict spread 2026-09-04 20:42:41 -07:00
mateo-berri
92e449d4d0 test(cost): cover reasoning nested in text_tokens beside audio output 2026-09-04 20:35:21 -07:00
devin-ai-integration[bot]
e7dd524a3c
feat(otel): stamp litellm.request.route on the LLM call span (#39698)
* feat(otel): stamp litellm.request.route on the LLM call span

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(otel): drop redundant comment on REQUEST_ROUTE

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(otel): Final-annotate route test locals, drop field comment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): read litellm.request.route off the server span

The LLM call span took the auth-normalized literal path from logging
metadata, which disagrees with the SERVER span wherever FastAPI matched a
template: on /engines/{model:path}/chat/completions the LLM span spelled the
model name while http.route carried the template, so the two spans grouped
into different buckets and the PR's premise did not hold.

Read the value off the span that already holds it. The request's root SERVER
span is anchored per request for parenting, and its attributes stay readable
after it ends, so request_root_http_route() answers from the async close
callback with the same http.route the SERVER span exports: the route template
on a normal route, the literal path where the passthrough hook rewrote it, and
the mount point on an MCP call. Nothing has to re-derive any of that, so the
two spans cannot drift apart.

The route the proxy recorded at auth stays as the backstop for a deployment
whose FastAPI instrumentation never mounted, where there is no server span to
disagree with. Off the proxy the attribute is omitted rather than empty.

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yucheng He <yucheng@berri.ai>
2026-09-05 03:33:44 +00:00
mateo-berri
d5bd788314 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_health_scoped_results_and_serialization 2026-09-04 20:32:45 -07:00
Mateo Wang
377b87c59c
Merge pull request #39306 from BerriAI/litellm_deflake_20260902
test: deflake JWT tamper, fuzzy picker, tag routing, liveliness, redis stall burst, and pre-commit interrupt tests
2026-09-04 20:30:50 -07:00
Mateo Wang
8d4a63d2a1
Merge pull request #39435 from BerriAI/litellm_debug_claude_session_report
feat(cli): add `lite debug claude` session report and /debug-lite slash command
2026-09-04 20:30:47 -07:00
mateo-berri
217cb7da65 fix(managed_files): resolve the creator org through the cached team lookup
Batch creation snapshotted the team's organization with a direct
litellm_teamtable query on every create. Go through get_team_object
instead, which serves the team auth already cached and only falls back
to the database when the team was never cached.
2026-09-04 20:30:33 -07:00
Mateo Wang
b77b7f158f
Merge pull request #39842 from BerriAI/litellm_lit_6949_unknown_model_spend_log_test
test(store_model_in_db): assert the 400 contract in the unknown-model spend log test
2026-09-04 20:29:56 -07:00
moe-berri
6385c7b3c5 fix(auto-router compression): close three review findings on the per-hop policy
Suppression state moves out of request metadata into a request-scoped ContextVar.
refresh_proxy_server_request_body_snapshot copies metadata into
proxy_server_request.body, which deployments persist to spend logs, so the marker
naming each suppressed guardrail was readable by the caller whose request produced
it. Recovering it was enough to replay {token}:{name} for any CustomGuardrail and
switch off a PII or content-filter guardrail, since the check never verified the
named guardrail was a compression one. Nothing is read from metadata now, so there
is no marker to forge and the per-process token is no longer needed.

Routing-side compression reads the live messages instead of a pre-guardrail copy.
arm_pre_call runs before the pre-call hook, so its snapshot held the prompt as it
was before any masking guardrail rewrote it, and messages_for_routing handed that
to a compression guardrail which POSTs it to an external service. Masked content
left the proxy anyway. The cost is one combination: when the model hop compressed
and the hops differ, routing now classifies on the compressed text, since no
uncompressed copy survives that a masking guardrail has already seen.

policy_for_model no longer falls back to a marker scoped to tags the request does
not carry, which applied an 'eu' policy to a 'us' request on config order alone.

Each fix carries a regression test; all three fail when the fix is reverted.
2026-09-04 20:28:17 -07:00
ryan-crabbe-berri
f73e683800 chore(ui): drop the fetch lookup note from the fetchClient docblock 2026-09-04 20:26:13 -07:00
mateo-berri
c4d09a31e0 fix(vector_stores): default the search-count debug log to a tuple so LIT002 stays within budget 2026-09-04 20:20:07 -07:00
mateo-berri
036d104533 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_govcloud_profiles_lit6421 2026-09-04 20:08:16 -07:00
mateo-berri
ff856080c5 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_gateway_rustls_provider 2026-09-04 20:08:06 -07:00
mateo-berri
e7aeb9199a Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_vector_store_surface_retrieval_failure 2026-09-04 20:07:55 -07:00
Mateo Wang
4af62a38c1
Merge branch 'litellm_internal_staging' into litellm_mistral_voxtral_tts_speech 2026-09-04 20:07:45 -07:00
Mateo Wang
9b64f8367e
Merge branch 'litellm_internal_staging' into litellm_batch_ui_logs 2026-09-04 20:07:32 -07:00
mateo-berri
4b0fa87b9e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5_r4 2026-09-04 20:06:56 -07:00
mateo-berri
b6a3cba25c Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_deflake_20260902 2026-09-04 20:03:51 -07:00
mateo-berri
9cde3d21b0 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_debug_claude_session_report 2026-09-04 20:03:49 -07:00
mateo-berri
79ca00aaa7 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_techdebt_20260903 2026-09-04 20:03:47 -07:00
mateo-berri
0c14777069 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit_5730_bedrock_batch_cancel_e2e 2026-09-04 20:03:47 -07:00
Mateo Wang
df3b8a699d
Merge pull request #39848 from BerriAI/litellm_datadog_guardrail_cost_by_unit_audit_field
fix(datadog_llm_obs): keep guardrail_cost_by_unit on redacted spans
2026-09-04 20:02:33 -07:00
mateo-berri
cb291b423e test(e2e): cover the reliability retry, cooldown, fallback, and routing-strategy cells 2026-09-04 20:00:52 -07:00
mateo-berri
41c8969f0a test(router): drive tag routing tests through acompletion until both deployments are seen 2026-09-04 19:54:40 -07:00