Commit graph

48019 commits

Author SHA1 Message Date
mateo-berri
fe242f8702 test(databricks): pin the streamed reasoning delta shape and alias 2026-09-09 13:16:57 -07:00
Mateo Wang
f529d6d6bd
Merge pull request #40444 from BerriAI/litellm_annotate_strict_budget_helpers
chore(lint): bring ANN202 and BLE001 back under the strict-rule budget
2026-09-09 13:11:05 -07:00
mateo-berri
f3a2844080 test(convert_dict_to_response): keep the regression test locals final and comment-free 2026-09-09 13:05:11 -07:00
kerry-berri
b23995ee29
Merge pull request #40372 from BerriAI/litellm_cli_skip_cost_map_fetch
fix(cli): skip remote model cost map fetch in lite CLI processes
2026-09-09 13:02:53 -07:00
mateo-berri
17ca562b6a test(llm_translation): expect an empty choices list to convert instead of raising 2026-09-09 13:01:11 -07:00
mateo
420282acb7 fix(registry): set max_output_tokens on vertex_ai/xai/grok-4.3 and grok-4.6
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 19:55:02 +00:00
jibanez-staticduo
fe9a4ffca6
fix(chatgpt): pin sideband budgets and preserve OAuth image identity 2026-09-09 21:48:44 +02:00
mateo-berri
261777d633 fix(databricks): keep top-level reasoning_content from OpenAI-compatible gateway models
The Databricks chat transformation only parsed reasoning out of FMAPI-style
reasoning content blocks, so external models behind Databricks AI Gateway that
return the OpenAI-style top-level reasoning_content string lost it, both in the
final message and in every streamed delta. Fall back to the shared OpenAI
reasoning helper when no reasoning block exists, and keep the delta's own
reasoning_content when streaming.
2026-09-09 12:40:29 -07:00
mateo-berri
35def27e7b fix(convert_dict_to_response): accept only a real list as choices and keep /v1/messages alive on an empty one
Narrows the no-choices guard so a dict, string, or None still raises the APIError while an empty list passes through,
guards the non-stream Anthropic bridge against indexing an empty choices list, and repairs test_completion_missing_role,
whose raw-response mock was patched in as the create() callable itself so the handler only ever saw a MagicMock
2026-09-09 12:37:29 -07:00
jibanez-staticduo
09f52d9aa4
fix(ci): update websocket fixtures and vulnerable development parser 2026-09-09 21:34:41 +02:00
mateo-berri
8bb6d8c120 chore(lint): bring ANN202 and BLE001 back under the strict-rule budget
Annotate the return types of dispatch_async and transform_then_dispatch in llm_http_handler and _send_batch in azure_sentinel, and mark four legitimate broad catches with the repo's noqa convention, so the promote PR's lint job passes the strict gate again. Supersedes #40328.
2026-09-09 12:34:08 -07:00
kerry
34d2d010c3 test(cli): drop lite e2e tests, the e2e runner does not install the package
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 19:33:59 +00:00
mateo-berri
c0d1fd45f3 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_mat228_pr40294 2026-09-09 12:31:55 -07:00
jibanez-staticduo
ef0a67c754
merge: integrate current upstream PR base 2026-09-09 21:21:32 +02:00
mateo
00e4380cb5 fix(registry): add Bedrock gpt-6-astra CRIS + mantle profiles, embed-v4/pegasus global profiles, gpt-image-2.5 entries; cap Vertex grok-4.1-fast output at 128k
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 19:15:44 +00:00
mateo
23ed208339 Merge litellm_internal_staging into rolling registry PR 2026-09-09 19:02:09 +00:00
jibanez-staticduo
511599d7b0
fix(chatgpt): preserve explicit gateway during provider resolution 2026-09-09 20:45:55 +02:00
devin-ai-integration[bot]
096984bfc2
fix(proxy): pin multi-root CA bundle to the server's root before handing it to Prisma (#40428)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 11:19:29 -07:00
ryan-crabbe-berri
f89e9ac749
Merge pull request #40425 from CaptainAni187/fix_aws_tests_ambient_ssl_cert_file
test: isolate bedrock aws tests from ambient SSL env vars
2026-09-09 11:06:14 -07:00
tin-berri
c82c9cbced
fix(router): strip encrypted reasoning on an auto-router tier change instead of a 503 (#40280)
A Responses API follow-up that replays reasoning.encrypted_content is pinned to the
deployment that minted it. Behind an auto-router the pre-routing hook rebinds the model
to the tier it picked before the candidate pool is built, so a turn that classifies into a
different tier never finds the origin and the affinity check raised its fail-fast 503,
whose text claims a cooldown that does not exist

When the deployment that minted the reasoning is not a member of the model group this turn
is routed to, strip the encrypted reasoning (keeping any readable summary, string or block
form) and dispatch to the routed group. Membership is tested by deployment id against the
candidate set the router itself resolved for the route (routing group, model_name, team,
and pattern alike), not by model-group name, so an alias, a provider-qualified spelling, a
team-public name, or a pattern route of the same group is not misread as a tier change.
An unknown origin (a removed deployment, or a forged/unauthenticated marker) is handled the
same as a cross-group one and its reasoning is stripped, so a real cross-group id and a
nonexistent id return the same response and cannot be used to enumerate deployment ids.
Unavailability within the origin's own group keeps the existing 429/503 fail-fast, so the
cooldown contract is unchanged

Resolves LIT-7195


Claude-Session: https://claude.ai/code/session_01KAumQbhzk6jdWWHFLA8Jar

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-09 11:04:11 -07:00
Animesh Kumar
6195fbf5f3 test: isolate bedrock aws tests from ambient SSL env vars
Nine cases assert the sts client is built with verify=True, but get_ssl_verify
reads SSL_CERT_FILE and SSL_VERIFY, so the argument depended on the ambient
environment. The published images set SSL_CERT_FILE, so the suite failed there
while passing in CI.

Fixes #40357
2026-09-09 23:24:12 +05:30
devin-ai-integration[bot]
c7163a80dd
perf(proxy): collapse per-worker SGR upserts into one statement per flush (#40362)
Each proxy worker flushed one Prisma upsert per active (date, category, route)
bucket every interval, so the Postgres primary saw workers x routes statements
per interval across the deployment. A flush now builds a single multi-row
INSERT ... ON CONFLICT DO UPDATE, and with use_redis_transaction_buffer on the
workers push snapshots to a Redis list that one lease-holding pod folds and
commits, so the whole deployment costs one statement per interval. The leader
keeps popping until the list is empty so a deployment wider than the dequeue
cap cannot build a backlog, and rows that fail both the commit and the Redis
re-queue fall back to the leader's own accumulator instead of being lost.

Resolves LIT-7371

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 10:48:06 -07:00
Clement
699ae63b2a
feat(router): support percentile-based TTFT routing (#40352)
* feat(router): support percentile-based TTFT routing

* fix(router): apply routing_strategy_args updates to the live selector

Runtime routing_strategy_args updates (config reload, update_settings)
only rebuilt the strategy selector when routing_strategy itself changed,
so a newly added ttft_percentile sat unused until the proxy restarted.

Also drops a comment that only restated the code it sat above.

Claude-Session: https://claude.ai/code/session_01PmqjhFYcUh6vA72d8W9gdB

* refactor(router): drop unreachable empty-samples guard in percentile latency

_percentile_latency is only called behind use_ttft, which already requires
a non-empty ttft sample list, so the early return was dead code and the one
line Codecov flagged as uncovered on this patch.

Claude-Session: https://claude.ai/code/session_01PmqjhFYcUh6vA72d8W9gdB

* test(router): cover the no-selector path of a routing_strategy_args update

simple-shuffle has no selector attribute to re-link, so the early return
guards a setattr with a None attribute name. Dropping the guard makes the
new test fail with "attribute name must be string, not 'NoneType'".

Claude-Session: https://claude.ai/code/session_01PmqjhFYcUh6vA72d8W9gdB

* fix(test): assert ValidationError on out-of-range ttft_percentile

pytest.raises(ValueError) tripped PT011 for being too broad. Pydantic
raises ValidationError for the gt/le constraint, so naming it satisfies
the rule and pins the assertion to the constraint under test.

Claude-Session: https://claude.ai/code/session_01PmqjhFYcUh6vA72d8W9gdB

* fix(router): drop Final from a per-deployment loop variable

basedpyright rejects "A Final variable cannot be assigned within a loop",
which pushed reportGeneralTypeIssues one over its budget. selected_latency
is rebound each iteration, so it matches its unannotated neighbours in the
same loop.

Claude-Session: https://claude.ai/code/session_01PmqjhFYcUh6vA72d8W9gdB

* test(router): exempt _apply_updated_routing_strategy_args from the name scan

The scan only reads test files with "router" in the filename, so it cannot
see the update_settings tests in router_strategy/test_lowest_latency.py.
Calling the private helper directly would test structure rather than
behaviour, so it joins the existing entries ignored for the same reason.

Claude-Session: https://claude.ai/code/session_01PmqjhFYcUh6vA72d8W9gdB
2026-09-09 10:47:34 -07:00
jibanez-staticduo
57e831d445
fix(chatgpt): authenticate sideband models and forward image headers 2026-09-09 19:44:19 +02:00
devin-ai-integration[bot]
7c6e33ef70
fix(proxy): ignore team_id="" on /key/update so team-less keys can be updated and imported (#40421)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 10:38:26 -07:00
devin-ai-integration[bot]
996ee5635a
perf(proxy): pipeline spend counter increments into one Redis call per request (#40371)
* perf(proxy): pipeline spend counter increments into one redis call

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): apply surviving spend increments before raising scope error

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(proxy): ruff format spend counter helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): settle inner spend counter gathers and fall back per key on pipeline failure

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): suppress BLE001 on pipeline fallback catch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): invalidate all batched spend counters on pipeline failure

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 10:35:50 -07:00
joshua-berri
ea0851de79
Merge pull request #40359 from BerriAI/litellm_fix_mcp_connection_errors_31318
fix(mcp): surface connection failures across transports
2026-09-09 10:22:25 -07:00
ryan-crabbe-berri
ff2f122846
Merge pull request #40174 from BerriAI/litellm_cost_estimate_cache_tokens
feat(proxy): price cache and reasoning tokens in /cost/estimate
2026-09-09 10:18:59 -07:00
ryan-crabbe-berri
1763aeff62
Merge pull request #40303 from mubashir1osmani/litellm_fix_delete_passthrough_ui
fix(ui): repair pass-through delete confirm dialog and disable delete for config endpoints
2026-09-09 10:15:38 -07:00
kerry
6e71b90a88 test(e2e): dedupe lite env setup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 17:07:34 +00:00
ryan-crabbe-berri
53d23b90ee Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_cost_estimate_cache_tokens
# Conflicts:
#	litellm/litellm_core_utils/litellm_logging.py
2026-09-09 10:05:48 -07:00
kerry
2ece873538 test(e2e): lite CLI never fetches the model cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 17:05:38 +00:00
kerry
3a54e5bcb9 Merge origin/litellm_internal_staging into litellm_cli_skip_cost_map_fetch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 17:00:37 +00:00
kerry-berri
bfcc6404d3
Merge pull request #40350 from BerriAI/litellm_cost_map_background_retries
fix(cost-map): keep first fetch blocking, run retries in background
2026-09-09 09:59:07 -07:00
ryan-crabbe-berri
36f3ca95d8 fix(proxy): report /cost/estimate rates from the call that billed them
The estimate looked the reported per-token rates up a second time, with the
provider this endpoint resolved rather than the one completion_cost infers.
The provider decides whether a token tier threshold is inclusive, so an
unrouted xai model sitting exactly on 200k billed at the tier rate and
reported the base rate, half of it.

completion_cost now hands back the rates its own lines were billed at, and
the endpoint reports those.

Claude-Session: https://claude.ai/code/session_01RLKy5DMi3XCBUJ37WzfNi1
2026-09-09 09:57:38 -07:00
yujonglee
1183b2abc6
fix(integrations): pass original request object to post-call guardrail hooks (#40414) 2026-09-09 09:46:37 -07:00
Joshua Valluru
dbf9490229 fix(mcp): expose shared SDK timeout normalization 2026-09-09 09:44:14 -07:00
mateo
6dbaf43ba2 fix(registry): drop OpenAI shutdown date from shared computer-use-preview entry
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 16:42:07 +00:00
Mateo Wang
9b6c7c8bf0
Apply suggestion from @greptile-apps[bot]
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-09-09 09:17:14 -07:00
ryan-crabbe-berri
e8140eb269
Merge pull request #39985 from BerriAI/litellm_lit_5858_jwt_team_grants
fix(proxy): apply team model aliases on the JWT auth path
2026-09-09 09:09:01 -07:00
kerry
1fc1aaabda test(cli): drop lite --version subprocess regression
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 16:00:35 +00:00
ryan-crabbe-berri
360fa65631 Merge branch 'litellm_internal_staging' into litellm_lit_5858_jwt_team_grants
Claude-Session: https://claude.ai/code/session_01Hn5E8Jz1LjGLFyiYxBRcBW
2026-09-09 08:58:50 -07:00
Joshua Valluru
a1588c2602 fix(mcp): preserve stream failures across transports 2026-09-09 08:28:18 -07:00
Joshua Valluru
0225a16f48 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_mcp_connection_errors_31318 2026-09-09 08:00:01 -07:00
mateo
2881b8cd45 fix(cost): carry output_cost_per_second_720p through model info
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 14:09:25 +00:00
mateo
b067729082 fix(registry): add xAI Imagine Video 720p per second rates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 13:51:40 +00:00
mateo
9f21ae395a fix(registry): correct eu Claude 3.5 Haiku Bedrock pricing, add Nova v1 tool_choice, Azure gpt-5.5 snapshot retirement
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 13:16:04 +00:00
mateo
f058d6a966 merge litellm_internal_staging 2026-09-09 13:05:02 +00:00
jibanez-staticduo
6e87b4b985
fix(chatgpt): resolve gateway URLs without touching token storage 2026-09-09 13:32:33 +02:00
jibanez-staticduo
96594e7b0e
fix(chatgpt): honor configured gateways across Codex transports 2026-09-09 13:30:28 +02:00