Mateo Wang
c25c098bc1
Merge pull request #41469 from BerriAI/litellm_bedrock_mantle_responses_drop_top_p
...
fix(responses): drop top_p for gpt-5 reasoning models when drop_params is set
2026-09-17 18:04:19 -07:00
Mateo Wang
706f69af50
Merge pull request #41665 from BerriAI/litellm_remove_dead_vertex_v1beta1_stub
...
refactor(vertex_ai): remove constant-False is_using_v1beta1_features stub and its dead call sites
2026-09-17 17:56:22 -07:00
kerry-berri
f51f01fb54
Merge pull request #41443 from BerriAI/litellm_remove_brittle_price_pinning_tests
...
test: delete unit-test assertions that pin cost-map prices, limits and deprecation dates
2026-09-17 17:52:42 -07:00
Mateo Wang
2fb502a556
Merge pull request #41702 from BerriAI/litellm_bedrock_invoke_tool_search_opus48_gen5
...
fix(bedrock): gate Invoke tool search on the model map for Opus 4.8 and gen 5 Claude
2026-09-17 17:39:34 -07:00
kerry
b97c2d6307
test: drop pinned bedrock invoke cost literals
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 00:31:11 +00:00
kerry
84d4d17bf8
Merge remote-tracking branch 'origin/main' into litellm_remove_brittle_price_pinning_tests
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
# Conflicts:
# tests/test_litellm/llms/bedrock/chat/test_converse_transformation.py
2026-09-18 00:29:52 +00:00
kerry
c4620170ca
test: delete assertions that pin vendor cost map facts
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 00:28:49 +00:00
mateo-berri
9ae5bde829
fix(bedrock_mantle): resolve gpt-5 sampling rules from the OpenAI catalogue entry
2026-09-17 17:25:59 -07:00
Mateo Wang
02a20fe264
Merge pull request #41699 from BerriAI/litellm_fireworks_minimax_m3_supports_vision
...
fix(fireworks_ai): restore supports_vision on minimax-m3 in the cost map
2026-09-17 17:15:56 -07:00
yucheng-berri
e99902c4dd
Merge pull request #41569 from BerriAI/litellm_azure_ptu_spillover_cost
...
fix(cost): price Azure PTU spillover requests at standard token rates
2026-09-17 17:09:40 -07:00
mateo-berri
99b83a2d52
fix(fireworks_ai): restore supports_vision on minimax-m3 in the cost map
2026-09-17 16:39:31 -07:00
kerry
91619376d2
test: restore azure ai cached-token billing coverage with derived rates
...
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 22:49:02 +00:00
Mateo Wang
fec8231b83
Merge pull request #41419 from BerriAI/litellm_bedrock_openai_no_cachepoint
...
fix(bedrock): never emit Converse cachePoint for OpenAI-family models
2026-09-17 15:43:04 -07:00
kerry
eb2be3758a
test: read cost expectations from the catalog row the code bills against
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 22:22:24 +00:00
kerry
ccff1fa95f
test: derive the remaining cost-map pins from the catalog entry
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 21:49:16 +00:00
mateo-berri
36b471ff24
Merge commit '1dd4c13815' into litellm_bedrock_openai_no_cachepoint
2026-09-17 14:36:33 -07:00
Yassin Kortam
bc9f4fec5b
Merge pull request #41542 from BerriAI/litellm_bedrock_realtime_sdk_0_11
...
fix(bedrock): support aws-sdk-bedrock-runtime 0.10/0.11 in Bedrock Realtime
2026-09-17 14:26:39 -07:00
Yassin Kortam
1b4739c415
Merge pull request #41493 from BerriAI/litellm_bridge_mid_conversation_system_turns
...
fix(anthropic-bridge): convert mid-conversation system turns to user turns on /v1/messages to chat completions
2026-09-17 14:25:36 -07:00
yassin
6e84ff0cb2
fix(bedrock): keep raw SDK import failure out of the realtime client error
...
Log the underlying ImportError server side and send the client only the installed
version, the supported range and the install hint
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:57:28 +00:00
kerry
f5940bcee4
Merge remote-tracking branch 'origin/main' into litellm_remove_brittle_price_pinning_tests
2026-09-17 20:52:32 +00:00
yassin
a56390ed09
Merge remote-tracking branch 'origin/main' into litellm_bedrock_realtime_sdk_0_11
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
# Conflicts:
# uv.lock
2026-09-17 20:32:59 +00:00
mateo
6a76ca0c72
refactor(vertex_ai): remove constant-False is_using_v1beta1_features stub and its dead call sites
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 20:16:13 +00:00
yassin
177e6a0a97
test(anthropic-bridge): bound role reads instead of wall-clock time in the long system run test
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:34:16 +00:00
kerry-berri
0f5bc0ffa9
Merge pull request #41627 from BerriAI/litellm_fireworks_minimax_m3_vision_tests
...
test(fireworks_ai): stop pinning vision support on minimax-m3
2026-09-17 12:28:23 -07:00
yassin
2fea3f53b7
perf(anthropic-bridge): reorder mid-conversation system runs in a single pass
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 19:06:06 +00:00
joshua-berri
1e7b03a6ed
Merge pull request #41619 from BerriAI/litellm_fix_mcp_guardrail_context_4889
...
fix(mcp): preserve request-selected guardrails during tool execution
2026-09-17 19:04:04 +00:00
kerry
6d20e68706
test(fireworks_ai): stop pinning vision support on minimax-m3
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:57:50 +00:00
Yujong Lee
b26935416a
Merge remote-tracking branch 'github/main' into litellm_rust_bridge_declarative_route_catalog
...
# Conflicts:
# tests/e2e/access_control/test_model_access_group_e2e.py
2026-09-17 11:08:08 -07:00
yassin
7c6fd3090f
Merge remote-tracking branch 'origin/main' into litellm_bridge_mid_conversation_system_turns
2026-09-17 18:07:06 +00:00
yassin
99667ad633
fix(anthropic-bridge): keep mid-conversation system turns when the target declares supports_mid_conversation_system
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 18:06:55 +00:00
Joshua Valluru
743684bdbe
fix(mcp): preserve request-selected guardrails during tool execution
2026-09-17 09:57:08 -07:00
yucheng
b237c185db
test(guardrails): type the recording guardrail logging_obj as the logging object
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:12:27 +00:00
yucheng
2d925e5dde
fix(guardrails): scope the logging_only reply scan with the request's own translation
...
The chat-shaped output handler now takes the input translation as its
request scoping, so the logged request is scoped exactly once and with
the pre-call semantics of the surface it arrived on. This drops the
unscoped chat_shaped_request_conversation detour from af312dc8 , which
made the Anthropic response scan remove in-sequence system turns under
skip_system while the request scan kept them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:40:44 +00:00
yucheng
a57483d1c8
fix(cost): price Azure PTU spillover requests at standard token rates
...
Azure PTU deployments carry zeroed per-token pricing because the reservation
is billed flat by the hour. When Azure spills a request onto pay-as-you-go
capacity it returns x-ms-is-spilled-over: true, and that traffic was still
priced at zero. The response cost calculator now detects the spillover header
on the result's hidden params or the logged provider response headers and
skips the zeroed custom pricing only for genuine PTU deployments while the
feature flag is on. Azure sync streaming now also records response headers on
the logging object, matching the async paths.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:24:05 +00:00
yucheng
87263cefca
Merge remote-tracking branch 'origin/main' into litellm_post_call_guardrail_context
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
# Conflicts:
# tests/test_litellm/llms/openai/responses/test_openai_responses_guardrail_handler.py
2026-09-17 05:08:02 +00:00
tin-berri
d18e06f736
Merge pull request #41508 from BerriAI/litellm_1789600151_discover_context_limits
...
feat(router): discover token limits for hosted OpenAI-compatible models
2026-09-16 20:29:57 -07:00
yassin
bb9ff8cb2c
fix(bedrock): keep realtime SDK error range inside websocket close reason
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:57:33 +00:00
yassin
0259e8c7d5
fix(bedrock): support aws-sdk-bedrock-runtime 0.10 and 0.11 in the realtime handler
...
The bedrock-realtime extra pinned aws-sdk-bedrock-runtime 0.7.x, whose Config and BedrockRuntimeClient surface is gone in 0.11. The handler now resolves AsyncBedrockRuntimeConfig, builds AsyncBedrockRuntimeClient with the awscrt duplex transport, closes the client when the session ends, and tells an absent SDK apart from an installed but unsupported version. Moves the pin to >=0.10.0,<0.12.0 with the awscrt extra
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 01:40:52 +00:00
mateo-berri
47b2479c94
fix(bedrock): gate Invoke tool search on the model map's supports_tool_search flag
...
The Bedrock InvokeModel transformations decided whether to send the
tool-search-tool-2025-10-19 beta from hardcoded model name lists (a pattern
list on the messages path, an "opus-4" substring on the chat path), so Opus 4.8,
Opus 5 and Sonnet 5 never got the beta on the messages path, Opus 5 and Sonnet 5
never got it on the chat path, Opus 4.1 got it without support, and
/v1/model/info reported supports_tool_search as unset for all three.
Both paths now read the model map through one shared helper: the Bedrock
entries for Opus 4.8, Opus 5 and Sonnet 5 carry supports_tool_search
explicitly, and a claude-tool-search fallback rule flags Claude 4.5 and newer
for unmapped ids, inference-profile ARNs and mapped entries with no opinion,
so the next Claude gets the beta with no code change. An explicit false on a
resolved entry still wins.
2026-09-16 18:23:14 -07:00
kerry-berri
2e4840ee18
Merge pull request #41509 from BerriAI/litellm_mantle_gpt5_verbosity
...
fix(bedrock_mantle): accept and forward verbosity on gpt-5.x chat completions
2026-09-16 17:20:33 -07:00
Moe Khalil
d0ed8145b0
chore: merge latest main into model info discovery branch
2026-09-17 00:19:16 +00:00
Mateo Wang
d4a22acb66
Merge pull request #41513 from BerriAI/litellm_internal_copy_31400
...
fix(bedrock): neutralize orphaned tool blocks instead of raising or injecting a dummy tool (internal copy of #31400 )
2026-09-16 17:12:35 -07:00
Moe Khalil
39286245b5
chore: merge main into model info discovery branch
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 00:10:00 +00:00
kerry
0a47fe160d
Merge remote-tracking branch 'origin/main' into litellm_mantle_gpt5_verbosity
2026-09-17 00:06:50 +00:00
Mateo Wang
913ef6ed49
Merge pull request #33856 from BerriAI/litellm_azure_ai_responses_native
...
fix(azure_ai): route Responses API to native /openai/v1/responses for Foundry Models
2026-09-16 17:06:29 -07:00
Mateo Wang
319f427c40
Merge pull request #41201 from BerriAI/litellm_gemini_37_38_flash_no_minimal_thinking
...
fix(gemini): map minimal thinking to low for Gemini 3.7 and 3.8 Flash
2026-09-16 16:52:08 -07:00
mateo-berri
ada0a1ad3a
fix(azure_ai): strip the azure_ai/ prefix when a Responses call is remapped to azure
...
A catalog OpenAI name on an .openai.azure.com host (or with AZURE_AI_API_BASE set to one) is remapped from azure_ai to azure before the Responses request is built, and the azure_ai/ prefix stayed in the wire model, so Azure answered DeploymentNotFound. The Azure Responses config now strips azure_ai/ next to responses/ and o_series/.
2026-09-16 16:49:29 -07:00
Mateo Wang
390ab45448
Merge pull request #37506 from BerriAI/litellm_dashscope_reasoning_effort
...
LiteLLM Rust / rust-wheel (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
fix(dashscope): forward reasoning_effort to the provider
2026-09-16 16:29:25 -07:00
kerry
15c18ad6cd
fix(bedrock_mantle): accept verbosity on gpt-5.x chat completions
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:27:07 +00:00
Moe Khalil
b7c6befb37
feat(router): discover token limits for hosted OpenAI-compatible models
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 23:19:54 +00:00