mateo-berri
c1e39810ec
fix(mistral): send reasoning_effort as a level the model accepts
...
Declare the live-verified reasoning_effort_levels on the Mistral cost-map
entries and round an undeclared request to the nearest declared level
(up to the weakest level at least as strong, down to the strongest when
the request exceeds the ceiling). Codex's default medium no longer 400s
on mistral-medium-latest, mistral-small-latest, or the vibe-cli family;
an entry that declares nothing keeps forwarding the value verbatim
2026-09-18 11:58:04 -07:00
kerry
fb8816d826
Merge remote-tracking branch 'origin/main' into litellm_lit_8128_off_peak_pricing_schema
2026-09-18 18:55:24 +00:00
mateo-berri
73fddb999e
fix(router): resolve stream_timeout before generic timeouts on the passthrough route
2026-09-18 11:53:26 -07:00
yassin
307df09792
test(mcp): assert the self-revoke response instead of echoing the delete mock
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 18:51:43 +00:00
ryan-crabbe-berri
839c040b53
Merge pull request #41695 from BerriAI/litellm_issue_classifier
...
ci: classify new issues into domain, provider, kind, priority and lift labels
2026-09-18 11:50:56 -07:00
Mateo Wang
adb5ee0881
Merge pull request #41275 from BerriAI/litellm_azure_strip_file_format
...
fix(azure): strip litellm format field from file and image content parts
2026-09-18 11:46:23 -07:00
yassin
3a5f286273
Merge remote-tracking branch 'origin/main' into litellm_fix_tpm_window_reset_sibling_counters
2026-09-18 18:44:33 +00:00
Mateo Wang
88c9dd1294
Merge pull request #41511 from BerriAI/litellm_foundry_a2a_entra_agents
...
feat(a2a): reach Microsoft Foundry agents with Entra auth and versioned card discovery
2026-09-18 11:44:14 -07:00
mateo-berri
cf08440557
fix(azure-ai): price FLUX.2 Flex by its resolved cost map key
2026-09-18 11:42:52 -07:00
mateo-berri
a4da989aa9
test(bedrock): type the STS recording helper
2026-09-18 11:42:52 -07:00
yassin
1e6b33ffab
Merge remote-tracking branch 'origin/main' into litellm_azure_speech_passthrough
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
# Conflicts:
# helm/litellm/templates/ingress.yaml
# litellm/proxy/_types.py
# litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py
# litellm/proxy/pass_through_endpoints/success_handler.py
# terraform/litellm/aws/locals.tf
# terraform/litellm/gcp/locals.tf
# tests/test_litellm/proxy/middleware/test_billable_request_metrics_middleware.py
2026-09-18 18:42:22 +00:00
Yassin Kortam
1653d132c5
Merge pull request #41506 from BerriAI/litellm_vertex_gcs_file_content_streaming
...
feat(vertex_ai): stream GCS batch output files from /v1/files/{id}/content
2026-09-18 11:42:14 -07:00
mateo-berri
42541a9233
fix(proxy): drop echoed cost-map pricing on a row's next save and build /model/info pricing stamps without mutation
2026-09-18 11:42:07 -07:00
yassin
aeca6ed7ba
chore: merge main into litellm_deepgram_listen_websocket_passthrough
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 18:39:39 +00:00
yassin
048aaad627
Merge remote-tracking branch 'origin/main' into litellm_mcp_admin_terminate_sessions_revoke_credentials
2026-09-18 18:39:16 +00:00
Yassin Kortam
0594dd7caf
Merge pull request #41539 from BerriAI/litellm_vault_login_secret_namespace
...
feat(vault): add separate login and secret namespaces for HashiCorp Vault
2026-09-18 11:38:49 -07:00
Yassin Kortam
84ae0805ba
Merge pull request #41620 from BerriAI/litellm_team_member_temp_budget_increase
...
feat(proxy): temporary budget increase for team members
2026-09-18 11:38:34 -07:00
ryan-crabbe-berri
294a9e15de
ci: classify new issues into domain, provider, kind, priority and lift labels
...
Every issue opened from now on is gated on the template headings, sent once
through the LiteLLM proxy with a strict JSON schema, and labelled from the
manifest in .github/labels.json. Old-template issues are not touched. The bug
template shrinks to Description, Config, LiteLLM Version and Steps to Repro,
both templates gain a domain dropdown, and the labelers that keyed off the old
component dropdown go away.
2026-09-18 11:37:55 -07:00
Yassin Kortam
f6d9b2552f
Merge pull request #41692 from BerriAI/litellm_mcp_gateway_sessions_by_client_user
...
feat(mcp): show live gateway sessions by AI client and user
2026-09-18 11:37:16 -07:00
Yassin Kortam
ca79c393d5
Merge pull request #41636 from BerriAI/litellm_per_key_end_user_default_budget
...
feat(proxy): per-key default budget for dynamically created customers
2026-09-18 11:36:59 -07:00
Yassin Kortam
f129e7e2d0
Merge pull request #41515 from BerriAI/litellm_transcribe_passthrough
...
feat(proxy): add Amazon Transcribe pass-through with completion-time job pricing
2026-09-18 11:35:06 -07:00
yassin
3219a875d0
fix(proxy): serialize in-memory window rollover and reset siblings once per expired window
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 18:20:18 +00:00
Yassin Kortam
d42f448e41
Merge pull request #41585 from BerriAI/litellm_lazy_fastapi_bpe_imports
...
perf: defer fastapi and tiktoken BPE imports out of import litellm
2026-09-18 11:19:13 -07:00
Yassin Kortam
660f3dfdd3
Merge pull request #41555 from BerriAI/litellm_max_parallel_requests_queue_size
...
feat(router): reject with 429 when a deployment's max_parallel_requests slots are all in use
2026-09-18 11:18:08 -07:00
kerry
4b215e2a60
fix(schema): require a schedule and well-formed windows in off_peak_pricing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 18:17:36 +00:00
mateo-berri
860b0203c0
test(bedrock): expect canonical session tags in the dynamic auth params propagation test
2026-09-18 11:17:34 -07:00
kerry
cbcb55af89
fix(schema): classify off_peak_pricing as a structured object in the model prices schema generator
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 18:11:17 +00:00
kerry
d68e20e586
Merge remote-tracking branch 'origin/litellm_e2e_cost_calculation_scripted_provider' into litellm_e2e_cost_calculation_scripted_provider
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
2026-09-18 18:05:20 +00:00
Mateo Wang
db756b9393
Merge pull request #40730 from BerriAI/litellm_responses_nested_additional_drop_params
...
fix(responses): honor nested additional_drop_params paths
2026-09-18 11:04:38 -07:00
kerry
e87830feba
Merge remote-tracking branch 'origin/main' into litellm_e2e_cost_calculation_scripted_provider
2026-09-18 17:56:49 +00:00
kerry-berri
fd58c31cc7
Merge pull request #41832 from BerriAI/litellm_lit_8111_cache_read_missing_rate
...
fix(cost): bill cache-read tokens at the input rate when the map has no cache-read rate
2026-09-18 10:56:25 -07:00
yassin
fa70e49b81
chore: merge main into litellm_transcribe_passthrough
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 17:54:59 +00:00
mateo-berri
b03957ba9c
fix(proxy): unpin cost-map pricing copied into model_info and report pricing overrides
...
A model_info blob that carries key next to pricing fields is a copy of a /model/info response (only litellm.get_model_info emits key), so those pricing fields are dropped when the row is loaded from the DB and on every Reload Price Data, and the deployment follows the current cost map again. Prices typed into litellm_params, or into model_info without key, stay as they are.
/model/info, /v1/model/info and /v2/model/info now report model_info.pricing_overrides, the pricing fields the deployment sets itself, and the Admin UI model page says whether a price follows the cost map or overrides it.
2026-09-18 10:49:36 -07:00
mateo-berri
2a76f54a5a
fix(bedrock): merge main and thread aws_session_tags through the auth struct
...
Merge origin/main (a9ee15372f ) into the typed AwsAuthParams refactor so the
session tags PR #40446 added land in the struct: resolve_credentials
canonicalizes aws_session_tags before STS, the realtime path forwards them,
and Files upload/download plus bodiless S3 signing now assume the role with
the tags instead of dropping them.
2026-09-18 10:41:52 -07:00
mateo-berri
b523ed9a2d
merge: origin/main into litellm_jwt_token_exchange_grant
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
2026-09-18 10:35:53 -07:00
ryan
8dfda93123
test(proxy): drop explanatory docstrings from routing_groups regression tests
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 17:32:31 +00:00
mateo-berri
f6ee046199
fix(proxy): read the user row past the recent-miss memo on a database-only lookup
...
get_user_object skipped the database for db_cache_expiry seconds after a miss on the same worker even when the caller asked for check_db_only, so the token exchange mint could answer no_active_key for a user JWT auth had just created. A database-only read now always reaches the database.
2026-09-18 10:25:45 -07:00
Devin AI
dba9ff801f
fix(proxy): reset sibling tpm/rpm counters when shared window rolls over
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 17:18:49 +00:00
ryan
c3bc55d18f
chore: merge main into litellm_routing_groups_atomic_validation
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 16:59:23 +00:00
yuneng-jiang
c4ab1d98e9
Merge pull request #41779 from BerriAI/litellm_settings_store_precedence
...
refactor(proxy): make the config file win over the database
2026-09-18 09:52:09 -07:00
kerry
af61d9247c
test(cost_tracking): expect cache reads without a map rate to estimate at the input rate
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 16:33:17 +00:00
kerry
88d1eb3b35
fix(cost): bill cache-read tokens at the input rate when the map has no cache-read rate
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 16:22:14 +00:00
ryan
817cdcef64
test(scim): drop redundant docstring from the valueless entitlements PUT test
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 16:00:54 +00:00
Mateo Wang
861f79797f
Merge pull request #41663 from BerriAI/litellm_remove_legacy_interactions_schema_flag
...
refactor(interactions): remove expired use_legacy_interactions_schema shim
2026-09-18 08:48:30 -07:00
ryan
b6410d563b
fix(scim): accept entitlements and roles entries without a value on SCIM user PUT
...
SCIMMultiValuedAttribute required value, so a PUT /scim/v2/Users/{id} that
carried an IdP-specific entitlements entry such as {"groups": [...]} failed
body validation with 422 and the suspend (active: false) never reached
update_user. value is now optional and unknown members are kept, so the
suspend is applied, keys are blocked, and the entries are stored as sent
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 15:46:33 +00:00
ryan-crabbe-berri
4f70b88a1f
Merge pull request #41707 from BerriAI/litellm_jwt_mapping_cache_evict_on_bulk_key_delete
...
fix(proxy): evict jwt key mapping cache on user, team, org, and bulk key deletion
2026-09-18 08:41:31 -07:00
kerry
5f3a86aee5
test(e2e): use TypeAlias over 3.12 type statements in e2e models
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 13:27:14 +00:00
kerry
dda7776346
test(e2e): make cost-calculation cases MECE by rate-key ownership with realistic fixtures
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 13:15:29 +00:00
yucheng
8fc74a78bc
fix(proxy): reject blank trusted_proxy_ranges entries before they are dropped
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 11:50:30 +00:00
yucheng
cacd12b87d
fix(proxy): treat a malformed trusted_proxy_ranges entry as an undeclared topology
...
A list with an entry that is not an address or CIDR range no longer switches
the per-source Admin UI sign-in limit on against the direct peer address, so a
typo cannot make a shared ingress address the bucket for every user behind it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 11:28:35 +00:00