Commit graph

19740 commits

Author SHA1 Message Date
yucheng
6464c1fd6c fix(proxy): make SettingsStore.clear() terminate when the config file owns a key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 20:18:48 +00:00
mateo-berri
c54c0b049d fix(vertex_ai): keep reasoning_effort unsupported on Vertex AI Mistral partner models
Vertex AI Mistral models reused MistralConfig, whose reasoning_effort
advertisement checks the mistral provider entry of the cost map, so
vertex_ai/mistral-medium-3 started advertising reasoning_effort and
drop_params stopped dropping it, turning a 200 into a Vertex 400.
VertexAIMistralConfig scopes that lookup to the vertex_ai provider, and
MistralConfig now reads the provider from its custom_llm_provider
property instead of a hardcoded "mistral".
2026-09-18 13:16:56 -07:00
Yassin Kortam
2e46b10320
Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split
perf(proxy): split aggregated usage query into key-free rollups and bounded top-N keys
2026-09-18 13:16:49 -07:00
yassin
1feffc3635 Merge remote-tracking branch 'origin/main' into litellm_mcp_admin_terminate_sessions_revoke_credentials 2026-09-18 20:16:22 +00:00
yassin
817aefc413 test(proxy): keep stored pass-through route tests off the shared FastAPI app so the allowlist coverage test stays order independent
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 20:15:48 +00:00
mateo-berri
1581c45615 fix(proxy): check and persist only the settings the request sent
POST /config/update compared litellm_settings after lowercasing the
callback list, so a config file spelling a callback in mixed case refused
the same list sent back, and it stored every general_settings model default
next to the keys the request set. Both now use the request as sent; only
the stored callback list is lowercased.

Also drops config_data from the router settings reload callers the previous
commit left behind and teaches the legacy MockProxyConfig the ownership
check.
2026-09-18 13:14:22 -07:00
ryan-crabbe-berri
6e3b6d6d03
Merge pull request #41830 from BerriAI/litellm_scim_multivalued_optional_value
fix(scim): accept entitlements and roles entries without a value on SCIM user PUT
2026-09-18 13:13:51 -07:00
Devin AI
04193745a6 fix(docs-test): restore line-anchored table regex in router settings check
A merge on this branch dropped the ^ anchor and MULTILINE flag from the
doc_key_pattern, so the unanchored match started swallowing key names
into captured fields whenever a row's description cell itself contains
a pipe (Literal unions and similar). The check then reported 27
long-documented keys as undocumented. Restores the exact pattern used
on main, verified against a live litellm-docs checkout: all 58 Router
init params resolve as documented.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 20:10:20 +00:00
mateo-berri
d1563e0b55 feat: honor eager_input_streaming on Bedrock and Anthropic Claude tools 2026-09-18 13:09:13 -07:00
ryan-crabbe-berri
d37320f8cf fix(proxy): report null cost for unpriced deployments on the /model/info id lookup
GET /model/info?litellm_model_id= went through _get_proxy_model_info, a copy of _enrich_model_info_with_litellm_data that missed the unpriced check, so the same deployment read as null in the list and as 0 on the id lookup. The id lookup now delegates to the shared helper, so the two paths cannot drift again
2026-09-18 13:08:55 -07:00
mateo-berri
74e9fb2323 fix(passthrough): validate only the winning timeout value in the resolver 2026-09-18 13:07:33 -07:00
yassin
b92820dcfe Merge branch 'litellm_usage_key_free_aggregate_split' into litellm_daily_global_spend_table 2026-09-18 20:06:47 +00:00
yassin
191ca14872 Merge remote-tracking branch 'origin/main' into litellm_usage_key_free_aggregate_split 2026-09-18 20:05:54 +00:00
Mateo Wang
57377c9249
Merge pull request #41860 from BerriAI/litellm_router_settings_doc_table_parse
test(docs): read only the first column of the router_settings reference table
2026-09-18 13:03:51 -07:00
Devin AI
fc186f5613 fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT models on Converse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 20:02:31 +00:00
Yucheng He
39a199d9a2 fix(proxy): let in-flight scheduled jobs finish before cancelling them at shutdown
Cancelling every in-flight job the moment shutdown reached the scheduler
dropped the rows a write job had already popped: flush_gateway_requests
drains its accumulator before committing and does not restore it on
CancelledError, and update_spend requeues its batch only after the
shutdown drain had already run.

Shutdown now waits up to JOB_FINISH_TIMEOUT_SECONDS for in-flight jobs
to finish on their own, cancels the ones still running, and does both
before the shutdown flushes so a requeued batch is still written. The
cleanup run never finishes inside the grace, so it is still cancelled
and still records outcome="aborted".

Resolves LIT-6990
2026-09-18 20:00:48 +00:00
Yucheng He
8ce8dd9b3b fix(proxy): pause the scheduler at shutdown start and keep cleanup progress per run
Review follow-ups on #41213:

- Pause the scheduler as the first shutdown step so a job whose fire time
  falls inside the shutdown window does not start only to be cancelled.
  Jobs already running keep the whole window and are cancelled and
  awaited before the database disconnects, as before.
- Keep the cleanup run's progress in a task-scoped ContextVar rather than
  on the cleaner instance, so two runs overlapping on one cleaner
  (APSCHEDULER_MAX_INSTANCES above 1 without a Redis lock) each report
  their own rows and batches on cancellation.
- Drop the module docstrings the repository comment policy does not
  allow; the rationale lives in the PR description.
2026-09-18 20:00:47 +00:00
Yucheng He
c1c566db87 fix(proxy): record aborted outcome when spend-log cleanup is cancelled at shutdown
cleanup_old_spend_logs only caught Exception, so a run cut short by
CancelledError recorded no outcome and logged nothing. Under uvicorn the
job was never cancelled at all: uvicorn re-raises the captured SIGTERM as
soon as the lifespan shutdown returns, before asyncio cancels outstanding
tasks, so an in-flight scheduler job simply died with the process.

The cleanup now handles CancelledError by logging elapsed time, rows
deleted and batch count at error level, recording outcome="aborted", and
re-raising. The lifespan shutdown stops the scheduler and awaits the jobs
it cancels while the database is still connected, so that handler runs
under uvicorn too, and the pod lock is released instead of orphaned.

Resolves LIT-6990
2026-09-18 20:00:19 +00:00
jesus
8549189cbb test(proxy): use local_model_cost_map fixture for alias listing test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 19:59:24 +00:00
ryan-crabbe-berri
6b54238083
Merge pull request #40807 from BerriAI/litellm_service_account_team_key_mgmt
feat(keys): let team service account keys use key management endpoints for their own team
2026-09-18 12:58:41 -07:00
Yassin Kortam
2edda5aec3
Merge pull request #41757 from BerriAI/litellm_typesafe_compaction_guardrail
feat(guardrails): add TypeSafe Jev relevance-based compaction guardrail
2026-09-18 12:56:52 -07:00
jesus
32e6a86319 fix(router): report null cost for unpriced deployments instead of 0
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 12:52:57 -07:00
yassin
cf22e4b171 fix(proxy): make SettingsStore.clear terminate when the config file owns a key
MutableMapping.clear pops items until the mapping is empty, but __delitem__
keeps config-owned keys, so clear spun forever on any store that had loaded a
config file. unittest.mock.patch.dict calls clear on exit, which is why the
proxy-infra and proxy-endpoints shards hung at 99 percent until the 20 minute
job timeout on every run since the store started refusing config-owned writes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 19:51:15 +00:00
Mateo Wang
a7a3f6802e
Merge pull request #41203 from clonylu/fix/anthropic-thinking-binding-beta-header
fix(anthropic): register thinking-binding-controls-2026-08-01 in beta headers config
2026-09-18 12:50:55 -07:00
mateo-berri
58beea2275 fix(passthrough): read the stream flag by truthiness in the timeout resolver 2026-09-18 12:45:54 -07:00
mateo-berri
8533cb9d59 test(proxy): drop the pass-through routes the reload tests add to the shared app
The two pass-through reload tests register routes on the shared app and
never take them off, so a later allowlist test in the same xdist worker
finds routes that no component exposes. Restore the app's route list
when each of those tests finishes.
2026-09-18 12:44:42 -07:00
Devin AI
4f904f69b6 test(proxy): restore app routes after pass-through reload tests
The two pass-through reload tests register real FastAPI routes on the
shared proxy app and only clean up the internal registry, so any test
running after them in the same worker sees stray /v1/kept-* and
/v1/deleted-* routes. test_component_allowlists counts those as
uncovered and fails. Snapshot app.routes and the registry up front and
restore both in a finally block.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 19:44:18 +00:00
mateo-berri
7cfec730d5 fix(proxy): refuse config-owned keys on POST /config/update
POST /config/update stored keys the config file owns and answered 200
while the file silently kept winning. Run the config-owned check for
general_settings, litellm_settings, and router_settings before the first
database read, the way /config/field/update already does, so a refused
request stores nothing.

The config reload re-loaded the merged settings as yaml settings, which
turned every saved router setting read-only after one tick. Read the
saved router settings row instead so database-owned values stay writable.

The integration harness seeds num_retries through /config/update instead
of the config file, which is what the effective-settings and observed
routing tests need to keep exercising a database-owned value.
2026-09-18 12:42:08 -07:00
mateo-berri
c0b0ba20b2 fix(azure-ai): keep FLUX.2 tolerant of OpenAI-only image params 2026-09-18 12:41:40 -07:00
yassin
fe8cf02823 feat(mcp): resolve the allowlisted client identity from the JWT claim or an opt-in header instead of clientInfo.name
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 19:27:43 +00:00
yassin
1e7c5400fd fix(proxy): canonicalize azure speech paths and bill uploaded short audio
Resolve dot segments in the /azure_speech endpoint path before the endpoint family and the admin-only batch guard are decided, so the guard and the forwarded upstream path agree. Bill short-audio requests for the longer of the uploaded audio duration and the recognized duration, so a NoMatch or silence response still charges for the audio Azure processed

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 19:27:25 +00:00
Yuneng Jiang
51f0620439
test(e2e): use deployed Azure model 2026-09-18 12:24:25 -07:00
mateo-berri
b4dc081c27 fix(mistral): never round reasoning_effort none onto the strength ladder 2026-09-18 12:22:34 -07:00
Devin AI
a774e71175 test(bedrock_mantle): document mantle in-region pricing invariant
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 19:18:45 +00:00
Devin AI
2299846b48 fix(bedrock_mantle): align gpt-5.6-sol pricing with the AWS in-region model card
Co-authored-by: kusumakarb <kusumakarbvss@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 19:18:02 +00:00
yassin
b5ad17e848 test(docs): read only the first column of the router_settings reference table
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 19:16:26 +00:00
yassin
649d299c44 Merge remote-tracking branch 'origin/main' into litellm_fix_tpm_window_reset_sibling_counters 2026-09-18 19:13:05 +00:00
yassin
afa8b8a903 test(docs): read only the first column of the router_settings reference table
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 19:11:39 +00:00
yassin
84f7adec2e fix(passthrough): match deepgram listen routes served under a path prefix
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 19:06:24 +00:00
Mateo Wang
a2626726a2
Merge pull request #40500 from BerriAI/litellm_bedrock_aws_auth_params
fix(bedrock): send aws_session_tags on every STS call via one typed auth struct
2026-09-18 12:01:52 -07:00
yassin
c5919a3c0e Merge remote-tracking branch 'origin/main' into litellm_mcp_client_allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/proxy/_experimental/mcp_server/server.py
#	litellm/proxy/proxy_server.py
2026-09-18 19:01:08 +00:00
ryan-crabbe-berri
ed9e666b54
Merge pull request #41351 from BerriAI/litellm_routing_groups_atomic_validation
fix(router): validate routing_groups at save time and keep invalid DB groups from blocking SSO load
2026-09-18 11:59:21 -07:00
mateo-berri
c1e39810ec fix(mistral): send reasoning_effort as a level the model accepts
Declare the live-verified reasoning_effort_levels on the Mistral cost-map
entries and round an undeclared request to the nearest declared level
(up to the weakest level at least as strong, down to the strongest when
the request exceeds the ceiling). Codex's default medium no longer 400s
on mistral-medium-latest, mistral-small-latest, or the vibe-cli family;
an entry that declares nothing keeps forwarding the value verbatim
2026-09-18 11:58:04 -07:00
kerry
fb8816d826 Merge remote-tracking branch 'origin/main' into litellm_lit_8128_off_peak_pricing_schema 2026-09-18 18:55:24 +00:00
mateo-berri
73fddb999e fix(router): resolve stream_timeout before generic timeouts on the passthrough route 2026-09-18 11:53:26 -07:00
yassin
307df09792 test(mcp): assert the self-revoke response instead of echoing the delete mock
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 18:51:43 +00:00
ryan-crabbe-berri
839c040b53
Merge pull request #41695 from BerriAI/litellm_issue_classifier
ci: classify new issues into domain, provider, kind, priority and lift labels
2026-09-18 11:50:56 -07:00
Mateo Wang
adb5ee0881
Merge pull request #41275 from BerriAI/litellm_azure_strip_file_format
fix(azure): strip litellm format field from file and image content parts
2026-09-18 11:46:23 -07:00
yassin
3a5f286273 Merge remote-tracking branch 'origin/main' into litellm_fix_tpm_window_reset_sibling_counters 2026-09-18 18:44:33 +00:00
Mateo Wang
88c9dd1294
Merge pull request #41511 from BerriAI/litellm_foundry_a2a_entra_agents
feat(a2a): reach Microsoft Foundry agents with Entra auth and versioned card discovery
2026-09-18 11:44:14 -07:00