ryan-crabbe-berri
faad94af94
Merge branch 'litellm_window_spend_writer' into litellm_window_spend_reader
2026-08-29 16:02:46 -07:00
ryan-crabbe-berri
b2a08100e1
Merge branch 'litellm_window_spend_schema' into litellm_window_spend_writer
2026-08-29 16:02:23 -07:00
tin-berri
36ea28b092
fix(anthropic): emit signature-only thinking blocks on the /v1/messages bridge ( #38809 )
2026-08-29 15:04:22 -07:00
Mateo Wang
20cfccaf5f
Merge pull request #37208 from BerriAI/litellm_managed_batches_observability
...
fix(batches): aggregate reasoning tokens and per-line pass/fail counts
2026-08-29 14:43:03 -07:00
Mateo Wang
c62c2afa09
Merge pull request #38234 from BerriAI/litellm_request_timeouts
...
fix(proxy): give every `requests` call a timeout so a silent server cannot hang the caller
2026-08-29 14:30:46 -07:00
Mateo Wang
2963b47cda
test: patch the Logging handler instead of the class in the poller error-file test
2026-08-29 14:09:33 -07:00
ryan-crabbe-berri
b5ec80d903
Merge commit 'a5f47a271a' into litellm_window_spend_reader
2026-08-29 13:55:53 -07:00
ryan-crabbe-berri
a5f47a271a
fix(proxy): re-queue budget window spend increments when the commit fails
...
Budget enforcement trusts a current LiteLLM_BudgetWindowSpend row without
reconciling it against LiteLLM_SpendLogs, so an increment dropped after a
failed commit let the entity spend past its window limit after the next
counter reseed. Failed increments now go back on the in-memory queue, or
back to the Redis buffer, and retry on the next scheduler tick like every
other spend category.
2026-08-29 13:55:36 -07:00
devin-ai-integration[bot]
645792955d
feat(proxy): cyberark conjur secret manager configuration via Admin UI ( #38445 )
...
* feat(proxy): CyberArk Conjur secret manager configuration via Admin UI
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): mock networking base-url helpers in AdminPanel test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): restore deployment CyberArk env config on delete and roll back on persist failure
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): reinit env-configured hashicorp vault manager after cyberark persist rollback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 13:36:08 -07:00
Mateo Wang
8dd9c4acb1
Merge pull request #30782 from emerzon/litellm_veo_31_lite
...
feat(vertex-ai): add veo 3.1 lite model metadata
2026-08-29 13:36:02 -07:00
Mateo Wang
306daf13b5
Merge pull request #38752 from BerriAI/litellm_deflake_20260829
...
fix: bound Hugging Face config fetch and keep embedding tests off the network
2026-08-29 13:33:06 -07:00
mateo-berri
4d5205c355
fix(proxy): give the remaining CLI clients a request timeout
...
The keys, credentials, models, model groups, and chat clients still sent
requests with no timeout, so a proxy that accepts the connection and
never answers pinned the caller forever. They now default to the same
30 seconds as their teams and users siblings, with chat on the OpenAI
SDK's 600 second default, and Client wires its timeout through to all of
them. S113 cannot see Session methods, so each client gets a
hanging-server regression test instead.
2026-08-29 13:32:39 -07:00
mateo-berri
8aba6e9203
Merge branch 'litellm_internal_staging' into litellm_request_timeouts
2026-08-29 13:32:33 -07:00
ryan-crabbe-berri
2c8efca0d3
Merge commit '56dd4e06ac' into litellm_window_spend_reader
2026-08-29 13:30:10 -07:00
ryan-crabbe-berri
56dd4e06ac
test(proxy): satisfy the test-quality gate for the window spend writer tests
2026-08-29 13:29:59 -07:00
Mateo Wang
38145c2082
test: undo the drive-by reformat below the poller error-file regression test
2026-08-29 13:17:43 -07:00
Mateo Wang
e0ed0a4c7a
Merge pull request #35017 from BerriAI/litellm_lit_4913_headroom_streaming_ccr
...
fix(headroom): resolve CCR retrieval on streaming /chat/completions
2026-08-29 13:01:35 -07:00
devin-ai-integration[bot]
f0340fef16
feat(mcp_gateway): add RFC 7662 introspection for gateway session tokens ( #38726 )
...
* feat(mcp_gateway): add RFC 7662 introspection for gateway session tokens
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp_gateway): omit Bearer token_type for refresh introspection and allow mcp-scoped keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp_gateway): cover introspection of RS256-signed session tokens
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp_gateway): load the discoverable router on a cold /introspect request
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(openapi): regenerate lazy snapshot and schema.d.ts for /introspect
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 12:56:04 -07:00
Mateo Wang
9ed7de6c02
Merge pull request #38670 from BerriAI/devin_ai_38659_cohere_embed_dispatch
...
fix(bedrock): route all cohere.embed models to the cohere embedding config
2026-08-29 12:55:32 -07:00
Mateo Wang
99884f0eaa
test: fake the provider file boundary in the poller error-file regression test
2026-08-29 12:53:52 -07:00
Mateo Wang
817bbe1dc6
Merge pull request #34440 from dan2k3k4/litellm_soniox_srt_cue_grouping
...
fix(soniox): align synthesized SRT/VTT cues to real speech timing
2026-08-29 12:49:57 -07:00
mateo-berri
2affd800ec
test(headroom): cover stream conversion after deployment-level compression
2026-08-29 12:45:56 -07:00
Mateo Wang
c453920f7a
Merge pull request #38285 from BerriAI/litellm_azure_v1_image_routes
...
fix(azure): use /openai/v1 image routes for v1, preview and latest api versions
2026-08-29 12:45:37 -07:00
mateo-berri
886d39c3a2
test(bedrock): expect cohere embed base64 encoding_format to normalize to float
2026-08-29 12:44:01 -07:00
yuneng-jiang
d435a62ce6
Merge pull request #38782 from BerriAI/litellm_/logs-reopen-shadcn-migration-2c9526
...
fix(ui): restore the reopen control for the log drawer's trace sidebar
2026-08-29 12:38:26 -07:00
ryan-crabbe-berri
8463cb901e
Merge remote-tracking branch 'origin/litellm_window_spend_writer' into litellm_window_spend_reader
2026-08-29 12:29:14 -07:00
ryan-crabbe-berri
2fac72392a
Merge remote-tracking branch 'origin/litellm_window_spend_schema' into litellm_window_spend_writer
2026-08-29 12:27:58 -07:00
mateo-berri
a3eac3f771
fix(bedrock): normalize encoding_format base64 to float for cohere embed models
2026-08-29 12:13:09 -07:00
Mateo Wang
4ef012627d
fix: count error-file failures in the batch cost poller path
2026-08-29 12:06:42 -07:00
mateo-berri
4d4cf40334
fix(headroom): delegate to the parent deployment hook so deployment-level configs still compress
2026-08-29 12:06:38 -07:00
mateo-berri
a007fa49e5
Merge branch 'litellm_internal_staging' into litellm_veo_31_lite
2026-08-29 12:04:42 -07:00
mateo-berri
4c42c01cb2
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_soniox_srt_cue_grouping
...
# Conflicts:
# litellm/llms/soniox/common_utils.py
2026-08-29 12:02:21 -07:00
Mateo Wang
002d0068f5
Merge pull request #38580 from BerriAI/devin_ai_fix_model_new_read_replica_lag_38556
...
fix(proxy): pin model reconcile read to the writer DB so /model/new does not 500 under read replica lag
2026-08-29 12:02:08 -07:00
Yuneng Jiang
08c83c12e9
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/logs-reopen-shadcn-migration-2c9526
2026-08-29 11:55:20 -07:00
Mateo Wang
c24f821652
Merge pull request #34849 from BerriAI/litellm_keyless_key_managed_resource_owner
...
fix(managed resources): let keys with no user_id or team_id read their own batches and files
2026-08-29 11:50:00 -07:00
ryan-crabbe-berri
e8994e8ce1
Merge branch 'litellm_window_spend_writer' into litellm_window_spend_reader
2026-08-29 11:49:42 -07:00
ryan-crabbe-berri
4e22a5ef5a
Merge branch 'litellm_window_spend_schema' into litellm_window_spend_writer
2026-08-29 11:47:37 -07:00
mateo-berri
f2f988bd57
test: cover forged _headroom_interception_converted_stream strip at the proxy boundary
2026-08-29 11:46:23 -07:00
Mateo Wang
5e60ec5c31
Merge pull request #38743 from BerriAI/litellm_techdebt_20260829
...
refactor: clean up tech debt that landed on 2026-08-29
2026-08-29 11:46:05 -07:00
mateo-berri
6d3e687ce4
fix(db): let the writer pin yield to the replica while the writer is degraded
2026-08-29 11:44:45 -07:00
mateo-berri
9e01bd1441
fix(azure): send the deployment name as the body model on v1 image routes
2026-08-29 11:41:58 -07:00
mateo-berri
68f891fd2b
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_keyless_key_managed_resource_owner
2026-08-29 11:33:08 -07:00
mateo-berri
a8c36e8307
Merge remote-tracking branch 'origin/litellm_internal_staging' into devin_ai_fix_model_new_read_replica_lag_38556
2026-08-29 11:29:07 -07:00
Mateo Wang
fee8619708
test: add batch request count keys to gcs pub sub spend logs fixture
2026-08-29 11:26:39 -07:00
yuneng-jiang
fa25ff2a2e
Merge pull request #38626 from BerriAI/litellm_ui_model_links_team_key_info
...
feat(ui): link team and key model chips to the models page filtered to that group
2026-08-29 11:26:15 -07:00
ryan-crabbe-berri
7745fe887f
Merge branch 'litellm_window_spend_writer' into litellm_window_spend_reader
2026-08-29 11:06:05 -07:00
ryan-crabbe-berri
041cae8280
fix(proxy): bound the window spend seed exclusion to the batch's own start time
...
request_id can be chosen by the client through x-litellm-call-id, so an
unbounded NOT (request_id = ANY(batch)) let a replayed old id drop that id's
historical LiteLLM_SpendLogs row from the one-time seed while its increment
still landed. The increment now carries the request start, the batch keeps
the earliest one, and the seed only excludes ids whose startTime is at or
after it.
2026-08-29 11:05:57 -07:00
tin-berri
2a5d09ee87
fix(policy): let the AI policy suggester drop sampling params its model refuses ( #38594 )
...
The suggester pins temperature=0.2 for tool-selection determinism and passed no
drop_params, so an operator-supplied reasoning model whose only accepted temperature is 1
made litellm raise UnsupportedParamsError and the whole suggestion fail. The default
gpt-4o-mini is unaffected; the failure needs the caller to name a model.
Every other internal LLM call the proxy makes on a user's behalf already opts in through
judge_acompletion, which sets drop_params=True on both dispatch paths. This was the one
caller outside that contract, so the sampling preference is now advisory here too and the
call degrades instead of dying.
Resolves LIT-6352
2026-08-29 10:58:27 -07:00
Mateo Wang
c3edb95e8d
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_managed_batches_observability
...
# Conflicts:
# tests/test_litellm/proxy/spend_tracking/test_spend_management_endpoints.py
2026-08-29 10:56:02 -07:00
mateo-berri
acb621c35b
Merge remote-tracking branch 'origin/litellm_internal_staging' into devin_ai_fix_model_new_read_replica_lag_38556
2026-08-29 10:54:46 -07:00