Commit graph

48142 commits

Author SHA1 Message Date
kerry
536a85b429 fix(cost-map): keep register_model url fetch to a single attempt
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 02:48:17 +00:00
kerry
aa0a9ab3ea refactor(cost-map): drop initial_outcome flag from retry loop
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 02:29:56 +00:00
kerry
7d52832692 merge: litellm_internal_staging into litellm_cost_map_background_retries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 02:17:30 +00:00
Mateo Wang
24ef3ec63b
Merge pull request #37781 from ZXT-zjbiliy/fix/build-base-response-empty-choices
fix(stream_chunk_builder): guard empty choices and missing role in build_base_response
2026-09-08 19:14:03 -07:00
mateo-berri
ddedb4867b fix: discard a streamed rewrite that drops or adds a tool call
A guardrail that removes or adds a tool call on an ended stream used to be
silently ignored: every handler substitutes the original list on a count
mismatch and the executor skipped its observer once the translation could
deliver rewrites. The executor now tracks the count change on the observer
and releases the original chunks with the discard warning on every
translation, matching what the merge base did for any tool call rewrite
2026-09-08 19:09:44 -07:00
Mateo Wang
f8e456d105
Merge pull request #40275 from BerriAI/litellm_lit6852_spend_attribution
fix(spend-tracking): recover key alias for session tokens from spend logs
2026-09-08 19:05:43 -07:00
kerry
0086b62b45 fix(cost-map): keep first fetch blocking, run retries in background
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 02:05:08 +00:00
Yuneng Jiang
1638473332
fix(bedrock): keep deletion response IDs in request context 2026-09-08 19:02:08 -07:00
tin-berri
1a9c6ce390
fix(mcp): preserve proxy logging and authorization coverage (#40337)
* test(mcp): exercise /mcp/proxy authorization against the real registry instead of patched manager methods

* fix(mcp): preserve proxy logging and authorization coverage

* test(mcp): respect the proxy FastAPI import boundary
2026-09-09 02:01:01 +00:00
yuneng-jiang
0d62970865
Merge pull request #40334 from BerriAI/litellm_/release-version-bump-787548
chore: bump litellm-enterprise 0.1.65 -> 0.1.66
2026-09-08 18:53:51 -07:00
mateo-berri
38a3de7641 fix(azure_ai): use the pixel cap the live MAI endpoint enforces 2026-09-08 18:46:18 -07:00
yuneng-jiang
86ee031217
Merge branch 'litellm_internal_staging' into litellm_/release-version-bump-787548 2026-09-08 18:44:54 -07:00
yuneng-jiang
1fbd1cb9ce
Merge pull request #40347 from BerriAI/litellm_fix_mcp_proxy_test_isolation
test(mcp): fix proxy fixture isolation after manager reload
2026-09-08 18:44:45 -07:00
Mateo Wang
d75aa4445d
Merge pull request #40179 from BerriAI/litellm_lit_2133_cost_map_provenance
feat(cost_map): report which revision of the price map the proxy is serving
2026-09-08 18:42:19 -07:00
Yuneng Jiang
f17632c036
chore: sync batch tests with latest staging 2026-09-08 18:36:15 -07:00
Yuneng Jiang
fb21852f7b
test(mcp): resolve current manager in proxy fixtures 2026-09-08 18:35:25 -07:00
tin-berri
314e573529
feat(auto-router): refresh family reasoning presets (#40341) 2026-09-08 18:28:46 -07:00
mateo-berri
84a99dfc31 fix(azure_ai): surface rejected MAI image params as 400 and drop comments 2026-09-08 18:24:51 -07:00
ryan-crabbe-berri
f5e4aa38ba
Merge pull request #40342 from BerriAI/litellm_prompt_cache_key_session_id
fix(anthropic): key the /v1/messages prompt cache on Claude Code's session_id only
2026-09-08 18:20:56 -07:00
ryan-crabbe-berri
634852a183 fix(anthropic): key the /v1/messages prompt cache on Claude Code's session_id only
The bridges derived prompt_cache_key as the first 64 chars of metadata.user_id.
Claude Code packs a JSON object into that field whose prefix is the per-install
device_id, so every session and subagent on one machine shared a single key,
and a plain end-user id pinned all of that user's conversations to one slot.

Parse the JSON and use session_id; send no key otherwise so the provider falls
back to its own prompt-prefix hashing. An explicit prompt_cache_key still wins.

Fixes #39145
2026-09-08 18:06:39 -07:00
yuneng-jiang
a99ecacffd
Merge pull request #40336 from BerriAI/litellm_extend_diskcache_deadline_oct1
chore(ci): extend diskcache scan exception to October 1
2026-09-08 18:05:00 -07:00
devin-ai-integration[bot]
43a1b2992a
fix(otel v2): restore the Datadog auth span and the last-wins callback merge (#40335)
* fix(otel v2): restore the Datadog auth span and the last-wins callback merge

Move @tracer.wrap() back onto user_api_key_auth so USE_DDTRACE=true emits the
auth span again, and let a failure entry's callback_vars take part in the
destination merge so the resolver picks the same account the runtime parser does

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel v2): drop docstrings from the two regression tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rerun proxy-infra after the flaky test_check_migration process-tree test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 01:00:26 +00:00
yuneng-jiang
9b186680c4
Merge branch 'litellm_internal_staging' into litellm_extend_diskcache_deadline_oct1 2026-09-08 17:55:38 -07:00
Yuneng Jiang
24ee66328a
test: bound Datadog read-back retries using reset headers 2026-09-08 17:55:21 -07:00
yuneng-jiang
c417e1084d
Merge branch 'litellm_internal_staging' into litellm_/release-version-bump-787548 2026-09-08 17:54:07 -07:00
yuneng-jiang
fb36c3c5a5
Merge pull request #40333 from BerriAI/litellm_fix_prisma_timeout_test_cleanup
test(proxy): fix Prisma timeout cleanup after subreaper tests
2026-09-08 17:53:56 -07:00
mateo-berri
529b8706ee test(spend-tracking): mock the spend-log scan through the transaction its caller now opens 2026-09-08 17:51:13 -07:00
mateo-berri
1133507565 refactor: type the buffered stream rewrite helpers without Any 2026-09-08 17:46:55 -07:00
Yuneng Jiang
810d48f28f
chore(ci): extend diskcache scan exception to October 1 2026-09-08 17:41:27 -07:00
Yuneng Jiang
94a81f003e
bump: litellm-enterprise 0.1.65 -> 0.1.66 2026-09-08 17:40:34 -07:00
mateo-berri
c6a0381595 fix(token_counter): bound concurrent HuggingFace encodes so a burst of large counts cannot exhaust memory 2026-09-08 17:38:01 -07:00
Yuneng Jiang
5a7919f3f5
test(proxy): isolate reaper state and reap Prisma fixture children 2026-09-08 17:35:43 -07:00
tin-berri
754a2afe12
feat(mcp): add schema discovery proxy mode (#40298) 2026-09-09 00:30:20 +00:00
Mateo Wang
599daea985
Merge pull request #36718 from BerriAI/litellm_fix_count_tokens_budget_reservation_leak
fix(budget_reservation): don't reserve budget on token counting routes
2026-09-08 17:29:27 -07:00
mateo-berri
359c26aa1d test(policy_engine): cover the legacy stream paths with no response and no rescan
A translation that hands the hook no assembled response leaves the stream
as it is, and a rewrite the translation cannot rescan is released as the
original stream with a warning.
2026-09-08 17:27:46 -07:00
moe-berri
6112274350
Merge pull request #40273 from BerriAI/litellm_non_reasoning_tier
feat(auto_router): opt-in NON_REASONING tier below SIMPLE
2026-09-08 17:27:35 -07:00
mateo-berri
268b944081 fix(spend-tracking): bound the spend-log scan with a statement timeout and name only unanimous alias, team, and owner 2026-09-08 17:24:22 -07:00
mateo-berri
5ea447e295 fix(azure): keep the model group segment exact and casefold only deployment names in the relay guard
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
2026-09-08 17:19:23 -07:00
mateo-berri
4fa34accb8 Merge remote-tracking branch 'origin/litellm_internal_staging' into HEAD
# Conflicts:
#	litellm/litellm_core_utils/litellm_logging.py
2026-09-08 17:17:31 -07:00
Mateo Wang
402351d980
Merge pull request #40268 from BerriAI/litellm_fireworks_responses_reasoning_instructions
fix(fireworks_ai): fold instructions and developer items into one leading system message on the Responses path
2026-09-08 17:16:52 -07:00
mateo-berri
9e67c08309 fix(policy_engine): keep tool-only legacy streams deliverable
A Messages stream that ends with only tool_use blocks reaches the legacy
step without a texts key, while the non-streaming rescan of the same
response sends an empty list. Compare both as empty and keep the stream
when the hook left the tool calls alone.
2026-09-08 17:16:39 -07:00
mateo-berri
260097b2dd refactor(hosted_vllm): drop redundant image edit docstrings 2026-09-08 17:16:34 -07:00
mateo-berri
a892e67c40 fix(azure): match deployment segments case-insensitively and keep relay helpers immutable 2026-09-08 17:13:53 -07:00
mateo-berri
831a2a13fb fix(azure): price azure_ai transcriptions at the azure_ai cost-map entry 2026-09-08 17:12:18 -07:00
mateo-berri
0c58346ba9 fix(proxy): warn when a deferred background policy was matched through a request tag
Retrieval re-matches only the key, team, and model scopes, so a post_call
policy that reached a pending background response through a request-body
tag does not govern the completed response. Log that at submit, next to the
deferral, and cover the retrieval re-match with tag-scoped tests.
2026-09-08 17:07:37 -07:00
mateo-berri
b456caa05e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit_7174_stream_tool_call_rewrites 2026-09-08 17:05:11 -07:00
Yuneng Jiang
b8be30219c
test: stop guardrail retries at the polling deadline 2026-09-08 17:02:00 -07:00
mateo-berri
88d2d77553 feat(hosted_vllm): add image edit support
Register a HostedVLLMImageEditConfig so hosted_vllm/<model> deployments route POST /v1/images/edits to the vLLM-Omni OpenAI-compatible endpoint instead of failing with 'image edit is not supported for hosted_vllm' before any request is sent
2026-09-08 16:59:55 -07:00
mateo-berri
0f224dd4ed fix: sign Bedrock guardrail, embedding, and Mantle requests off the event loop
Bedrock Guardrails resolved credentials and signed inline on the loop in
its three async paths, Titan embeddings signed each batch item inline in
the async loop, and Bedrock Mantle requests slipped past the off-loop
gate because BedrockMantleAuthMixin composes a BaseAWSLLM instead of
inheriting from it. Introduce the SignsRequestsWithAWS marker that both
BaseAWSLLM and the Mantle mixin carry so sign_request_off_loop_if_aws
covers Mantle, and move the guardrail and embedding signing into
asyncio.to_thread. Every new test fails at the previous tip.
2026-09-08 16:59:34 -07:00
devin-ai-integration[bot]
075655c7ee
test(azure_sentinel): pin batch_size as a per-request bound under concurrent events (#40320)
* test(azure_sentinel): pin batch_size as a per-request bound under concurrent events

Adds a regression test to the mapped Azure Sentinel test file for the concurrency scenario from LIT-6920: 40 records logged concurrently at batch_size=5 while each ingestion request is still in flight. Asserts no request carries more than batch_size records, every record arrives exactly once in order, and the queue is empty afterwards. Runs for both the standard log queue and the audit log queue.

The test fails on the tree before #39880 (whole shared queue serialized per threshold send, then cleared) and passes on current staging. It is independent of the size-split coverage that #39880 added for LIT-5899.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_sentinel): gate the first send on events so later records provably arrive while it is in flight

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-08 16:55:40 -07:00