Commit graph

47905 commits

Author SHA1 Message Date
Joshua Valluru
7b6ef9206e test(mcp): cover server notifications during tool listing 2026-09-08 22:14:47 -07:00
Joshua Valluru
15392e7b3a fix(mcp): surface connection test failures safely 2026-09-08 21:26:04 -07:00
kerry
9a721abf0d test(cost-map): clear LITELLM_LOCAL_MODEL_COST_MAP in register_model url test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 03:50:49 +00:00
mateo-berri
c66ae07e3b fix(hosted_vllm): reject image edit params vLLM-Omni ignores 2026-09-08 20:11:01 -07:00
Mateo Wang
ee7c7e14f3
Merge pull request #40189 from BerriAI/litellm_lit_3157_azure_ai_catalog_models
fix(azure_ai): price seven Foundry catalog names and charge the model router fee once
2026-09-08 20:08:40 -07:00
mateo-berri
94f9230d13 fix(proxy): type the new pipeline tests and keep tag values out of the deferral warning
Every test this PR adds now annotates its fixture and parametrize
parameters. The submit-time warning for a tag-matched deferred policy
names only the policies, since a wildcard attachment pattern would let
caller-provided tag text reach the log.
2026-09-08 19:59:09 -07:00
kerry
4179f086e0 refactor(cost-map): share local-fallback and remote-accept paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 02:58:51 +00:00
tin-berri
902dd7b2b6
fix(mcp): log proxy tool dispatch exceptions (#40351) 2026-09-08 19:56:21 -07:00
kerry
536a85b429 fix(cost-map): keep register_model url fetch to a single attempt
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 02:48:17 +00:00
kerry
aa0a9ab3ea refactor(cost-map): drop initial_outcome flag from retry loop
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 02:29:56 +00:00
kerry
7d52832692 merge: litellm_internal_staging into litellm_cost_map_background_retries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 02:17:30 +00:00
Mateo Wang
24ef3ec63b
Merge pull request #37781 from ZXT-zjbiliy/fix/build-base-response-empty-choices
fix(stream_chunk_builder): guard empty choices and missing role in build_base_response
2026-09-08 19:14:03 -07:00
Mateo Wang
f8e456d105
Merge pull request #40275 from BerriAI/litellm_lit6852_spend_attribution
fix(spend-tracking): recover key alias for session tokens from spend logs
2026-09-08 19:05:43 -07:00
kerry
0086b62b45 fix(cost-map): keep first fetch blocking, run retries in background
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 02:05:08 +00:00
Yuneng Jiang
1638473332
fix(bedrock): keep deletion response IDs in request context 2026-09-08 19:02:08 -07:00
tin-berri
1a9c6ce390
fix(mcp): preserve proxy logging and authorization coverage (#40337)
* test(mcp): exercise /mcp/proxy authorization against the real registry instead of patched manager methods

* fix(mcp): preserve proxy logging and authorization coverage

* test(mcp): respect the proxy FastAPI import boundary
2026-09-09 02:01:01 +00:00
yuneng-jiang
0d62970865
Merge pull request #40334 from BerriAI/litellm_/release-version-bump-787548
chore: bump litellm-enterprise 0.1.65 -> 0.1.66
2026-09-08 18:53:51 -07:00
yuneng-jiang
86ee031217
Merge branch 'litellm_internal_staging' into litellm_/release-version-bump-787548 2026-09-08 18:44:54 -07:00
yuneng-jiang
1fbd1cb9ce
Merge pull request #40347 from BerriAI/litellm_fix_mcp_proxy_test_isolation
test(mcp): fix proxy fixture isolation after manager reload
2026-09-08 18:44:45 -07:00
Mateo Wang
d75aa4445d
Merge pull request #40179 from BerriAI/litellm_lit_2133_cost_map_provenance
feat(cost_map): report which revision of the price map the proxy is serving
2026-09-08 18:42:19 -07:00
Yuneng Jiang
f17632c036
chore: sync batch tests with latest staging 2026-09-08 18:36:15 -07:00
Yuneng Jiang
fb21852f7b
test(mcp): resolve current manager in proxy fixtures 2026-09-08 18:35:25 -07:00
tin-berri
314e573529
feat(auto-router): refresh family reasoning presets (#40341) 2026-09-08 18:28:46 -07:00
ryan-crabbe-berri
f5e4aa38ba
Merge pull request #40342 from BerriAI/litellm_prompt_cache_key_session_id
fix(anthropic): key the /v1/messages prompt cache on Claude Code's session_id only
2026-09-08 18:20:56 -07:00
ryan-crabbe-berri
634852a183 fix(anthropic): key the /v1/messages prompt cache on Claude Code's session_id only
The bridges derived prompt_cache_key as the first 64 chars of metadata.user_id.
Claude Code packs a JSON object into that field whose prefix is the per-install
device_id, so every session and subagent on one machine shared a single key,
and a plain end-user id pinned all of that user's conversations to one slot.

Parse the JSON and use session_id; send no key otherwise so the provider falls
back to its own prompt-prefix hashing. An explicit prompt_cache_key still wins.

Fixes #39145
2026-09-08 18:06:39 -07:00
yuneng-jiang
a99ecacffd
Merge pull request #40336 from BerriAI/litellm_extend_diskcache_deadline_oct1
chore(ci): extend diskcache scan exception to October 1
2026-09-08 18:05:00 -07:00
devin-ai-integration[bot]
43a1b2992a
fix(otel v2): restore the Datadog auth span and the last-wins callback merge (#40335)
* fix(otel v2): restore the Datadog auth span and the last-wins callback merge

Move @tracer.wrap() back onto user_api_key_auth so USE_DDTRACE=true emits the
auth span again, and let a failure entry's callback_vars take part in the
destination merge so the resolver picks the same account the runtime parser does

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel v2): drop docstrings from the two regression tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rerun proxy-infra after the flaky test_check_migration process-tree test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 01:00:26 +00:00
yuneng-jiang
9b186680c4
Merge branch 'litellm_internal_staging' into litellm_extend_diskcache_deadline_oct1 2026-09-08 17:55:38 -07:00
Yuneng Jiang
24ee66328a
test: bound Datadog read-back retries using reset headers 2026-09-08 17:55:21 -07:00
yuneng-jiang
c417e1084d
Merge branch 'litellm_internal_staging' into litellm_/release-version-bump-787548 2026-09-08 17:54:07 -07:00
yuneng-jiang
fb36c3c5a5
Merge pull request #40333 from BerriAI/litellm_fix_prisma_timeout_test_cleanup
test(proxy): fix Prisma timeout cleanup after subreaper tests
2026-09-08 17:53:56 -07:00
mateo-berri
529b8706ee test(spend-tracking): mock the spend-log scan through the transaction its caller now opens 2026-09-08 17:51:13 -07:00
Yuneng Jiang
810d48f28f
chore(ci): extend diskcache scan exception to October 1 2026-09-08 17:41:27 -07:00
Yuneng Jiang
94a81f003e
bump: litellm-enterprise 0.1.65 -> 0.1.66 2026-09-08 17:40:34 -07:00
Yuneng Jiang
5a7919f3f5
test(proxy): isolate reaper state and reap Prisma fixture children 2026-09-08 17:35:43 -07:00
tin-berri
754a2afe12
feat(mcp): add schema discovery proxy mode (#40298) 2026-09-09 00:30:20 +00:00
Mateo Wang
599daea985
Merge pull request #36718 from BerriAI/litellm_fix_count_tokens_budget_reservation_leak
fix(budget_reservation): don't reserve budget on token counting routes
2026-09-08 17:29:27 -07:00
moe-berri
6112274350
Merge pull request #40273 from BerriAI/litellm_non_reasoning_tier
feat(auto_router): opt-in NON_REASONING tier below SIMPLE
2026-09-08 17:27:35 -07:00
mateo-berri
268b944081 fix(spend-tracking): bound the spend-log scan with a statement timeout and name only unanimous alias, team, and owner 2026-09-08 17:24:22 -07:00
Mateo Wang
402351d980
Merge pull request #40268 from BerriAI/litellm_fireworks_responses_reasoning_instructions
fix(fireworks_ai): fold instructions and developer items into one leading system message on the Responses path
2026-09-08 17:16:52 -07:00
mateo-berri
260097b2dd refactor(hosted_vllm): drop redundant image edit docstrings 2026-09-08 17:16:34 -07:00
mateo-berri
831a2a13fb fix(azure): price azure_ai transcriptions at the azure_ai cost-map entry 2026-09-08 17:12:18 -07:00
mateo-berri
0c58346ba9 fix(proxy): warn when a deferred background policy was matched through a request tag
Retrieval re-matches only the key, team, and model scopes, so a post_call
policy that reached a pending background response through a request-body
tag does not govern the completed response. Log that at submit, next to the
deferral, and cover the retrieval re-match with tag-scoped tests.
2026-09-08 17:07:37 -07:00
Yuneng Jiang
b8be30219c
test: stop guardrail retries at the polling deadline 2026-09-08 17:02:00 -07:00
mateo-berri
88d2d77553 feat(hosted_vllm): add image edit support
Register a HostedVLLMImageEditConfig so hosted_vllm/<model> deployments route POST /v1/images/edits to the vLLM-Omni OpenAI-compatible endpoint instead of failing with 'image edit is not supported for hosted_vllm' before any request is sent
2026-09-08 16:59:55 -07:00
devin-ai-integration[bot]
075655c7ee
test(azure_sentinel): pin batch_size as a per-request bound under concurrent events (#40320)
* test(azure_sentinel): pin batch_size as a per-request bound under concurrent events

Adds a regression test to the mapped Azure Sentinel test file for the concurrency scenario from LIT-6920: 40 records logged concurrently at batch_size=5 while each ingestion request is still in flight. Asserts no request carries more than batch_size records, every record arrives exactly once in order, and the queue is empty afterwards. Runs for both the standard log queue and the audit log queue.

The test fails on the tree before #39880 (whole shared queue serialized per threshold send, then cleared) and passes on current staging. It is independent of the size-split coverage that #39880 added for LIT-5899.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_sentinel): gate the first send on events so later records provably arrive while it is in flight

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-08 16:55:40 -07:00
moe-berri
9d7e09e4e8 fix(ui): release the plan-mode floor when the non-reasoning tier is cleared
Turning the tier off, or switching to a classifier that cannot emit it, dropped
the flag and the pool but left plan_mode_min_tier naming a tier that is no longer
active. The backend rejects that on save, and the switch is disabled after a
classifier change, so the operator had no way to clear it.

Both paths now release the floor when it points at the cleared tier. An orphaned
keyword rule is left alone on purpose: getKeywordTierRulesError already names it
at the save gate, which is how a removed custom tier behaves.
2026-09-08 16:53:52 -07:00
Yuneng Jiang
253600fc61
test: wait for requested guardrail propagation 2026-09-08 16:52:56 -07:00
yuneng-jiang
54dc1d7644
Merge pull request #40323 from BerriAI/litellm_merge_main_into_staging
chore(ci): merge main into internal staging
2026-09-08 16:44:52 -07:00
mateo-berri
7a6c0cbf08 fix(spend-tracking): drop the owner of a digest shared by several users and back off failed scans 2026-09-08 16:43:28 -07:00