Commit graph

17245 commits

Author SHA1 Message Date
joshua-berri
d44e52a3cb
Merge pull request #40453 from BerriAI/litellm_fix_mcp_auth_challenge_6635
fix(mcp): challenge and scope gateway-owned server authentication
2026-09-09 21:51:33 -07:00
joshua-berri
17e9b2329e
Merge pull request #40454 from BerriAI/litellm_fix_mcp_debug_auth_resolution_6892
fix(mcp): report resolved upstream authentication in debug headers
2026-09-09 21:51:22 -07:00
Joshua Valluru
95f6e96ef9 test(mcp): type auth diagnostics regression parameters 2026-09-09 20:57:33 -07:00
Mateo Wang
69245feff4
Merge pull request #39863 from BerriAI/litellm_lit_7022_azure_ai_passthrough_config
fix(azure_ai): add passthrough config so router-model relays reach the deployment's own endpoint
2026-09-09 20:56:57 -07:00
mrinal
05e68fb7a7 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_health_check_db_storm 2026-09-10 03:48:34 +00:00
Joshua Valluru
a0bb021f02 fix(mcp): preserve gateway-owned resource scope through consent 2026-09-09 20:41:56 -07:00
Joshua Valluru
935611d25d fix(mcp): respect optional discovery capabilities and quiet unsupported methods 2026-09-09 20:21:19 -07:00
mateo-berri
b4919d9bd7 fix(bedrock): walk the output location on an unfiltered files list
Some checks failed
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
A list without a purpose covered the input bucket only, so a deployment
with a separate s3_output_bucket_name never saw its batch outputs unless
the caller passed purpose=batch_output. The listing now follows the input
location to its last page and then walks the output location whenever it
differs from the input one, in bucket or in prefix, so the unfiltered list
matches what OpenAI returns
2026-09-09 19:53:57 -07:00
joshua-berri
083ddefa92
Merge pull request #40498 from BerriAI/litellm_fix_mcp_edit_tool_preview_7135
fix(mcp): refresh tool previews when editing connection settings
2026-09-09 19:48:00 -07:00
mateo-berri
d71f4aeff9 fix(cost): layer deployment OCR rates over the cost map field by field 2026-09-09 19:47:00 -07:00
Mateo Wang
8ec2f00955
Merge pull request #40451 from BerriAI/litellm_responses_bridge_replay_encrypted_reasoning
fix(anthropic): replay OpenAI encrypted reasoning byte for byte through the /v1/messages bridge
2026-09-09 19:45:11 -07:00
mateo-berri
966a58d1c8 Merge branch 'litellm_internal_staging' into litellm_lit_7022_azure_ai_passthrough_config 2026-09-09 19:42:28 -07:00
Mateo Wang
f0fac55fe1
Merge pull request #40114 from BerriAI/litellm_fix_background_polling_disconnect_guard
fix(responses): keep background polling alive after the client disconnects
2026-09-09 19:41:19 -07:00
mateo-berri
a6681950e8 Merge branch 'litellm_internal_staging' into litellm_lit_7022_azure_ai_passthrough_config 2026-09-09 19:35:55 -07:00
Mateo Wang
c005431cad
Merge pull request #40489 from BerriAI/litellm_lite_claude_apikeyhelper_conflict
fix(cli): let the apiKeyHelper supply Claude Code's key under lite claude
2026-09-09 19:33:59 -07:00
Mateo Wang
26626348f8
Merge pull request #40270 from BerriAI/litellm_bedrock_sign_request_off_loop
fix(bedrock): sign requests off the event loop on every async path
2026-09-09 19:31:58 -07:00
Mateo Wang
261aa2f16f
Merge pull request #40186 from BerriAI/litellm_lit_5546_count_tokens_offload
fix(token_counter): release the GIL for HuggingFace counts and cap exact counting per string
2026-09-09 19:29:17 -07:00
mateo-berri
5a1be56426 fix(router): drop bridge-tagged reasoning blocks whole when the minting deployment is not in the routed group
The cross-group branch of EncryptedContentAffinityCheck removed only the signature from
Anthropic-shaped thinking blocks, which left unsigned thinking blocks that Anthropic and
Bedrock reject (thinking.signature: Field required). Drop the whole block, the way #40280
drops undecryptable Responses input items, so the routed request carries the conversation
text with no reasoning item for those turns
2026-09-09 19:27:54 -07:00
mateo-berri
942b647d42 merge: bring litellm_internal_staging into the Bedrock files delete and list fix
Converge on staging's delete plumbing (_S3DeleteContext read from the logging call's additional_args, _sign_s3_request_without_body, the credential-stripping delete_data in the managed-files hook) and keep this PR's listing support, the 400 mapping for out-of-bucket file ids, the proxy-admin-only raw cloud id rule, and the OpenAI FileDeleted delete response.

Two staging tests move to this PR's contract: an out-of-bucket delete raises BedrockError 400 instead of ValueError, and deleting a stored provider output returns FileDeleted rather than the stored file object.
2026-09-09 19:27:40 -07:00
Mateo Wang
9839bfcdbd
Merge pull request #40465 from BerriAI/litellm_e2e_spend_rows_join_virtual_key
test(e2e): every spend row a virtual key writes joins its token across all write paths
2026-09-09 19:25:41 -07:00
mateo-berri
763270e875 fix(cost): treat annotation-only deployment pricing as custom OCR pricing 2026-09-09 19:24:55 -07:00
mateo-berri
69c8c70547 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_background_polling_disconnect_guard 2026-09-09 19:21:47 -07:00
mateo-berri
50215c8717 test(files): cover the admin-only raw cloud id rule for Vertex GCS ids 2026-09-09 19:17:09 -07:00
mateo-berri
3dcc09b26c fix(bedrock): keep the signing pool to AWS configs on the async-transform path
The async-transform path in the shared handler sent every provider's sign_and_log to the AWS pool, so Ollama, Snowflake, and watsonx queued behind Bedrock refreshes there. Only SignsRequestsWithAWS configs take run_aws_signing now, the rest keep the default-executor hop they had. The executor isolation test also runs on its own loop instead of pinning a one-thread default executor on the session-scoped pytest loop
2026-09-09 19:16:11 -07:00
Mateo Wang
2611f6420a
Merge pull request #33409 from BerriAI/litellm_fix_bedrock_deepseek_thinking_leak
fix(bedrock): stop leaking Anthropic `thinking`/`reasoning_effort` into DeepSeek Converse requests
2026-09-09 19:13:16 -07:00
mateo-berri
433b436371 Merge remote-tracking branch 'origin/litellm_internal_staging' into pr36609_fork
# Conflicts:
#	litellm/cost_calculator.py
2026-09-09 19:13:15 -07:00
Mateo Wang
1d18b61e3e
Merge pull request #40074 from mihidumh/fix/mai-image-unsupported-params
fix(azure_ai): reject unsupported n and size params on MAI image models
2026-09-09 19:11:11 -07:00
mateo-berri
8627576c9c fix(azure_ai): keep a deployment named like the first native path segment in the relayed URL 2026-09-09 19:11:04 -07:00
mateo-berri
ab9dc75411 merge: origin/litellm_internal_staging into litellm_lite_claude_apikeyhelper_conflict
Resolves the conflicts with the lite configure claude work from #40319: every persistent
writer and reader of Claude Code's settings file now resolves it through CLAUDE_CONFIG_DIR,
the lite up backup check only guards the default file, and each settings file keeps its own
undo receipt (the default file keeps ~/.litellm/claude_configure_state.json, any other file
gets ~/.litellm/claude_configure_state/<sha256 of its resolved path>.json).
2026-09-09 19:09:52 -07:00
mateo-berri
45fad445eb fix(cost): bill request-level OCR pricing on direct SDK calls 2026-09-09 19:07:16 -07:00
mateo-berri
263c34fa26 Merge branch 'litellm_internal_staging' into litellm_fix_bedrock_deepseek_thinking_leak 2026-09-09 19:01:39 -07:00
mateo-berri
d93baa3f2c fix(bedrock): sign on a dedicated executor instead of the shared default one
asyncio.to_thread puts every Bedrock signing on the loop's default executor, the same pool every provider's async entry point hops through, so signings parked on botocore's refresh lock queued unrelated providers behind Bedrock. run_aws_signing runs them on a 16-thread pool only AWS signing uses
2026-09-09 18:52:00 -07:00
Mateo Wang
db7ca65b69
Merge pull request #40462 from BerriAI/litellm_fix_responses_stream_named_tool_choice
fix(responses): echo a named tool_choice in the Responses API shape on the chat-completions bridge
2026-09-09 18:48:32 -07:00
Mateo Wang
aefa1040a0
Merge pull request #40249 from BerriAI/litellm_fix_responses_bridge_reasoning_effort
fix: keep reasoning_effort for mode: responses bridge deployments
2026-09-09 18:46:58 -07:00
mateo-berri
36c1e5e17d refactor(compression): build the protected index set without mutation 2026-09-09 18:45:22 -07:00
mateo-berri
71cf1c7901 test(e2e): drop the unused batch list client and fail loudly on a user teardown miss
The failed-batch redesign left list_batches, BatchList, BatchListQuery and
the batch object's metadata and created_at fields with no caller, and they
duplicated the batches suite's own client. delete_user discarded its result,
so a user that outlived the class fixture went unnoticed; unwrap turns that
into a teardown error like delete_key already does.
2026-09-09 18:34:15 -07:00
mateo-berri
708381948b fix(cli): compare the resolved settings path when deciding whether the lite up and autoroute backups guard it 2026-09-09 18:34:07 -07:00
mateo-berri
a25eccafc9 fix(azure_ai): drop the api_base path segments a relay already repeats 2026-09-09 18:32:13 -07:00
mateo-berri
4716b46c24 fix(router): read the affinity pin from the messages argument and strip bridge reasoning on a cross-group route
The encrypted_content_affinity check only read the Anthropic history from request_kwargs["messages"], so a caller that passes it through the callback's messages argument alone skipped the pin. Read the argument first and fall back to the kwargs.

When the minting deployment is not a candidate of the routed group, the base already strips the Responses input's encrypted reasoning; do the same for the bridge-tagged thinking blocks in Anthropic messages so the routed deployment gets the readable thinking text instead of ciphertext it cannot decrypt.
2026-09-09 18:31:08 -07:00
mateo-berri
4adf99557f fix(responses): echo "auto" for a tool_choice the bridge cannot express instead of failing after the provider call 2026-09-09 18:27:12 -07:00
mateo-berri
4d4906e94e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_responses_bridge_replay_encrypted_reasoning
# Conflicts:
#	litellm/router_utils/pre_call_checks/encrypted_content_affinity_check.py
#	tests/test_litellm/router_utils/pre_call_checks/test_encrypted_content_affinity_check.py
2026-09-09 18:23:07 -07:00
Joshua Valluru
3169c80252 fix(mcp): preserve IPv6 hosts when comparing preview origins 2026-09-09 18:21:55 -07:00
mateo-berri
b3f2de058b fix(cli): only guard the default settings.json with the lite up and autoroute backups on login --config-claude 2026-09-09 18:19:56 -07:00
Mateo Wang
b22ca7ac6d
Merge pull request #40464 from BerriAI/litellm_internal_copy_35771
fix(azure): respect DEFAULT_MAX_RETRIES in initialize_azure_sdk_client (internal copy of #35771)
2026-09-09 18:19:08 -07:00
mateo-berri
c1ca963d75 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit_5546_count_tokens_offload 2026-09-09 18:18:17 -07:00
mateo-berri
47eb7f4349 fix(token_counter): give each event loop its own count limiter
A single process-wide anyio.CapacityLimiter wakes waiters with the
asyncio.Event of whichever loop created it, so a fifth loop in another
thread (asyncio.run per request under Celery, gunicorn sync workers, or
run_async_function from function_with_fallbacks) waited forever once four
counts were in flight, and building it at import time raised
AsyncLibraryNotFoundError on anyio below 4.2. The limiter now lives in a
RunVar and is created on first use inside the running loop.

TOKEN_COUNTER_MAX_CONCURRENT_COUNTS reads from the environment like
TOKEN_COUNTER_MAX_EXACT_CHARS, and the proxy's tokenizer lookup runs on the
shared thread pool so a Hub download no longer holds a count slot.
2026-09-09 18:17:32 -07:00
mateo-berri
76c9342e9d fix(proxy): run migrations through python -m prisma when the prisma console script is not on PATH
The proxy probed for the Prisma CLI by spawning the bare `prisma` console
script, so a launcher whose PATH lacked the interpreter's bin directory
printed "prisma package not found", skipped every migration and served
traffic against an empty schema. Every Prisma command now falls back to
`python -m prisma` when the console script is not on PATH, and the boot
probe checks for the console script or the importable package instead of
spawning anything.
2026-09-09 18:17:12 -07:00
Mateo Wang
21ae29b759
Merge pull request #40284 from BerriAI/litellm_legacy_hook_streaming_pipeline_step
feat(guardrails): run legacy post-call hooks as streaming pipeline steps
2026-09-09 18:14:37 -07:00
Mateo Wang
53f3e70f02
Merge pull request #40461 from BerriAI/litellm_lit_7373_responses_custom_tool_call_guardrail
fix(guardrails): scan and rewrite Responses custom_tool_call output items
2026-09-09 18:13:00 -07:00
Joshua Valluru
493c98f9bf fix(mcp): normalize default ports in preview origins 2026-09-09 18:11:52 -07:00