Commit graph

18968 commits

Author SHA1 Message Date
yuneng
3f4fe7db82 docs(tests): define the tier contract for unit, integration and e2e
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 06:24:06 +00:00
yassin
92d3a1d87d test(proxy): expect 422 for per-model budget rejections on cursor route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 06:10:16 +00:00
yassin
f9244749e0 fix(proxy): return 422 instead of 429 for BudgetExceededError
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 06:00:06 +00:00
kerry
9f24699e4c fix(fal_ai): reject empty image lists in image edit requests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 05:44:56 +00:00
mateo-berri
468f74c628 ci(e2e): fix the stage-mirror batch reds and keep a redacted pytest log
The changed-test gate booted its stage-mirror stack without files_settings
or finetune_settings, so every raw upload with a custom_llm_provider hit a
500, and it exported the whole provider env into the gateways, so the
AWS_ROLE_NAME the assume-role test needs made the GovCloud deployment run
an AssumeRole with its static keys. The gate also deleted its pytest output,
so a red run left nothing to read. The mirror config now carries the
openai, azure, and vertex_ai file settings, gateways start without
AWS_ROLE_NAME, and the workflow uploads the pass logs and junit files with
every secret value, every field of a JSON-valued secret, and their
XML-escaped forms replaced before the raw files are removed.
2026-09-19 22:37:07 -07:00
kerry
cc7dce6a21 fix(fal_ai): accept every FileTypes image input and derive gpt-image qualities from pricing rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 05:33:06 +00:00
kerry
62c215be19 fix(fal_ai): accept every FileTypes image input and derive gpt-image qualities from pricing metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 05:31:27 +00:00
kerry
365dc9a3b5 feat(fal_ai): add gpt-image-2.5 flare/sunburst, flux/dev and image edits
Route openai/gpt-image-2.5/{flare,sunburst}/text-to-image through the existing GPT Image config with the xhigh and max quality tiers, add a dedicated fal-ai/flux/dev config, and add a Fal image-edit config so /v1/images/edits works for the gpt-image-2.5 and gpt-image-2 edit endpoints. Add flat and quality-by-size keyed pricing rows so spend is non-zero

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 05:14:15 +00:00
Mateo Wang
58065d46fd
Merge pull request #42071 from BerriAI/litellm_remove_dead_telemetry_flag 2026-09-19 21:48:02 -07:00
mateo-berri
0c68c58eb1 test(proxy): expect bedrock_mantle in the anthropic header provider list
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
2026-09-19 20:39:53 -07:00
Mateo Wang
4011367b39
Merge pull request #40986 from BerriAI/litellm_lit_7346_multi_choice_stream_guardrails
fix(guardrails): scan each choice's tool-call arguments apart on n>1 streams and log why a rewrite was discarded
2026-09-19 20:34:05 -07:00
Mateo Wang
9d7f77988a
Merge pull request #41974 from BerriAI/litellm_fix_startup_view_creation_race
fix(proxy): wait for the spend-log table before creating startup views
2026-09-19 20:33:49 -07:00
mateo-berri
3ffe6272c9 fix(router): hash the prompt caching affinity prefix off the event loop
Offload the per-block hashing through offload_token_count on both the pre-call
read and the success-event write, hash raw bytes as base64 instead of raising,
drop the unused serialize_object helper, and bind the chained digest, the
message envelope, and the bytes path in the regression tests
2026-09-19 20:22:28 -07:00
mateo-berri
19d77e2442 test(guardrails): type the recorder hook's request_data as a Mapping 2026-09-19 20:21:37 -07:00
mateo-berri
f24208f9ca fix(bedrock_mantle): price region-prefixed Claude responses from the bare Bedrock row 2026-09-19 20:17:55 -07:00
kerry
66078f4834 Merge remote-tracking branch 'origin/main' into litellm_azure_ai_mai_image_2_5_pro
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 03:15:05 +00:00
mateo
b2e123da43 fix(proxy): drop legacy telemetry key from persisted WORKER_CONFIG before initialize
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 03:14:11 +00:00
mateo-berri
0f0c0fe499 fix: drop a blank anthropic-beta header before it reaches the provider 2026-09-19 20:13:59 -07:00
mateo-berri
c13dcb0abf fix(proxy): forward a client's anthropic-beta and anthropic-version headers to bedrock_mantle 2026-09-19 20:11:14 -07:00
Mateo Wang
ef7da9b49f
Merge pull request #42072 from BerriAI/litellm_mcp_cold_worker_tools_call
fix(mcp): tools/call no longer 404s on a worker that has not served tools/list
2026-09-19 20:04:33 -07:00
mateo-berri
327447bc10 Merge remote-tracking branch 'origin/main' into litellm_fix_startup_view_creation_race
# Conflicts:
#	tests/test_litellm/proxy/test_proxy_server.py
2026-09-19 20:02:33 -07:00
mateo-berri
517fff5bb7 fix(router): keep prompt caching affinity when the breakpoint moves
The prompt_caching pre-call check keyed a deployment pin on a hash of the
whole cacheable prefix, cache_control markers included. Agent clients
such as Claude Code move the marker to the newest user turn on every
request, so the key changed every turn, the pin never matched, and a
multi-turn session drifted across deployments and lost its provider
cache.

Hash the prefix per content block with the markers stripped, chained so
every block position has a key, and write the pin at the breakpoint
block. Lookup walks back over the last PROMPT_CACHE_LOOKBACK_POSITIONS
positions (a run of tool_use or tool_result blocks counting as one), the
same window the provider probes for a cached prefix, in one batch cache
read. Both sides hash the prefix after base64 truncation so a request
carrying raw image bytes derives the keys the success event stored.
2026-09-19 19:52:22 -07:00
Joshua Valluru
4fda0092d3 fix(mcp): explain missing public client dependencies 2026-09-19 19:52:13 -07:00
joshua-berri
8df260a13d
Merge pull request #42051 from BerriAI/litellm_mcp_oauth_e2e_3467_rework
test(e2e): restore MCP OAuth happy-path coverage (LIT-3467)
2026-09-20 02:50:14 +00:00
Joshua Valluru
124196cbaa fix(auth): preserve scope-admin email policy during status lookup 2026-09-19 19:48:53 -07:00
mateo
fdd91a347a test: drop narration comment from telemetry flag test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 02:44:52 +00:00
Mateo Wang
b4447096e4
Merge pull request #42067 from BerriAI/litellm_genai_adapter_response_schema_tool_params
fix(google_genai): forward response schema and tool parameters through the generateContent adapter
2026-09-19 19:43:32 -07:00
Mateo Wang
79e25d1e96
Merge pull request #42045 from BerriAI/litellm_lit_8201_notfound_retry_policy
fix(router): add NotFoundErrorRetries so a retry policy can pin 404 retries
2026-09-19 19:37:45 -07:00
mateo
8e530cf819 fix: keep --telemetry as a hidden no-op so existing start commands still parse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 02:36:59 +00:00
Joshua Valluru
6ea74d70f1 fix(auth): enforce SCIM status for admin JWTs and refresh SCIM caches 2026-09-19 19:33:47 -07:00
mateo-berri
a3e9ed34fe fix(mcp): gate the pre-call listing per tool, not per server
A tools/call on a cold worker listed the target server once and then never
again, so a later caller whose credentials expose a wider upstream catalog
got 404 for tools the first caller never had. Gate the pre-call listing on
whether this worker already exposes the requested tool, so callers with
different catalogs no longer mask each other. Removing the per-server guard
also drops the empty-listing case that re-listed on every call.
2026-09-19 19:29:33 -07:00
yuneng-jiang
09e14a485c
Merge pull request #42061 from BerriAI/litellm_fix_gcs_pub_sub_autorouter_golden
test(logging): add autorouter estimate keys to the GCS pub/sub spend-log golden
2026-09-19 19:22:41 -07:00
mateo-berri
5eb967d925 fix(google_genai): drop non-object tool parameters instead of forwarding them 2026-09-19 19:19:33 -07:00
Mateo Wang
5417abd586
Merge pull request #42011 from BerriAI/litellm_scrub_default_master_key
docs: stop advertising sk-1234 as the master key in shipped configs and examples
2026-09-19 19:05:31 -07:00
Mateo Wang
b6dd3d932c
Merge pull request #42019 from BerriAI/litellm_master_key_boot_enforcement
feat(proxy)!: refuse to start with an unset, empty, or publicly known master key
2026-09-19 19:04:02 -07:00
mateo-berri
325d17aca9 fix(litellm): keep a function tool without a body on the chat route
A tools entry of only {"type": "function"} has nothing for the Responses
bridge to convert, and the bridge raised a 500 for it where the chat
route returns the provider's own 400. The gate now counts a tool as a
function tool only when it carries a function body or a top-level name,
on every provider the gate serves
2026-09-19 18:58:08 -07:00
ryan-crabbe-berri
ecf17513fb refactor(proxy): rename the local development override to dangerously_permit_weak_or_unset_master_key so the name says exactly what it permits 2026-09-19 18:53:14 -07:00
mateo-berri
2e83871d54 test(guardrails): type the recorder hook's logging_obj as object 2026-09-19 18:52:28 -07:00
mateo-berri
47ebfa10a0 fix(google_genai): reuse the shared key filter for Gemini-only schema keys 2026-09-19 18:52:00 -07:00
mateo-berri
92ff54f134 fix(mcp): list a never-listed server before its first tools/call
The startup tool-name fill skips servers whose upstream wants the caller's
own token (true_passthrough, OAuth discovery), and mcp 2 no longer runs the
list handler before an uncached tools/call, so every uvicorn worker that had
not served tools/list answered 404 "Tool not found" for prefixed tools/call
and the REST server_id route on those servers.

On a resolution miss, execute_mcp_tool now lists the prefix-matched (or
server_id-requested) server once, with the caller's credentials, through the
existing tools/list path, then resolves as before. Listing failures fall
through to the existing 404, a worker that already listed the server never
re-lists it, and a server outside the caller's allowed set is never listed.
2026-09-19 18:49:26 -07:00
mateo-berri
827d1c99a0 test: type the cache hook test helpers
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
2026-09-19 18:48:39 -07:00
mateo
f820472488 chore: remove the dead telemetry flag from the SDK, proxy CLI and configs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 01:44:10 +00:00
mateo-berri
3772993032 fix(anthropic_messages): only Mantle consumes get_llm_provider's api_base
The /v1/messages handler passed the api_base get_llm_provider resolved to every
provider's native messages config, which shadowed DEEPSEEK_ANTHROPIC_API_BASE and
TENCENT_ANTHROPIC_API_BASE with the chat default and changed the azure_ai
precedence. Messages configs now opt in through uses_get_llm_provider_api_base(),
true only for Bedrock Mantle, whose region-prefixed model must resolve to a
region host before the prefix is stripped. Also registers
BedrockMantleAnthropicMessagesConfig in the lazy import registry.
2026-09-19 18:42:05 -07:00
mateo-berri
2c3fc4cbff test: drop narrating docstrings and wrap long lines in the cache hook tests 2026-09-19 18:34:23 -07:00
mateo-berri
875f015e24 fix(token_counter): count replayed redacted_thinking blocks so prompt_caching keeps pinning
A conversation that replays a redacted_thinking block (Anthropic redacted reasoning, or the
/v1/messages bridge's stand-in for a reasoning item that carries no summary) made
_count_content_list raise, is_prompt_caching_valid_prompt swallowed that to False, and the
prompt_caching pre-call check neither recorded nor pinned the serving deployment, so the
conversation bounced across the group and paid a cache write on every deployment. The block
now counts like a thinking block with no text: zero tokens for the encrypted payload.
2026-09-19 18:33:56 -07:00
mateo-berri
fba179f2c0 fix(google_genai): forward response schema and tool parameters through the generateContent adapter 2026-09-19 18:31:09 -07:00
joshua-berri
daecea3eb8
Merge pull request #42050 from BerriAI/litellm_mcp_scoped_regressions_4506_rework
test(mcp): restore scoped execution and credential isolation regressions
2026-09-20 01:29:44 +00:00
Tin Chi Lo
94b2fd827b feat(ui): show prompt caching requests and net savings 2026-09-19 18:27:38 -07:00
Mateo Wang
93e39d5042
Merge pull request #42062 from BerriAI/litellm_pr38499_batch_retrieve_model_group
fix(router): stamp model_group when retrieving a batch, so batch tokens are attributable (internal copy of #38499)
2026-09-19 18:23:06 -07:00
mateo-berri
b0971ee0ba fix: count extra_body tools and cache_control in place of the direct ones 2026-09-19 18:22:41 -07:00