Commit graph

50471 commits

Author SHA1 Message Date
Yujong Lee
96baeb8b04 refactor(rust): remove gateway, config, router, realtime, and trace-parity infrastructure
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:00:07 +00:00
Yuneng Jiang
1972a30def
revert(e2e): unmount Gemini, its api_base means two things
Build 227 mounted Gemini and turned TestGeminiFiles::test_gemini_file_upload
red. litellm's two Gemini endpoints disagree about what api_base means.
Chat composes {api_base}/models/{model}:{endpoint} and defaults api_base to
https://generativelanguage.googleapis.com/v1beta, so the version lives
inside it. File upload composes {api_base}/upload/v1beta/files and defaults
to the host root, so the version lives outside it. A single api_base cannot
satisfy both, and a registration carries no signal about which endpoint the
deployment will be used for, so the edge cannot route one and not the other.

Backing it out rather than working around it. The cache must never turn a
passing test red, which is the same rule the Bedrock model allowlist
follows, and Gemini was 7 of roughly 1030 edge calls in that build. Anyone
pointing litellm's Gemini provider at an AI gateway or a corporate proxy
hits this too, so the fix belongs in litellm; mounting Gemini is one line
once it lands.

This reverts commit 8a553ceb58.
2026-09-16 08:37:25 -07:00
Yuneng Jiang
c447c3312d
feat(e2e): separate a provider error from a body that failed its rule
Build 226's 62 Bedrock rejections are the question this is trying to
answer, and "incomplete" would have covered both candidate causes at once.
Replaying the completeness rules over eight streams captured from live
Bedrock, covering tool use, extended thinking and a max-tokens stop on
both streaming endpoints, accepts every one of them, so a rule that is too
strict is the less likely half. A provider that answered 429 or 5xx and
was retried out of sight is the other, and it now counts as
rejected_error_status rather than being folded in with a grammar failure.
2026-09-16 08:13:04 -07:00
Joshua Valluru
fdb8e3533b fix(mcp): validate credentials in existing request paths 2026-09-16 08:07:10 -07:00
Joshua Valluru
87190604f6 chore: integrate current main for MCP security compatibility 2026-09-16 07:54:13 -07:00
Yuneng Jiang
7006da9cde
feat(e2e): say why a response was not recorded
Build 226 routed Bedrock streaming for the first time and rejected 62 of
220 misses on that mount, and the counters could not say why. A flat
rejected count covers three unrelated things with opposite fixes: the
consumer walking away mid-capture, a body that arrived whole and failed
its endpoint's rule, and a provider that could not be reached. Each now
also counts its own reason.

A consumer that walks away was counting nothing at all. Abandoning the
capture generator raises GeneratorExit at its yield, so neither branch of
the old accounting ran and the miss simply vanished from the report, which
is also why misses could exceed writes plus rejected with nothing to
explain the gap. The decision moves into settle() so the generator's
finally owns the accounting and an abandoned capture is counted like any
other rejection.
2026-09-16 07:47:31 -07:00
Yuneng Jiang
8a553ceb58
feat(e2e): mount Gemini on the provider cache
Gemini needs none of the machinery Bedrock needed. litellm composes
{api_base}/models/{model}:{endpoint} from a custom api_base, so a plain
path-prefixed mount reaches it, and the credential travels as a static
x-goog-api-key header that no host rewrite invalidates. Nothing is
re-signed and nothing leaves the cache key, so a recording still cannot
cross credentials.

A finished turn names a finishReason on every candidate and reports
usageMetadata. The reason is read as a string rather than compared to
STOP: MAX_TOKENS and the safety reasons end a turn just as finally, and
rejecting them would send every one of them upstream forever. Streaming
is the half worth care. Gemini repeats usageMetadata on every chunk and
names a finishReason only on the last, so the terminator is the final
event rather than any event, and a stream the connection cut short ends
on a chunk carrying usage and no reason.

The mount's upstream base carries the API version, so the path the rules
see is /v1beta/models/..., not the one the proxy sent. The first version
of this anchored the rule at the start of that path, which passed every
test against a stub with no version prefix and would have cached nothing
at all in a real run. Caught by replaying the rules over responses
captured from live gemini-2.5-flash, which is also why the tests now
mount their stub under the version prefix.

Vertex stays unmounted and is a separate provider here: litellm grafts
the default Vertex path onto an api_base only when that api_base has no
path of its own, so Vertex needs a root-mounted edge on its own port.
2026-09-16 07:38:44 -07:00
Devin AI
e64552867a refactor(guardrails): drop stale buffering comment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 14:21:36 +00:00
Devin AI
4d596082de feat(guardrails): release buffered stream chunks after each passing scan
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 14:20:32 +00:00
Devin AI
e0dd1350f4 fix(images): build the merged edit form in one comprehension
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
ai-gateway image / ai-gateway release image (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 13:48:26 +00:00
joshua-berri
9cd787386e
Merge pull request #41314 from BerriAI/litellm_fix_mcp_jwt_oauth_persistence
fix(mcp): authorize JWT OAuth credential persistence
2026-09-16 06:47:40 -07:00
Yuneng Jiang
30a691ed55
feat(e2e): cache Bedrock streaming responses
The Claude Code compat cells drive the real CLI, which always streams, so
converse-stream and invoke-with-response-stream were most of the suite's
Bedrock traffic and all of it bypassed the edge.

AWS frames those as binary vnd.amazon.eventstream rather than SSE, so
botocore's own parser reads the frames and validates both CRCs, and each
endpoint is then held to its terminal grammar. Two details drove the rule.
A ConverseStream ends with metadata, not with messageStop, and metadata is
what carries the token usage litellm prices the call from, so a stream cut
between the two names a stop reason but would replay as a free call. And a
dropped connection is invisible to the parser: it yields the frames it did
receive and silently discards a trailing partial one, so a stream cut one
byte short parses clean. The body is checked against the frame lengths it
declares to catch that.

The invoke stream carries the ordinary Anthropic event grammar inside its
chunk frames, so it shares the completeness rule with the SSE mounts.

Validated against three real Bedrock eventstream captures, and the tests
build their own frames rather than pasting a capture, with one test holding
that framing to botocore's parser.
2026-09-16 06:46:19 -07:00
Devin AI
a295f3f0c9 chore: merge main into litellm_fix_image_edits_bracketed_alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 13:33:05 +00:00
Devin AI
ba820783ac fix(tests): import Bedrock usage types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 13:25:14 +00:00
Devin AI
1674a3d767 fix(bedrock): never emit Converse cachePoint for OpenAI-family models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 13:21:17 +00:00
Devin AI
642525f571 fix(registry): add azure gpt-image-2.5 entries and together/azure deprecation dates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 13:19:53 +00:00
berriai-litellm-provider-info-sync[bot]
0443605c40
chore(prices): sync Together AI prices: 4 models, 4 deprecated
together_ai/google/gemma-4-31B-it: deprecation_date
together_ai/intfloat/multilingual-e5-large-instruct: deprecation_date
together_ai/openai/gpt-oss-20b: deprecation_date
together_ai/thinkingmachines/Inkling-Small: deprecation_date
2026-09-16 13:16:12 +00:00
Devin AI
4fd69b25e1 Merge remote-tracking branch 'origin/main' into litellm_registry_audit_2026_09_14 2026-09-16 13:05:28 +00:00
yassin
a8fff5b091 fix(content_filter): refuse a trim that splits a conditional word across the cut
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 12:45:58 +00:00
yassin
4fbe631146 test(content_filter): annotate streaming test locals as Final and type the logging metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 12:17:47 +00:00
Yuneng Jiang
bd1c2d6f07 fix(e2e): keep the tool-continuation echo-back test on the live path
The key normalizes a unique marker so two builds match, which is the whole
point, but it makes this test's identity collide with an earlier run's: it mints
a fresh receipt, sends it through a tool result, and asserts the model echoes it
back verbatim, so a stale recording matched and answered with the old receipt.
Build 223 is where that surfaced, once the corpus was full enough for the first
call to hit. A test that asserts a provider echoed this run's own unique value
belongs on the live path.
2026-09-16 05:09:04 -07:00
yassin
60642e875b perf(content_filter): back off refused streamed buffer cuts by one context length
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 11:51:06 +00:00
Yuneng Jiang
7c2234be3a test(e2e): enforce the cross-region invariant on the Bedrock allowlist
The allowlist rejects an unlisted model before the region resolver runs, so the
two negative cases that used to cover the resolver were passing for the wrong
reason and two mutations of it survived. Answering an env-referenced region with
the default mount is only sound because every allowlisted model is a `us.`
profile that fans out across the US regions, so assert that on the list itself
and drop the per-call branch it made unreachable.
2026-09-16 04:41:29 -07:00
yassin
7e429dee87 fix(content_filter): keep exception phrases and open conditional sentences in the streamed buffer
Trimming the streamed buffer to the retained tail could drop a category
exception phrase that suppresses a later keyword, or the identifier word
of an unfinished sentence that a conditional category pairs with a later
block word. Refuse the cut while either would leave the buffer so the
bounded scan masks and blocks exactly like a scan of the full text

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 11:18:10 +00:00
yassin
52f06906fe fix(content_filter): widen the streamed scan tail to the longest configured keyword
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 10:58:30 +00:00
Yuneng Jiang
ebf34cd880 docs(e2e): say plainly that Bedrock streaming is not cached 2026-09-16 03:35:00 -07:00
Yuneng Jiang
c7246adc1d fix(e2e): route only the Bedrock models the runner role can invoke
The edge re-signs with the run pod's identity, whose IAM policy is an explicit
per-model allowlist. Matching on the `anthropic.` infix instead routed every
Anthropic-on-Bedrock model, so a model outside the policy came back 403 from
Bedrock with no fallback, taking the whole claude_code Bedrock matrix red.
An unlisted model now keeps its direct path and loses only caching.
2026-09-16 03:29:02 -07:00
runjivu
4a70bc3ba3 fix: re-check budget on router fallback targets
Budget is enforced once during auth, against the requested model group.
`_is_model_cost_zero` waives every budget check for a zero-cost group, and the
router then picks a fallback target afterwards, inside `run_async_fallback`,
where nothing re-checks budget. A free model with a paid fallback therefore
bills with no budget gate at all.

Add `fallback_budget_check`, the budget sibling of the existing
`fallback_access_check`: a predicate awaited per fallback target that skips
targets the caller cannot pay for. The primary attempt is untouched, so a
zero-cost model is never blocked by budget and only the paid fallback is
refused.

Counter reads pass `max_budget` so `get_current_spend` verifies against
authoritative recorded spend, matching the auth-time key and user checks; a
counter restored from an older snapshot reads as a hit rather than a clean
miss, so without it a stale-low value would keep admitting paid fallbacks.

A zero-cost fallback target is always allowed, and a team key does not inherit
the key owner's personal budget unless `apply_user_budget_to_team_keys` is set,
matching `_PROXY_MaxBudgetLimiter`.

Scope is key and user budgets. Team, team-member, end-user, org, global and
per-model budgets are not covered yet: those auth-path functions enforce rather
than report, so reusing them would fire threshold alerts and take spend
reservations for a target that is then skipped. Two limitations of that scope
are documented in the module docstring: the check reads the spend counter
rather than reserving against it, so concurrent fallbacks can cross a cap
together; and a request reaching the router without
`metadata["user_api_key_auth"]` is not restricted. Both are shared with
`fallback_model_access.py`.

Opt-in via `general_settings.enforce_fallback_budget`.

Relates to #41344

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 19:06:38 +09:00
Yuneng Jiang
30c6241e3a
fix(e2e): a null error field is not an error
Every OpenAI Responses body carries `error: null` at the top level, and the
completeness check tested the key's presence rather than its value, so it
rejected every single one. The cost was silent: nothing failed, the endpoint
simply never cached, which is exactly the outcome the endpoint was added for.

Found by driving the edge against the real providers rather than the synthetic
fixtures, which carried no error key at all. Reading the value instead of the
key is also more accurate for chat completions and messages, where a real error
body carries a populated error object.
2026-09-16 03:01:07 -07:00
yassin
62ecb11ab9 perf(content_filter): scan a bounded window per streamed chunk
The streaming post-call hook rescanned the whole accumulated choice buffer on every chunk, so scan cost grew quadratically with output length. Keep a bounded per-choice buffer instead: once it exceeds twice the scan context, drop the head when masking the head and tail separately yields the same output as masking the whole buffer, so no pattern, phrase or exception straddles the cut. Detections from the dropped head are kept and merged, deduplicated, into the final log row

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 09:59:53 +00:00
yassin
6b7cafe92b refactor(proxy): share one typed increment for key spend and total_spend writes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 09:58:36 +00:00
Yuneng Jiang
aebfcf7da3
fix(e2e): route Bedrock deployments whose region only the proxy can resolve
Almost every Bedrock deployment in the suite declares
aws_region_name="os.environ/AWS_REGION". The mount resolver treated that
string as a region name, produced a mount nothing serves, and left the whole
Anthropic-on-Bedrock surface on its direct path, which is the one thing
mounting Bedrock was for.

The run pod does not share the proxy's environment, so the harness genuinely
cannot resolve that reference. A `us.` inference profile fans out across the US
regions and is reachable from any of them, so those route to the default mount
whatever the proxy resolved. A model that is not cross-region and declares its
region that way keeps its direct path rather than being sent to a region it may
not exist in.
2026-09-16 02:45:06 -07:00
yassin
6cf35ed71b feat(proxy): expose lifetime total_spend on virtual keys
Adds a persistent total_spend column to LiteLLM_VerificationToken and LiteLLM_DeletedVerificationToken, incremented in the same write as spend and left alone by budget resets. Surfaces it on /key/info, /key/list and the Admin UI Virtual Keys table and key detail view

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 09:42:20 +00:00
Yuneng Jiang
b68e60f706
feat(e2e): cache the responses and embeddings endpoints behind the edge
Chat completions and messages were the only cacheable paths. The suite also
drives /v1/embeddings and /v1/responses through the same OpenAI mount, so both
now cache, each with its own completeness rule: a chat response's `choices`
check would reject a perfectly good embedding, and a Responses run that never
reached `response.completed` must stay out of the cache the same way a
truncated stream does.

Vertex and Gemini stay off the edge. litellm's `_check_custom_proxy` rewrites a
path-prefixed vertex api_base into `{api_base}:{endpoint}`, dropping project,
location and model, so a mount under a path prefix cannot work without a
root-mounted edge on its own port or a change in litellm. Shipping an
unvalidated URL guess would have been worse than saying so in PROVIDER_CACHE.md.

Also finishes the MountPolicy move: a mount now carries its signer and its
unkeyed headers together instead of a bare signer map.
2026-09-16 02:34:09 -07:00
Yuneng Jiang
2d40254b57 feat(e2e): key the provider cache per test and mount Bedrock behind it
The exact-request cache reused 5% of routed traffic (build 218: 19 hits,
350 misses) because every test salts its prompt with a fresh unique_marker(),
so the same test could never match itself across builds. It also routed only
openai and anthropic, while the week's flakiness was Bedrock.

Key is now HMAC(test id + method + URL + headers + body, with every
unique_marker() token replaced by a placeholder, + FIFO slot index). The slot
index is what keeps two marker-only-different calls in one test on two
recordings and therefore two provider response ids, so spend rows still
reconcile one per invocation. A call outside any test is not cacheable.

Bedrock gets a region-qualified mount and SigV4 re-signing, since the edge
rewrites the Host the proxy signed. Signature headers are excluded from the
key for signing mounts only, because x-amz-date would otherwise make every
Bedrock request a permanent miss; every other mount still keys on its
credentials whole. Only Anthropic-on-Bedrock chat deployments route:
embeddings, image generation, rerank and realtime keep their direct path, and
so do deployments carrying their own aws_role_name or static keys, whose whole
point is to prove the product's assume-role chain rather than the runner's.
The two eventstream actions bypass the cache and go live, still signed.

Counters are now attributed per mount as well as in total, so a build can
report a per-provider hit rate instead of one number.
2026-09-16 02:15:46 -07:00
shivam
cd594f104a fix(otel): drop stale trace headers before injecting passthrough trace context
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 08:34:53 +00:00
yucheng
3628025aae fix(proxy): keep caller metadata.trace_id ahead of the OTel fallback on litellm_metadata routes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 08:30:54 +00:00
yucheng
9974cf4bf8 fix(proxy): honor key-level disable_fallbacks after first pre-call pass
Some checks failed
ai-gateway image / ai-gateway release image (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Key metadata disable_fallbacks only lands on data during add_key_level_controls,
so the local rate-limit fallback retry now rechecks it post pre-call. Also use a
real UserAPIKeyAuth in the skip pre-call test since the path reads router_settings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 08:17:58 +00:00
yucheng
b82de95625 refactor(passthrough): bind trace-enriched headers to a Final local
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 08:04:27 +00:00
yucheng
898fbd37a7 test(proxy): mark locals Final in the OTel trace id fallback tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 08:02:04 +00:00
yucheng
391da46e2c test: type the v3 limiter rig and otel key helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 07:59:33 +00:00
yucheng
4cb4493fa7 fix(proxy): let the OTel trace id fallback fill a null litellm_trace_id
A body that serializes litellm_trace_id as null or an empty string carries no identity, so it must not
block the server span fallback. Also mark the nested metadata write as an out-param store

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 07:48:58 +00:00
yucheng
49417d4fa2 fix(proxy): ignore non-span parent_otel_span when deriving litellm_trace_id
UserAPIKeyAuth.parent_otel_span is Any at runtime (opentelemetry is an optional extra), so the OTel
trace-id fallback must only format an int trace id, otherwise an object that merely quacks like a span
turns the whole request into a 500

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 07:41:48 +00:00
yucheng
8899583d06 docs(otel): tighten inject_trace_context docstring
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 07:33:39 +00:00
yucheng
05ededf8a0 fix(otel): parent passthrough trace propagation on the legacy request span
Pass user_api_key_dict.parent_otel_span into the outgoing W3C injection so the
legacy otel callback propagates its litellm_request span, falling back to the
otel_v2 request root span and then the ambient span. Extend the mapped unit
tests to assert the propagated trace and span ids over real captured headers
for HTTP and WebSocket passthrough with forwarding on and off.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 07:24:52 +00:00
yucheng
f1fd1c8996 fix(proxy): default litellm_trace_id to the OTel server span trace id
When the otel callback is enabled and the client sends no trace or session identity, the request now inherits the W3C trace id of the proxy's server span as litellm_trace_id and metadata.trace_id. The missing_session_id policy and SpendLogs then persist that value as session_id, so a trace in the OTel backend and its row in the Logs UI carry the same id. Explicit x-litellm-trace-id, traceparent, body metadata.trace_id and litellm_trace_id keep priority.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 07:22:03 +00:00
yucheng
9faaf7f4d4 refactor(proxy): assign the fallback model on the fresh snapshot instead of building a dict literal
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 07:15:13 +00:00
yucheng
2b6184d768 fix(proxy): resolve rate-limit fallbacks after model normalization and retry from a client-request snapshot
The fallback retry in _pre_call_with_fallbacks re-entered common_processing_pre_call_logic with data already enriched by the first pass, so add_litellm_data_to_request deep-copied a metadata dict holding the live OTel span and the request failed with a 500 (cannot pickle '_thread.RLock') instead of the intended 429 or fallback. Capture the configured fallbacks and a snapshot of the client request before the first pass, look up the fallback chain by the normalized model group after the limiter raises, and run each fallback attempt on a fresh copy of that snapshot. Replaces the mock-heavy tests with a rig that runs the real v3 limiter and a live OTel span through the proxy_logging_obj seam

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 07:08:24 +00:00
yucheng
2e699914e1 Merge remote-tracking branch 'origin/main' into litellm_lit_7470_rate_limit_fallback_pristine_data 2026-09-16 06:57:08 +00:00
shivam
7ba073aa26 test(otel): cover websocket trace propagation with forwarding on and off
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 06:55:06 +00:00