Commit graph

46263 commits

Author SHA1 Message Date
Mateo Wang
982d3a5476
Merge pull request #38496 from BerriAI/litellm_fix_exception_type_unbound_local
fix(exception_mapping_utils): map unmapped exceptions when model and provider are unset
2026-08-27 10:01:46 -07:00
Mateo Wang
86365263aa
Merge pull request #38484 from BerriAI/litellm_techdebt_20260827
refactor: clean up fresh tech debt from 2026-08-27 window
2026-08-27 09:50:07 -07:00
mateo-berri
a0689f04c4 fix(model_prices): cap ministral-3-3b at Mistral API's 131072 and mirror Anthropic family flags on new DeepInfra Claude rows 2026-08-27 09:48:19 -07:00
Mateo Wang
98d231c09b
Merge pull request #38398 from daniel-meismer-zocdoc/litellm_mcp_bearer_scheme_refresh
fix(mcp): canonicalize bearer scheme on bridge egress
2026-08-27 09:41:33 -07:00
samtsai15
d4904ad1e7 docs(guardrails): correct two comments that outran what the code guarantees
_ImageFetchBudget's class docstring still described claims as optimistic, taking
"the largest cap it could use" — the behaviour claim() had before the previous
commit made it all or nothing. Two docstrings in one class contradicting each
other is worse than either being absent.

_file_backed_image_count argued that structured_messages being narrowed by the
skip and scope flags is why it does not compare counts against inputs["images"].
That is true and it is half the picture: the same narrowing means a file source in
a message the scope excluded is not seen here at all. `images` is extracted from
every message while structured_messages holds only the scoped subset -- different
lists in guardrail_translation/handler.py, neither of which this PR changes. The
blind spot is now named, along with why reading the unscoped list would trade it
for a worse answer: refusing content the operator's skip flags deliberately took
out of scanning.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:30:06 +08:00
Mateo Wang
86ef1fb08b
Merge pull request #38391 from BerriAI/litellm_toggle_internal_health_check_logs
feat(ui): toggle internal health check visibility in request logs
2026-08-27 09:29:45 -07:00
samtsai15
ed991c4f6c test(guardrails): an oversized remote image under allow drops, it does not derail
The previous commit's all-or-nothing grant left one line uncovered: the return
after the PayloadTooLargeError arm. Under `block` that arm raises, so nothing was
reaching its fall-through.

The gap is worth a test rather than a shrug. An operator who sets
on_unscannable_image=allow asked for the image to go unscanned, not for the whole
request to die on a url that streams past the cap. The test asserts the text
alongside it still reaches the scan, and that the transfer was cut off rather than
read in full first.

The existing allow-policy oversize test uses an inline data URI, which is checked
after decoding and never touches this arm.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:20:16 +08:00
samtsai15
05488171b3 fix(guardrails): grant a whole image's budget or none, so the rejection reads true
The download budget handed out whatever was left. With a few hundred bytes
remaining that became the per-image cap, and a perfectly ordinary 50 KB image was
refused with

    remote image over ApplyGuardrail's 4 MB limit: remote body exceeds 100 bytes

which blames AWS for this request having spent its own budget, and contradicts
itself in the same sentence. This PR argues that a guardrail's failure modes have
to be legible; that message is not.

claim() is all or nothing now. "Over the per-image limit" and "this request is out
of download budget" stay separately diagnosable, at the cost of up to one image's
worth of the 80 MB going unused at the tail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:13:31 +08:00
Devin AI
5d26ae0fcd fix(model_prices): absorb Databricks/Z.AI and xAI registry PRs, add Together and Azure deprecation dates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-27 13:16:12 +00:00
Devin AI
4b3e82b8a3 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_bedrock_sol_anthropic_1hr 2026-08-27 13:03:14 +00:00
mateo-berri
abaaf8b210 fix(proxy): leave the deferred-logging flush parameter untyped
Typing logging_obj as the concrete Logging class makes basedpyright see the
_enqueue_deferred_logging reset as a protected access from outside the class,
pushing reportPrivateUsage one over the tree-wide budget. The tests pin that
attribute on a mock, so the reset has to stay a direct write on the passed
object. Restore the parameter annotation this file already had.
2026-08-27 11:18:01 +00:00
mateo-berri
5226401644 fix(proxy): keep the direct deferred-logging attribute write
The Any-reduction pass routed the deferred-logging teardown through a new
Logging.clear_deferred_logging_enqueue() helper. Callers in the request path
hand this function a MagicMock in tests, which absorbs the method call and
leaves _enqueue_deferred_logging set, so three deferred-guardrail-logging
assertions failed. Restore the direct attribute write and drop the helper.
2026-08-27 11:05:11 +00:00
mateo-berri
663aaa02a4 chore(lint): re-ratchet budgets against the new merge base
The merge brought 326 upstream commits, which moved every gate's base
count. Re-runs the ratchet once so the limits track the cleared headroom
at cd63c7e5a7 instead of the pre-merge base.

basedpyright -1955 across 48 rules, ruff-strict -310, LIT -106.
2026-08-27 10:46:47 +00:00
mateo-berri
2483a34dbc Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5_0826
# Conflicts:
#	basedpyright-code-budget.json
#	ruff-strict-budget.json
#	type-discipline-budget.json
2026-08-27 10:39:56 +00:00
mateo-berri
b74e696615 refactor(types): replace Any with real types across 178 backend files
Types provider request and response bodies at their boundaries with
TypedDicts and Protocols instead of dict[str, Any], so the untyped-to-typed
crossing is paid once per boundary rather than once per field read.

Removes 1,700 reportAny/reportExplicitAny errors and 1,953 basedpyright
errors overall, plus 310 ruff strict-rule and 106 LIT-rule violations.
No cast, type: ignore, noqa, or new Any annotations anywhere in the diff.

Ratchets the basedpyright, ruff-strict, and type-discipline budgets to the
new counts so the cleared headroom cannot silently grow back.
2026-08-27 10:14:58 +00:00
samtsai15
e54f69c48d test(guardrails): keep the file-image scan walking past an unreadable entry
Codecov left one line uncovered: the `continue` that skips a structured_messages
entry which is not a mapping.

Worth a test rather than a shrug. structured_messages comes from the caller, so
its shape is not guaranteed, and the loop has to keep going past an entry it
cannot read. Stopping or throwing there would let one junk element hide a file
image sitting after it, turning the guard into the bypass it exists to prevent --
so the test puts the file image last, behind a string, an int, and a message whose
content is not a list.

Added lines in bedrock_guardrails.py are back to full coverage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 17:54:49 +08:00
Devin AI
09ff5bf6cd Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_deflake_20260821 2026-08-27 09:17:37 +00:00
mateo-berri
ae95acfb05 fix(exception_mapping_utils): map unmapped exceptions when model and provider are unset 2026-08-27 02:08:18 -07:00
samtsai15
1930ae1634 refactor(guardrails): move the file-image refusal out of apply_guardrail
The previous commit added two branches to apply_guardrail and pushed it from 14 to
16 against ruff-strict's max-complexity of 15, taking the codebase C901 total from
312 to 313 with the budget at 312. It was one branch under the ceiling, so the
check had to live somewhere else.

Both conditions now sit in _refuse_file_backed_images: input_type, because images
are a request-side concern, and the count. apply_guardrail is back to 14 and the
call site reads as one statement.

Worth recording why this took two tries to see. `ruff check --select C901` uses
pyproject.toml, where max-complexity is 10, and apply_guardrail was already over
that both before and after -- so a before/after comparison showed no change and I
read the failure as a local artifact problem. The gate runs ruff with
--config ruff-strict.toml, where the ceiling is 15, and that is the only threshold
the budget is counted against.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 16:51:20 +08:00
samtsai15
4368765374 fix(guardrails): refuse a file-backed image instead of documenting that it passes
An Anthropic `{"type": "image", "source": {"type": "file", "file_id": ...}}` block
carries no bytes, so the extractor yields nothing for it while the provider still
forwards the file to the model. The docstring called that a known gap. Documented
is not the same as safe, and a silent pass is the exact failure this whole path
exists to remove, so it is now refused: blocked by default, allowed only where the
operator sets on_unscannable_image.

Resolving the file id would need a Files API client this guardrail does not have,
which is a larger change than the one this PR is making.

The check reads inputs["structured_messages"], which on /v1/messages carries the
raw Anthropic blocks, so the file source is visible to the guardrail without
altering GenericGuardrailAPIInputs. Putting a marker in inputs["images"] was the
alternative and would have reached five other guardrails that consume that field,
trading one gap for four new unknowns.

It detects the file shape specifically rather than comparing an image count against
inputs["images"]. structured_messages is already narrowed by the skip and scope
flags, so a mismatch is not by itself evidence of a dropped image, and refusing a
legitimate request would be worse than the gap being closed. A test pins that
base64 and url sources are still accepted.

Placed before the incremental and latest-message-only shortcuts, so a file-backed
image cannot be skipped by them either.

The first draft of the refusal test passed for the wrong reason: without the check
the request died on absent AWS credentials rather than on the bypass. The Bedrock
call is now stubbed, so removing the check makes it fail with DID NOT RAISE -- the
silent pass itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 16:40:39 +08:00
mateo-berri
bcb6a0a998 test(together_ai): assert fail-open supported params for models missing from the registry 2026-08-27 01:27:32 -07:00
samtsai15
c830ee204c fix: type the forwarded kwargs so basedpyright can check the httpx call
CI's basedpyright budget moved with the base and surfaced 12 new
reportArgumentType errors, ten of them on one line.

`request_kwargs` was typed `dict[str, object]` to keep an earlier gate happy, but
`object` does not unpack into httpx's typed parameters: `stream()` names ten of
them (content, data, files, params, headers, cookies, auth, follow_redirects,
timeout, extensions) and every one was an error. `dict[str, Any]` is what the
values actually are -- whatever `async_safe_get`'s caller passed through -- and it
lets the call be checked instead of merely tolerated.

The other two are the OUTPUT branch's `enumerate(filtered_messages)`. Making
skip_scan conditional on `image_urls` removed pyright's narrowing there, since
skip_scan used to be the only way that name could still be None. It cannot be None
at runtime -- images exist on the request side only, so a response scan returns
early exactly as before -- but the branch now says so with `or ()` rather than
resting on that indirection.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 16:26:31 +08:00
mateo-berri
b4c6e01fcc feat(together_ai): add zai-org/GLM-5.3-Flash to the model registry
Adds pricing (0.15/0.50 per 1M tokens, 0.03 cached read), the 1M context window, and capability flags (tools, parallel tools, tool choice, response schema, reasoning, vision) for Together AI's zai-org/GLM-5.3-Flash, mirrored into the backup cost map, with exact-value regression tests.
2026-08-27 01:12:12 -07:00
samtsai15
e84522118c style: satisfy LIT010 on the lines this PR touches
The base gained LIT010 (assignment without a Final declaration) while this PR was
open, so code that passed the gate at the old base now trips it.

Three are genuinely never rebound and take Final: `per_message`, `chunks`, and the
awaited fetch result, which is now bound separately so the unpack below has room
for its reason.

Two are real rebinding and say so. `total` is the running byte count the cap is
measured against. `mime_type` is overwritten further down when the caller passes
`format`, so it cannot be Final; the reason sits on the unpack line itself, since
a reason on the preceding line does not count.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 16:11:44 +08:00
samtsai15
5008f1f97e Merge litellm_internal_staging into fix/bedrock-guardrail-image-input
The base moved 478 commits while this PR was open. One conflict, in
test_litellm_core_utils_prompt_templates_factory.py: #38169 added
`from typing import Final` to the import block where this branch added
`import socket`. Both kept; the tests each side appended do not overlap.

Merged rather than rebased so the commit SHAs referenced in this PR's
description and review replies keep resolving.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 16:05:00 +08:00
Devin AI
63d7920f8b refactor: dedupe server_tool_use web search reads and type fresh test locals
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-27 08:01:52 +00:00
samtsai15
ef71da890e test: cover the capped fetch in the suites that own it
The capped path was only ever exercised through the Bedrock guardrail, so the two
shared files it lives in were the last uncovered lines on this PR. Both now have
tests where the code does, not only where its first caller does.

url_utils: _underlying_httpx_client's TypeError had no test at all. Every existing
one goes through AsyncHTTPHandler, whose `.client` is a real httpx client, so the
isinstance guard always held. It exists because `cast` is banned here, and without
it the declared return type would be a lie. Also covered: aborting mid-transfer,
asserting on how many bytes were pulled rather than only that it raised, and the
rebuilt response dropping content-encoding and content-length -- aiter_bytes
yields decoded bytes, so carrying those over would describe the body wrongly.

factory: get_image_details_async's body never ran under test. The model-path tests
stub the whole method out, and the guardrail tests fake the transport underneath
it. One test drives the real method with a cap; the other pins the reason the
previous commit exists, calling process_image_async with no cap against a stub
written to the old one-parameter signature.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 16:00:31 +08:00
samtsai15
fb5c886542 test(guardrails): pin the default-deny arms this PR relies on
The PR's position is that anything it cannot recognise must not reach the model
unscanned, but nothing asserted that. Every unrecognised shape landed on an
untested `return None`, so a refactor letting one fall through to the text branch
would have been reported to the operator as scanned content.

One parametrized test walks the shapes a caller can actually put on the wire --
absent content, content that is not a list, a part that is not a mapping, an image
part with no url, a url that is neither string nor mapping, a part carrying neither
image nor text -- and asserts none of them becomes a content item. Stating the
contract once beats restating each branch.

The rest cover paths that are behaviour rather than defence:

- `image_url` as a bare string, which OpenAI accepts alongside `{"url": ...}` and
  which reaches the model as an image either way
- a bare base64 payload whose format cannot be sniffed, left to
  on_unscannable_image instead of guessed at. It goes through apply_guardrail,
  the only caller that normalizes before decoding
- an oversized inline image under `allow`. Under `block` that branch raises and
  never returns, so this is the only path reaching its fall-through

_get_image_url and _retained_image_bytes are exercised directly. Both are reached
only through callers that already checked the shape, so their guards are otherwise
unreachable and would be dropped in a refactor without anything noticing.

Added lines in bedrock_guardrails.py now measure at full coverage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 15:30:23 +08:00
tin-berri
cd63c7e5a7
feat(ui): put the auto-router savings hero on a spend rail and a four-tile row (#38470)
The savings card carried four numbers in two stacked halves: the headline
saving with its delta on the left over the two spend rows, and avg saved per
session on the right. Give the headline the whole left half, move the two
spend rows into a rail on the right, and drop avg saved per session into the
metric row below as its first tile, with the session count as an inline hint.

Each spend row stays a description list so assistive tech keeps the label to
value association, with the shadcn Separator between the two rows. Both hero
columns are minmax(0,1fr) so a large total wraps instead of overflowing the
card, which also fixes the clipping the old 1fr columns already had. Metric
grows one optional hint slot so the new tile reuses the same presenter as its
three siblings.
2026-08-27 00:22:38 -07:00
yatishgoel
f0c336cccb fix(router): apply model renames to the in-memory deployment list 2026-08-27 12:42:41 +05:30
samtsai15
66cddac504 fix(guardrails): bound a request's image downloads in total, not just per image
Capping each image at 4 MB does not bound a request.
_create_bedrock_input_content_request gathers over every message and
_build_input_content_items then gathers over every part, so the fetches all start
together, and _MAX_IMAGES_PER_APPLY_GUARDRAIL_CALL is not consulted until
_bin_pack_bedrock_content runs on items that are already resident. 200 urls at
4 MB is 800 MB, chosen by the caller.

Two request-scoped bounds. A byte budget of _MAX_IMAGE_BYTES times
_MAX_IMAGES_PER_APPLY_GUARDRAIL_CALL, held for one content-request build and
passed down rather than kept on the guardrail, which is a callback instance shared
by every request. And a semaphore of 4, so the reserved bytes are not all in
flight at once.

Each fetch reserves a cap and refunds what a usable image did not take. A response
that decoded to nothing is charged in full: the transfer happened, and refunding it
would let a url serving megabytes of junk be repeated down the whole list for free,
which is the exhaustion being guarded against. An earlier draft refunded the whole
reservation and bounded only in-flight bytes; the regression test caught that the
decoded images still accumulated without limit.

Inline base64 draws on neither budget nor gate. Those bytes arrived in the request
body the proxy already accepted, so charging them to a download quota would refuse
inline images for no reason.

The total is what a single ApplyGuardrail call would accept anyway (20 images at
4 MB), so no request the API would take in one call is refused. A conversation
chunked across several calls can exceed it, and images past the budget are then
unscannable and left to on_unscannable_image, which blocks by default.

Without the budget the regression test fetches 209,715,200 bytes for one request.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 15:08:35 +08:00
samtsai15
355a444944 fix: forward max_bytes only when the caller set one
`process_image_async` passed `max_bytes=` on every call. The parameter has a
default, so the signature stayed compatible, but the *call* did not: an override
or test stub written against the previous signature now gets an unexpected
keyword and raises.

test_url_with_format_param caught it. It stubs `get_image_details_async` with a
one-parameter fake, so the model call died on TypeError and the provider mock it
asserts on was never reached.

Omitting the keyword when there is nothing to cap leaves the model-call image
paths byte-for-byte as they were, which is what the parameter being additive was
supposed to mean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 15:08:18 +08:00
yuneng-jiang
586e3d8de5
Merge branch 'litellm_internal_staging' into litellm_e2e_deflake_fallback_cache 2026-08-27 00:05:40 -07:00
yuneng-jiang
192ccaaf02
Merge pull request #38469 from BerriAI/litellm_e2e_gemini_chat_thinking_budget
fix(e2e): disable thinking on the gemini chat cost test instead of racing its budget
2026-08-27 00:02:59 -07:00
yuneng-jiang
2e2d68e869
Merge pull request #38468 from BerriAI/litellm_e2e_deflake_cacheable_prefix_size
fix(e2e): size the mid-conversation-system cache prefix above the minimum deterministically
2026-08-27 00:02:53 -07:00
tin-berri
81dc8dba1c
fix(ui): carry a preset's per-tier litellm_params through the prefill (#38453)
* fix(ui): carry a preset's per-tier litellm_params through the prefill

buildPresetPrefill rebuilt the complexity router config field by field and
never emitted tier_model_params, so a bundled preset that declares per-model
litellm_params (reasoning_effort, for instance) lost them before the create
form ever saw them. Both halves of the round trip already existed:
hydrateTierModelParams reads either storage shape, and serializeTierModelConfigs
writes them back on submit.

Hydrating alone is not enough. Tier entries get rewritten to the caller's
registered model spelling, which can differ from the preset's literal string by
version-separator punctuation, while the params stay keyed on what the preset
spelled. serializeTierModelConfigs then drops any param whose key is not in the
tier, silently. The param keys go through the same resolver as the tier entries.

* test(ui): catch a preset spelling the same model two ways in one tier

buildPresetPrefill resolves every model reference through normalizeModelName,
so two spellings of the same model in one tier (e.g. "claude-sonnet-4-5" and
"claude-sonnet-4.5") collapse to one key. For tier_model_configs that means one
model's litellm_params silently overwrites the other's - flagged by Greptile
on #38453 (P2, confirmed real via a throwaway repro, not a regression: on the
merge base both param sets were already dropped).

Nothing else validates preset authoring, and these are trusted, checked-in
JSON, so the fix is a static test over the bundled data rather than runtime
code. Exports normalizeModelName so the test exercises the actual resolution
rule instead of a hand-rolled copy of it. Verified the test fails when a
preset is mutated to spell one model two ways, and passes clean on the real
bundled presets.
2026-08-26 23:48:04 -07:00
mateo-berri
edde2e50ef fix(openai): drop tool_reference parts from tool messages at the chat boundary
OpenAI's chat completions API rejects tool_reference content parts in
role tool messages, so a mixed text plus reference tool result carried
through the Anthropic adapter turned a previously working request into
a 400 on chat-routed OpenAI and Azure deployments. Strip the reference
parts there, keeping a reference-only result as an empty-text tool
message so the preceding tool_call stays answered, mirroring the
Responses bridge skip.
2026-08-26 23:43:41 -07:00
yucheng-berri
8ebcb3e181
feat(newrelic): per-team cost and usage metrics via team callbacks (#37610)
* feat(newrelic): per-team cost and usage metrics via team callbacks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(newrelic): retry transient 429/408 metric posts instead of dropping

* fix(newrelic): drop only records queued when the drain began, not mid-drain arrivals

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-26 23:42:02 -07:00
Yuneng Jiang
0ec2d95506
fix(e2e): disable thinking on the gemini chat cost test instead of racing its budget
`test_gemini_chat_returns_content_and_logs_cost` asks gemini-2.5-flash to
"reply with the single word pong" under `max_tokens=32`, and has been seen
returning no content at all:

    completion_tokens=29, reasoning_tokens=29, content=None

gemini-2.5-flash defaults to dynamic thinking, and `max_tokens` maps to
`maxOutputTokens`, which on the 2.5 family counts thinking tokens as well as
visible output. So the model is free to spend the entire budget on thoughts and
emit nothing, which is exactly what the usage above shows.

Raising the limit alone does not fix this. Dynamic thinking on 2.5 Flash is
documented up to 24576 tokens, so no budget small enough to be reasonable for a
one-word smoke test is safe. The fix is to take thinking out of the picture:
`reasoning_effort="none"` maps to `thinkingConfig.thinkingBudget=0` for the 2.5
family, so the whole limit is available to visible output. Verified against this
checkout:

    get_optional_params(model="gemini-2.5-flash", custom_llm_provider="gemini",
                        max_tokens=32)
    -> {'max_output_tokens': 32}                       # no thinkingConfig at all

    get_optional_params(model="gemini-2.5-flash", custom_llm_provider="gemini",
                        max_tokens=64, reasoning_effort="none")
    -> {'max_output_tokens': 64,
        'thinkingConfig': {'thinkingBudget': 0, 'includeThoughts': False}}

This mirrors what the OpenAI tool tests in this same file already do with
gpt-5.6 for the same failure mode. `max_tokens` goes to 64 for headroom; with
thinking disabled that is ample for a one-word answer.

Neither `covers` claim changes: the call still exercises the gemini chat
translation path and still produces a costed SpendLogs row.
2026-08-26 23:38:43 -07:00
Yuneng Jiang
2e55fa1411
fix(e2e): size the mid-conversation-system cache prefix above the minimum deterministically
`_cacheable_system_block` embedded the per-run marker in all 300 paragraphs, so
the block's token count moved with the marker's own tokenization. Measured over
40 random markers the size ranged 3611-5408 tokens (median 4509): 15% of runs
landed under the 4096-token minimum cacheable prefix of Haiku 4.5, despite the
docstring claiming the prompt was comfortably above it.

When the system block is under the minimum, no cache entry is written at the
system breakpoint. The entry at the second breakpoint still gets written,
because system + first user turn clears the minimum -- which is why the failures
report a large cache_creation with cache_read stuck at 0
(`cache_creation_input_tokens=5610 cache_read_input_tokens=0`, and 5610 is the
whole prefix, not the user turn's share). `_prime_prompt_cache` rotates the user
turn on every attempt, so that second entry never prefix-matches the next
attempt either. Every attempt re-creates the full prefix, cache_read never rises
above 0, and the loop burns its 60s deadline:

    prompt cache never became readable in full within 60.0s

That is the single most frequent flake in the e2e suite, 9 of 38 runs, and it
hits all three provider classes identically because they share this helper.

Move the marker out of the repeated paragraph so it appears once, and size the
block at 1500 paragraphs. The prefix is now 8056-8060 tokens across markers --
spread 4 tokens instead of 1797, and 1.97x the minimum in the worst case. The
same marker-per-repetition pattern in `_first_turn_user_text` is fixed the same
way. Both copies of the helpers stay byte-identical.
2026-08-26 23:34:19 -07:00
yuneng-jiang
807ee7f232
Merge pull request #38454 from BerriAI/litellm_e2e_vertex_live_model
fix(e2e): move the vertex realtime suite off the retired Live preview model
2026-08-26 23:27:31 -07:00
mateo-berri
e26ea0bc95 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit6103_tool_reference_passthrough 2026-08-26 23:07:11 -07:00
Yuneng Jiang
8741e8a1ad
test(e2e): create the key under the dashboard session key
The test claiming mgmt.key.generate.happy_path signed in and then only read
/key/list, so nothing proved the session key an admin's sign-in mints is
actually accepted on /key/generate. It now does what an admin filling in
Create New Key does: POST /key/generate under the session key, read the new
key back from /key/info, see it in the dashboard's own /key/list, and drive
real traffic through it to confirm its model scope is enforced.

Adds ManagementClient.generate_key with the same caller_key seam update_key
and key_list already use, so the suite can call the route as the master key
or as a virtual key. Also wraps the over-long models import.

Refusing the dashboard session key on /key/generate turns only this test red;
the master-key generate, the key edit, and regenerate stay green.
2026-08-26 22:52:44 -07:00
tin-berri
166694948f
fix(ui): show custom technical keywords on every router whose scorer runs (#38451)
The keywords feed the scorer's technical dimension, so they change tier decisions
on any router that scores. The control rendered only for classifier_type
'heuristic', while the scoring knobs right below it already gated on
heuristicScoringRole(value) !== 'never'. The two disagreed, so an operator could
edit boundaries and weights on a router whose keywords they could neither see nor
set.

That hid the control on an LLM classifier using the default heuristic fallback,
and on heuristic_first, which runs the scorer on every request to decide whether
to short-circuit. Both now read the same predicate as the panel below them.
2026-08-26 22:37:37 -07:00
mateo-berri
99af9ad9eb fix(anthropic): carry tool_reference tool results through the guardrail translation round trip 2026-08-26 22:35:46 -07:00
mateo-berri
ab71807985 test(streaming): drop redundant docstrings from flex-tier regression tests 2026-08-26 22:23:17 -07:00
mateo-berri
f7228a4670 fix(streaming): preserve parsed-chunk provider_specific_fields so Vertex flex streams bill at flex rates 2026-08-26 21:47:15 -07:00
Mateo Wang
3b50819468
Merge pull request #38439 from BerriAI/litellm_messages_tool_usage_cost_header
Some checks are pending
Postgres Tests / proxy-behavior (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / misc (push) Waiting to run
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
fix(anthropic_adapter): carry web search cost into /v1/messages breakdown headers
2026-08-26 21:18:10 -07:00
mateo-berri
449c091391 fix(realtime): carry audio output tokens into response.done usage so Gemini Live native audio bills at the audio rate 2026-08-26 21:10:12 -07:00
Mateo Wang
aedaf4d0b0
Merge pull request #38434 from BerriAI/litellm_propagate_prompt_deletes
fix(prompts): propagate prompt deletes to every worker and pod
2026-08-26 21:03:00 -07:00