Commit graph

44402 commits

Author SHA1 Message Date
mateo-berri
a631bf730a fix(anthropic): forward Claude Code safeguards and dangerous-tool-use beta to Bedrock Invoke and Vertex on /v1/messages
Backport of #42288 to stable/1.99.x.
Cherry-picked from merge commit fc82f6e8fa (litellm_safeguards_bedrock_vertex_messages).
The line has no bedrock_mantle beta-header mapping and no Mantle /v1/messages route, so the Mantle mapping, its test file, and the bedrock_mantle test parameter are left out.
2026-09-22 10:16:54 -07:00
Yassin Kortam
9f6a6f89f1 fix(anthropic): forward safeguards and anthropic-beta unchanged on native /v1/messages
Backport of #42152 to stable/1.99.x.
Cherry-picked from merge commit e912ebe999 (litellm_claude_code_safeguards_passthrough).
2026-09-22 10:16:25 -07:00
yuneng-jiang
3c08f26b6b
Merge pull request #41323 from BerriAI/litellm_backport_1_99_x_41230_0915
chore(release): backport #41230 to stable/1.99.x and cut 1.99.2
2026-09-15 17:26:38 -07:00
Yuneng Jiang
d8eae9d924
chore: refresh uv.lock for 1.99.2 2026-09-15 16:03:38 -07:00
Yuneng Jiang
d29c63927a
bump: version 1.99.1 → 1.99.2 2026-09-15 16:03:19 -07:00
Yuneng Jiang
325cd75ace
chore(deps): bump pypdf to 6.16.1 2026-09-15 16:02:11 -07:00
Yuneng Jiang
40c50fbe0f
chore(deps): bump gitpython to 3.1.59 2026-09-15 16:02:11 -07:00
Yuneng Jiang
6befc61f33
chore(deps): bump tornado to 6.5.8 2026-09-15 16:02:05 -07:00
Yassin Kortam
1f8d1dc885
Merge pull request #41230 from BerriAI/litellm_client_timeout_408_skips_cooldown
fix(router): stop counting caller-set timeout 408s toward deployment cooldown

(cherry picked from commit 140229bc4a)
2026-09-15 15:59:20 -07:00
yuneng-jiang
10f4033437
Merge pull request #39179 from BerriAI/litellm_backport_1_99_x_otelcache
chore(release): backport #38716 to stable/1.99.x and cut 1.99.1
2026-09-01 14:18:02 -07:00
Yuneng Jiang
3ebb415b54
chore: refresh uv.lock for 1.99.1 2026-09-01 12:56:04 -07:00
Yuneng Jiang
d72131846b
bump: version 1.99.0 → 1.99.1 2026-09-01 12:55:43 -07:00
Yuneng Jiang
ddc4d8f9d9
chore(deps): bump restrictedpython to 8.5 2026-09-01 12:51:06 -07:00
devin-ai-integration[bot]
ad75134fce
fix(otel): emit cache token counts on OTel v2 LLM spans (#38716)
* fix(otel): emit cache token counts on OTel v2 LLM spans

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): trim comment in LLMUsage adapter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): drop casts in LLMUsage cache token adapter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(deps): bump restrictedpython to 8.3 for GHSA-ffg3-p8fm-mjx2

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 40edeaaecb)
2026-09-01 12:50:35 -07:00
yuneng-jiang
fa647f742d
Merge pull request #39048 from BerriAI/litellm_/docker-build-failure-9f6247
fix(docker): pin apk python to 3.13 on rc/1.99.0 (cherry-pick #38917)
2026-08-31 18:27:11 -07:00
Mateo Wang
5892e30d42
fix(docker): pin apk python to 3.13 in migrations image
Cherry-pick of #38973 onto rc/1.99.0, same failure mode as #38917:
Wolfi's python3 roll to 3.14 breaks the uvloop 0.21.0 sdist build.

(cherry picked from commit 0f84a7053a)
2026-08-31 18:20:07 -07:00
Mateo Wang
91f1a57100
fix(docker): bump wolfi-base for glibc 2.44 and pin apk python to 3.13
Cherry-pick of #38917 onto rc/1.99.0 so the release Docker builds stop
picking up Wolfi's python 3.14 roll, which has no uvloop 0.21.0 wheel
and fails the sdist build.

(cherry picked from commit f814945d4c)
2026-08-31 18:12:18 -07:00
yuneng-jiang
d0c86678ed
Merge pull request #38862 from BerriAI/litellm_/litellm-e2e-rc-1-99-0-11cfd8
test(e2e): backport the select-anchoring and router-fallback spec de-flakes
2026-08-29 19:38:20 -07:00
Yuneng Jiang
da1d292dce
test(e2e): measure the select popup after it settles instead of mid-flight
Both anchoring tests read the trigger's box before the click and the popup's box the
instant it turns visible. Base UI places the popup asynchronously and opening it can
shift the trigger, so both boxes could be sampled before the layout settled. The run
on 1eedaa3a43 missed by 4.2px (expected >= 446.015, got 441.799) on a tree with no UI
changes at all, having passed on 21092d633b, which differs only in a deleted python
test and a budget json.

Each assertion now re-reads both boxes under expect.poll. The conditions themselves
are unchanged: the popup must sit at or below the trigger's bottom edge in the first
test and must not overlap it in the second. Polling cannot mask a genuinely misplaced
popup, since one that never lands correctly still fails when the poll times out.

(cherry picked from commit d8679508d4)
2026-08-29 19:11:27 -07:00
Yuneng Jiang
a2aecfea7f
fix(e2e): only the negative fallback assertion needs every replica to agree
litellm-e2e-ui 68 failed the test this PR was meant to stabilise: "fallback
never took effect", streak 4 of a required 5, 60s timeout. Requiring a
consecutive streak of 200s after the fallback is set was wrong. It asserts that
the fallback path succeeds five times running, which is a reliability claim the
test never intended to make, and the path is inherently retry-ish because the
broken primary is attempted first on every call. One intermittent non-200
resets the streak, so a mostly-working fallback never converges.

The two directions are not symmetric:

  before the write  proving NO replica serves it   -> needs every replica
  after the write   proving the fallback serves it -> one success is the claim

So the control keeps a multi-sample window and the success assertion goes back
to polling for a first sighting, on the wider 60s budget rather than the
original 30s that expired on litellm-e2e-ui 63.

Also drops the two local rebinds Greptile flagged against the repo's
no-reassignment convention: the streak counter is gone with the helper it lived
in, and the cache-round loop is now a lazy generator consumed by next().

(cherry picked from commit b95801172c)
2026-08-29 19:11:27 -07:00
Yuneng Jiang
ab1da05e59
test(e2e): de-flake the cost-header cache read and the router fallback control
Two e2e tests fail on timing rather than on litellm behaviour. Measured over the
last ~35 litellm-e2e / litellm-e2e-ui runs:

  routerSettings.spec.ts:254  9/35 runs (7 flaky-on-retry, 2 hard failures)
  test_cost_headers_e2e.py    1/29 runs it appeared in

Router fallback control
-----------------------
The e2e stack runs replicaCount 2 with proxy_config_reload_interval_seconds 7,
and every request is routed independently, so an observation of the new config
only proves the replica that served it reloaded. patchRouterSettings returns as
soon as /config/update returns, and clearBrokenFallback never waits at all, so a
retry's one-shot control assertion could be answered by a sibling replica still
holding the previous attempt's fallback. That is exactly the observed pair of
errors: "fallback never took effect" on the first attempt and "broken primary
unexpectedly succeeded on its own" on the retry.

Both assertions now poll for a consecutive streak spanning more than one reload
cycle, mirroring the PROPAGATION_TIMEOUT / settle_propagation doctrine the Python
suite already applies in e2e_config.py.

Cost-header cache read
----------------------
The prime and measure calls fired back to back with no gap, and each retry threw
away the prefix it had just paid to prime in favour of a fresh one. OpenAI
publishes a primed prefix asynchronously and routes cache lookups by
prompt_cache_key, so the test was rerolling the least likely path to a hit.

Each round now pins a prompt_cache_key and re-reads the same primed prefix up to
CACHE_REREADS times before rotating, so a fresh prefix is spent only after the
primed one has genuinely failed to become readable.

No production code changes; prompt_cache_key is added to the e2e ChatBody model,
which serializes exclude_none and so is inert for every other caller.

(cherry picked from commit 84dfc18f6b)
2026-08-29 19:11:26 -07:00
yuneng-jiang
1411852684
Merge pull request #38855 from BerriAI/litellm_/backport-38251-rc-1-99-0
[Backport rc/1.99.0] fix(anthropic): translate tool_result document blocks in the /v1/messages bridge
2026-08-29 17:51:18 -07:00
mateo-berri
f0ee067aaa
fix(anthropic): translate tool_result document blocks in the /v1/messages bridge
(cherry picked from commit 17845b4fb0)
2026-08-29 17:46:57 -07:00
yuneng-jiang
4b281e7892
Merge pull request #38854 from BerriAI/litellm_backport_model_test_dialog_stray_text
fix(ui): drop stray text next to Close in the model connection test dialog
2026-08-29 17:45:06 -07:00
Yuneng Jiang
2689e16765
chore: update Next.js build artifacts (2026-08-30 00:39 UTC, node v24.19.0) 2026-08-29 17:39:09 -07:00
yuneng-jiang
339b713f8e
Merge pull request #38852 from BerriAI/litellm_/model-test-connection-artifact-805be5
fix(ui): drop stray text next to Close in the model connection test dialog

(cherry picked from commit 8e45522117)
2026-08-29 17:36:17 -07:00
yuneng-jiang
3d14b70ec0
Merge pull request #38848 from BerriAI/litellm_e2e-rc-1-99-0-fixture-backports
fix(e2e): backport the vertex realtime and vision image fixture fixes to rc/1.99.0
2026-08-29 16:57:15 -07:00
Yuneng Jiang
4717cb37b8
test(e2e): serve the vision image from our own fixture
The two vision tests pointed at a Wikipedia-hosted cat photo, so every run
depended on upload.wikimedia.org staying up and unthrottled. It throttled,
and the 429 surfaced as a bedrock APIConnectionError, which reads as a
gateway failure rather than what it was.

The image is now a fixture in the repo, passed as a data URL. That also puts
the two providers on the same bytes: litellm downloads the image itself for
bedrock, while openai is handed the link and fetches it from its own servers,
so the hosted URL quietly meant the two tests were not testing the same thing.

The image was generated for this repo rather than borrowed, so nothing here
carries a third-party license. Also drops a stale comment about openai prompt
caching that sat above the vision helper; no caching test uses it.

(cherry picked from commit ff418ffb9c)
2026-08-29 16:54:13 -07:00
Yuneng Jiang
f1e7d9eef7
fix(e2e): move the vertex realtime suite off the retired Live preview model
Google withdrew gemini-live-2.5-flash-preview-native-audio-09-2025 from the
Vertex Live API. Every session dies at setup:

  received 1007 (invalid frame payload data)
  gemini-live-2.5-flash-preview-native-audio-09-2025 is not supported in the live api.

The client sees session.created (the proxy synthesizes it on connect) and then
nothing, so both vertex_ai realtime tests time out waiting for session.updated.

Confirmed by probing the Vertex Live endpoint directly with the e2e stack's own
credentials:

  gemini-live-2.5-flash-preview-native-audio-09-2025 -> 1007, not supported
  gemini-live-2.5-flash-native-audio                 -> setupComplete

so this swaps to the non-preview sibling, which is the same native-audio class
and is what the cost map already carries for vertex_ai.

Not a litellm regression. The suspicion fell on #38395 because it removed the
native-audio speechConfig strip, but the setup payload this suite sends is
byte-identical either side of that change: the strip only fires when a client
sends a voice, and the e2e SessionConfig has no voice field. Google's rejection
names the model, not a field.

The gemini (Google AI Studio) provider keeps the -09-2025 id, which still works
there; only the Vertex endpoint dropped it.

(cherry picked from commit a215ecaf3d)
2026-08-29 16:54:13 -07:00
yuneng-jiang
59658e1db1
Merge pull request #38821 from BerriAI/litellm_/shadcn-migration-bug-fixes-e122ae
fix(ui): backport the shadcn migration regression fixes onto 1.99.0
2026-08-29 16:13:20 -07:00
Yuneng Jiang
47d3340eb5
chore: update Next.js build artifacts (2026-08-29 23:08 UTC, node v24.19.0) 2026-08-29 16:08:53 -07:00
yuneng-jiang
fdd424a45a
Merge pull request #38830 from BerriAI/litellm_paginated_select_deletion_query
fix(ui): keep a deleted-from search query instead of blanking the box

(cherry picked from commit b3235fa786)
2026-08-29 16:01:53 -07:00
yuneng-jiang
eb67432174
Merge pull request #38782 from BerriAI/litellm_/logs-reopen-shadcn-migration-2c9526
fix(ui): restore the reopen control for the log drawer's trace sidebar

(cherry picked from commit d435a62ce6)
2026-08-29 15:15:11 -07:00
yuneng-jiang
16c65387d8
Merge pull request #38778 from BerriAI/litellm_/json-readability-logs-661421
fix(ui): make the logs JSON viewer follow the theme in dark mode

(cherry picked from commit 207893db9c)
2026-08-29 15:14:59 -07:00
ryan-crabbe-berri
756e03c903
Merge pull request #38771 from BerriAI/litellm_/api-reference-dark-mode-4ec67f
fix(ui): make code blocks follow the theme in dark mode

(cherry picked from commit e585aab3ea)
2026-08-29 15:14:59 -07:00
yuneng-jiang
47ce15027b
Merge pull request #38588 from BerriAI/litellm_/dark-mode-logo-strategy-7b99f2
feat(ui): make provider logos readable in dark mode

(cherry picked from commit 733d0b5af5)
2026-08-29 15:14:45 -07:00
ryan-crabbe-berri
b23632f628
Merge pull request #38601 from BerriAI/litellm_ui_navbar_papercuts
fix(ui): one-click theme toggle and matching Docs/Blog styling in the top bar

(cherry picked from commit 3beb02e512)
2026-08-29 15:14:45 -07:00
ryan-crabbe-berri
31321883e7
Merge pull request #38574 from BerriAI/litellm_combobox_server_search_hardening
fix(ui): stop server-searched comboboxes from clobbering picks and queries

(cherry picked from commit 3c41392893)
2026-08-29 15:14:41 -07:00
yuneng-jiang
79d7bebc4d
fix(ui): let the paginated search select keep what the user types (#38475)
* fix(ui): let the paginated search select keep what the user types

The combobox handed Base UI a freshly built option object for the current
selection every time a page of results came back. Base UI answers a changed
value by rewriting the input with that option's label, so every search response
wiped the query mid-typing and the list never narrowed. Once a user had been
picked in the Usage page filter box, no other user could be reached.

The component now owns the input text. It holds the query while the list is
open, falls back to the selected option's label once the list closes, and
remembers the picked option so its label survives later pages that no longer
carry it, the way the multi-select sibling already does.

* refactor(ui): name the paginated select's search state instead of commenting it

* fix(ui): start a fresh query when typing lands on the selected label

Focusing the filter box without clicking it leaves the caret at the end of the
selected option's label, so the next keystroke extended that label into a query
no server could match. Only a click cleared the box first.

A keystroke that arrives while the box is showing a label is now read as the
start of a new query, wherever in the label it landed.

(cherry picked from commit 3746ba58d7)
2026-08-29 15:13:41 -07:00
tin-berri
9ca168d930
fix(ui): open select popups below the trigger instead of over it (#38554)
The shared SelectContent wrapper defaulted alignItemWithTrigger to true,
which puts Base UI's positioner into item-aligned mode and places the
popup so the active item sits on top of the trigger. In that mode the
side and sideOffset the wrapper passes two lines above are ignored, and
the popup reports data-side="none".

The overlap only becomes visible once the items are tall enough to
matter, which is why the autorouter Template picker shows it clearly:
its options are three-line cards, so the popup covers both the select
box and its own label.

No call site in the dashboard asked for item-aligned mode. 21 of them
across 15 files already passed alignItemWithTrigger={false} by hand to
undo the default, and the remaining 127 inherited the bug. Flipping the
default makes side and sideOffset live, so collision handling works and
a select with no room below now flips above the trigger rather than
covering it. The 21 hand-written opt-outs are deleted as redundant.

(cherry picked from commit 71449b9c55)
2026-08-29 15:13:41 -07:00
yuneng-jiang
cbac4921bb
Merge pull request #38366 from BerriAI/litellm_fix_add_model_public_name_focus
fix(ui): keep focus in the add model public name input while typing

(cherry picked from commit f066b01b0a)
2026-08-29 15:13:41 -07:00
ryan-crabbe-berri
c1585be98a
Merge pull request #38273 from BerriAI/litellm_fix_flow_builder_guardrail_dropdown_zindex
fix(ui): stack policy flow builder below the popup layer so guardrail options render

(cherry picked from commit 50a42ba38e)
2026-08-29 15:13:41 -07:00
yuneng-jiang
c4867ccd18
fix(ui): make playground chat bubbles theme-aware (#37978)
The playground message bubble painted its fill, border and avatar circle from
inline hex values, so in dark mode both bubbles stayed near-white while the text
inherited the dark foreground: the message body was unreadable. The MCP-events
placeholder bubble in ChatUI carried the same three fills.

They move onto the tokens the rest of the sweep already uses, so the assistant
surface is bg-card over border-border and the user surface is the info tint at
the same weight the other selected-state surfaces take. Light mode keeps the
same colour family it had.

The regression test asserts the token classes and that no inline style survives
on either surface, which is the exact shape the bug took.

(cherry picked from commit 3fb1009f81)
2026-08-29 15:13:41 -07:00
yuneng-jiang
6b29016987
fix(ui): restore the public model name tooltip layout in the add model flow (#37986)
The tooltip popup is an inline-flex row, so the four sibling blocks passed as a fragment laid out side by side in four columns. Wrap them in a single flex-col container instead.

The inline code samples also used bg-muted, which is defined against the page surface, not the inverted tooltip surface, so they rendered as near-white chips carrying near-white text. Tint them from the popup's own token instead.

(cherry picked from commit 6db5a5d660)
2026-08-29 15:13:40 -07:00
yuneng-jiang
5c3fa165c8
fix(ui): theme the created-key box so it follows dark mode (#37985)
The virtual key shown after creating a key sits in a div with a
hardcoded #f8f8f8 inline background, so in dark mode the box keeps
the light background while the key text inherits the light foreground
color, leaving the key nearly unreadable. Swap the inline styles for
the bg-muted and text-foreground tokens, which resolve per theme.

(cherry picked from commit 5f56be3294)
2026-08-29 15:13:40 -07:00
yuneng-jiang
947dbbf029
Merge pull request #37913 from BerriAI/litellm_internal_staging
Some checks failed
Helm unit test / unit-test (push) Has been cancelled
Scorecard supply-chain security / Scorecard analysis (push) Has been cancelled
Code Quality Checks / code-quality (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Unit Tests: Documentation Validation / documentation (push) Has been cancelled
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Has been cancelled
Unit Tests / core-utils (push) Has been cancelled
Unit Tests / enterprise-routing (push) Has been cancelled
Unit Tests / integrations (push) Has been cancelled
Unit Tests / All Other Providers (push) Has been cancelled
Unit Tests / Vertex AI (push) Has been cancelled
Unit Tests / misc (push) Has been cancelled
Unit Tests / proxy-auth (push) Has been cancelled
Unit Tests / proxy-endpoints (push) Has been cancelled
Unit Tests / proxy-infra (push) Has been cancelled
Unit Tests / proxy-server (push) Has been cancelled
Unit Tests / responses-caching-types (push) Has been cancelled
Unit Tests: Proxy DB Operations / auth-checks (push) Has been cancelled
Unit Tests: Proxy DB Operations / budgets (push) Has been cancelled
Unit Tests: Proxy DB Operations / custom-logging (push) Has been cancelled
Unit Tests: Proxy DB Operations / db-and-spend (push) Has been cancelled
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Has been cancelled
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Has been cancelled
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Has been cancelled
Unit Tests: Proxy DB Operations / key-generation (push) Has been cancelled
Unit Tests: Proxy DB Operations / logging-misc (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-runtime (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-server-core (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-utils (push) Has been cancelled
chore(ci): promote internal staging to main
2026-08-22 15:24:27 -07:00
Mateo Wang
75613bf22f
test: add regression coverage for twelve closed issues (#37974)
* test: add regression coverage for twelve closed issues

Adds targeted regression tests for behavior that was fixed but left ungated,
so the fixes cannot silently regress:

- #33772 openai cache_write_tokens cost
- #34309 Responses API cache cost_breakdown
- #35363 /v1/responses batch spend
- #36619 auto-router api_base/api_key leak on a shared model name
- #35359 batch fallbacks within the owning model group
- #36523 passthrough streamed Responses spend log
- #36646 passthrough embeddings spend log
- #37147 non-object metadata on create_batch is a 400
- #35362 unscoped list files reads the managed-file store
- #33221 gpt-5.6 bridges to Responses on function tools alone
- #34487 LLM complexity classifier runs for every caller metadata shape
- #35124 streamed /v1/messages emits success logging on both bridges

Cost assertions read rates from litellm.model_cost rather than hardcoding
dollar amounts, so they do not drift on repricing.

* fix: stop the new regression tests polluting and tripping over shared global state

Two shard failures, both from global state the new tests share with their
neighbours rather than from the behaviour under test.

test_main.py's local_cost_map pinned litellm.model_cost but left the
get_model_info lru_cache warm, so completion_cost billed at whatever prices
were cached earlier in the process while the assertions read the pinned map.
Clear the cache on both sides of the fixture, matching the local_model_cost_map
fixture in tests/test_litellm/conftest.py.

The anthropic messages streaming tests called GLOBAL_LOGGING_WORKER.flush()
on whatever queue happened to be around. A queue left non-empty by an earlier
test is still bound to that test's loop, so join() either hangs or raises
"bound to a different event loop". Rebind to the running loop before the call
and wait for the captured payload instead of a fixed sleep.
2026-08-22 22:24:05 +00:00
yuneng-jiang
ba07340964
chore: update Next.js build artifacts (2026-08-22 22:11 UTC, node v24.19.0) (#37976) 2026-08-22 15:20:50 -07:00
Mateo Wang
11cbe472ac
Merge pull request #36355 from harryzhou2000/fix/responses-bridge-preserve-reasoning-input-items
fix(responses-bridge): preserve reasoning input items and signed thinking blocks
2026-08-22 15:07:02 -07:00
mateo-berri
19a3fe1b66 fix(responses-bridge): fall back to summary text when content carries none
An empty content list, or one holding only opaque blocks, still lets the
provider-bound branch replay the summary text. The inspection path treated
any non-None content as final, so that replayed text stayed invisible to
guardrails and token counting.
2026-08-22 14:51:38 -07:00