Commit graph

45043 commits

Author SHA1 Message Date
Mateo Wang
ea6ac3dd99
Merge pull request #38283 from BerriAI/litellm_together_regression_tests
test(together_ai): regression suite across chat, responses, and messages surfaces
2026-08-25 17:58:14 -07:00
Mateo Wang
d0c527f3bb
Merge pull request #36245 from BerriAI/litellm_fix_headroom_stream_leak
fix(anthropic): buffer streamed responses carrying server-fulfilled tools so retrieval tool calls never reach the client
2026-08-25 17:48:31 -07:00
mateo-berri
602bf1c0ac test(together_ai): close the injected async client after the streaming test 2026-08-25 17:45:54 -07:00
mateo-berri
7c3da429c7 test(together_ai): regression suite across chat, responses, and messages surfaces
Adds streaming, async, /v1/responses, and /v1/messages coverage for the
Together AI overhaul (#38233, #38248, #38230, #38265, #38275), plus the
legacy api.together.xyz host and TOGETHER_AI_API_BASE through
litellm.completion. Each new test fails under a one-line mutation of the
merged code.
2026-08-25 17:40:41 -07:00
ryan-crabbe-berri
89d2d34453 fix(ui): lint z-index utilities behind arbitrary Tailwind variants
The variant prefix pattern only understood word variants, so
data-[side=top]:z-50 or [&>*]:z-[5] slipped past the rule. Parse the
utility as everything after the last top-level colon (brackets and
parens nest) and also strip the important marker. Formats the two files
prettier flagged in CI
2026-08-25 17:34:45 -07:00
mateo-berri
8d877980e0 fix(caching): carry has_buffered_provider_output through the anthropic stream cache writer 2026-08-25 17:31:46 -07:00
Mateo Wang
99e1eaaca0
Merge pull request #38205 from BerriAI/litellm_decrease_anys_opus5_round2
refactor(repositories): type prisma table access with one generic protocol
2026-08-25 17:19:57 -07:00
Matthew Lapointe
418012aac5 fix(bedrock): type GPT-5 reasoning field and update capability test
Type the GPT-5.x reasoning payload with a ReadOnly TypedDict so the dict
literal satisfies the type-discipline budget, and drop the now-redundant
thinking pop (the thinking mapping is already skipped for these models).
Update the cross-region capability test to expect reasoning_effort offered
and thinking/output_config withheld for GPT-5.x on Converse.
2026-08-25 20:06:53 -04:00
ryan-crabbe-berri
6385b6d801 refactor(ui): replace hand-picked z-index values with one named scale and lint it
Regression LIT-6143 (the policy Flow Builder painting its guardrail dropdown
underneath a position: fixed shell at z-index 1000) was one instance of a
class of bug: pages picking their own z-index numbers above the portalled
popup layer. This removes the class.

- globals.css defines the only z-index values in the dashboard as Tailwind
  utilities: z-raised, z-chrome, z-sticky, z-sticky-pinned, z-floating,
  z-overlay, z-popup; every numeric, arbitrary and inline z-index across
  src is migrated onto them and tailwind-merge learns the tokens
- new local/no-ad-hoc-z-index ESLint rule bans z-<n>, z-[...], z-(...) and
  inline zIndex everywhere, and reserves z-popup for the portalled
  primitives in components/ui (and the DataTable menus)
- the Flow Builder renders in the dashboard content area instead of as a
  fixed full-screen overlay, so it has no stacking level at all
- the guardrail content-filter Add keyword / Add pattern / Custom pattern
  dialogs drop the leftover z-[1100] (renamed from ABOVE_ANTD_MODAL when
  antd was removed) that hid their own Action select and pattern combobox
  behind the dialog, the same bug as LIT-6143
2026-08-25 17:02:35 -07:00
mateo-berri
64d8b24c10 fix(router): deliver the hold-back retrieval error to the client instead of falling back 2026-08-25 17:00:00 -07:00
ryan-crabbe-berri
434add755a
Merge pull request #38274 from BerriAI/litellm_test_lint_os_environ
test: gate the test tree on B003 so a test cannot swap os.environ for a plain dict
2026-08-25 16:51:45 -07:00
mateo-berri
7965cfdd2b fix(proxy): keep listing a keyless user's own models when their user row is missing
Keys minted by /key/generate get no LiteLLM_UserTable row, so /v2/model/info?user_models_only=true for such a user hit the new None guard and returned 400 where the merge base returned the user's own models. Skip the team-model merge for a missing row instead of raising, since a user with no row belongs to no team
2026-08-25 16:51:03 -07:00
mateo-berri
98b04de455 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5_round2
# Conflicts:
#	litellm/responses/litellm_completion_transformation/session_handler.py
2026-08-25 16:48:41 -07:00
ryan-crabbe-berri
50a42ba38e
Merge pull request #38273 from BerriAI/litellm_fix_flow_builder_guardrail_dropdown_zindex
fix(ui): stack policy flow builder below the popup layer so guardrail options render
2026-08-25 16:43:53 -07:00
mateo-berri
1aca5a93d4 Merge remote-tracking branch 'origin/litellm_fix_headroom_stream_leak' into litellm_fix_headroom_stream_leak
# Conflicts:
#	litellm/proxy/common_request_processing.py
#	tests/test_litellm/proxy/test_budget_reservation.py
2026-08-25 16:40:05 -07:00
mateo-berri
1b37397633 fix(router): forward the hold-back keepalive ping live and carry the withheld-output flag through the stream wrappers
The router's pre-content ping filter dropped AgenticAnthropicStreamingIterator's
hold-back keepalive, so a held-back turn sent the client nothing until the buffer
settled. A ping that no lifecycle frame precedes is now forwarded live, since a
fallback's message_start can still follow it without overlapping lifecycles

The proxy's cancel-refund guard checked isinstance against the iterator, but the
proxy only ever sees it behind FallbackAwareAnthropicMessagesStream and
AnthropicMessagesStreamingResponse, so a disconnect during hold-back refunded the
budget reservation anyway. Both wrappers now forward a duck-typed
has_buffered_provider_output flag, and the router wrapper follows a fallback
source so the flag tracks the stream actually being consumed
2026-08-25 16:39:14 -07:00
yucheng-berri
ba8d8b6e14
fix(logging): redact tool call arguments to valid JSON and preserve null content (#38182)
* fix(logging): redact tool call arguments to valid JSON and preserve null content

Resolves LIT-6102

* refactor(logging): centralize redacted tool-call arguments constant and satisfy test-quality gate

* fix(responses): drop Final annotations on loop-assigned locals flagged by basedpyright

* fix(responses): skip custom tool calls in redacted-arguments normalizer

* fix(logging): keep the redaction sentinel in stored tool-call arguments and preserve null output text
2026-08-25 16:38:18 -07:00
mateo
1bc441ca2a fix(proxy): detect held-back provider output through the anthropic streaming response wrapper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-25 23:33:47 +00:00
Matthew Lapointe
9cc276a96e fix(bedrock): never forward Anthropic thinking for OpenAI GPT-5.x Converse
Stop advertising thinking/output_config as supported for OpenAI GPT-5.x and
skip the thinking mapping for these models, so a request combining thinking
with reasoning_effort can no longer leak a thinking block into
additionalModelRequestFields regardless of parameter order, which Bedrock
rejects with unknown_parameter.
2026-08-25 19:32:54 -04:00
mateo-berri
7aa8efcf47 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_together_structured_outputs 2026-08-25 16:29:05 -07:00
yucheng-berri
75bf9f9452
fix(router): persist attempted_fallbacks and original_model_group into spend logs metadata (#38107) 2026-08-25 16:22:01 -07:00
mateo-berri
d16afd2027 fix(memory): return 404 when the memory row vanishes before delete
prisma-client-py delete() returns None when the row is already gone, so the
handler reported deleted: true for a row this caller never removed. Surface
the same 404 the read path uses instead
2026-08-25 16:20:17 -07:00
Matthew Lapointe
74e86d3c0d fix(bedrock): route reasoning_effort to reasoning.effort for OpenAI GPT-5.x on Converse
OpenAI GPT-5.x models on Bedrock Converse expect reasoning effort under
additionalModelRequestFields as {"reasoning": {"effort": ...}}. They were
falling into the Anthropic branch and emitting a `thinking` block, which
Converse rejects with unknown_parameter.

The bedrock_converse gpt-5.6 entries were also missing supports_reasoning,
so reasoning_effort was dropped before mapping. Setting the flag lets the
existing config-driven supported-params path accept it, rather than adding
another model-name branch.
2026-08-25 19:20:00 -04:00
mateo-berri
4582496c8a Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5_round2
# Conflicts:
#	basedpyright-code-budget.json
#	litellm/proxy/auth/user_api_key_auth.py
#	litellm/proxy/management_endpoints/team_endpoints.py
#	litellm/proxy/management_helpers/utils.py
#	ruff-strict-budget.json
#	tests/test_litellm/proxy/management_endpoints/test_team_endpoints.py
#	type-discipline-budget.json
2026-08-25 16:17:29 -07:00
Mateo Wang
1ff615c335
Merge pull request #38275 from BerriAI/litellm_together_chat_template_kwargs
fix(together_ai): strip internal thinking fields from outbound messages, keep reasoning_content
2026-08-25 16:15:50 -07:00
Deepanshu Lulla
6684bb343b
perf(streaming): add shared JSONFragmentAccumulator for Vertex and Anthropic (#36610)
* perf(streaming): add shared JSONFragmentAccumulator for Vertex and Anthropic

Vertex's handle_accumulated_json_chunk and Anthropic's
_handle_accumulated_json_chunk each independently accumulated SSE fragments
into a JSON envelope with self.accumulated_json += fragment. Because the
attribute holds a live reference, CPython copies the whole prior buffer on
every fragment, making buffer assembly O(n^2) in total payload size.
Anthropic additionally had no completeness heuristic at all and called
json.loads on the whole buffer after every fragment, and could wedge forever
on two concatenated envelopes.

Add JSONFragmentAccumulator in litellm_core_utils/: fragments append to a
list in O(1), a could_close_json heuristic lets callers skip the join+parse
entirely until a value could plausibly be complete, and pop_next_value peels
one JSON value off the front of the buffer at a time using
json.JSONDecoder().raw_decode, keeping any unconsumed remainder instead of
failing on concatenated values. Migrate both providers onto it; Anthropic's
__next__/__anext__ end-of-stream handlers now delegate to
_handle_accumulated_json_chunk(is_final=True) instead of duplicating the
parse-and-reset logic inline.

Fixes #31861.

* test(streaming): close diff-coverage gaps in JSONFragmentAccumulator migration

Codecov flagged 11 uncovered lines in the migration: the accumulated_json
setter, __next__/__anext__'s end-of-stream drain branches in both providers,
and the pop_next_value "not found" path when a buffer's newest fragment ends
in "}" but is genuinely incomplete (an inner object closed, the outer one
didn't). Add targeted tests for each.

* fix(streaming): make JSONFragmentAccumulator's completeness heuristic O(1)

could_close_json rescanned every preceding blank fragment on each call, so a
hostile upstream that sends malformed JSON (never closing) followed by many
blank keepalive fragments could drive that scan, and the join+parse it
gates, to O(n^2) total. Track the last non-blank fragment's trailing byte
incrementally in append/pop_next_value/set instead of rescanning the buffer.

Reported by automated review on PR #36610.

* fix(streaming): make JSONFragmentAccumulator.pop_next_value O(1) per call

pop_next_value previously rebuilt the full remaining string and sliced a
new remainder on every call, so draining N concatenated JSON values
already sitting in one buffer cost O(n^2) total. Replace the rebuild-and-
slice with a materialize-once cursor: pending fragments are joined into
the buffer only when new ones have arrived since the last pop, and
consumed values are dropped by advancing an offset instead of copying
the remaining string.

* test(streaming): make JSONFragmentAccumulator drain regression test CI-stable

The 80k-value drain test used an absolute ms budget that flaked on a
busier CI runner (233.5ms vs a 150ms budget calibrated on a quiet
machine). Replace it with a doubling-ratio check: draining twice as many
concatenated values should take roughly 2x as long for O(n), not
~4x for O(n^2), and that ratio holds regardless of machine speed.

* test(streaming): suppress TQ002 on the append-laziness spy test

A test-quality gate (TQ002: don't assert only that a mock was called)
landed upstream since this branch's last rebase and now flags
test_append_never_calls_raw_decode. The test verifies append() defers
all decoding to pop_next_value, which has no caller-observable proxy
other than spying on the stdlib call it must avoid making.

* test(vertex): spy on raw_decode instead of json.loads in accumulator regression tests

Post-migration to the shared JSONFragmentAccumulator, Vertex's decode
path goes through json.JSONDecoder.raw_decode, not json.loads. The two
O(n^2)/partial-fragment regression tests still patched json.loads, which
that path never calls, so both passed unconditionally regardless of
whether the underlying implementation regressed. Verified by simulating
an eager-reparse regression: both tests now fail against it and pass
against the correct implementation.

* fix(streaming): widen JSONFragmentAccumulator's whitespace skip to match str.strip()

pop_next_value's whitespace skip only matched json.decoder.WHITESPACE's
ASCII set, narrower than str.strip() (Unicode-aware) which the O(1)-cursor
rewrite replaced. A non-ASCII separator like U+00A0 between two
concatenated JSON values on one SSE line made raw_decode fail on it, and
the buffer never advanced past that byte again, permanently stranding
everything after it for the rest of the stream. Use str.isspace() to
match str.strip()'s tolerance.

---------

Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
2026-08-25 16:15:47 -07:00
mateo-berri
4ab654cbc5 merge: litellm_internal_staging into litellm_fix_headroom_stream_leak 2026-08-25 16:04:33 -07:00
mateo-berri
8a61b286d1 test(together_ai): annotate preserved-thinking test helper and capture list 2026-08-25 16:03:43 -07:00
devin-ai-integration[bot]
bb22742025
fix(rerank): emit latency and cost headers on /rerank (#35419)
* fix(rerank): emit latency and cost headers on /rerank

Thread the logging object into the rerank httpx calls and pass hidden_params through to get_custom_headers, so x-litellm-overhead-duration-ms, x-litellm-response-duration-ms, x-litellm-response-cost, x-litellm-call-id and the LITELLM_DETAILED_TIMING x-litellm-timing-* headers show up on rerank like they do on chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rerank): keep zero response cost in the /rerank cost header

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: assign the new rerank endpoint tests to the proxy-endpoints shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: suppress TQ008 on the rerank header tests with reasons

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: milan <milan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
2026-08-25 15:54:25 -07:00
mateo-berri
1ba66b1555 fix(together_ai): strip internal thinking fields from outbound messages, pin chat_template_kwargs passthrough 2026-08-25 15:48:58 -07:00
ryan-crabbe-berri
8298cbe1f7 test: gate the test tree on B003 so a test cannot swap os.environ for a plain dict 2026-08-25 15:43:26 -07:00
Daniel Meismer
31f7b9409a fix(router): resolve hidden aliases for explicit lookup
Co-Authored-By: Codex
2026-08-25 18:41:23 -04:00
Daniel Meismer
c673bcd970 fix(mcp): preserve provider access token lifetime
Co-Authored-By: Codex
2026-08-25 18:41:16 -04:00
Devin AI
3ebd18b057 fix(ui): stack policy flow builder below the popup layer so guardrail options render
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-25 22:37:02 +00:00
mateo-berri
5ab9e63628 fix(together_ai): fail open on response_format instead of dropping it for unregistered models 2026-08-25 15:24:58 -07:00
Mateo Wang
04818a3554
Merge pull request #38267 from BerriAI/litellm_fix_responses_user_document_drop
fix(anthropic): carry user-content document blocks through the /v1/messages responses bridge
2026-08-25 15:20:30 -07:00
Mateo Wang
e4ff44f623
Merge pull request #38265 from BerriAI/litellm_together_tools_passthrough
fix(together_ai): pass tools through for models missing from the registry
2026-08-25 15:12:08 -07:00
Mateo Wang
dc40377959
Merge pull request #38261 from BerriAI/litellm_fix_responses_tool_result_document_drop
fix(anthropic): carry tool_result document blocks through the /v1/messages responses bridge
2026-08-25 15:09:23 -07:00
Mateo Wang
296ebc1103
Merge pull request #38252 from BerriAI/litellm_pr_template_caveat_severity
docs(pr-template): split Caveats bullets into severity tiers and call for plain engineering language
2026-08-25 15:04:37 -07:00
mateo-berri
61ee2f5ecb fix(anthropic): carry user-content document blocks through the /v1/messages responses bridge 2026-08-25 15:00:57 -07:00
Yassin Kortam
6c0c91c5ad
fix(team): serialize member_add, member_delete, and delete under the team's advisory lock (#37969)
* fix(proxy): make /team/member_delete's four cleanups atomic

The team roster update, the user.teams update, the team membership
delete, and the team-scoped verification token delete ran as four
sequential writes with no transaction around them, so a failure
between any two left the removal half applied. Thread a single
prisma transaction through all four writes, following the same
tx.<table> pattern /team/member_add and /team/member_update already
use, so either all four land or none do.

* fix(team): serialize member_add, member_delete, and delete under the team's advisory lock

/team/member_add validated a team exists and then wrote the user's teams array and
a membership row without holding anything across that gap, so a /team/delete could
commit its reference sweeps in between and leave a member pointing at a team id that
no longer exists. The write path already re-read members_with_roles under a row lock
before this change, but SELECT ... FOR UPDATE can deadlock with the access-group
endpoints, which lock an access group and then a team.

member_add now takes pg_advisory_xact_lock(hashtext(team_id)) before re-reading the
team and only writes if it is still there, so a delete that already committed is
visible before any write happens. delete_team takes the same lock around its own
row delete and reference sweep, so the two requests can never interleave: whichever
acquires the lock first runs to completion before the other's read can proceed.

Dropping the row lock from member_add's read also dropped the incidental protection
it gave against a concurrent member_delete, which still wrote from the snapshot it
validated against, unlocked, and could silently overwrite whatever member_add had
just committed. member_delete now takes the same advisory lock and re-reads the
roster under it before computing its own write, so it can never resurrect a member
by overwriting from stale data.

Resolves LIT-5544

* fix(team): run member writes on the advisory lock's transaction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(team): keep member writes on the lock holder's connection after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(team): keep the transactional member create an upsert on user_id

The transaction path was creating the email-identified user row outright, where the
regular client path upserts on user_id. Share one upsert helper between both member
paths so the create stays idempotent on the lock holder's connection.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(team): read member_delete's user and key rows on the lock-holding transaction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-25 21:55:11 +00:00
Mateo Wang
9962f4fba8
Merge pull request #38251 from BerriAI/litellm_fix_tool_result_document_drop
fix(anthropic): translate tool_result document blocks in the /v1/messages bridge
2026-08-25 14:54:01 -07:00
Mateo Wang
a40157c55d
chore: make it more concise 2026-08-25 14:53:24 -07:00
Mateo Wang
d61ccd0f58
chore: make it more clear 2026-08-25 14:51:53 -07:00
Mateo Wang
99c1b33fd3
chore: humanize the CLAUDE.md 2026-08-25 14:50:14 -07:00
mateo-berri
ec365a3c12 fix(together_ai): pass tools through for models missing from the registry
Tools passed for Together models with no supports_function_calling entry in
model_prices_and_context_window.json were rejected with UnsupportedParamsError
by default and silently dropped under drop_params, which made models emit tool
calls as plain text. Fail open instead: pass tool params through with a warning
and let Together validate. Models the registry explicitly marks as not
supporting function calling keep the loud contract: raise by default, drop with
a warning under drop_params.
2026-08-25 14:38:34 -07:00
Deepanshu Lulla
75c4565dde
fix(cerebras): add max_retries and extra_headers to get_supported_openai_params (#36601)
Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
2026-08-25 14:10:56 -07:00
Mateo Wang
d8d384eada
Merge pull request #38230 from BerriAI/litellm_together_models_backfill
feat(models): add missing Together AI serverless models to the cost map
2026-08-25 14:02:58 -07:00
mateo-berri
c3c9e3a919 test(anthropic): pin dropped file-id and empty-url document sources to string output 2026-08-25 13:58:37 -07:00
mateo-berri
e1a0920445 refactor(anthropic): bind tool_result file parts once instead of seeding an empty tuple 2026-08-25 13:54:34 -07:00