Commit graph

45730 commits

Author SHA1 Message Date
yuneng-jiang
a4a0386717
Merge pull request #38637 from BerriAI/litellm_/red-tests-review-fbb718
test: refresh the suites that drifted from langfuse and OpenAI's retired Assistants API
2026-08-28 09:58:44 -07:00
mateo-berri
b8f91235a6 fix(cost): only let a WxH request size drive image pricing 2026-08-28 09:52:31 -07:00
Mateo Wang
ecf84a2c2a
Merge pull request #38657 from BerriAI/litellm_count_tokens_fallback_tools_system
fix(proxy): count tools, system, and Anthropic image and document blocks in the count_tokens fallback (internal copy of #36671)
2026-08-28 09:52:14 -07:00
tin-berri
721db0f03e
feat(ui): edit the auto-router tier set with custom classifier-defined tiers (#38603)
* feat(ui): edit the auto-router tier set with custom classifier-defined tiers

The editor over the model layer beneath it. An Edit tiers button turns the tier
list into an editor: a tier takes a name, a classifier definition, and models,
between two and eight rows. Restore defaults resets to the built-in four rather
than stacking them on top. Keyword rules follow a rename, an orphaned rule
blocks the save, and both forms dry-run the exact payload against the backend
validator before writing.

The edit modal hydrates a stored custom set into rows, and an untouched
open-and-save round-trips byte-identically, per-model reasoning efforts
included. A form that never opens the editor submits the same bytes as before.

The cost-optimization tier chart renders arbitrary tier names: the guard that
returned no models for a non-built-in name is gone, and the fixed four-color
array gives way to the shared cycle.

* refactor(ui): extract tier editor sections to clear new lint warnings

* test(ui): drop narration comments per repo convention

* fix(ui): default editingTiers so the build's type check passes

* fix(ui): restore the mid-dry-run submit guard and its regression tests
2026-08-28 09:45:11 -07:00
mateo-berri
bf047148fc fix(cost): price image generations from the requested quality when the response omits it 2026-08-28 09:38:21 -07:00
Yuneng Jiang
2b0932f930
test(access-groups): give each xdist worker its own fixture ids
Every test in the file seeds, reads and deletes the same fixed group and team
ids, and auth_ui_unit_tests runs pytest with -n 2. Two tests landing on the two
workers at once tread on each other: one worker's _clean_db DELETE wipes rows
the other just seeded, and its sync writes land in the other's read.

Both shapes showed up on 13a0976bb6, a commit that renames a passthrough test
and nothing else. test_reconcile_is_idempotent... read back an empty table, and
test_reconcile_handles_a_null_array_column read the idempotent test's team on
its own second group.

Scoping the ids to PYTEST_XDIST_WORKER keeps each worker in its own rows. Tests
on one worker still run in sequence, so no isolation is lost.

Reproduced against a local Postgres: -n 2 failed 6 out of 6 runs before, passed
6 out of 6 after, and serial runs are green either way. Stripping the COALESCE
guard from the mirror's SQL still fails the suite, so the ids are all that
changed.
2026-08-28 09:35:12 -07:00
mateo-berri
b05ac5fefd test(registry): allow /v1beta/interactions in the supported_endpoints schema 2026-08-28 09:29:09 -07:00
Yuneng Jiang
0ec619e7d1
test: close mutation-testing gaps in container, skills and openai-like config factories
Mutation testing surfaced three factory functions whose tests ran against them
but asserted nothing that a mutation could break, so every planted bug survived.

- litellm/llms/litellm_proxy/skills/code_execution.py: the OpenAI and Anthropic
  tool schemas were unpinned (the Anthropic one was not reached by any test at
  all) and the handler's default fallbacks were unchecked
- litellm/containers/endpoint_factory.py: the endpoints.json contract, the
  generated sync/async function set and the response-type mapping were unpinned
- litellm/llms/openai_like/dynamic_config.py: the generated Responses API config
  class had no coverage of auth header, URL resolution or the store override

The openai_like tests clear _responses_config_cache around each test. Without
that, the module-level cache hands back a class built before the mutation and
the tests pass against mutated code.

Verified by re-running mutmut per scope:
  llms/litellm_proxy  45.2% -> 62.8%  (70 mutants newly killed)
  containers          36.8% -> 84.3%  (45 mutants newly killed)
  llms/openai_like    55.7% -> 66.9%  (34 mutants newly killed)
2026-08-28 09:28:14 -07:00
mateo-berri
1aa214e69c fix(registry): align gemini-omni-flash-preview limits with the models API
The models API reports 131072 input / 65536 output for the preview model
and the Interactions API accepts 100k tokens but rejects 130k, so the
1,048,576 input limit copied from the docs was wrong.
2026-08-28 09:19:22 -07:00
ryan-crabbe-berri
5476f91b4a
Merge pull request #38662 from BerriAI/litellm_fix_playground_model_group_info_llm_api_key 2026-08-28 09:18:57 -07:00
mateo-berri
02f787308d fix(registry): correct gemini omni, grok-4.20 multi-agent, kimi-k2.7-code entries
Gemini omni 1.1 flash and omni flash preview only answer on the Interactions
API, so both now list /v1beta/interactions as their endpoint and 1.1 flash
gets the 131072 / 65536 limits the models API reports.

grok-4.20-multi-agent and -latest now match the dated entry (mode responses,
/v1/responses only), and all three drop function calling and tool choice
since the API rejects client-side tools outside a beta.

kimi-k2.7-code gets the capability flags kimi-k2.6 carries (tools, reasoning,
JSON mode, image and video input) plus max_output_tokens.

grok-imagine-image-2.0 gets a low quality tier at $0.04 so quality=low is
not billed at the $0.06 default.
2026-08-28 09:17:46 -07:00
Yuneng Jiang
13a0976bb6
test(passthrough): name the test for what it now covers
The target is a local server, so the OpenAI host check already excludes it
and the docstring's claim about the route path no longer holds.
2026-08-28 09:11:48 -07:00
Yuneng Jiang
cddf1f6c03
chore(lint): ratchet the TQ ceilings for the passthrough isolation 2026-08-28 09:05:56 -07:00
Yuneng Jiang
021e03fe90
test(passthrough): serve the passthrough target locally instead of calling OpenAI
Greptile flagged this test as coupled to OpenAI's availability. The coupling was
not the status assertion it pointed at, and it predates this PR: the test it
replaced called /v1/assistants live the same way, and pass_through_endpoints gates
success logging on `response.status_code < 400`, so an upstream outage has always
meant no log fires and the payload assertions fail regardless.

The target is now a local HTTP server on an ephemeral port, so the test is offline
either way. It still exercises the generic passthrough handler, since
_is_supported_openai_endpoint does not claim a 127.0.0.1 URL any more than it
claimed /v1/moderations, and it now also asserts what the upstream actually
received rather than only what came back.

respx was the obvious approach and does not work here: it patches httpx transports,
and the passthrough issues its request through the custom aiohttp transport, so the
call went to the real api.openai.com and returned 401 while respx sat unused.

Mutation checked: gating off the success enqueue fails the test, and tampering with
the logged response body fails it.
2026-08-28 09:03:57 -07:00
Devin AI
7f3ff3b47f fix(bedrock): route all cohere.embed models to BedrockCohereEmbeddingConfig
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-28 15:31:11 +00:00
mateo-berri
3bbd2fac33 fix(batches): fill a managed batch page past rows that will not parse
The managed batch listing fetched one page of rows, derived has_more from
that raw fetch, then dropped every row whose stored blob would not parse.
last_id came from the survivors, so a page of corrupt or legacy rows came
back as data [], last_id null, has_more true, and a client following
last_id could not advance. The OpenAI SDK's auto-paginator, which cursors
off the last item in data, stopped silently and returned a truncated list.

Read chunks until page_size + 1 batches survive parsing and file-id
resolution or the caller's rows run out, the way the managed file listing
already does, so a page carries data and a usable cursor while parseable
rows remain and has_more only says true when another one exists. The first
chunk keeps the old page_size + 1 size so a healthy page still costs one
query; a scan that has to continue widens to the file listing's
continuation chunk and stops resolving rows once the page is full.
2026-08-28 07:54:18 -07:00
yassin
418b820af4 fix(proxy): let llm_api virtual keys read /model_group/info
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-28 14:42:22 +00:00
mateo-berri
24d226c6c2 chore(token_counter): drop docstrings and test prose that restated the count_tokens branches 2026-08-28 06:28:21 -07:00
Devin AI
2b02c95ff2 fix(registry): add zai/glm-5.3-flash, databricks-glm-5-3-flash, moonshot/kimi-k2.7-code; xai grok-imagine-image-pro deprecation date
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-28 13:13:38 +00:00
Andrew Mattie
f5fbde9151 fix(streaming): align assembled provider model 2026-08-28 08:13:14 -05:00
mateo-berri
83ab87091b fix(proxy): only attach tools to the count_tokens fallback when counting messages 2026-08-28 06:10:30 -07:00
Devin AI
f153203ebe Merge remote-tracking branch 'origin/litellm_internal_staging' into devin/1787857843-registry-audit-rolling 2026-08-28 13:04:58 +00:00
mateo-berri
70ba0bb973 fix(proxy): count tools, system, and Anthropic document blocks in the count_tokens fallback 2026-08-28 05:41:38 -07:00
Fazeel Usmani
7cd3a27d60 update test signature for Anthropic image block handling 2026-08-28 05:35:06 -07:00
Fazeel Usmani
fe28781dfd fix(token-counter): enhance handling of Anthropic image blocks in token counting 2026-08-28 05:35:05 -07:00
Fazeel Usmani
1cce589aa0 fix(token-counter): count Anthropic native image content blocks
`_count_content_list` accepted text, image_url, tool_use, tool_result,
thinking and tool_reference, and raised on anything else, so an
Anthropic-native `{"type": "image", "source": {...}}` block aborted the
whole count. That is the documented Anthropic image format and exactly
what /v1/messages receives.

Three user-visible effects. /v1/messages/count_tokens and
/utils/token_counter return 500, and the router's context-window
pre-call check swallows the ValueError and returns every deployment
unfiltered, so an oversized prompt carrying an image is dispatched to
the provider instead of being rejected locally with a 400.

Prices the block through the existing image path: a base64 source
becomes a data URI, a url source passes through, and a file source
falls back to the default image token count. Blocks nested inside
tool_result.content are covered too, because _count_anthropic_content
recurses back into _count_content_list.

Fixes #36604
2026-08-28 05:35:05 -07:00
Devin AI
cb70941660 fix(tests): drain the global logging worker in RAG aquery billing tests instead of polling
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-28 09:58:22 +00:00
Devin AI
d990f24b9d chore(techdebt): type new signatures and drop slop comments from the last 24h
Removes restating comments added with the Teams alerting destination and the
lazy OpenAPI snapshot refactor, types three signatures that shipped untyped or
with bare dict, and ratchets the strict and basedpyright budgets down.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-28 07:59:32 +00:00
Yuneng Jiang
4cc6119d08
chore(lint): ratchet the TQ ceilings for the batch completions fix 2026-08-28 00:27:12 -07:00
Yuneng Jiang
e8b9f3675b
test(batch): make the upstream-failure tolerance actually reachable
batch_completion collects per-request failures into its result list rather than
raising them; its own source says "return exceptions if any". So the test's
`except Timeout` and `except litellm.InternalServerError` arms could never fire for
the case they were written for. An upstream 500 instead reached
`response.choices`, raised AttributeError on the exception object, and fell through
to the bare `except Exception` that calls pytest.fail. That is what CircleCI hit.

The tolerance now reads the returned values, which is where the failures actually
are. The same two exception types are tolerated as before, nothing broader.

Checked against four injected outcomes: three InternalServerErrors pass, three
Timeouts pass, an AuthenticationError fails, and a response whose content is None
fails. So it is not tolerating its way to a vacuous green.
2026-08-28 00:25:53 -07:00
tin-berri
ca0b951a43
feat(spend): report prompt caching savings as total and gateway-attributed (#38134)
* feat(spend): report prompt caching savings as total and gateway-attributed

`prompt_caching_savings_spend` credited every cached request, including caching a
client asked for with its own `cache_control` and caching a provider does implicitly,
so the number overstated what the gateway had any hand in.

Gating that column in place would have fixed the overstatement by changing what the
column means, leaving rows written before the change saying "all caching savings" and
rows after saying "gateway-injected only" with nothing to tell them apart, and forcing
a decision about rewriting history. It also breaks the cache-leakage estimate on the
dashboard, whose numerator would be gated while its denominator, the cached token
counts, would not, so the rate it extrapolates from would be quietly diluted.

Report both instead. `prompt_caching_savings_spend` keeps meaning every net dollar
caching saved, which is what a customer means by "what did caching save me", and the
new `gateway_injected_caching_savings_spend` carries the subset litellm caused by
injecting the breakpoints itself. Both are derived from the same marker, so this
changes what is done with it rather than how it is obtained.

The attributed figure is normally the smaller of the two, being a subset of the same
requests, but not always: a request that writes cache it never reads has negative net
savings, and excluding such a request can lift the attributed figure above the total.

Also stops the marker riding into a fallback leg. The fallback rebuild spread the
failed attempt's metadata forward, so a deployment that injected nothing inherited the
marker and was credited anyway, which silently restored the very overstatement this
separates out.

* fix(bedrock): credit gateway caching where the tool cachePoint is placed (#38478)

The savings marker records breakpoints litellm placed, and a tool_config
injection point becomes one only in the converse transform, and only when the
request carries tools. The prompt hook cannot see either condition, so marking
on the point's presence credited request shapes that cached nothing, while
Bedrock tool caching the gateway did cause went uncredited.

Record it at the placement site instead. The marker's reader also resolves its
bucket by value now: litellm_params declares litellm_metadata as None on every
request, so asking the shared name resolver named a bucket that was not there
and the mark was dropped.
2026-08-28 00:19:06 -07:00
Yuneng Jiang
d8679508d4
test(e2e): measure the select popup after it settles instead of mid-flight
Both anchoring tests read the trigger's box before the click and the popup's box the
instant it turns visible. Base UI places the popup asynchronously and opening it can
shift the trigger, so both boxes could be sampled before the layout settled. The run
on 1eedaa3a43 missed by 4.2px (expected >= 446.015, got 441.799) on a tree with no UI
changes at all, having passed on 21092d633b, which differs only in a deleted python
test and a budget json.

Each assertion now re-reads both boxes under expect.poll. The conditions themselves
are unchanged: the popup must sit at or below the trigger's bottom edge in the first
test and must not overlap it in the second. Polling cannot mask a genuinely misplaced
popup, since one that never lands correctly still fails when the poll times out.
2026-08-28 00:15:02 -07:00
Yuneng Jiang
0c5c96dbf6
test(e2e): unskip four tests whose blockers no longer hold
/v1/batches now rejects a missing input_file_id with a 400 through
raise_if_required_body_param_missing, so the contract negative that was
skipped for "500s instead of 400" passes as written. Verified against a
live proxy.

The three Datadog MCP tests were skipped because each one sent a
`telemetry` argument that search_datadog_logs rejects with "unexpected
additional properties". That argument was never a documented Datadog
parameter and no assertion reads it, so it is dropped and the tests run
again unchanged otherwise.
2026-08-28 00:01:51 -07:00
Yuneng Jiang
1eedaa3a43
test(passthrough): drop the Azure assistants test, retired upstream on 2026-08-26
Azure now answers the create call with 410 and code assistants_api_deprecated:
"The Assistants API has been retired. Follow the migration guide to update your
workloads." Microsoft retired it on the same day OpenAI retired theirs, which is
why this landed with the OpenAI ones rather than before them.

This job runs pytest with -x, so the test was also hiding everything after it in
tests/pass_through_tests.

test_pass_through_file_operations stays. It only asks /v1/files for
purpose="assistants", which still answers 200 on both upload and delete.
2026-08-27 23:58:23 -07:00
Yuneng Jiang
120971dacc
chore(lint): ratchet the TQ ceilings down to what these test fixes reached 2026-08-27 23:53:14 -07:00
Yuneng Jiang
21092d633b
test(opik): stop the batching test racing its own 5-second flush timer
test_opik_logging_http_request asserted "nothing has been POSTed yet" roughly one
second into a window governed by OpikLogger's 5-second periodic flush. On a loaded
CI worker the five preceding acompletion calls eat that budget, the periodic flush
fires, and the assertion flips. Reproduced with no product changes at all: letting
5.5 seconds pass before the assertion drains the queue and sets mock_post.called,
which is exactly the failure CircleCI reports.

The test now pins flush_interval past anything the test can reach, so the two
batching assertions measure batching instead of wall clock, and drives the flush
path explicitly at the end rather than sleeping the interval. That last phase used
to be near-vacuous, since the size-triggered flush had already emptied the queue.

Assertions now match only calls to Opik's own /traces/batch and /spans/batch.
get_async_httpx_client caches one client per special provider, so the mock is
process-wide and any other logging callback's POST would otherwise count.

Dropped the teardown that closed that shared client, which broke every later test
in the same worker that logs through it, and the try/except that turned assertion
failures into a pytest.fail with no traceback.

Mutation checked: flushing on every event and never flushing on size both fail the
test.
2026-08-27 23:43:38 -07:00
Yuneng Jiang
a47c2f76ab
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/red-tests-review-fbb718 2026-08-27 23:37:37 -07:00
Yuneng Jiang
feffb62266
test: refresh the suites that drifted from langfuse and OpenAI's retired Assistants API
Two unrelated causes, both leaving staging red with tests that no longer describe
anything true.

#38264 gave LangFuseLogger a langfuse_environment argument and started carrying it
in the credentials dict. The handler test's fake logger did not accept the new
keyword, so constructing it raised TypeError, and four cases in
test_langfuse_unit_tests rebuilt the cache key by hand from three fields and missed
on the four-field key production now writes. Caching itself was never broken: the
handler sets and gets with the same dict. The fake now takes the argument and
asserts it is forwarded, and the cache assertion issues a second identical request
and expects the same logger back, which is the behaviour that matters and cannot
rot the next time a credential field is added.

OpenAI has retired the Assistants API. /v1/assistants and /v1/threads both answer
404 with a valid key, where every live route answers 401, so nothing calling them
can pass again. test_custom_logger_passthrough covered generic passthrough logging
and only used assistants because it is a route with no provider-specific handler;
it moves to /v1/moderations, which is still unclaimed by
_is_supported_openai_endpoint, so the same generic branch is exercised. The two
tests there asserted the same thing against different dead routes, so they collapse
into one. The Ruby suite existed solely to drive assistants, threads, messages and
runs, so it goes along with the RVM and bundler steps that were installed only to
run it, and the two dead OpenAI assistants cases leave
test_openai_assistants_passthrough.

The Azure assistants case in that file stays. Azure runs its own lifecycle and I
could not reach the CI deployment to check whether that API is still there.
2026-08-27 23:24:36 -07:00
yuneng-jiang
2b10dc5a7a
Merge pull request #38304 from BerriAI/litellm_fix_staging_ci_regressions
test: fix staging CI regressions from #38182, #38144, #38265, #37962, and #37969
2026-08-27 23:21:37 -07:00
Yuneng Jiang
74bc1efe65
Merge branch 'litellm_internal_staging' into litellm_fix_staging_ci_regressions
Resolve tests/llm_translation/test_together_ai.py in favor of staging:
bcb6a0a998 already landed the fail-open assertion for models missing from
the registry, so both models now list response_format and tools. This
branch's narrower gating of response_format no longer matches behavior.
2026-08-27 23:12:42 -07:00
tin-berri
02c1c45b53
fix(complexity_router): route client housekeeping calls to the cheapest tier (#38598)
A coding agent names each conversation by quoting the whole session and asking
for a title. The classifier rated the quoted session rather than the request, so
the cheapest call the client makes routed to the most expensive tier: 11 of 17
title generations in one day of real traffic came back COMPLEX.

Recognize those prompts by literal sentinel on the newest ask and route them to
the cheapest configured tier without classifying them, so the call costs nothing
to route. The placement is scoped to the one request that carries the sentinel:
it never displaces an operator's classifier plugin, the bandit cannot reach above
the tier as raised, it never becomes the session pin, and the sentinel that
matched is recorded on the routing decision. Detection reads the newest ask
alone, so a title request quoted into a later turn cannot cheapen the work that
follows it, and a keyword rule, an escalation keyword or the plan-mode floor all
still decide over it.

Regenerates the lazy OpenAPI snapshot, which was already stale on the base for an
unrelated Presidio guardrail field and failed the schema check on every PR.

Resolves LIT-6349
2026-08-27 22:55:09 -07:00
tin-berri
b3322e30a7
fix(anthropic): handle per-level reasoning_effort flags without supports_reasoning (#38618)
Some checks failed
Unit Tests / caching-local (push) Has been cancelled
Unit Tests / core-utils (push) Has been cancelled
Unit Tests / enterprise-package (push) Has been cancelled
Unit Tests / enterprise-routing (push) Has been cancelled
Unit Tests / integrations (push) Has been cancelled
Unit Tests / All Other Providers (push) Has been cancelled
Unit Tests / Vertex AI (push) Has been cancelled
Unit Tests / misc (push) Has been cancelled
Unit Tests / proxy-auth (push) Has been cancelled
Unit Tests / proxy-endpoints (push) Has been cancelled
Unit Tests / proxy-extras (push) Has been cancelled
Unit Tests / proxy-infra (push) Has been cancelled
Unit Tests / proxy-server (push) Has been cancelled
Unit Tests / responses-caching-types (push) Has been cancelled
GitHub Actions Security Analysis / zizmor (push) Has been cancelled
Postgres Tests / schema-migration (push) Has been cancelled
Postgres Tests / proxy-behavior (push) Has been cancelled
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Has been cancelled
Unit Tests: Proxy DB Operations / auth-checks (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-utils (push) Has been cancelled
Unit Tests: Proxy DB Operations / budgets (push) Has been cancelled
Unit Tests: Proxy DB Operations / custom-logging (push) Has been cancelled
Unit Tests: Proxy DB Operations / db-and-spend (push) Has been cancelled
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Has been cancelled
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Has been cancelled
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Has been cancelled
Unit Tests: Proxy DB Operations / key-generation (push) Has been cancelled
Unit Tests: Proxy DB Operations / logging-misc (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-runtime (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-server-core (push) Has been cancelled
* fix(anthropic): handle per-level reasoning_effort flags without supports_reasoning

When a model has only per-level flags (e.g. supports_minimal_reasoning_effort: true)
but no explicit supports_reasoning flag, treat it as implicitly reasoning-capable.
This fixes gpt-5-search-api which declares minimal support but was incorrectly
degraded to low/minimal floor due to missing explicit supports_reasoning flag.

Test: verify per-level flag enables resolution path even without supports_reasoning.

Note: This change indirectly causes 20 azure deployments to forward max/xhigh
instead of degrading to high when requested, as these models now correctly
resolve their supported efforts through declared capability flags. This is
intended behavior (avoiding unnecessary degradation) but silent; operators
seeing increased latency/cost should check reasoning effort changes in logs.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(anthropic): explicit supports_reasoning=False wins over per-level flags

Greptile P1: the implicit-True branch bypassed the operator's explicit
supports_reasoning: false escape hatch when per-level flags were present
or inherited through the bare-twin lookup. Return () first on explicit
False, then apply the per-level implication only when the flag is unset.

Also drops a test comment that restated the test name (P2).

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-27 22:19:22 -07:00
tin-berri
39e5b0c2d1
fix(anthropic): drop and self-heal empty thinking blocks on /v1/messages (#38625)
* fix(anthropic): drop and self-heal empty thinking blocks on /v1/messages

* test(anthropic): pin early-signature carry across the blank thinking chunk skip
2026-08-27 21:41:56 -07:00
Andrew Mattie
134a4cd9fd fix(streaming): preserve provider model for cost calculation 2026-08-27 23:26:00 -05:00
yuneng-jiang
bd7e9c1997
Merge pull request #38624 from BerriAI/litellm_/circleci-regressions-review-f0a542
fix(exceptions): keep a refused connection an APIConnectionError
2026-08-27 21:16:37 -07:00
ryan-crabbe-berri
546a9aa39e
Merge pull request #36514 from ansh-agrawal/feature/enforce-model-rpm-tpm-on-create
feat(proxy): opt-in flags to require rpm/tpm on model and project create
2026-08-27 21:14:12 -07:00
ryan-crabbe-berri
56a80c8125 feat(ui): link team and key model chips to the models page filtered to that group
Model chips on the team info and virtual key info pages (overview and settings tabs) now link to /models-and-endpoints?model_group=<name>, which the All Models tab reads through a new nuqs-backed model group filter. Grant sentinels such as all-proxy-models stay plain badges. The selected group is also sent as the server-side search so the matching deployments are fetched even when they are beyond the first page.
2026-08-27 21:08:33 -07:00
ryan-crabbe-berri
3beb02e512
Merge pull request #38601 from BerriAI/litellm_ui_navbar_papercuts
fix(ui): one-click theme toggle and matching Docs/Blog styling in the top bar
2026-08-27 21:06:41 -07:00
ryan-crabbe-berri
a96555593e test(proxy): narrow pytest.raises to HTTPException/ProxyException
Satisfies the test-tree ruff gate (PT011, B017)
2026-08-27 21:05:48 -07:00
Yuneng Jiang
a4049b730c
fix(exceptions): keep a refused connection an APIConnectionError
#38318 taught exception_type to map upstream status codes for providers with
no branch of their own. It reads the status code off the exception, but
_handle_error stamps 500 onto every failure that never carried one, so a
refused connection reached the mapper wearing a status code nothing upstream
had sent, and came back as InternalServerError instead of APIConnectionError.

The two are not interchangeable to a caller: a 5xx says the provider answered
and failed, which the router treats as a reason to cool the deployment down,
while a connection error says the request never landed.

BaseLLMException now records whether its status code was received or
synthesized, _handle_error sets that when it invents the 500, and the status
mapper declines to act on a code litellm made up, so those failures fall
through to the APIConnectionError the branch was always meant to produce.

Genuine upstream 5xx responses are untouched, which the second test pins.
The search transformation assertion #38318 had loosened to InternalServerError
goes back to APIConnectionError for the same reason.
2026-08-27 21:00:24 -07:00