Commit graph

16473 commits

Author SHA1 Message Date
devin-ai-integration[bot]
6d0367ce35
feat(prometheus): expose per-key and per-team rate limit allowed and used gauges (#39236)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 18:03:18 -07:00
mateo-berri
6d8c18d518 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_gateway_injection_scope 2026-09-01 18:02:02 -07:00
tin-berri
81277252e1
fix(datadog_llm_obs): send tool calls, tool results and cache tokens in DD's own fields (#39222)
The LLM Obs callback copied litellm's OpenAI-shaped objects into the span
verbatim, so every field Datadog names differently landed somewhere it does
not read: tool calls kept their nested `function` wrapper instead of DD's
name/arguments/tool_id, tool messages carried no result linking them to their
call, the request's tools were never sent, and prompt-cache counts sat inside
meta.metadata rather than the span metrics its cache dashboards chart.

One rule governs the message mapper: add the fields Datadog declares, and never
destroy content it did not understand. Content collapses to its text only when
it has text, so a content list carrying tool or image blocks rides along
unchanged, and absent messages map to an empty input rather than a fabricated
turn. Tool calls and results are read from both dialects, the OpenAI
`tool_calls` / `role: tool` shape and the Anthropic `tool_use` / `tool_result`
content blocks, so /v1/messages sessions gain tool linking they never had.

Cache counts come from the same owners the savings dashboard uses, so every
provider spelling resolves through one place rather than a second local guess.
The three cache metrics partition the input count: litellm's normalized prompt
total includes both cache categories, as the cost calculator's pricing helper
documents, so the non-cached residual subtracts reads AND writes. Counting a
primed prefix as ordinary input had inflated non-cached usage by exactly the
cache-write count on every priming request.

Correlating a result to its call reads ids and names structurally and parses no
arguments, so a tool call's arguments are decoded once per span rather than
once per pass, and arguments past a size bound ship as the raw string instead
of paying a decode that multiplies memory on hostile compact JSON.

The flat `output_tool_calls.*` metadata copies go away with this: they were a
second representation of a fact that now has its own field on the same span.
2026-09-01 18:01:13 -07:00
mateo-berri
c9435b5ff3 fix(guardrails): key stream rewrites by choice index and scan delta-only responses buffers
Chat streaming write-backs now match chunks by the choice's index field
instead of its list position, delivering rewrites to the right choice on
n>1 streams; an ended-stream rewrite on a multi-choice buffer fails
closed since stream_chunk_builder collapses the choices. The Responses
fallback joins output_text.delta events when delivery is expected, so a
delta-only buffer is guardrail-checked instead of released raw.
2026-09-01 18:00:42 -07:00
tin-berri
48dd06e841
fix(bedrock): gate Converse cachePoint emission on model prompt caching support (#39210)
Bedrock rejects requests carrying cachePoint blocks for models whose entry in the cost map does not declare supports_prompt_caching (403 "You invoked an unsupported model or your request did not allow prompt caching"). Clients like Claude Code attach cache_control to every request, so any such model behind the gateway failed on every call. The new bedrock_model_accepts_cache_points predicate drops cachePoint emission for map-known non-caching models at all three emission funnels, keeps emitting for unmapped ids (application inference profile ARNs), and skips the gateway injection credit when the tool_config point is not placed.
2026-09-01 18:00:31 -07:00
devin-ai-integration[bot]
c001975152
fix(aiohttp_transport): map transport-internal CancelledError to a retryable ConnectError (#39240)
aiohttp shields its DNS resolution task; when the connector closes it cancels
that child, so the request task sees CancelledError without ever being
cancelled itself. map_aiohttp_exceptions() only caught Exception, so the
BaseException skipped transport mapping, router retries and proxy error
handling, and /v1/responses answered 500 "No response returned".

Catch CancelledError in the mapper, re-raise when the current task is really
being cancelled (Task.cancelling() > 0), and otherwise map it to
httpx.ConnectError so the usual retry, fallback and error mapping apply.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 17:55:47 -07:00
devin-ai-integration[bot]
93219a9257
fix(docker): install bedrock-realtime extra in monolith proxy images (#39223)
The Dockerfile, docker/Dockerfile.non_root and docker/Dockerfile.database uv sync stages never passed --extra bedrock-realtime, so aws-sdk-bedrock-runtime was absent from the image venv and Bedrock Nova Sonic /v1/realtime sessions failed with 'Missing aws_sdk_bedrock_runtime'. gateway/Dockerfile already had the extra (PR #34426).

Adds a static check over every uv sync in the proxy Dockerfiles and an image-level import probe that the image-scan workflow runs against the built root, non-root and gateway images.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 17:51:55 -07:00
devin-ai-integration[bot]
47b9d838aa
perf(scim): resolve group members with one user table read per member (#39228)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 17:49:40 -07:00
devin-ai-integration[bot]
04a25083a6
fix(cost-map): retry transient boot fetch failures and recover config deployments dropped by a stale cost map (#39230)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 17:49:06 -07:00
mateo-berri
c614d68d11 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_dashscope_rerank_endpoint 2026-09-01 17:48:40 -07:00
devin-ai-integration[bot]
69029c139e
fix(mcp): report per-server outcomes in aggregate REST tools/list (#39232)
GET /mcp-rest/tools/list without server_id returned only the tools of
the servers that answered and silently dropped any server whose listing
failed (for example an OAuth-protected server without credentials), so
clients could not tell a partial listing from a complete one.

The aggregate response now carries a server_outcomes map keyed by server
alias with the same classified outcome (ok/auth_required/forbidden/...)
that the MCP protocol path already puts in _meta. Healthy tools and the
HTTP 200 status are unchanged.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 17:48:01 -07:00
moe-berri
e3a61c82da test(router): register indirect session routing coverage 2026-09-01 17:46:41 -07:00
mateo-berri
ac19d0dbdf fix(spend): keep every-deployment scope on gateway cache-injection marks
The caching-savings marker litellm_gateway_injected_cache credits gateway-earned
prompt-caching savings to the deployment it names, or to every deployment via
the empty-string sentinel. Two paths lost that scope:

- the router prompt-management factory stamps a provisional deployment's
  model_info into kwargs before the prompt pass runs, so an injection recorded
  there named that provisional pick and a differently-billed deployment lost
  the credit
- record_gateway_injection overwrote on every positive delta, so a per-leg
  stamp (the Bedrock converse tool_config one included) downgraded an
  existing every-deployment mark and the leg billed after a failover lost
  the credit

record_gateway_injection now takes injected_for_every_deployment, the two
pre-choice callers declare it, and an every-deployment mark is never narrowed
by a later per-leg stamp. Per-leg marks still overwrite each other. Spend
amounts are untouched; only the savings attribution is affected.

Also unblocks make lint at the staging tip: tests/e2e/test_junit_properties.py
landed three basedpyright reds via an e2e-only PR whose lint job skipped, now
suppressed as the deliberate duck-typed double they are.
2026-09-01 17:44:29 -07:00
moe-berri
82d046c428 fix(router): make Claude session cleanup best effort 2026-09-01 17:43:48 -07:00
mateo-berri
8c7fe00d80 fix: compare stream event types by equality so typed completed events keep their usage 2026-09-01 17:37:07 -07:00
moe-berri
1cd99a036e fix(router): route Claude Code subagents through session router 2026-09-01 17:29:58 -07:00
mateo-berri
4fbe4ce2e2 fix(guardrail_translation): deliver stream rewrites on incomplete and failed responses terminals 2026-09-01 17:29:52 -07:00
mateo-berri
269f2336ec fix(dashscope): remap chat-shaped api_base to the live rerank route
DashScope serves rerank at /compatible-api/v1/reranks, but the chat default
and DASHSCOPE_API_BASE both hand rerank a /compatible-mode/v1 base, which
turned into the dead /compatible-mode/v1/reranks route. Chat-shaped
.aliyuncs.com bases now remap to the same host's rerank route, preserving
the mainland/intl region that API keys are scoped to. Explicit rerank
api_base values and the *_API_BASE_RERANK env vars still win. Applies to
dashscope, qwencloud, and qwen_ai_platform.
2026-09-01 17:19:14 -07:00
mateo-berri
7ed2e8acff Merge remote-tracking branch 'origin/litellm_internal_staging' into HEAD
# Conflicts:
#	tests/test_litellm/test_router.py
2026-09-01 17:17:24 -07:00
mateo-berri
fcd9052179 feat(proxy): honor model_info.display_name in the Anthropic-shaped /v1/models listing 2026-09-01 17:10:52 -07:00
mateo-berri
c4982ca407 fix(mcp): error instead of silent empty tools when scoped MCP access is denied; grant agent MCP servers from the UI 2026-09-01 17:04:48 -07:00
mateo-berri
ec677e5e74 Merge remote-tracking branch 'origin/litellm_fix_post_call_policy_pipeline' into litellm_post_call_pipeline_stream_rewrite
# Conflicts:
#	litellm/llms/anthropic/chat/guardrail_translation/handler.py
#	litellm/llms/openai/chat/guardrail_translation/handler.py
2026-09-01 17:01:42 -07:00
Mateo Wang
3dac3f7a36
Merge pull request #35816 from BerriAI/litellm_anthropic_stream_model_alias
fix(proxy): report requested model on Anthropic streaming message_start
2026-09-01 16:52:41 -07:00
tin-berri
59da6e75a5
feat(router): fall back on anthropic safeguard refusals on /v1/messages (#39157)
* fix(router): resolve fallbacks against the tier a pre-routing hook selected

A complexity or auto router picks a tier behind the router group name, but
fallback lookup kept using kwargs["model"], which is still the router name. The
tier's configured chain never ran, so a provider failure on its first hop went
straight back to the client with "No fallback model group found for original
model_group=smart-router".

The hook assigns the selected model to a local only, and fallback resolution runs
on an outer kwargs dict that **kwargs already copied, so writing it there is not
visible. Record the selection in the metadata bucket instead, which is a nested
dict shared by reference across those copies and is how the router already
carries values back up, then key fallback lookup off it when present.

Applies to the generic, context-window, content-policy and weighted-failover
lookups. Reporting keeps using the router name, since that is what the caller
asked for.

Fixes #38832

* fix(router): annotate the recorded-selection helper with a read-only mapping

record_pre_routing_selection only reads the request kwargs, writing into the
nested metadata bucket it finds there, so Mapping states what it actually needs
and clears the LIT001 mutable-annotation budget without a suppression.

* test(router): assert the no-kwargs path leaks nothing

The tolerated-None case called the helper without checking anything, which the
test-quality gate counts as a test with no assertion. Assert that a fresh mapping
still reads back empty, so the case proves the call is a no-op rather than only
that it does not raise.

* fix(router): stop declaring loop-assigned locals Final in the selection helpers

Both helpers annotated a loop-assigned local as Final, which reassigns a Final on
every iteration and cost three basedpyright errors. Read the buckets through a
generator instead, so the write path iterates a for-target and the read path
resolves in one shot with next(), which also matches the functional style the
type-discipline rules ask for.

* style(router): apply ruff format to the selection helpers

* fix(router): derive the pre-routing tier fresh on every fallback hop

The metadata buckets also carry whatever the caller sent, so an inbound
pre_routing_selected_model let a client pick which fallback chain its
request fell into. A fallback hop also inherited the previous hop's tier,
so the second hop keyed its own failure off the tier that already failed
and never ran its own chain.

Clear the key at the top of async_function_with_fallbacks. Every hop
re-enters there, so only the hook that routed that hop can set it.

* fix(router): drop the cast at the fallback-hop clear call site

* feat(router): fall back on anthropic safeguard refusals on /v1/messages

---------

Co-authored-by: Priyansh Nandwana <nandwana.priyansh103@gmail.com>
2026-09-01 16:50:12 -07:00
Devin AI
acc65b27d2 fix(router): pass container create/list through when model names no deployment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 23:39:42 +00:00
ryan-crabbe-berri
8f56dbe7a3
Merge pull request #39209 from BerriAI/litellm_e2e_junit_source_property
test: record each e2e test's source location in the JUnit report
2026-09-01 16:27:41 -07:00
Devin AI
4a68abfd49 fix(proxy): route container create and list through model_list deployments
Container create and list requests had no container ID to decode, so the
router called the provider handler directly and the OpenAI transformation
fell back to the global OPENAI_API_KEY. Proxies configured only with
model_list credentials sent Authorization: Bearer None. Route through
_ageneric_api_call_with_fallbacks when the caller passes a model, expose
the list endpoint's model query param to the router, and encode the
managed container ID on the async create path so follow-up calls route
to the same deployment.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 23:26:23 +00:00
ryan-crabbe-berri
fc1a5fd7f9
Merge pull request #39206 from BerriAI/litellm_lit_3925_clear_team_key_create
fix: stop a cleared Team field from blocking personal key creation
2026-09-01 16:21:52 -07:00
devin-ai-integration[bot]
3888a85045
fix(budget): reject known estimates over remaining budget under fail_closed_budget_enforcement (#39214)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 16:13:26 -07:00
ryan-crabbe-berri
c14cf9d173 fix(access_groups): reconcile update deltas against the derived team set
Team membership deltas on PUT now start from the teams that really carry
the group, so a team the mirror column missed can be detached. Read
endpoints go through a typed TeamRepository instead of the untyped db
handle, and the where clause always carries both OR arms.

Claude-Session: https://claude.ai/code/session_01QvQzYztinxj8ZuD5YxbVdL
2026-09-01 16:03:22 -07:00
ryan-crabbe-berri
0c2d4c5773 Refuse a source path carrying a colon
`path:line` cannot represent a path that itself contains a colon, and the
one way pytest produces one is a Windows absolute location: separator
normalization turns `C:\app\e2e\a2a\test_x.py` into `C:/app/...`, which
slipped past the leading-slash check and composed the nonsense repo path
`tests/e2e/C:/app/e2e/a2a/test_x.py`.

Reject the colon itself rather than special-casing a drive letter: it is
the character the format reserves, so no path containing one was ever
linkable.

Claude-Session: https://claude.ai/code/session_017dTKXwJkzhtVLzDhePHsKG
2026-09-01 16:02:54 -07:00
ryan-crabbe-berri
00e40c0afe Record each e2e test's source location in the JUnit report
The JUnit report is the only thing that leaves the e2e run, and it says
where a test's results came from but never where its code lives. A reader
looking at `test_cell_claimed_only_by_a_skipped_test_is_uncovered` on the
status page has a name and nothing else -- no file, no line, no way to
reach the source short of grepping the repo by hand.

Pytest knows the location; the report format loses it. The `xunit1` family
wrote `file=` and `line=` onto every `<testcase>`, and the `xunit2` default
this suite runs on drops both. Switching families back would change the
document for every consumer of the same XML -- the Buildkite Test Engine
upload and the Loki pipeline included -- so add the location the way this
suite already adds `package` and `covers`: as a `<property>`, which is
purely additive.

`source` is repo-relative and one-based (`tests/e2e/a2a/test_x.py:41`), so
a consumer can build a link without knowing how pytest was started. That
takes normalizing the two launch shapes -- the runner image runs from its
own copy at /app/e2e, a developer runs from the repo root -- which is the
same normalization `package_from_nodeid` was already doing in reverse, now
factored into `suite_parts` so the two cannot drift apart. Paths that
escape the suite, and tests pytest reports no line for, emit an empty
string: a test with no link beats a link that 404s.

Claude-Session: https://claude.ai/code/session_017dTKXwJkzhtVLzDhePHsKG
2026-09-01 16:02:54 -07:00
ryan-crabbe-berri
346efa0c33
Merge pull request #39197 from BerriAI/litellm_e2e_reliability_retry_context_window
test(e2e): cover retry-on-timeout and the context-window fallback
2026-09-01 15:51:01 -07:00
ryan-crabbe-berri
6a469c2159 fix(access_groups): derive attached teams from the team table and reject unknown team ids
GET /v1/access_group and GET /v1/access_group/{id} (and the /v1/unified_access_group aliases) used to
return the assigned_team_ids column verbatim. That column is a denormalized mirror of
LiteLLM_TeamTable.access_group_ids and can be stale or hold ids of teams that no longer exist, so the
Attached Teams view drifted from reality.

The read path now runs one team find_many per request, unioning teams whose access_group_ids carry any
group in the response with teams listed in the stored columns. Only real team rows come back, so ghost
ids drop out and teams the mirror missed are added. The stored order is kept for ids that survive and
newly discovered teams are appended.

Create and update now resolve the requested assigned_team_ids inside the transaction and answer 400
with the missing ids before anything is written, instead of silently storing ids that point nowhere.

Refs LIT-6593

Claude-Session: https://claude.ai/code/session_01QvQzYztinxj8ZuD5YxbVdL
2026-09-01 15:46:09 -07:00
Yujong Lee
5799a32cdd fix(vector-store): route pre-call searches through router 2026-09-01 15:37:50 -07:00
Yujong Lee
6805d01709 fix(vector-store): preserve aliases with embedding config 2026-09-01 15:33:18 -07:00
Yujong Lee
5635811726 fix(vector-store): route embeddings through router 2026-09-01 15:33:18 -07:00
Yujong Lee
bfa5eac76b fix(vector-store): resolve embedding aliases for search 2026-09-01 15:33:18 -07:00
ryan-crabbe-berri
9c417ba08b
Merge pull request #38969 from emerzon/litellm_strict_order_fallback
fix(router): keep order fallback on the requested order level
2026-09-01 15:21:59 -07:00
ryan-crabbe-berri
55d638412b fix: stop a cleared Team field from blocking personal key creation
Clearing the Team combobox in the Create Key modal left team_id set to an
empty string, so /key/generate treated the request as team key generation
and failed with a team-not-found error for non-admin members.

TeamDropdown now emits null on clear, and GenerateKeyRequest normalizes an
empty team_id to None so the request runs the personal key path.
2026-09-01 15:17:00 -07:00
ryan-crabbe-berri
45fa78470d test(router): inject the upstream client instead of mutating litellm.aclient_session
The text-completion wire test set litellm.aclient_session, which the
test-quality gate (TQ005) flags as a process-wide global write. Pass an
AsyncOpenAI client through the router's client kwarg instead, so the test
owns its transport and needs no cache flush or global restore.

Claude-Session: https://claude.ai/code/session_01XKkTFa6g7Rmd6vtHL91GMn
2026-09-01 15:14:56 -07:00
devin-ai-integration[bot]
97dbd8efcb
fix(docker): add public Wolfi apk repo to runtime image (#39033)
* fix(docker): add public Wolfi apk repo to runtime image

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(docker): accept quote variants in Wolfi repo assertion

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 15:11:15 -07:00
devin-ai-integration[bot]
846900320e
feat(alerting): slack alerts for per-user daily/monthly spend thresholds and spend anomaly detection (#38438)
* feat(alerting): slack alerts for per-user daily/monthly spend thresholds and spend anomaly detection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(alerting): use specific ValidationError matches in config rejection test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): tolerate mocked slack alerting args when scheduling user spend scan

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(alerting): reject non-finite values in user spend alert settings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 15:09:03 -07:00
devin-ai-integration[bot]
5988d93fed
fix(logging): guarantee max_parallel_requests slot release when streaming logging fails (#39093)
* fix(logging): guarantee max_parallel_requests slot release when stream logging fails

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(logging): cover guardrail branch of streaming logging hook failure isolation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 15:08:11 -07:00
devin-ai-integration[bot]
f9c6eda909
fix(cli): quote the Claude Code apiKeyHelper for cmd.exe on Windows (#39174)
* fix(cli): quote the Claude Code apiKeyHelper for cmd.exe on Windows

lite up and lite login --config-claude wrote the helper command with
POSIX shlex quoting, so a backslashed Windows install path came out
wrapped in single quotes that cmd.exe and PowerShell take literally.
Quote every token with the cmd.exe rules already used for agent shims
when running on Windows, and keep the POSIX output unchanged elsewhere.

Resolves LIT-6627

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cli): split the Windows apiKeyHelper with cmd.exe and C runtime rules

The invocation test pulled tokens back out with a regex, which cannot see
the doubled quotes or the percent guard quote_for_cmd emits. Model the two
parsers that read the helper on Windows instead and check argv round
trips for backslashed, spaced, metacharacter, percent and quoted tokens

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 15:08:00 -07:00
ryan-crabbe-berri
0b89c59be2 fix(router): consume _target_order at deployment selection so it never reaches a provider
Reading _target_order with .get left it in the request kwargs after selection, and only
nine provider boundaries stripped it. _atext_completion and _aadapter_completion spread
the raw kwargs, so an order-2 hop on /completions sent _target_order upstream, which
real providers reject as an unknown argument. Popping at selection strips it for every
path in one place; the PR's retry-keeping test already passed with pop because each retry
hands the callee its own kwargs copy.

Claude-Session: https://claude.ai/code/session_01XKkTFa6g7Rmd6vtHL91GMn
2026-09-01 15:02:01 -07:00
ryan-crabbe-berri
af11db9fe5 test(e2e): cover retry-on-timeout and the context-window fallback
Two P0 rows in the reliability coverage registry had no test.

reliability.retry.timeout.succeeds_within_retries gets a new file. The model
group is a pair: an always-timing-out deployment holding all of the group's
shuffle weight, and a healthy backup at weight 0. The weighted pick always opens
on the timing-out one, its first Timeout benches it via an allowed_fails_policy
of TimeoutErrorAllowedFails 0, and the retry falls through to the only
deployment left, so the outcome is a completion plus a reported retry with no
random first pick in the middle.

reliability.fallback.context_window.routes_to_fallback joins the existing
fallbacks spec. It registers a genuinely small-context OpenAI deployment, sends
a prompt past its limit so the provider refuses it on length, and reroutes with
context_window_fallbacks, which is the setting that handles that refusal rather
than plain fallbacks.

Both drive real provider calls through router_settings_override, so no config
change and no second proxy is needed. Reliability & Performance goes 16/36 to
18/36.

Claude-Session: https://claude.ai/code/session_01QvQzYztinxj8ZuD5YxbVdL
2026-09-01 14:51:15 -07:00
mateo
4da12795fc fix: filter deployment default API key limits
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-01 21:50:45 +00:00
mateo-berri
a38dfecd96 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_stream_modify_response_chunks 2026-09-01 14:49:06 -07:00
mateo-berri
d59fcda8af fix(rerank): adopt declared authenticating providers in arerank instead of resolving them
get_llm_provider runs the OAuth device flow for github_copilot and chatgpt,
so calling it on the event loop before the executor dispatch let an
authenticated caller block the loop for the length of the polling window.
Adopt the declared provider via declared_authenticating_provider, matching
the metadata callers in utils.py, and only resolve for everything else.
2026-09-01 14:47:59 -07:00