Commit graph

42067 commits

Author SHA1 Message Date
ryan-crabbe-berri
7745fe887f Merge branch 'litellm_window_spend_writer' into litellm_window_spend_reader 2026-08-29 11:06:05 -07:00
ryan-crabbe-berri
041cae8280 fix(proxy): bound the window spend seed exclusion to the batch's own start time
request_id can be chosen by the client through x-litellm-call-id, so an
unbounded NOT (request_id = ANY(batch)) let a replayed old id drop that id's
historical LiteLLM_SpendLogs row from the one-time seed while its increment
still landed. The increment now carries the request start, the batch keeps
the earliest one, and the seed only excludes ids whose startTime is at or
after it.
2026-08-29 11:05:57 -07:00
ryan-crabbe-berri
e7d99079de Merge branch 'litellm_window_spend_writer' into litellm_window_spend_reader 2026-08-29 10:51:50 -07:00
ryan-crabbe-berri
83cedfd2b2 refactor(proxy): type the budget window reset_at passed to the window spend enqueue 2026-08-29 10:51:38 -07:00
ryan-crabbe-berri
124d08d592 perf(proxy): read budget-window spend from the maintained window table
Per-window budget enforcement aggregated LiteLLM_SpendLogs on every cold
counter and on every 5s authoritative floor check. SpendLogs has no index on
api_key or team_id, so each check range-scanned the highest-volume table.

Reads now hit the LiteLLM_BudgetWindowSpend row by primary key and only fall
back to the aggregate when no row exists for the window being enforced. A row
is current when its window_start is at or past the caller's expected start, so
a pod holding a stale reset_at trusts a window another pod already rolled
instead of summing the previous window back in.

window_duration is threaded from the budget_limits entry through to the read
rather than parsed back out of the counter key. The reader never writes rows.
2026-08-04 18:56:32 -07:00
ryan-crabbe-berri
187e0fab60 feat(proxy): maintain per-window budget spend rows in the spend writer
Multi-window budgets (budget_limits on keys and teams) enforce off Redis
counters with a 60s TTL. Every cold counter, and every authoritative floor
check, aggregates LiteLLM_SpendLogs with a range scan over startTime; that
table has no index on api_key or team_id, so the scan lands on the highest
volume table in the schema and saturates the connection pool (#35766).

Every other budget feature reads a maintained running spend value instead.
This gives windows the same shape by keeping LiteLLM_BudgetWindowSpend up to
date: one row per configured window, whose window_start rolls forward in
place.

The cost callback already iterates a key's and a team's windows with the
window start computed and the actual cost in hand, so it enqueues there, onto
a new WindowSpendUpdateQueue. Increments are enqueued even when the cache
increment is skipped for a reserved counter, since the reservation only
pre-charged an estimate and the row still owes the actual cost. Windows with
no reset_at slide with wall clock and cannot be represented by a single row,
so they are left to the read path's existing aggregate.

The queue flushes alongside the daily spend queues, through the Redis buffer
when one is configured (only the pod-lock winner commits) and directly
otherwise. A flush selects the primary keys that already exist, seeds the
ones that do not from LiteLLM_SpendLogs so a new row cannot undercount spend
that predates it, and applies one batch of upserts ordered by primary key. An
increment at or behind the stored window_start adds into the row, matching how
in-flight requests carry into a window after a reset; a newer one rolls the
window and starts from that increment. The conflict arm adds only the
increment, never the seeded base, so two pods seeding the same new window
cannot double count it.

The seed excludes the requests its own batch is about to apply. Spend logs are
drained by a separate monitor that fires on a ~2s poll whenever anything is
queued, while window increments flush on the much slower batch tick, so by
seed time the batch's log rows are normally already in the table; counting
them in the aggregate and again in the increments made a fresh row land at
exactly twice the true spend. Each increment therefore carries the
LiteLLM_SpendLogs request_id it was recorded under, which update_database now
returns rather than having the callback re-derive it (a cache hit appends
time.time() to that id, so a second derivation would not match).

The reset job rolls each expired window's row alongside the counter it zeroes,
conditional on the stored window_start still being behind the new one so a
pod that already rolled it is not clobbered.
2026-08-04 18:55:48 -07:00
ryan-crabbe-berri
3b1ab1908c feat(proxy): add LiteLLM_BudgetWindowSpend table for per-window budget spend
Multi-window budgets (budget_limits on keys/teams) currently keep window
spend only in cache. Every cold or expired counter recomputes the window
by aggregating LiteLLM_SpendLogs, which has no usable index for that
query and saturates the DB on large tables (#35766).

This adds a LiteLLM_BudgetWindowSpend table holding one row per
configured window, keyed (entity_type, entity_id, window_duration),
with window_start identifying the period the spend belongs to.
Follow-up PRs maintain these rows from the spend update writer and move
window budget enforcement reads onto them.
2026-08-04 16:56:37 -07:00
tin-berri
ffcb54b06d
feat(ui): reorder Add Auto Router into name + template, with a collapsible detailed config (#35746)
* feat(ui): add template picker to the Add Auto Router flow

Add Auto Router now opens straight into name + an optional Template
dropdown (Anthropic/OpenAI model-family presets or Custom). A preset
prefills the full complexity-router config and collapses the Detailed
Configuration section to a one-line tier summary; choosing Custom (or
nothing yet) leaves it expanded, and a caller can toggle it manually
at any point. A preset option greys out with the specific missing
model(s) named when the caller lacks a model it needs, or while the
model list is loading or failed to load.

Prefill and submit-gating logic live in testable pure functions
(buildPresetPrefill, getReferencedModelsError) rather than inline in
the component, per the dashboard's own testing guidance.

* refactor(ui): memoize presetAvailability

Consistency with the other memoized derived values it closes over
(availableModelSet, presets). Negligible perf impact with two
presets today, but keeps the pattern uniform as more get added.

* refactor(ui): drop pointless useMemo around getAllPresets()

getAllPresets() already returns a stable module-level array
reference; wrapping it in useMemo added React machinery for
something that can't change.

* refactor(ui): hoist presets to module scope

getAllPresets() was still being called from inside the component
body on every render even after dropping the useMemo wrapper.
Resolving it once at module load, alongside PRESETS' own
module-level initialization in autorouter_presets.ts, is the
actually-clean version of the previous fix.

* fix(ui): collapse Detailed Configuration by default

It was defaulting to expanded before any template was chosen, so
the modal still opened onto the full tier/classifier form instead
of just Name + Template. Custom still auto-expands it, and a
preset still collapses it after prefilling.

* fix(ui): list Custom Configuration last in the Template dropdown

Custom is the escape hatch, not the headline choice, so the bundled
presets now come first with Custom listed after them.

Also lets the collapsed Detailed Configuration summary wrap onto
its own line(s) instead of sharing a line with the section label
and truncating mid-model-name.

* feat(ui): match preset models across "-"/"." version separators

Admins spell version numbers inconsistently (claude-sonnet-4-5 vs
claude-sonnet-4.5), so a preset's hardcoded name and a caller's
registered one can refer to the same model while differing only in
that punctuation. getMissingModels (and therefore presetAvailability
and the submit-blocking check) now treats the two as equivalent.

Applying a preset writes the caller's actual registered spelling
into the tiers, not the preset's literal string, since the caller
may only have the dotted (or hyphenated) form and never the other
one - buildPresetPrefill now takes the available-models set for
this rewrite. Two different model names never collide; only the
separator within one version number does.

* fix(ui): re-check referenced models inside submitRecommendedRouter

submitBlockedReason disables the button for a stale/missing model
reference, but Form's onFinish (wired to the same handler) fires on
a real form submission regardless of the button's own disabled
state. The other four blocking checks already re-validate inside
submitRecommendedRouter for this exact reason; this one was missing
it, so a router could still be created referencing a model no
longer in availableModelSet.

Found by Bugbot.

* Update autorouter_presets.json
2026-08-04 23:24:18 +00:00
Mateo Wang
1798d9d2a9
Merge pull request #35834 from BerriAI/litellm_fix_cursor_variant_budget_bypass
fix(proxy): enforce per-model budgets against resolved cursor model variants
2026-08-04 16:14:06 -07:00
ryan-crabbe-berri
f4538679c0
fix(proxy): apply key_alias/key_hash filters to all /key/list visibility branches (#35840)
* fix(proxy): apply key_alias/key_hash filters to all /key/list visibility branches

The filters previously lived only in the own-keys OR branch, so a team admin's admin-team branch matched every team key and the Key Alias filter in the Virtual Keys UI appeared broken. Both filters are now global AND conditions alongside team_id/project_id/access_group_id/agent_id, narrowing every visibility branch while leaving unfiltered visibility unchanged.

* chore: drop new explanatory comments flagged by review

* chore: restore schema.d.ts to base enum order
2026-08-04 23:07:57 +00:00
yucheng-berri
58ead7f653
fix(azure_storage): honor AZURE_STORAGE_ENDPOINT_SUFFIX for sovereign clouds (#35806)
The azure_storage logging callback and the azure blob files backend built every
storage URL against the hardcoded commercial host, so an Azure Government account
was unreachable with no way to override it.

Read AZURE_STORAGE_ENDPOINT_SUFFIX (default core.windows.net) once in
AzureBlobStorageLogger and derive the Data Lake and Blob hosts from it, so all
seven previously hardcoded sites follow the configured cloud. Parse stored blob
URLs with urlparse instead of matching the commercial host, so URLs persisted
before the suffix was configured still resolve, and pin the resulting
host-validation boundary with tests.
2026-08-04 16:06:41 -07:00
Deepanshu Lulla
e2950a8995
fix(router): eagerly fetch Vertex AI deferred stream to surface HTTP errors in _acompletion fallback path (#34627)
* fix(router): eagerly fetch deferred stream to surface HTTP errors in fallback path

Providers like Vertex AI and Bedrock defer their HTTP call until the first
__anext__ on the returned CustomStreamWrapper (completion_stream=None,
make_call set). Errors raised inside __anext__ (e.g. 429, 503) escape the
_acompletion try/except block, so fail_calls is never incremented, deployment
cooldown does not fire, and the standard fallback chain is bypassed.

Call fetch_stream() on the wrapper before delegating to
_acompletion_streaming_iterator when completion_stream is None and make_call
is set. Any HTTP error now propagates through _acompletion's except block,
increments fail_calls, and enters the normal retry/fallback chain.

Strip Content-Length, Transfer-Encoding, Content-Encoding, and Content-Type
from exception headers at the same point to prevent HTTP framing mismatches
when LiteLLM builds its own error response body.

Add a re-raise guard in _acompletion_streaming_iterator (async and sync paths)
so MidStreamFallbackError with already-generated content re-raises to the
caller instead of silently injecting a continuation prompt into a fresh request
to a fallback model.

Apply logging cleanup in async_function_with_fallbacks_common_utils: use
%s-style formatting and exc_info=True instead of f-strings with
traceback.format_exc().

* fix(router): undo success_calls on deferred-stream fetch failure; broaden header strip

* fix(router): extract header-strip helper to keep _acompletion under strict C901 threshold

* test(router): add unit tests for _strip_http_framing_headers to satisfy router coverage gate

* test(router): add sync _completion_streaming_iterator re-raise test for mid-chunk MidStreamFallbackError

* fix(router): restore Fallbacks context in no-fallback log; document update_team mcp_rpm_limit

The log and debug message when no fallback model group is found was missing
the Fallbacks list, making it hard to understand why routing failed.

Also adds the missing mcp_rpm_limit documentation to update_team to fix
the documentation_test_api_docs CI check.

* fix(router): preserve original traceback in deferred stream fetch error re-raise

Using bare `raise` instead of `raise fetch_err` keeps the full inner
traceback from fetch_stream() intact so the error origin is visible in
logs and debuggers without being anchored to this line.

* style(test): restore black-style formatting in test_router.py

An earlier commit on this branch collapsed the file's pre-existing
multi-line formatting into single lines while adding the deferred-stream
tests, producing a diff full of unrelated reformatting noise. Restores
the untouched code to its original formatting; the actual new/changed
test content is unaffected (verified via AST comparison).

* fix(router): re-raise mid-stream fallback on any generated content, not just text

The re-raise guard added for MidStreamFallbackError only checked
generated_content, which tracks text deltas alone. A stream that emitted a
tool-call or reasoning-only chunk before failing had generated_content=""
despite already streaming to the client, so the router silently retried
and the client saw duplicated/inconsistent output. The guard now also
inspects the wrapper's raw chunks for tool_calls/reasoning_content.

Also moves the deferred-stream HTTP-framing-header stripping out of
Router._acompletion into the proxy's _handle_llm_api_exception: Router is
used directly as an SDK as well as by the proxy, and stripping headers
there dropped legitimate provider metadata (content-type,
proxy-authenticate) for direct SDK callers who never see the proxy's own
response construction.

schema.d.ts regenerated via make pre-commit; unrelated to this change.

* test(router): add direct coverage for _stream_chunks_have_generated_content

CI's router_code_coverage check flags any router.py function never referenced
by name in a test file; the new helper was only exercised indirectly through
the mid-stream re-raise guard tests.

* revert(ui): drop incidental schema.d.ts regeneration

Committing router.py/common_request_processing.py touched
pre_commit_lint.sh's litellm/proxy trigger for the API-type-sync check,
which force-regenerated schema.d.ts even though neither file changes any
route or model. The regenerated ordering of two unrelated Union/enum
fields (stream_timeout, user_role) isn't stable across process
invocations even against completely unmodified backend code (confirmed
by regenerating twice against the pre-existing committed code and getting
the same diff both times), so this reverts to the original committed
file rather than chase non-deterministic output.

* fix(proxy): strip framing headers on the pre-existing ProxyException branch too

_handle_llm_api_exception filtered framing headers into a local `headers`
dict, but for an exception that's already a ProxyException, it merged
{**e.headers, **headers}: the original e.headers came first, so a framing
header present there but absent from the filtered `headers` (because it
was just stripped) was never overwritten and survived into the response
unfiltered. Filters the merged result instead of relying on the merge
order to do it implicitly.

* chore: retrigger CI (no GitHub Actions check-suite was created for the previous two pushes)

* fix(router): detect thinking_blocks as generated content in mid-stream guard

Greptile flagged that a thinking-only delta (Anthropic extended thinking,
Delta.thinking_blocks) wasn't recognized as already-streamed content, so
a stream that emitted only thinking blocks before failing could still
restart via fallback and append an unrelated response after content the
client already received.

* fix(proxy): strip browser-facing security headers from provider exceptions too

veria-ai flagged that the framing-header denylist still let a malicious or
misconfigured provider set browser-facing headers (Access-Control-Allow-Origin,
Content-Security-Policy, Clear-Site-Data, etc.) on the proxy's own error
response. Adds a dedicated _BROWSER_SECURITY_HEADERS set alongside the
existing framing one and strips both wherever provider exception headers
reach the client response.

* refactor(router): address maintainer review mechanicals

- List[ModelResponseStream] -> list[ModelResponseStream] in
  _stream_chunks_have_generated_content (ruff UP006 strict-budget gate)
- drop _strip_http_framing_headers and its 3 tests: the proxy inlines the
  filter directly now, so the helper has had no production caller since
  the header-stripping was moved out of Router
- move HTTP_FRAMING_HEADERS/BROWSER_SECURITY_HEADERS/
  UNSAFE_PROXY_RESPONSE_HEADERS from router.py into litellm/constants.py,
  removing the router.py <-> proxy import path the two CodeQL
  cyclic-import alerts were pointing at
- move the eager fetch_stream() call before success_calls/logging/
  _track_deployment_metrics instead of incrementing then compensating
  with a manual decrement on failure
- fix a dead assert message: `mock_fallback.assert_not_called(), "..."`
  built a tuple, not an assert-with-message; assert_not_called() already
  raises on its own so this just drops the inert string

* revert(router): pull mid-stream continuation-removal out of this PR

Removing the continuation-prompt fallback (retrying with the partial
response as a prefixed assistant message) so a stream failing after
partial content always re-raises instead was a scope decision beyond
what this PR's title/issue (#31874) describe, and it directly conflicts
with #30242/#30743, which are already fixing the same code path for
Anthropic's removal of assistant-message prefill on Sonnet 4.6+/Opus
4.6+. Landing this PR's version first would delete the branch those PRs
are patching; landing theirs first would have this PR undo their fix on
rebase.

Restores the original prefill-based continuation-resume behavior
(including the is_pre_first_chunk guard already in litellm_internal_staging)
in both _acompletion_streaming_iterator and _completion_streaming_iterator,
and removes _stream_chunks_have_generated_content along with the tests
that only existed to cover the guard. This PR now only touches the
deferred-stream eager-fetch fix and the header-stripping fixes; the
non-text-content re-raise idea becomes a follow-up PR built on top of
whichever of #30242/#30743 lands.

* fix(proxy): re-filter unsafe headers after the response-headers hook merge

_handle_llm_api_exception filtered provider/framing headers once, then
merged in post_call_response_headers_hook's return value afterward
without re-filtering. The ProxyException branch happened to re-filter
after its own header merge, but the HTTPException/httpx.HTTPStatusError/
generic-exception branches passed the post-hook headers straight through
unfiltered, so a callback hook (any custom guardrail/logging plugin)
returning an unsafe header would bypass the strip entirely for those
paths. Filters once, right after the hook merge, so every branch gets
the same guarantee.

* Revert "revert(router): pull mid-stream continuation-removal out of this PR"

This reverts commit c5ca101f61746a9b12a480c4bc48d95fc0c69f8d.

* fix(router): detect reasoning_items as generated content in mid-stream guard

Greptile flagged that a structured reasoning-only delta (Delta.reasoning_items,
the OpenAI Responses-API-style reasoning item) wasn't recognized as
already-streamed content by _stream_chunks_have_generated_content, alongside
the existing thinking_blocks/tool_calls checks, so a stream that emitted only
reasoning_items before failing could still restart via fallback.

* fix(router): annotate _stream_chunks_have_generated_content with Sequence, not list

The type_discipline_gate LIT001 check flags mutable-collection parameter
annotations. chunks is only iterated, never mutated, so Sequence is the
correct read-only annotation and clears the ratcheted budget ceiling.

* fix(router): surface original provider exception, not the internal wrapper, when mid-stream fallback gives up

When content has already streamed and MidStreamFallbackError carries
original_exception (e.g. RateLimitError), both the async and sync
streaming iterators bare-re-raised the wrapper itself, so the client
lost the specific error type/code/provider_specific_fields instead of
seeing the real provider error. The fallback-failure path a few lines
below already unwraps to original_exception for the same reason; apply
the same pattern here.

Also extend _stream_chunks_have_generated_content to recognize audio,
images, and annotations deltas as generated content, matching
is_chunk_non_empty's existing annotations check and Delta's treatment
of audio/images as first-class content fields — a stream carrying only
one of these before failing was not recognized as already-streamed,
so the router could still restart it via fallback after the client had
received real content.

* chore: retrigger CI (frontend-lint cancelled, schema.d.ts flake)

frontend-lint's check-run shows conclusion=cancelled on 70e47f4897 with
no superseding run, and this PR touches no UI files. Verify schema.d.ts
matches the proxy OpenAPI spec is on the previously diagnosed
stream_timeout/user_role Union-ordering nondeterminism (e9fc5e5063).
Empty commit to force a fresh CI run for both rather than a manual
rerun, which requires repo admin rights this fork PR doesn't have.

---------

Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com>
2026-08-04 22:44:43 +00:00
ryan-crabbe-berri
9ea5cfce0e
fix(proxy): persist periodic reload schedule state so status survives restarts and fires without store_model_in_db (#35165)
* fix(proxy): persist periodic reload schedule state so status survives restarts and fires without store_model_in_db

The model cost map and Anthropic beta headers reload schedules kept their
last-run time in a per-pod module global, so GET /schedule/*/status reported
last_run null after any restart and the Admin UI showed the reload as never
having run. The reload check also only ran from the add_deployment job, which
is registered only when store_model_in_db is true, so config-file deployments
stored a schedule that never fired.

Persist last_run_at and reload_requested_at as dedicated columns on
LiteLLM_Config, owned by the reload job and manual reload endpoints, while the
schedule endpoints own the param_value JSON (interval_hours); no writer can
clobber another's fields. Serve status entirely from the row. Register the
check as its own periodic_reload_job outside the store_model_in_db gate.
Replace the force_reload boolean with a reload_requested_at timestamp each pod
compares against its own in-memory last reload, so a manual reload reaches
every pod exactly once instead of being cleared by the first poller. Run the
blocking fetches via asyncio.to_thread, and stamp last_run_at with update_many
so a schedule cancelled mid-poll is not resurrected.

* fix(proxy): compare reload requests against pod data age seeded at boot

A pod that had never reloaded kept its in-memory clock at None, and with no
interval configured nothing ever set it, so every manual reload request was
ignored by every pod except the one serving the click (Greptile P1 on the
previous commit). Seed the per-pod timestamp at boot as the time its data was
loaded and reload whenever a request or the interval is older than that, which
also removes both None special cases from the due predicate. A schedule whose
row has no last_run_at fires on the next tick so the first run does not wait a
full interval.

* fix(proxy): scope reload persistence to the model cost map and seed the pod clock from the actual load time

Revert the Anthropic beta headers reload path to its previous JSON-flag
implementation so this PR only changes the price data reload; the beta headers
path keeps working exactly as before and can migrate to the shared module in a
follow-up. The unused columns on its config row are inert.

Seed model_cost_map_loaded_at from the timestamp get_model_cost_map records at
the actual import-time fetch instead of ProxyConfig construction time, closing
the startup window where a manual reload request stamped between the fetch and
the constructor compared as older than the pod's data and was skipped
(Greptile P1 on the previous commit).

* refactor(proxy): drop the legacy force_reload backfill from the reload tracking migration

The backfill only carried over a manual reload clicked in the seconds before an
upgrade, and every upgrade restarts the pods, which re-fetch the cost map at
import and so already deliver what that request asked for. Removing it makes
the migration schema-only, so prisma db push and prisma migrate deploy leave
the database in the same state instead of diverging on a data statement that
only one of them runs.

* fix(proxy): stamp reload timestamps at the precision they are stored at

Postgres stores these columns as TIMESTAMP(3) while Python stamps microseconds,
so a pod comparing its in-memory clock against the persisted copy of the same
instant read as newer and skipped the reload request it had just recorded.
Truncate every stamp to milliseconds at the source, and floor the boot seed the
same way, so the in-memory value and its persisted copy compare exactly.

* fix(proxy): identify manual reloads by revision instead of comparing timestamps

Comparing a request timestamp against each pod's data age made correctness depend
on clock resolution: Postgres stores TIMESTAMP(3) while Python stamps microseconds,
and two events inside the same millisecond are indistinguishable no matter how the
comparison is written.

Replace reload_requested_at with a reload_revision counter the manual reload
endpoint increments atomically in the database. Each pod records the revision it
last applied and reloads whenever the row's differs, so a request reaches every pod
exactly once regardless of clock skew or precision, and concurrent requests publish
distinct revisions instead of overwriting one another. A pod adopts the current
revision on its first poll, since data it loaded at boot already satisfies any
earlier request. Interval reloads still key off the pod's own data age, where hour
scale comparisons make precision irrelevant.

* fix(proxy): seed the applied reload revision at startup

A pod adopted whatever revision it found on its first poll, so a manual reload
published while the pod was starting was marked applied without ever being
served and the pod kept the prices it fetched at import. Read the row once at
startup instead, right after that fetch, and treat a missing row as revision 0

* style(tests): revert incidental reformatting of test_proxy_server.py

An earlier ruff format run reflowed the whole file from its 88-column
formatting, adding ~1150 lines of churn unrelated to this PR. Replay only
the real test changes onto the original formatting

* fix(proxy): serve an outstanding reload request on a booting pod

Seeding the applied revision at startup left a window: a manual reload
published after the import-time cost map fetch but before startup read the
row was marked applied without ever being fetched, stranding that pod on
stale prices when no interval was configured. A pod now starts unapplied and
serves any outstanding request on its first poll, which costs one redundant
fetch per boot and removes the window along with the seeding step

* fix(proxy): accept a reload interval still encoded as JSON text

param_value is written with safe_dumps, and a raw row read can return it
decoded or as a string depending on the driver. Strict validation rejected
the string, so the schedule read as disabled and an admin's configured
reloads silently stopped. Mirrors the guard ConfigRepository.get_param
already carries for the same column

* fix(proxy): cancel a reload schedule without resetting the revision

* fix(proxy): null the interval in JSON so cancelling keeps the revision

prisma rejects a null literal for a Json? column, so update_many writes an
interval-less object instead. The fake config table now rejects the same input
the database does, which is what the live run caught and the mock did not.

Also records the run before adopting the revision, so a failed status write
leaves the request unserved for the next poll rather than reporting a run that
never landed.

* fix(ui): match the CI-generated user_role union order in schema.d.ts
2026-08-04 15:42:57 -07:00
mateo-berri
334e10470b refactor(proxy): make the cursor responses-path body single-assignment 2026-08-04 15:21:35 -07:00
Yassin Kortam
d1ca826ff6
docs(helm): replace the classic chart's 128Mi resource example with the documented 4Gi sizing (#35830)
The litellm-helm values file shipped the stock helm create boilerplate for
resources: an empty default plus a commented 100m/128Mi example it invites
operators to uncomment. 128Mi is roughly 32x below what the proxy needs at
DB-connected steady state, and it was the only sizing figure this chart ever
showed, so operators who followed it were sized for OOMKills.

Point the example at the documented 1 CPU / 4Gi per worker instead, link the
production sizing guidance, and note why the default stays unset. The
migration job's commented block carried the same trap with a 100m/100Mi
example; drop those numbers rather than substitute proxy figures that do not
transfer to a job that migrates and exits.

The defaults are deliberately left at {} so no existing release changes shape
on upgrade; rendered output is unchanged.
2026-08-04 15:17:12 -07:00
yuneng-jiang
e4fd790f1c
Merge pull request #35835 from BerriAI/litellm_/elated-margulis-7f300f
refactor(ui): route MCP session tokens through the shared storage helper
2026-08-04 15:03:29 -07:00
yuneng-jiang
e64536c425
test(e2e): retry provider-transient statuses at the transport with bounded backoff (#35824)
* test(e2e): retry provider-transient statuses at the transport with bounded backoff

The Anthropic passthrough cost test failed a full-suite run on a real 529
overloaded_error. Passthrough routes forward provider responses verbatim
and bypass the router's num_retries, so provider blips reach the harness
only on those paths. Following standard practice, the retry is scoped to
the dependency boundary instead of rerunning tests: only the enumerated
transient statuses (500/502/503/504/529, the set production SDKs retry by
default) are retried, with bounded exponential backoff and a printed line
per retry so flakiness stays visible in run logs.

429 is deliberately excluded: the quota suites assert the proxy's own
rate-limit and budget 429s, and a transport that absorbed them would break
those tests. Network errors and timeouts are not retried either, so a hang
surfaces as a hang. request_with_retry takes injected callables, and the
new harness tests pin the contract with protocol fakes, no monkeypatching

* test(e2e): narrow the transport retry to 529, the one status the proxy cannot emit

Greptile's review is right that status-only classification could absorb an
intermittently failing proxy: at the transport a 500/502/503/504 from the
proxy is indistinguishable from one it relayed, and the proxy is the system
under test. 529 is the only status litellm provably never originates
(Anthropic's overload signal, forwarded verbatim on passthrough) and the
only transient observed across the full-suite runs, so the set shrinks to
exactly that. The canary tests now also pin 500/502/503/504 as never
retried
2026-08-04 14:57:42 -07:00
Mateo Wang
05204795cf
Merge pull request #35828 from BerriAI/litellm_zero_local_basedpyright_headroom
chore(lint): zero out basedpyright headroom for purely local rules
2026-08-04 14:51:49 -07:00
Mateo Wang
4395e974db
Merge pull request #35825 from BerriAI/litellm_claude_md_em_dash_order
docs(CLAUDE.md): prefer commas over semicolons when replacing em dashes
2026-08-04 14:47:32 -07:00
mateo-berri
1cd481d4f2 fix(proxy): enforce per-model budgets against resolved cursor model variants 2026-08-04 14:36:04 -07:00
mateo-berri
ae54f0c95d docs(CLAUDE.md): add colon to em dash replacement list 2026-08-04 14:34:01 -07:00
Mateo Wang
27885076e7
Merge pull request #35738 from BerriAI/litellm_bedrock_tool_choice_parallel_conflict
fix(bedrock): drop conflicting tool_choice.type when toolConfig.toolChoice is set
2026-08-04 14:30:10 -07:00
mateo-berri
bd7d270e17 docs(CLAUDE.md): weight punctuation variety instead of defaulting to comma 2026-08-04 14:24:29 -07:00
mateo-berri
98d4f9151c chore(lint): zero out basedpyright headroom for purely local rules 2026-08-04 14:20:55 -07:00
Mateo Wang
98fed43ae7
chore: make it more concise 2026-08-04 14:19:56 -07:00
Yuneng Jiang
560a4ac891
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/elated-margulis-7f300f 2026-08-04 14:16:02 -07:00
Yuneng Jiang
f1383f16fa
refactor(ui): route MCP session tokens through the shared storage helper
mcpTokenStore was the only OAuth path writing straight to window.sessionStorage;
useMcpOAuthFlow, useToolsOAuthFlow, the callback page and the edit-screen UI state
all already go through secureStorage. Align it so the OAuth surface has one storage
format instead of two.

The stored payload also carried a refresh_token that nothing ever read back. All
three read sites take access_token only, and nothing reads the mcp-session-token:
keys directly, so the field was write-only. Drop it from the store and from the four
callers that populated it. The client-forwarded modes (true_passthrough and
oauth_delegate) re-authorize rather than refresh, and authorization_code is
unaffected because it persists through storeMCPOAuthUserCredential on the backend,
which keeps its own refresh token.

Entries written before this change decode to null and are treated as absent, which
surfaces the normal Authorize prompt; they are session-scoped and expire in an hour.

Add two regression tests that decode the stored value before asserting, so neither
can pass merely because the payload is no longer plain text.
2026-08-04 14:15:29 -07:00
yuneng-jiang
5045a576ad
bump: litellm-proxy-extras 0.4.81 -> 0.4.82, litellm 1.96.0 -> 1.97.0 (#35810) 2026-08-04 21:14:36 +00:00
Mateo Wang
5159cba6b0
Merge pull request #35807 from BerriAI/litellm_enforce_final_variables
feat(lint): enforce Final on locals and freeze function parameters (LIT010, LIT011)
2026-08-04 14:13:46 -07:00
yuneng-jiang
e86f2209a4
test(e2e): move load/perf testing out of the main suite and drop the vllm passthrough test (#35820)
The Locust throughput SLO test is a different testing category from
functional e2e (variance-driven, historically flaky, currently
skip-annotated against LIT-5119) and erodes trust in the suite as a
release gate; it comes out of the default collection along with its
exclusive plumbing (locustfile, load-mock registration fixtures,
run_chat_load). Re-implementation as its own pipeline is tracked in
LIT-5163. The weekly session-anomaly test never ran in the suite (opt-in
via E2E_WEEKLY_ANOMALY, driven by its own workflow) and stays, as do the
markerless aggregation unit tests.

The vllm passthrough test read-times-out (60s) against the shared
vllm-cpu backend in every run on the per-SHA e2e stack; it is removed
until LIT-5164 establishes whether that is backend capacity or a
passthrough defect. Its registry cells return to the gap list, which is
the honest state
2026-08-04 14:12:25 -07:00
mateo-berri
09388532d2 docs(CLAUDE.md): prefer commas over semicolons when replacing em dashes 2026-08-04 14:10:03 -07:00
Mateo Wang
8445cf158b
Merge pull request #35555 from BerriAI/devin/1785632264-gemini-robotics-er-2
feat(gemini): add gemini-robotics-er-2-preview and gemini-robotics-er-1.6-preview
2026-08-04 14:08:59 -07:00
mateo-berri
258d154f18 chore(lint): zero out reportGeneralTypeIssues headroom 2026-08-04 13:52:00 -07:00
mateo-berri
72983e28c0 fix(lint): exempt the runtime-settable config surface in litellm/__init__.py from LIT010
Module-level names in litellm/__init__.py are the SDK's documented config
surface: users assign litellm.api_key and friends directly, and the proxy
rebinds them via setattr from litellm_settings. The package ships py.typed,
so the Final sweep made every such documented assignment a mypy error
("Cannot assign to final name") in downstream codebases. Strip Final from
the module scope of that file, keep it on function locals, and teach LIT010
that the config surface's module scope is exempt so the gate stays green
without suppression comments
2026-08-04 13:49:58 -07:00
mubashir1osmani
ad79b314c5
test(e2e): cover legacy text /completions endpoint (#34431)
* test(e2e): cover legacy text /completions endpoint

The /completions (and /v1/completions) text-completion route had zero e2e
coverage despite being the second-busiest endpoint in production; everything
'completions' in the suite was chat. Add a text-completion endpoint test that
registers an OpenAI instruct deployment, drives /v1/completions through the
gateway, and asserts real generated text. Adds text_completions() + the
completion request/result models to EndpointsClient, the 'completions' endpoint
to the coverage registry vocab, and the registry cell.

* test(e2e): assert /v1/completions choices shape, not just joined text

Assert the response carries a choices array and the first choice has real text,
so a malformed response (no choices) and a clean-but-empty completion are
distinct failures. Drop the unused text property / id / model fields (model only
what the test reads).
2026-08-04 13:48:07 -07:00
mateo-berri
050a8bdd09 fix(gemini): mark gemini-robotics-er-1.6-preview as supporting prompt caching 2026-08-04 13:34:01 -07:00
yuneng-jiang
49eb19c39f
chore(deps): upgrade cryptography to 50.0.0 (#35803)
Moves the proxy extra's cryptography floor from 48.0.1 to 49.0.0 and widens the
ceiling to <51, then holds the lock at 50.0.0 with a uv override

mlflow caps cryptography at <50 even in its newest release, so publishing a
plain >=50.0.0,<51.0 range would make `pip install "litellm[proxy,mlflow]"`
unresolvable for downstream consumers. Publishing >=49.0.0,<51.0 keeps that
combination installable (it resolves to 49.0.0), while the
override-dependencies entry, which is a uv workspace setting and never reaches
published metadata, keeps our own lock and Docker images on 50.0.0

mlflow only uses PBKDF2HMAC, AESGCM, Fernet and InvalidTag from cryptography;
none of those are affected by the 49 or 50 breaking changes, so overriding its
ceiling is safe in practice

Lock delta is cryptography 48.0.1 -> 50.0.0, the mlflow trio 3.14.0 -> 3.15.0
and msal 1.36.0 -> 1.37.0

cryptography 49 dropped its x86_64 macOS and 32-bit Windows wheels. Linux CI
and the Docker images are unaffected; developers on Intel Macs will build from
source
2026-08-04 13:32:30 -07:00
Mateo Wang
c1450e9fa9
chore: fix formatting 2026-08-04 13:24:15 -07:00
mateo-berri
741aa901cd docs(claude): state the Final-binding and frozen-parameter conventions 2026-08-04 13:20:59 -07:00
mubashir1osmani
dcb4e5033c
test(e2e): vendor API strategy coverage across endpoints (#34649)
* test(e2e): cover vendor strategy gaps for chat contract, image edits, auth, team activity

Resolves the first slice of LIT-4778 (vendor API testing strategy): image edits happy path, chat multi-turn + validation + sanitization, LLM-route auth header matrix, and /team/daily/activity structure

* test(e2e): expand vendor API strategy coverage across endpoints

Adds validation cases on existing endpoint suites, plus vector stores, search,
bedrock native, realtime HTTP secrets/calls, responses retrieve, files/batches
contract, and chat stream SSE. Registers coverage cells for LIT-4778

* test(e2e): finish vendor strategy open items

Audio transcription negatives, vector-store file attach/poll/search,
OpenAI moderation category matrix across chat/messages/responses, and
smoke model matrix for chat (LIT-4778)

* test(e2e): harden vendor strategy suite against live env edges

Fix stream [DONE] tracking, XSS no-crash contract, realtime model routing,
vector store list/search models, responses validation, and provider-denied
Bedrock paths so the suite is stable against a live proxy

* test(e2e): rename suites, drop vendor_contract, fix greptile gaps

Move shared status helpers into e2e_http, rename chat auth headers and
chat security suites, remove vendor_contract and dev_config files_settings,
and tighten transcription validation plus vector-store search assertions

* test(e2e): route bedrock stream disconnects through e2e_http

Catch mid-stream RequestException in the shared harness so bedrock native
tests do not import requests directly

* fix(e2e): address greptile and veria review on vendor strategy suite

Store search tool keys as os.environ refs and resolve them in SearchAPIRouter.
Tighten validation helpers and assertions so 5xx/empty/unrelated failures no longer pass coverage cells

* fix(e2e): drop search_api_router os.environ expansion from vendor suite

Keep the PR test-only. Search tools register without an api_key so the
proxy falls back to its own PERPLEXITY/TAVILY env, same pattern as a2a.

* test(e2e): drop search e2e suite from vendor strategy PR

Remove the /v1/search coverage file and its registry rows so this PR
no longer carries search endpoint testing.
2026-08-04 20:19:34 +00:00
mateo-berri
2708620d6a feat(lint): enforce Final on locals and freeze function parameters (LIT010, LIT011) 2026-08-04 12:54:39 -07:00
yuneng-jiang
487074f602
chore(build): move the Admin UI toolchain to Node 24 (#35801)
* chore(build): move the Admin UI toolchain to Node 24

Node 18 and Node 20 both reached end of life (2025-04-30 and 2026-04-30), and
the release images along with every CI lane were still building on them. Node 24
is the current LTS through 2028-04-30, so this moves the four UI build images,
the CircleCI lanes, and the four GitHub Actions workflows onto it

Node 24 also ships npm 11.17, which is the first line that implements the
min-release-age setting this repo already carries in its .npmrc files. On npm 10
the key is parsed and discarded, so the release-age gate has had no effect
regardless of its value. Tightening the dashboard's engines range and turning on
engine-strict makes an unsupported npm fail loudly rather than skip the gate
quietly, and a new step in the UI build workflow probes an impossible cooldown
so an inert setting cannot pass unnoticed again

Node 24's bundled undici tightened its brand check on RequestInit.signal, which
rejects the AbortSignal jsdom installs and broke the two cases in
src/lib/http/api.test.ts that rebase a request onto a runtime base url. Under
jsdom the Request global comes from Node while AbortSignal comes from jsdom;
tests/jsdomFetchEnv.ts delegates to the jsdom environment and then restores
Node's native AbortController and AbortSignal so both come from one realm.
Upgrading jsdom does not address this, as jsdom still does not own Request

The workflows now read ui/litellm-dashboard/.nvmrc instead of repeating a
literal, so the Node version has a single source of truth, and ui/Dockerfile is
pinned by digest to match the other three build images. The lockfile changes are
npm 11 normalising the engines range and dropping optional peer entries it no
longer records

* fix(build): point every Admin UI build script at .nvmrc

The enterprise Docker path was left on Node 18. docker/build_admin_ui.sh runs
only when enterprise/enterprise_ui/enterprise_colors.json is present, which it
never is in the OSS tree, so neither CI nor a default image build reaches it;
it pinned nvm to v18.17.0 and then built the dashboard, which now requires Node
24, so a customized enterprise image would have failed EBADENGINE

All three UI build scripts now resolve the version from
ui/litellm-dashboard/.nvmrc rather than carrying their own pin, so the Node
version has a single home across Docker, CI, and local builds. build_ui.sh was
on v20 and build_ui_custom_path.sh on v18.17.0

Also drops the dependency-cooldown probe from the UI build workflow. The
engines floor plus engine-strict already fails an unsupported npm loudly at
install time, so the probe was redundant, and treating any nonzero exit from a
live registry call as proof of enforcement made it unsound besides
2026-08-04 12:36:07 -07:00
devin-ai-integration[bot]
355ae9989b
fix(proxy): propagate user_email and bind api_key on JWT auth attribution paths (#34331)
* fix(proxy): propagate user_email and bind api_key on JWT auth paths

Standard JWT auth built UserAPIKeyAuth with user_id but never user_email, and the first auto-registered request early-returned a key with token set but api_key unset, so spend-log attribution logged user_api_key_user_email and user_api_key_hash as null. Bind api_key to the token hash on the auto-registered key, copy user_email from the resolved user object on both the standard and auto-register JWT paths, and warn when enable_jwt_auth/litellm_jwtauth are placed at the config top level where they are silently ignored.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover misplaced top-level JWT config warning

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan <ryan@berri.ai>
2026-08-04 19:05:39 +00:00
Mateo Wang
cbeaf86c8d
Merge pull request #34029 from BerriAI/litellm_lit4395_cursor_agent
fix(proxy): make /cursor/chat/completions work with Cursor agent mode
2026-08-04 10:33:20 -07:00
mateo-berri
9eeff06263 Merge origin/litellm_internal_staging into litellm_lit4395_cursor_agent 2026-08-04 10:20:03 -07:00
yuneng-jiang
5ac1edcd59
fix(e2e): make spend-counter redis connection env-driven for non-cluster deployments (#35732) 2026-08-04 09:47:01 -07:00
yuneng-jiang
6b3d4f2380
feat(ui): add admin-configurable user banner (#35729)
* feat(ui): add admin-configurable user banner

Proxy admins can publish a markdown announcement that renders as a
dismissible banner on every dashboard page for all authenticated users,
editable from Admin Settings > UI Settings without a redeploy. Backed by
new /get/user_banner and /update/user_banner endpoints persisting to the
existing LiteLLM_UISettings table

* fix(ui): re-surface dismissed banner on identical republish

Stamp a server-side revision on every banner update and fold it into
the client dismissal signature, so unpublishing and republishing the
same message reaches users who dismissed the earlier run

* fix(ui): stamp banner revision as an opaque uuid instead of a counter

Two overlapping admin updates could read the same prior revision and
both persist the same incremented value, letting an identical republish
collide with a previously dismissed signature. A server-generated uuid
per update makes every publication identity unique by construction with
no read-modify-write

* refactor(ui): drop the server-side banner cache

Reads go straight to the single-row table; the dashboard already
throttles fetches client-side, so the cache only added staleness
windows under concurrent updates and multiple workers

* refactor(ui): move banner storage behind a domain repository and drop the store_model_in_db gate

UserBannerRepository owns the row shape instead of the endpoint
reaching through the generic .table bridge, and publishing no longer
depends on the unrelated STORE_MODEL_IN_DB flag; a connected database
remains the only requirement
2026-08-04 09:24:29 -07:00
Ahmed N
368dd0be5b
fix(groq): translate web_search_options to the browser_search tool (#34971) 2026-08-04 09:08:11 -07:00
tin-berri
956d5177d1
fix(proxy): log the model cost map reload failure lazily (#35750)
The reload-failure warning built its message with an f-string, so the
interpolation ran on every failed reload whether or not the warning level was
enabled. `test_logging_calls_do_not_build_their_message_eagerly` scans the whole
litellm package and asserts zero offenders, so this one call has been reddening
`misc / Run tests` on litellm_internal_staging for every branch cut from it

Passing the reason as a %-style argument defers the interpolation to
`record.getMessage()`, which only runs once the record passes the level check
2026-08-03 23:41:21 -07:00
devin-ai-integration[bot]
a625d1e1ca
feat(otel): stamp service tier attributes on inference spans (#35679)
* feat(otel): stamp service tier attributes on inference spans

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): bound requested service tier to known values

The requested tier is caller-controlled and reaches the span verbatim, so an
arbitrary string lands on every litellm_request span on success and on failure.
A 100k character value was stamped uncapped; safe_set_attribute does not
truncate and no span limits are configured.

Apply KNOWN_REQUEST_SERVICE_TIERS in get_requested_service_tier so both the
span attribute and the Prometheus label bound the value the same way. The
served tier stays unrestricted since it comes from the provider, so a tier a
provider adds later is still reported.

Prometheus label behavior is unchanged.

* fix: derive known service tiers from the ServiceTier enum

The allowlist omitted "fast", which litellm models as a real tier and prices
through the priority cost key, so a request naming it resolved to no tier on
the span and no Prometheus label.

Deriving the set from ServiceTier keeps the two in sync, so a tier added there
for cost calculation cannot go missing here.

Behavior change: a request with service_tier "fast" now carries the tier on the
span and on the Prometheus service_tier label, where it previously resolved to
none. Every other value resolves as before.

* refactor: build the known service tiers without a mutable intermediate

The set comprehension and set literal tripped LIT002, which bounds mutable
collections. Concatenating tuples keeps the derivation from ServiceTier while
every intermediate stays immutable; the resulting frozenset is unchanged.

---------

Co-authored-by: milan <milan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yucheng Zhu <yucheng@berri.ai>
2026-08-03 23:10:01 -07:00