Commit graph

13832 commits

Author SHA1 Message Date
elinacse
833670f7db fix(batch): track cost for managed batches with no attributable key/user/team
LiteLLM_ManagedObjectTable only stores created_by (user_id) and team_id,
never the raw API key hash. A batch created with the master key or a
team-less key has both null, so CheckBatchCost's synthetic logging_obj
for the completed batch carried no attributable key/user/team/end-user.
_should_track_cost_callback silently skipped the DB write in that case
(by design, to avoid tracking truly anonymous requests), with no error
or warning: batch_processed still became true, but no LiteLLM_SpendLogs
row was ever written despite real, already-incurred provider cost.

Extend the same allowance already made for unauthenticated pass-through
requests to aretrieve_batch's cost event, and pass job.team_id through
so a batch's team gets real attribution when one exists.
2026-08-02 12:20:46 +05:30
yuneng-jiang
491eda319c
Merge pull request #35572 from BerriAI/litellm_e2e_budget_spend_poll_to_deadline
Some checks are pending
Unit Tests: LLM Provider Transformations / All Other Providers (push) Waiting to run
Unit Tests: MCP, Secrets, Containers & Misc / misc (push) Waiting to run
Unit Tests: Proxy Auth & Key Management / proxy-auth (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Proxy API Endpoints / proxy-endpoints (push) Waiting to run
Unit Tests: Proxy API Endpoints / proxy-server (push) Waiting to run
Unit Tests: Proxy Infrastructure / proxy-infra (push) Waiting to run
Unit Tests: Proxy Legacy Tests / auth-and-jwt (push) Waiting to run
Unit Tests: Proxy Legacy Tests / key-generation (push) Waiting to run
Unit Tests: Proxy Legacy Tests / proxy-config (push) Waiting to run
Unit Tests: Proxy Legacy Tests / proxy-response-and-misc (push) Waiting to run
Unit Tests: Proxy Legacy Tests / proxy-server (push) Waiting to run
Unit Tests: Proxy Legacy Tests / proxy-server-extras (push) Waiting to run
Unit Tests: Proxy Legacy Tests / proxy-token-counter (push) Waiting to run
Unit Tests: Proxy Legacy Tests / proxy-user-auth-and-spend (push) Waiting to run
Unit Tests: Proxy Legacy Tests / proxy-utils (push) Waiting to run
Unit Tests: Responses, Caching & Types / responses-caching-types (push) Waiting to run
test(e2e): poll key spend to a deadline in budget reset advances tests
2026-08-01 19:58:58 -07:00
ryan-crabbe-berri
bf873dffe4 test(e2e): skip the strict-priority and throughput SLO tests pending LIT-5118 / LIT-5119
The strict-priority e2e (added with the zero-increment limiter fix) can
never pass on stage: the proxy there does not run the
dynamic_rate_limiter_v3 callbacks + priority_reservation settings the
module requires, confirmed by zero limiter log lines across every
gateway and backend pod during the 2026-08-02 run. Config lives in the
infra repo; LIT-5118 tracks adding it.

The throughput SLO test failed the same run with 65.9% of requests dying
at the ELB as 502/503 before reaching a pod. The per-replica SLO rework
fixed the RPS-floor assertion but cannot help when stage idles at one
warm gateway replica; LIT-5119 tracks pre-scaling the fleet for the load
phase.

Both skips name their ticket, and the coverage registry returns the two
cells to the gap list while they are in place.
2026-08-01 19:42:10 -07:00
ryan-crabbe-berri
ebdd854b98 test(e2e): poll key spend to a deadline in budget reset advances tests
A single read of key_info.spend races the batched spend writer: deltas
earned before a reset flush to the DB up to ~60s later
(proxy_batch_write_at) and land on the row after the reset zeroed it.
The stage runs on Jul 30 and Aug 2 failed
test_key_budget_reset_at_advances_after_window exactly this way, with
spend back at the driven total while budget_reset_at had advanced and
calls flowed again.

Replace the single reads in rung 3 (spend zeroed after reset) and rung 4
(roomy window keeps spend) with _poll_key_spend, which re-reads to a 90s
deadline covering one full flush-plus-reset cycle. A reset that never
zeroes the row keeps spend pinned and still times out, so the regression
guard keeps its teeth.
2026-08-01 19:29:56 -07:00
Devin AI
81a80b8c63 fix(anthropic): preserve speed=fast in usage for /v1/messages and pass-through
Fast mode is priced with a provider-specific multiplier applied off usage.speed, but only chat completions kept that field. The Messages route rebuilt usage with empty optional params, stream reassembly dropped speed and inference_geo, and the pass-through handler never read speed off the request body, so fast-mode spend was logged at the standard rate.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-02 01:02:27 +00:00
mateo-berri
f25f1d2921 feat(proxy): resolve Cursor thinking/fast model-name suffixes on /cursor/chat/completions
Cursor appends -thinking-<level> and -fast to custom model names when the
user picks a thinking level or fast mode, so a model configured as
claude-opus-5 arrives as claude-opus-5-thinking-xhigh-fast and fails
routing with no healthy deployments. When the raw name is not servable by
the router but the suffix-stripped base name is, rewrite the body to the
base model and carry the thinking level into reasoning_effort (chat
bodies) or reasoning.effort (Responses bodies), never clobbering an
effort the client already sent. Explicitly configured aliases keep
winning because the raw-name servability check runs first.
2026-08-01 17:42:03 -07:00
Devin AI
c5c5a27679 fix(files): enforce require_managed_files on file retrieve, content and delete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-02 00:19:42 +00:00
yuneng-jiang
ceaf556b2e
Merge pull request #35523 from BerriAI/litellm_ui_login_no_mcp_landing
fix(ui): land general login on the keys dashboard, send MCP consent to /ui/connect
2026-08-01 17:17:18 -07:00
mateo-berri
23d26d5e64 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit4395_cursor_agent
# Conflicts:
#	litellm/completion_extras/litellm_responses_transformation/transformation.py
#	litellm/litellm_core_utils/llm_response_utils/convert_dict_to_response.py
#	litellm/litellm_core_utils/prompt_templates/common_utils.py
#	litellm/litellm_core_utils/streaming_chunk_builder_utils.py
#	litellm/llms/openai/chat/gpt_transformation.py
#	litellm/main.py
#	litellm/proxy/response_api_endpoints/endpoints.py
#	ruff-strict-budget.json
#	type-discipline-budget.json
2026-08-01 17:06:53 -07:00
yucheng-berri
7c3b578e77
fix(team-callbacks): report API-registered callbacks from GET /team/{team_id}/callback (#35512)
* fix(team-callbacks): report API-registered callbacks from GET /team/{team_id}/callback

POST /team/{team_id}/callback writes metadata["logging"] while the GET read
metadata["callback_settings"], so every team configured through the API or the
Admin UI got back an empty list. c620d76fe4 migrated the writer to the new key
and left this reader on the old one.

Resolve the read the same way request-time resolution does in
_get_dynamic_logging_metadata: a logging slot that is present wins outright and
callback_settings stays as the deprecated fallback, so the endpoint reports what
a request would really do rather than the union of both shapes. An empty logging
list therefore reports no callbacks, matching a request that fires none.

Decrypt callback_vars for the response and mask the credential keys. Ciphertext
would be unusable to the caller, and a value encrypted under a key that is no
longer classified as sensitive would otherwise come back as a raw blob.

Resolves LIT-5093

* Update litellm/proxy/management_endpoints/team_callback_endpoints.py

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(team-callbacks): mask callback vars that fail to decrypt

decrypt_callback_vars passes a value through untouched when it cannot be
decrypted, which happens to existing rows after a salt-key rotation. Under a
key that is not classified as sensitive that blob reached the caller as opaque
ciphertext it could not use or tell apart from a real value, so mask anything
still carrying the encrypted prefix.

Raised by Greptile on the first commit.

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-01 17:02:40 -07:00
mateo-berri
85f367b545 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_up035_abc_imports
# Conflicts:
#	litellm/llms/bedrock/base_aws_llm.py
#	litellm/proxy/proxy_cli.py
#	litellm/proxy/proxy_server.py
#	litellm/repositories/model_repository.py
#	litellm/router.py
2026-08-01 16:18:04 -07:00
Tin Chi Lo
7194cafbc1 fix(ui): land general login on the keys dashboard, send MCP consent to /ui/connect
A keyless internal user signing in to the Admin UI was redirected off the
post-login landing to /ui/connect, which renders nothing but the MCP apps panel,
so a plain gateway sign-in ended on an MCP OAuth surface the user never asked
for. The landing now renders the keys dashboard for every role. The key lookup
that existed only to make that routing decision goes with it, along with the
useKeys enabled flag it was the sole caller of and the role-hydration hold that
guarded its one-frame dashboard flash

The gateway DCR consent flow moves the other way. Its /authorize handed the
browser to /ui/chat/integrations, whose layout hard-blocks when enable_chat_ui
is off, which is the default, and client-side redirects to /ui/ without the
query string; that destroys the connect_flow handle and strands the MCP client
until the 600s flow cookie expires. It now lands on /ui/connect, which reads
connect_flow and connect_client, mounts the consent banner and puts the apps
panel in connect mode. /ui/chat/integrations keeps its connect-mode handling
this release so flows sealed before the deploy still finish

Resolves LIT-5104
Resolves LIT-4911
2026-08-01 16:09:30 -07:00
yuneng-jiang
68b1e035fb
Merge pull request #35518 from BerriAI/litellm_/windows-litellm-import-fail-4acb08
fix(deps): move pydantic-settings into the base dependencies
2026-08-01 15:45:39 -07:00
mateo-berri
b604e2b20c refactor(lint): apply every safe ruff autofix and zero 28 strict-rule budgets
About 35,000 fixes ruff marks safe across 32 rules (UP006/UP045/UP007
modern annotations, UP032 f-strings, SIM114/SIM118, RET501, and
friends), removal of the 1,296 typing imports the rewrite orphaned, and
hand fixes for what the fixers could not see: five star-import
freeloaders of typing names, two F823 late-import annotations, the
/get/config/list introspection crash on types.UnionType, redundant
function-local RoleMappings imports in ui_sso.py that shadowed the
module-level name once the annotation lost its quotes, and one FURB168
tautology.

B009/B010/PIE804/RUF019 are excluded on purpose: their safe fixes
rewrite getattr/setattr/**-splat/key-in-dict escape hatches into forms
basedpyright then rejects (283 new errors measured), so their budgets
stay at base values.

ruff-strict-budget.json drops by 39,579 this commit (39,968 across the
branch) with 28 rules at an actual 0 and 9 more sharply down.
type-discipline-budget.json ratchets LIT002/LIT006/LIT009 down; LIT001
moves to the now-honest total: the checker matches the spelling `set`
but not the alias `Set`, so the 160 typing.Set annotations rewritten to
set[...] were always mutable-set annotations and only now count.
2026-08-01 15:43:29 -07:00
tin-berri
11ad3ff939
refactor(complexity_router): drop the tier-rubric override, close the rubric on the window it was given (#35504)
Two changes to the classifier's system role, both narrowing it rather than adding to it

classifier_tier_rubric let an operator replace the tier definitions. It shipped in
#35471 alongside the assistant-turn context window, but the two answer different halves
of the same report and only the context window was asked for. The override carried a
composed prompt, an overridable and a non-overridable half, a blank-is-unset rule, a
length-warning validator and a pair of dashboard controls. All of it goes

The rubric then closes on one of two lines, chosen by classifier_context_window_size.
At 0 no conversation is quoted, so the line is the original one, byte for byte: a
deployment that sends no context is told to classify the current message and nothing
else, which is what it could see all along. Above 0 the turns are quoted, and the
original line told the model to disregard them, which is how a request whose difficulty
was established in an earlier turn came back SIMPLE on the word "yes". There the line
instead says to classify the current message using the quoted turns as context, and to
rate what a short reply approves rather than the reply

The choice keys on the window and not on classifier_context_include_assistant_turns.
Whether the quoted turns are the user's alone or include the assistant's replies does
not change what the model needs told, and whose turn is whose is already on the turns.
Keying it on the assistant toggle would put the default deployment back on the original
line, which is the configuration the report was raised against

Folds in #35508, which built the window-dependent framing on top of the override this
removes; that PR is closed in favour of this one
2026-08-01 15:38:44 -07:00
yuneng-jiang
4861f0cd26
Merge pull request #35505 from BerriAI/litellm_/test-config-failures-745045
test(proxy): assert _delete_deployment's still-desired id set instead of a delete count
2026-08-01 15:38:35 -07:00
yuneng-jiang
22c3aae84e
Merge pull request #35506 from BerriAI/litellm_/gcs-pubsub-test-metadata-4685a9
test(logging): pin routing_decision and internal_call_origin in the gcs pubsub spend log fixture
2026-08-01 15:38:19 -07:00
yuneng-jiang
a9e1ab2e81
Merge pull request #35507 from BerriAI/litellm_/team-member-add-authz-tests-3948e1
test(proxy): separate the member_add permission gate from the provisioning gate
2026-08-01 15:35:22 -07:00
Yuneng Jiang
7447f9babc
fix(deps): move pydantic-settings into the base dependencies
`import litellm` reaches litellm/integrations/otel/model/config.py via
litellm_core_utils/litellm_logging.py, so pydantic-settings is needed at import
time. It was declared only in the `proxy` extra, which left a plain
`pip install litellm` unimportable on every platform.

Adds tests/base_sdk_tests/check_base_sdk_install.py and a base_sdk_install
CircleCI job that builds the wheel, installs it into a clean venv with no extras,
and smoke-checks the import, a mock completion, a mock embedding, the bundled
pricing metadata and the token counter. The check is stdlib-only on purpose;
installing pytest into that venv would add packaging, pluggy and iniconfig and
could mask the class of undeclared dependency it exists to catch.

Previously the Windows job was the only one installing without extras, so this
class of break was caught by accident rather than by design.
2026-08-01 15:25:26 -07:00
Yuneng Jiang
86312da3be
fix(ci): let the E2E proxy accept the mock testing params its suite sends
Gating the mock testing request params behind
general_settings.dangerously_allow_mock_testing_request_params (#35423) turned
every fallback, retry and timeout drill in tests/test_fallbacks.py into a 400:
the build_and_test job mounts proxy_server_config.yaml, which never opted in.

Opt that config in. It is the config the CI proxy runs with, and the suite it
serves exists to drive synthetic failures.

Add a unit test that ties the two together: it scans the top-level tests/test_*.py
files build_and_test globs for gated param names and fails if the config they run
against has not opted in, so the next change to either side is caught in a fast
lint-tier job rather than a Docker E2E.
2026-08-01 14:57:42 -07:00
Yuneng Jiang
ff4a50c768
test(proxy): separate the member_add permission gate from the provisioning gate
Adding a team member by a user_id with no user row is now proxy-admin-only,
so the /team/member_add authz matrix, which targeted a never-seeded user_id,
started 403ing every non-proxy-admin caller. Seed the member as a real user
row so the matrix reads _validate_team_member_add_permissions alone; leaving
it unseeded and relaxing the expectations to 403 would have left all 18 rows
green with that gate deleted outright.

Cover the new gate at the HTTP boundary, where only the helper was pinned
before: a team admin and an org admin both clear the permission check on the
same team and are still refused an unprovisioned user_id, with no user row
left behind. Pin the escape hatch that refusal names too, so closing the
email-invite path for non-proxy-admins cannot pass silently.

Promote the user seeder the member-info pins had kept private to conftest,
and reclaim invited users by their scratch-prefixed email, since an invite
allocates the user_id server-side.
2026-08-01 14:39:47 -07:00
Yuneng Jiang
67dd8924ee
test(proxy): assert _delete_deployment's still-desired id set instead of a delete count
_delete_deployment stopped returning a count of evictions in #35400 and now returns
the frozenset of ids the db and config still want, so a caller judging its own reload
can tell a deliberate eviction from a deployment that went missing. These two tests in
tests/local_testing were left comparing that frozenset against an int and have been
failing since; the directory is only referenced by .circleci/config.yml, which no
longer reports checks on PRs, so nothing caught them.

The eviction behavior itself is unchanged, so the fix is on the assertions: compare
against the expected id set, and pin the router's surviving ids so a mutation that
evicts the wrong deployment is caught rather than passing a bare length check.
2026-08-01 14:36:30 -07:00
Yuneng Jiang
aaf619c270
test(logging): pin routing_decision and internal_call_origin in the gcs pubsub spend log fixture 2026-08-01 14:34:04 -07:00
Yassin Kortam
33eda22386
fix(docker): honor USE_DDTRACE in the componentized gateway and backend images (#35490)
The componentized images exec uvicorn directly, so ddtrace-run never wraps the
interpreter. USE_DDTRACE is not inert there; the proxy lifespan still runs
patch_all and litellm's own manual spans still emit. What never gets installed
is ddtrace's ASGI TraceMiddleware: starlette builds its middleware stack lazily
on the first __call__, which is the lifespan scope, so patching from inside the
lifespan body is already too late and no root request span is ever created.

Route both entrypoints through a shared docker/component_entrypoint.sh that
mirrors the monolith's prod_entrypoint.sh contract, including the
DD_TRACE_OPENAI_ENABLED=False export that keeps ddtrace's openai integration
from double-reporting calls litellm instruments itself.

Co-authored-by: Yassin Kortam <yassin.kortam@gmail.com>
2026-08-01 14:12:59 -07:00
Mateo Wang
66a3f3780b
Merge pull request #35406 from BerriAI/devin_ai_fix_batch_output_file_id_encoding_lit4964
fix(batches): encode public model group on background-created output file ids
2026-08-01 14:12:43 -07:00
Yassin Kortam
72c8888563
fix(docker): bake prisma offline in the componentized migrations image (#35485)
The migrations image ran `prisma migrate deploy` against a bake anchored in
$HOME with no node in the runtime stage, so prisma-client-py fell through to
nodeenv and tried to download a Node runtime on first start. In an
egress-restricted cluster that fails outright, and under an arbitrary uid the
uid-specific cache path is unreadable, so the job never applies a migration.

Move the bake to /opt/prisma with world-readable modes, install node in the
runtime stage, and pin PRISMA_BINARY_CACHE_DIR / PRISMA_CLI_PATH /
PRISMA_OFFLINE_MODE so the migration entrypoint runs the cached CLI directly.
This is the same treatment the root, non_root and database images already
carry.

Resolves LIT-4727
2026-08-01 14:12:31 -07:00
Yassin Kortam
3d35eee560
test(e2e): derive the throughput SLO per replica and surface locust's error breakdown (#35494)
The floor was an absolute fleet number, so it asserted replicas x per-replica rate
and went red on how many gateway pods happened to be warm rather than on the request
path. The test now measures one replica first, with a short serial pass that only ever
occupies a single pod, and requires the concurrent phase to reach at least that rate.
A serial latency budget carries the request-path assertion the floor used to imply,
and both hold at one replica or seven.

Zero-error runs that "sustained 16.7 RPS" were queueing, not slow requests: the load
model is a mock_response deployment with no upstream, a single-worker replica serves
it in about 57ms, and 100 closed-loop users against 1/0.057 RPS of capacity sit at
6s each by Little's law.

The runner also kept locust's --json summary and threw away everything else, so a run
where 93% of requests failed said nothing about what they got. It now passes --csv,
reads the failure breakdown back, and reports locust's own generator-saturation
warnings, both folded into the assertion messages.

Resolves LIT-5054
2026-08-01 14:08:19 -07:00
Mateo Wang
9161e3ba67
Merge pull request #35436 from BerriAI/litellm_redis_pubsub_config_sync
feat(proxy): push config sync to pods via redis pub/sub
2026-08-01 14:06:59 -07:00
tin-berri
9bed4955d4
feat(complexity_router): let the classifier see assistant turns and rate what a short reply approves (#35471)
The LLM classifier's context window carried user turns only, so a conversation
whose difficulty was stated by the model rather than by the user was classified
without it. Asked to find events, the assistant answers "here is the plan, it is
complex, should I execute?", the user answers "yes", and the router rates the
word "yes" and picks the cheapest tier

Two independent causes, so two changes that are each provable on their own

classifier_context_include_assistant_turns adds assistant turns to the window.
It is off by default because turning it on shifts tier decisions, and therefore
spend, for an already-deployed router, and because assistant text is net-new
egress to the classifier deployment. With it on, classifier_context_window_size
counts the last N turns across both roles, which is what makes the assistant's
own statement of difficulty land in the window

Assistant text reaches the classifier payload and nothing else. The window is
read only by _build_classifier_user_payload, while keyword_tier_rules, escalation
matching, the heuristic scorer and the semantic embedding all read the human ask
through _iter_human_asks_newest_first. Those are substring and vector matchers,
so an assistant echoing an escalation keyword back to a user would choose the
model, and the spend, with nobody having asked. Rather than widen the shared
iterator, _iter_context_turns_newest_first is separate and feeds the window
alone, which makes the boundary structural instead of a rule to remember

The rubric ended "Classify only the current message", and the classifier applied
it literally: a request whose difficulty was established earlier came back SIMPLE
because the message being rated was the word "yes". A context window the rubric
then tells the model to disregard buys nothing, so the wording now asks it to
rate the work the current message approves, judged in the conversation it
continues, while still forbidding it to rate a quoted section as if that section
were the request

classifier_tier_rubric lets an operator replace the tier definitions. The
trust-boundary paragraph is appended and cannot be replaced: it defends the
operator against their own callers, so an operator writing tiers without that
threat in mind would otherwise hand every keyholder the top tier by omission.
Blank reads as unset so an empty form field falls back rather than sending a
rubric with no tiers in it

Turns are labelled by role only when assistant turns can appear, so the prompt of
every deployment that never asked for this is unchanged byte for byte
2026-08-01 13:59:35 -07:00
Yassin Kortam
14dd98cd5f
fix(bedrock): cache AssumeRole credentials per attributed identity (#35467)
The explicit AssumeRole branch of BaseAWSLLM.get_credentials returned without
touching the process-wide IAM cache, so every model request issued a fresh
sts:AssumeRole, and on ECS/EC2 an uncached sts:GetCallerIdentity ahead of it.
Route the whole role branch through _get_or_set_cached_credentials with the TTL
_auth_with_aws_role already computed and discarded. The cache key is the same
aws_* argument snapshot the other flows use, taken before the session-name
default is filled in, so each aws_session_name keeps its own STS session and no
attributed identity can be served another's credentials.

Credential fetches now single-flight behind striped locks. Without that, a burst
of concurrent misses on one key each issued their own STS call, which is the
same thundering herd the cache exists to prevent, moved to the miss window.
2026-08-01 13:54:42 -07:00
Yassin Kortam
669bfd6c60
feat(proxy): bound DB statement and lock time via general_settings (#35496)
A daily-spend batch upsert that outlives prisma-client-py's 30s HTTP read
timeout keeps running server side after the client gives up, holding its
row locks for as long as the database takes. Every later flush cycle
queues behind those locks, which is how one slow batch cascaded into
exhausted database sessions.

The query engine's own transaction timeout cannot end that wait: it
cannot interrupt a statement that is already executing. Measured against
real Postgres, a batch wrapped in db.tx(timeout=60s) still held its locks
for the full 90s the statement ran. Only a Postgres-side statement_timeout
bounded it.

database_statement_timeout and database_lock_timeout (seconds) are now
first-class general_settings keys, emitted as libpq
options=-c statement_timeout=<ms> on DATABASE_URL. They are opt-in, so an
unset config keeps today's behavior, and they are never applied to
DIRECT_URL, which serves migrations that legitimately run long.

Resolves LIT-4718

Co-authored-by: Yassin Kortam <yassin.kortam@gmail.com>
2026-08-01 13:54:17 -07:00
Yassin Kortam
4178a857ae
fix(spend): bound each spend-log write statement by payload bytes (#34956)
The Prisma query engine is a separate Rust process whose resident memory is
a high-water mark: it grows with the payload of the largest single statement
it executes and glibc never returns that memory to the OS, so a pod's memory
floor ratchets up to its worst-ever write and stays there for the life of
the worker. Memory-based autoscaling then reads a number that reflects the
largest write the pod has ever done rather than what it is doing now.

The spend-log flush handed Prisma a fixed 1000 rows per create_many. With
store_prompts_in_spend_logs enabled a single row carries the full prompt and
response, so one statement can be tens of megabytes and permanently costs
hundreds of megabytes of RSS. Row counts cannot express that budget: the
same 1000 rows range from well under a megabyte to tens of megabytes.

Split each flush into statements bounded by encoded payload size
(SPEND_LOG_WRITE_BATCH_MAX_BYTES, default 2MB) on top of the existing
1000-row cap. What is measured is the encoded statement, so the budget
counts what actually goes on the wire: the JSON escaping of quotes and
newlines, multibyte characters at their encoded width, the field names and
separators a 25-column row carries, and the brackets and row separators the
rows carry as one collection. Deployments that do not store prompts keep one
statement per 1000 rows and are unaffected; prompt-carrying flushes get
several small statements instead of one huge one. A row larger than the
budget is still written on its own rather than dropped, and a row the
serializer refuses counts as zero rather than raising out of the flush and
dropping every row queued behind it.

Splitting a flush must not multiply what a poison-row flood costs, so the
poison-isolation allowance is threaded through every statement of a 1000-row
group instead of being handed out fresh per statement. That is only safe
because the allowance now counts failed inserts rather than every insert:
the one insert a statement needs when nothing is poisoned is not charged, so
a healthy flush never runs the allowance down however many statements it
splits into, and a statement reached after the allowance is spent is still
attempted so clean rows behind a flood still persist. Failed inserts for a
group are bounded by the allowance plus one baseline insert per statement,
which restores the constant-per-group ceiling the single-statement path had.

Resolves LIT-4765
2026-08-01 13:53:36 -07:00
Yassin Kortam
a8cc6a921a
fix(router): honor request-level num_retries over a deployment's litellm_params value (#35483)
A failing deployment stamps its own litellm_params.num_retries onto the raised
exception, and async_function_with_retries adopted that value unconditionally. So a
model_list num_retries outranked both the x-litellm-num-retries header and the request
body, inverting the documented precedence to model_list > header > body >
litellm_settings.

The router could not tell a request-level value from its own default because the entry
points filled num_retries in with self.num_retries whenever the caller omitted it,
collapsing "the request asked for N" and "nobody asked". Drop that pre-fill from
_update_kwargs_before_fallbacks and from the six entry points that also did it a line
above their own call to it (image generation sync and async, adapter completion, file
create, batch create, batch cancel), all of which reach async_function_with_retries,
where the router/global default is already resolved. Leaving them would have made the
request value never None on those routes and permanently suppressed a deployment
num_retries there.

The sync text_completion pre-fill stays. That path resolves a deployment and calls
litellm.text_completion directly, never entering the retry loop, so no request-versus-
deployment ranking happens there and there is nothing to fix; removing the line would
only change which value is forwarded to litellm.text_completion, a behaviour change this
bug does not call for.

async_function_with_retries then adopts the deployment's value only when the request
carried none. Precedence is now header > body > model_list > litellm_settings, with the
deployment value still beating litellm_settings when the request is silent, on every
entry point that retries.

Resolves LIT-4772
2026-08-01 13:51:29 -07:00
Shivam Rawat
e204e629e0
Merge pull request #35422 from BerriAI/litellm_fix_tpm_only_dynamic_rate_limit
fix(rate-limit): enforce token limits when the pre-call increment is zero
2026-08-01 13:02:11 -07:00
Yassin Kortam
704b9da8ab
fix(a2a): keep config-defined agents registered and accept the documented agents: key (#35163)
The public A2A guide tells users to declare agents under a top-level
`agents:` key, but the proxy only ever read `agent_list:`, so the
documented config was silently ignored and GET /v1/agents returned an
empty list. Accept `agents` as the documented spelling and keep
`agent_list` working for anyone who found it by reading the source.
Selection is by key presence, so an explicitly empty `agents: []` is not
overridden by leftover legacy entries.

Config-defined agents were also dropped on any database-backed gateway:
the periodic reload rebuilt the registry from the DB rows plus a module
global that was declared and never assigned. The registry now remembers
the agents it loaded from config.yaml and replays them on every rebuild.
A database row wins a name collision, mirroring how config-declared MCP
servers are unioned under the database registry, so name lookups and
deregistration keep addressing exactly one agent.

Resolves LIT-4978
2026-08-01 12:39:13 -07:00
mubashir1osmani
4d43080a74 fix(pricing): apply OpenAI's gpt-5.6 terra/luna cut to Azure cost map
OpenAI cut Terra 20% and Luna 80% on 2026-07-30; openai and bedrock_mantle
entries already match. Azure global and us/eu data-zone terra/luna rows still
used the pre-cut rates, so spend tracking over-billed those Azure deployments.
Sol is unchanged. Cache-read, priority, and long-context fields scale with the
same multipliers already used for azure gpt-5.6.
2026-08-01 12:34:25 -07:00
yuneng-jiang
1e7b39d15c
Merge pull request #35473 from BerriAI/litellm_/ai-api-allowlist-model-info-6ba7db
feat(proxy): let AI API keys read /model/info
2026-08-01 12:04:45 -07:00
Shivam Rawat
334805990c fix(rate-limit): block check-only counters at the limit and scope tpm reservation to tokens
Aligns the fix with the constraints in LIT-4800. A zero-increment
counter now blocks at current >= limit, matching RPM's semantics; the
previous current > limit let a pool sitting exactly at its reservation
admit one extra request. reserve_tpm_tokens rebuilds its descriptors
with only tokens_per_unit so the requests dimension stays out of the
reservation pass, which deliberately leaves RPM to the separate
should_rate_limit check.
2026-08-01 11:56:18 -07:00
devin-ai-integration[bot]
a23aa47e1e
feat(prometheus): add global exclude_metrics and exclude_labels options (#34201)
* feat(prometheus): add global exclude_metrics and exclude_labels options

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(prometheus): apply global exclude_labels to hard-coded metric labels

Metrics built with hard-coded labelnames lists (guardrail, provider budget,
callback, managed file/batch, batch cost) bypassed prometheus_exclude_labels
because only labels resolved via get_labels_for_metric were filtered. Route
every metric through a factory that strips excluded labels at construction and
proxies labels() so excluded labels are dropped at emission too. Add the
non-enum hard-coded labels (guardrail_name, status, error_type, hook_type,
purpose, file_type, result) to exclude-config validation so they are accepted
instead of raising ValueError at logger init.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(prometheus): simplify exclude-label factory to the kwargs labelnames path

All metric definitions pass labelnames as a keyword argument, and the only
metrics that pass it positionally resolve their labels through
get_labels_for_metric, which already drops excluded labels, so they never carry
an excluded label into the factory. Drop the unreachable positional
reconstruction branch.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: re-trigger CI (flaky unrelated bedrock agentcore test)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(prometheus): use immutable constructions to satisfy LIT002 budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-01 11:52:27 -07:00
yuneng-jiang
20874f8d32
Merge pull request #35435 from BerriAI/litellm_team_member_mgmt_0731
fix(proxy): align team member add with existing user provisioning rules
2026-08-01 11:52:01 -07:00
Yuneng Jiang
7d6ee2a9ca
feat(proxy): let AI API keys read /model/info
Keys created with key_type=llm_api get allowed_routes=["llm_api_routes"],
which covered /v1/models but not /v1/model/info, so a client could list model
names but not read pricing, mode, or max_tokens without a second key.

Adds both /model/info and /v1/model/info (same handler) to llm_api_routes only.
Membership there is not the same as RouteChecks.is_llm_api_route(), which is
what gates DISABLE_LLM_API_ENDPOINTS, global/virtual-key budget enforcement,
enforce_user_param and JWT team attachment; /guardrails/apply_guardrail already
sits in the group the same way. /v2/model/info stays out: it is the paginated
Admin UI listing, not model metadata a caller needs at request time.

public_routes moves from set([...]) to a frozenset literal to keep the LIT002
and ruff-strict ceilings from rising; both budgets ratchet down by one.
2026-08-01 11:43:36 -07:00
Tin Chi Lo
c27f1b7b6d fix(tools): classify custom tool calls by one shared rule and make envelope payload extraction total
A chat tool-call dict was classified custom-vs-function with four different
spellings: the non-streaming parser required type == "custom", the streaming
Delta coercion also accepted a custom payload without type, and the stream
assembler required a type that later chunks never carry. The same payload could
be a custom tool call mid-stream, a TypeError on the completed message, and
silently dropped from the assembled message. is_custom_tool_call_dict() is now
the single discriminator (explicit custom type, or a custom payload present)
used by both parsers, and the assembler classifies from the accumulated custom
payload, matching how the deltas it consumes were classified.

The tool envelope converter picked one exclusive payload source: the nested
dict when present, else the top level. An empty nested envelope therefore
shadowed top-level fields and the normalized tool lost its name. Payload
extraction is now total over both locations, nested first, and an envelope
with no name anywhere passes through unchanged instead of being emitted
stripped.
2026-08-01 11:26:09 -07:00
Tin Chi Lo
4139f548da fix(bridge): resolve the effective OpenAI base once, shared by gate and chat handler
The gpt-5.4+ responses-bridge gate classified the endpoint from the call-level
api_base alone, while the OpenAI chat handler resolves arg > global > env >
default. A custom base configured via litellm.api_base or OPENAI_BASE_URL/
OPENAI_API_BASE was therefore invisible to the gate: it read blank as the
default OpenAI endpoint and bridged a request the custom backend has no
/responses route for.

Extract that resolution into one _resolve_openai_api_base() and have both the
gate and _complete_custom_openai() call it, so the gate can never classify an
endpoint the request won't hit. The gate compares the resolved base against the
default (import litellm seeds OPENAI_BASE_URL to the default, so "override is
non-None" is not a safe custom-endpoint signal); whitespace collapses to the
default as before. reasoning_effort="none" remains the escape hatch.
2026-08-01 11:26:09 -07:00
Tin Chi Lo
c7c656e8a9 fix(cursor): convert tools and tool_choice through one envelope rule
Chat Completions nests a named tool_choice under its tool type while the
Responses API keeps it flat; ChatCompletionNamedToolChoiceParam and
ChatCompletionNamedToolChoiceCustomParam both mark the nested key required. The
messages arm normalized tool definitions but forwarded tool_choice at whatever
level Cursor sent it, so a flat {"type": "custom", "name": "ApplyPatch"} reached
OpenAI unchanged and was rejected while the tool defs beside it nested correctly

A tool definition and a named tool_choice carry the same envelope, so both now
convert through a single _convert_tool_envelope, and _normalize_tool_dialect
moves tools and tool_choice together on each arm. That covers all four cells of
{tool def, tool_choice} x {to chat, to responses} and removes the shape where
one field can be converted while the other is missed, replacing three helpers
with two and cutting 24 lines

Also restores the end-to-end assertion that a flat tool_choice reaches
chat_completion nested, which had been flipped to pin the passthrough behavior
2026-08-01 11:26:09 -07:00
tin
56cc475c80 refactor(cursor): trim LOC in cursor byok tool normalization
- share one _CustomToolCallAccess mixin across the 4 new custom-tool
  classes instead of hand-rolling dict access on each
- inline the single-use _nest_flat_chat_tools / _flatten_chat_tools_for_responses
  list wrappers at their call sites
- drop _nest_flat_chat_tool_choice: it rewrote object-form chat tool_choice
  into {type,custom:{name}}, a shape OpenAI rejects; real Cursor never sends
  tool_choice on the messages arm, so pass it through unchanged
2026-08-01 11:26:09 -07:00
Tin Chi Lo
cc00650fec fix(litellm): treat a blank api_base as the default OpenAI endpoint in the bridge gate
A blank api_base (empty or whitespace) resolves to the default OpenAI
base downstream but is not None, so the constraint-enforcing-endpoint
check misclassified it as a custom backend and skipped the unset-effort
auto-bridge, leaving gpt-5.4+ function-tool requests to 400 at OpenAI.
The check now treats None, empty, and whitespace api_base alike; a real
custom base still opts out. Verified with get_llm_provider, which passes
a blank api_base through while resolving the provider to openai
2026-08-01 11:26:09 -07:00
Tin Chi Lo
d516a72c05 fix(litellm): scope the unset-effort responses bridge to constraint-enforcing endpoints
Chat-only OpenAI-compatible backends registered under the openai
provider with custom api_base and gpt-5.4+ model names served
tools-without-reasoning fine and have no /responses route, so the
unset-effort arm added for real OpenAI would have silently rerouted
previously working deployments. The arm now fires only when api_base is
unset (default OpenAI endpoint) or the provider is azure; an explicit
reasoning_effort keeps its pre-existing bridging behavior on any
api_base. Flagged lines also modernized to PEP 604
2026-08-01 11:26:09 -07:00
Tin Chi Lo
a5ba1caac5 test(helicone): stub the anthropic module unconditionally
An import probe proves nothing about the real SDK: it may be absent (it
lives in the proxy-runtime extra) and the tests/test_litellm/llms/anthropic
test package can shadow it once collection puts that path on sys.path,
which made the test order-sensitive across collection sets
2026-08-01 11:26:09 -07:00
Tin Chi Lo
e9d16bc35c fix(litellm): make the responses bridge and cursor routing total over the surfaces they now serve
Three gaps from the bridge becoming a mainstream path for chat traffic.
The chat to responses message converter only mapped function tool_calls,
so history carrying the native custom tool calls this PR introduced
raised "tool call not supported" on follow-up turns; custom entries now
map to custom_tool_call items and their results to
custom_tool_call_output. The stream translator returned an empty delta
for output_item.done on tool items, which left the responses guardrail
handler's tool extraction permanently empty (dead on staging too, where
the built chunk was discarded); stateless callers now receive the
complete tool call while per-stream callers keep the suppressed delta
that prevents client-side duplication. Cursor routing keyed on the
presence of a messages key, so a null or empty stub next to a real
agent-mode input array picked the chat arm; routing now keys on
messages content
2026-08-01 11:26:09 -07:00
Tin Chi Lo
bbba450301 fix(litellm): honor dict-form reasoning_effort in the bridge escape hatch and serialize custom tool calls in helicone and lunary logs
The bridge gate compared reasoning_effort against the string "none", so
litellm's dict form ({"effort": "none"}) wrongly bridged; the gate now
reads the effort value from either form and treats a summary inside the
dict as Responses-only regardless of effort. Helicone and lunary
previously skipped custom tool calls entirely; both now serialize them
(helicone as a tool_use block from the custom payload, lunary with the
custom name and input in its function fields, keeping type custom), with
new mapped tests for both integrations
2026-08-01 11:26:09 -07:00