Commit graph

4669 commits

Author SHA1 Message Date
ryan-crabbe-berri
d2440639d5 feat(budgets): enforce shared budgets on model access groups
A model access group could gate which models a caller reaches but never how
much that group of callers could spend in total. Capping a shared pool meant
setting a per-entity budget on every key by hand, which caps each key
separately and still leaves no way to read what the group cost.

Spend is attributed to a group only when the group's name appears on an
allowlist the caller was granted (key, team, team-member scope, project or
org) and that group serves the requested model. Asking for a model that
merely belongs to a group attributes nothing, because nothing about the
caller named the group. Levels are unioned rather than ranked, so a team
granted "*" whose member is scoped to one group still counts as gated by
that group.

Enforcement runs on both paths tags already use: a reservation counter on
the pre-call path and a read-time max_budget check inside the existing
concurrent budget gather, so the ceiling still holds under
disable_budget_reservation.

Adds LiteLLM_ModelAccessGroupBudgetTable, which is the only place a group is
ever a row: the groups themselves stay free-text strings in
model_info.access_groups, so a row exists only once someone gives that group
a budget. GET, PUT and DELETE /access_group/{name}/budget manage it, and
/access_group/{name}/info now carries the spend and budget alongside the
models.
2026-08-29 12:13:55 -07:00
mateo-berri
4d4cf40334 fix(headroom): delegate to the parent deployment hook so deployment-level configs still compress 2026-08-29 12:06:38 -07:00
Mateo Wang
002d0068f5
Merge pull request #38580 from BerriAI/devin_ai_fix_model_new_read_replica_lag_38556
fix(proxy): pin model reconcile read to the writer DB so /model/new does not 500 under read replica lag
2026-08-29 12:02:08 -07:00
ryan-crabbe-berri
e8994e8ce1 Merge branch 'litellm_window_spend_writer' into litellm_window_spend_reader 2026-08-29 11:49:42 -07:00
ryan-crabbe-berri
4e22a5ef5a Merge branch 'litellm_window_spend_schema' into litellm_window_spend_writer 2026-08-29 11:47:37 -07:00
mateo-berri
f2f988bd57 test: cover forged _headroom_interception_converted_stream strip at the proxy boundary 2026-08-29 11:46:23 -07:00
Mateo Wang
5e60ec5c31
Merge pull request #38743 from BerriAI/litellm_techdebt_20260829
refactor: clean up tech debt that landed on 2026-08-29
2026-08-29 11:46:05 -07:00
mateo-berri
6d3e687ce4 fix(db): let the writer pin yield to the replica while the writer is degraded 2026-08-29 11:44:45 -07:00
mateo-berri
a8c36e8307 Merge remote-tracking branch 'origin/litellm_internal_staging' into devin_ai_fix_model_new_read_replica_lag_38556 2026-08-29 11:29:07 -07:00
yuneng-jiang
fa25ff2a2e
Merge pull request #38626 from BerriAI/litellm_ui_model_links_team_key_info
feat(ui): link team and key model chips to the models page filtered to that group
2026-08-29 11:26:15 -07:00
ryan-crabbe-berri
7745fe887f Merge branch 'litellm_window_spend_writer' into litellm_window_spend_reader 2026-08-29 11:06:05 -07:00
ryan-crabbe-berri
041cae8280 fix(proxy): bound the window spend seed exclusion to the batch's own start time
request_id can be chosen by the client through x-litellm-call-id, so an
unbounded NOT (request_id = ANY(batch)) let a replayed old id drop that id's
historical LiteLLM_SpendLogs row from the one-time seed while its increment
still landed. The increment now carries the request start, the batch keeps
the earliest one, and the seed only excludes ids whose startTime is at or
after it.
2026-08-29 11:05:57 -07:00
tin-berri
2a5d09ee87
fix(policy): let the AI policy suggester drop sampling params its model refuses (#38594)
The suggester pins temperature=0.2 for tool-selection determinism and passed no
drop_params, so an operator-supplied reasoning model whose only accepted temperature is 1
made litellm raise UnsupportedParamsError and the whole suggestion fail. The default
gpt-4o-mini is unaffected; the failure needs the caller to name a model.

Every other internal LLM call the proxy makes on a user's behalf already opts in through
judge_acompletion, which sets drop_params=True on both dispatch paths. This was the one
caller outside that contract, so the sampling preference is now advisory here too and the
call degrades instead of dying.

Resolves LIT-6352
2026-08-29 10:58:27 -07:00
Mateo Wang
c3edb95e8d Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_managed_batches_observability
# Conflicts:
#	tests/test_litellm/proxy/spend_tracking/test_spend_management_endpoints.py
2026-08-29 10:56:02 -07:00
mateo-berri
acb621c35b Merge remote-tracking branch 'origin/litellm_internal_staging' into devin_ai_fix_model_new_read_replica_lag_38556 2026-08-29 10:54:46 -07:00
mateo-berri
2457e60cfc Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit_4913_headroom_streaming_ccr 2026-08-29 10:54:24 -07:00
ryan-crabbe-berri
08118e6246 fix(proxy): let the exact model= filter match team BYOK public names
Team-scoped deployments keep the internal model_name_{team_id}_{uuid} routing key and expose the public name in model_info.team_public_model_name. The dashboard links team model chips with the public name, so the exact filter now matches either name via the existing helper.
2026-08-29 10:52:15 -07:00
ryan-crabbe-berri
9beb5ead4d fix(proxy): keep the exact model= DB predicate within the type-discipline budget
The where clause now uses the exact name string directly and skips the DB query when the typed search cannot occur in that name, so no new mutable literals are added (LIT002 gate).
2026-08-29 10:40:39 -07:00
ryan-crabbe-berri
3e99ee8d0e fix(proxy): scope the DB-side model search by the exact model= filter
With model=<group>&search=<term>, the router list was narrowed to the group but the DB query only matched the substring, so other groups' rows leaked into the page and total_count.
2026-08-29 10:32:29 -07:00
devin-ai-integration[bot]
30efcfd684
feat(mcp): support asymmetric (RS256) signing for MCP gateway session tokens (#38728)
* feat(mcp): support RS256 signing for MCP gateway session tokens

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: ruff format session token modules

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): enforce key strength on rotated public keys and unique kids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 10:18:59 -07:00
devin-ai-integration[bot]
0de1825450
fix(health): honor allow_requests_on_db_unavailable in readiness probe (#37640)
* fix(health): honor allow_requests_on_db_unavailable in readiness probe

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(health): bound readiness DB check and pass reconnect timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(health): bound whole readiness DB check with one deadline

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(health): keep readiness deadline fallback within lint budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(health): suppress TQ008 for proxy-global readiness patches

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(db): release reconnect lock when a waiting reconnect is cancelled

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: milan <milan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
2026-08-29 10:17:12 -07:00
Mateo Wang
cb7d41a5c6
Merge pull request #38739 from BerriAI/litellm_fix_tag_routing_reads_merged_metadata_tags
fix(proxy): tag routing misses proxy-merged tags when chat requests carry litellm_metadata
2026-08-29 10:14:05 -07:00
Rāna(Bass Ver.)
4f630411f0
Merge branch 'litellm_internal_staging' into fix/34379-unblock-customer 2026-08-30 00:17:26 +08:00
mateo-berri
c37260a2bd Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5
# Conflicts:
#	basedpyright-code-budget.json
#	enterprise/litellm_enterprise/proxy/audit_logging_endpoints.py
#	litellm/_lazy_imports.py
#	litellm/a2a_protocol/litellm_completion_bridge/transformation.py
#	litellm/integrations/bitbucket/bitbucket_client.py
#	litellm/integrations/compression_interception/handler.py
#	litellm/integrations/prometheus_helpers/prometheus_api.py
#	litellm/litellm_core_utils/model_response_utils.py
#	litellm/litellm_core_utils/url_utils.py
#	litellm/llms/anthropic/experimental_pass_through/context_management/dispatcher.py
#	litellm/llms/anthropic/experimental_pass_through/responses_adapters/handler.py
#	litellm/llms/anthropic/skills/transformation.py
#	litellm/llms/azure/files/handler.py
#	litellm/llms/bedrock/realtime/handler.py
#	litellm/llms/chatgpt/chat/streaming_utils.py
#	litellm/llms/compactifai/chat/transformation.py
#	litellm/llms/oci/chat/cohere.py
#	litellm/llms/vertex_ai/vector_stores/rag_api/transformation.py
#	litellm/proxy/agent_endpoints/agent_registry.py
#	litellm/proxy/client/cli/commands/credentials.py
#	litellm/proxy/client/cli/commands/teams.py
#	litellm/proxy/common_utils/get_routes.py
#	litellm/proxy/db/routing_prisma_wrapper.py
#	litellm/proxy/guardrails/guardrail_hooks/custom_code/sandbox.py
#	litellm/proxy/guardrails/guardrail_hooks/hiddenlayer/hiddenlayer.py
#	litellm/proxy/guardrails/guardrail_hooks/llm_as_a_judge/__init__.py
#	litellm/proxy/guardrails/guardrail_hooks/promptguard/promptguard.py
#	litellm/rust_bridge/responses_websocket.py
#	litellm/secret_managers/secret_manager_handler.py
#	ruff-strict-budget.json
#	type-discipline-budget.json
2026-08-29 06:37:10 -07:00
mateo-berri
8d4620649f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5
# Conflicts:
#	basedpyright-code-budget.json
#	litellm/caching/valkey_semantic_cache.py
#	litellm/integrations/compression_interception/handler.py
#	litellm/integrations/custom_logger.py
#	litellm/llms/custom_httpx/container_handler.py
#	litellm/llms/infinity/rerank/transformation.py
#	litellm/proxy/agent_endpoints/agent_registry.py
#	litellm/repositories/base_repository.py
#	litellm/repositories/credentials_repository.py
#	litellm/repositories/team_repository.py
#	ruff-strict-budget.json
#	type-discipline-budget.json
2026-08-29 06:03:33 -07:00
Devin AI
adf5637fdf chore: merge litellm_internal_staging into litellm_techdebt_20260829
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 11:31:44 +00:00
Mateo Wang
6a3333d3c8
Merge pull request #38747 from BerriAI/litellm_aws_partition_helper
fix(aws): build every AWS endpoint and ARN from the region's partition (aws-cn, aws-us-gov)
2026-08-29 04:10:59 -07:00
mateo-berri
229970c500 fix(guardrails): unwrap HiddenParamsAsyncIteratorWrapper before deferred dispatch class sniffing 2026-08-29 02:30:34 -07:00
mateo-berri
1947c65081 test(aws): type the new partition test parameters 2026-08-29 02:27:18 -07:00
mateo-berri
664697133b Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_guardrail_stream_audit 2026-08-29 02:24:59 -07:00
mateo-berri
f60ccf6234 fix(guardrails): match deferred stream dispatch shape per stream owner and defer passthrough logging until guardrail eos 2026-08-29 02:23:12 -07:00
mateo-berri
64eec53fd8 fix(guardrails): surface post-flush stream blocks as in-stream error frames and keep guardrail_information in spend logs
A guardrail block or failed scan that fires after SSE chunks have been
flushed can no longer set an HTTP status, so raising HTTPException there
silently truncated the stream. _emit_streaming_http_error now routes
post-flush failures through the endpoint translation's
build_stream_error_items, emitting the surface-correct error frame on
chat completions (data: {error}), /v1/messages (event: error), and
/v1/responses (ErrorEvent with the next sequence number). Pre-flush
blocks still raise with a real HTTP status.

Successful flags-on scans also logged metadata.guardrail_information as
null: the chat handler planted litellm_metadata on a route whose bucket
is metadata, flipping the bucket for every later write, and responses
streams fired their spend log before the eos scan ran. The chat handler
now merges user_api_key metadata through get_or_create_metadata_bucket,
and deferred stream-complete logging is armed for aresponses like it
already was for anthropic_messages.
2026-08-29 01:25:08 -07:00
mateo-berri
ad8c1457d1 fix(aws): build every AWS endpoint and ARN from the region partition
Adds litellm/litellm_core_utils/aws_partition.py mapping a region to its
AWS partition (aws, aws-cn, aws-us-gov, and the iso partitions), its DNS
suffix, and its ARN prefix, and uses it at every AWS host and ARN build
site: bedrock (runtime, agent, agentcore, legacy client, batches, files,
realtime), sagemaker, polly, secrets manager, s3 log uploads, bedrock
passthrough routes, and rag ingestion. ARN detection now accepts
arn:aws-cn: and arn:aws-us-gov: prefixes.

STS region resolution now falls back to the configured aws_region_name
after the aws_sts_endpoint host and the AWS_REGION/AWS_DEFAULT_REGION env
vars, so cn and gov role assumption no longer silently signs against
us-west-2.

A partition sweep test walks every endpoint builder with cn regions and
asserts no amazonaws.com host or arn:aws: prefix comes out, plus an AST
guard that fails on any new f-string hardcoding either literal.
2026-08-29 01:21:59 -07:00
mateo-berri
5d34fb20ff fix(bedrock): map real batch record counts and guard zero-count retire 2026-08-29 01:05:36 -07:00
Devin AI
9bfb332904 refactor: replace fresh getattr/setattr and test type-ignores with typed access
Same-day debt cleanup on code that landed in the last 24 hours. No behavior change.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 07:59:27 +00:00
mateo-berri
8fcfaf519c test: drop duplicated callback assertion 2026-08-29 00:53:24 -07:00
mateo-berri
817b5f44f6 fix(anthropic_endpoints): serialize dict-detail HTTPExceptions on /v1/messages like sibling surfaces 2026-08-29 00:48:14 -07:00
mateo-berri
f58c3c0868 fix(proxy): fold litellm_metadata into metadata on chat routes so tag routing sees merged tags 2026-08-29 00:42:43 -07:00
mateo-berri
7eb757a49b fix(bedrock): route streamed responses-API output through the unified guardrail
Streamed /v1/responses returned 500 whenever a Bedrock post_call guardrail
was enabled: the hook fed responses-API events into stream_chunk_builder,
which only understands chat-completions chunks, and the wrapped KeyError
surfaced as litellm.APIError before any ApplyGuardrail scan ran.

Delegate responses-API routes to UnifiedLLMGuardrails, whose translation
layer scans the assembled response at end of stream and only then releases
the buffered events, so flagged content never reaches the client.
2026-08-29 00:01:02 -07:00
Mateo Wang
ae7e50f096
Merge pull request #38607 from BerriAI/litellm_fix_files_pre_call_hook
fix(proxy): trigger async_pre_call_hook on POST /v1/files uploads
2026-08-28 23:58:05 -07:00
mateo-berri
e486e43d0f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_guardrail_stream_audit 2026-08-28 18:16:00 -07:00
Mateo Wang
fef5d3d0f9
Merge pull request #38713 from BerriAI/litellm_fix_messages_stream_guardrail_logging_race
fix(guardrails): record post_call scans on native /v1/messages streams
2026-08-28 17:44:55 -07:00
mateo-berri
a90ad5fe5c feat(bedrock): honor streaming buffer/sampling config for unbuffered post_call scans 2026-08-28 17:35:15 -07:00
devin-ai-integration[bot]
592518202c
feat(terraform): add litellm_jwt_key_mapping resource (#38714)
* feat(terraform): add litellm_jwt_key_mapping resource

Adds a Terraform resource for the proxy's JWT to virtual key mappings, so a
JWT client identified by a claim such as client_id, azp or sub maps to a
virtual key and inherits its models, budgets, rate limits and spend tracking.

Covers the four mapping endpoints: /jwt/key/mapping/new, /info, /update and
/delete. is_active is applied through a follow-up update because the create
endpoint always starts a mapping active, a dropped description is sent as an
empty string because the update endpoint ignores absent fields, changing the
mapped key rotates it in place, and changing the claim name or value forces
replacement since the update endpoint cannot change them.

* fix(terraform): revert key on failed jwt_key_mapping update

Classic SDKv2 persists a failed Update's diff-applied values to state
regardless of the error, so a rejected key rotation left the new key in
state while the proxy kept the old one and the next plan falsely converged.
Revert key via GetChange and resync description/is_active/computed fields
from a post-failure Read, since Read alone can't recover key (the proxy
never returns it).

Also drop the case-insensitive "mapping not found" body match: the proxy
raises 404 for all three not-found paths (info, update, delete), so
checking the status code alone is sufficient.

Clarify the docs: referencing a litellm_key resource's write-only key is
not a null-then-400 situation, it's a static "Missing required argument"
error at plan time, in every apply ordering.

* fix(terraform): stop leaving an active mapping behind on failed cleanup

Two issues flagged by review:

- Create has no way to ask the proxy for an inactive mapping, so an
  is_active=false mapping is briefly active while the follow-up
  deactivation runs. If that deactivation call itself fails, the mapping
  used to stay active and untracked. It's now deleted instead, closing
  the exposure rather than leaving it open indefinitely.
- On a failed update, only `key` was reverted before the recovery read.
  If that read also failed, description/is_active kept the rejected
  values, so a later plan could report false convergence. Now all three
  are reverted before the read runs.

Both come with regression tests, mutation-verified against the pre-fix
code.

* fix(deps): bump restrictedpython to 8.5 for GHSA-ffg3-p8fm-mjx2

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: retrigger ci

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): stub anthropic judge credentials in funnel seeding test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* revert(deps): keep uv.lock unchanged to keep the PR terraform-only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Fabrice Pont <fabrice.pont@doctolib.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-28 17:10:44 -07:00
yucheng-berri
470eb9620a
test(shadow_eval): configure the anthropic sdk judge in the funnel-seed test (#38717)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 16:59:22 -07:00
mateo-berri
1a5a856e3e fix(guardrails): defer native /v1/messages stream logging until post_call scans finish 2026-08-28 16:05:45 -07:00
tin-berri
4e48d74455
feat(shadow_eval): measure both arms' cost so a job reports what the router would have saved (#38631)
The attempt row now prices the real arm (the payload's response_cost plus its own
routing classifier when it routed) beside the shadow arm (completion plus the
classifier cost the routing decision writes back), and flags turns litellm's
response cache served. A per-leg funnel table counts the eligible requests that
produced no row (lost the sampling dice, unjudgeable shape, concurrency shed),
so results can weigh judged rows against the traffic they stand for. Job results
gain per-slice and overall arm spends plus the coverage counts, the budget gates
charge the shadow arm's classifier spend against max_budget, and the dashboard
shows the measured cost comparison beside the win rate

Resolves LIT-6358

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 15:13:19 -07:00
Yassin Kortam
671f89d8bc
fix(proxy): reset a key's budget-window counters on spend reset (#38686)
* fix(proxy): reset a key's budget-window counters and broadcast the reset cross-pod

/key/{id}/reset_spend already reset the key's lifetime spend counter in
Redis, but a key with its own budget_limits (an extra time-windowed cap,
e.g. a daily budget layered on top of the lifetime max_budget) kept its
window counter untouched, so the key stayed 429'd on
"ExceededBudget: Key over <duration> budget" even after the admin action
reported spend back to $0.

Force-expire each window on reset: zero its Redis counter and restart the
window from now (window_start is derived as reset_at - budget_duration,
so reset_at must float to now + duration, not the next calendar boundary
get_budget_reset_time gives key creation - that boundary can still be
in the past relative to the spend that triggered the block).

Also close a second, narrower race: _delete_cache_key_object evicted the
cached key object only on the handling pod, so another pod could keep
serving the stale pre-reset object (and re-derive the pre-reset spend
counter via its own floor-marker cache) until its own TTL expired. It now
broadcasts the eviction, matching the pattern already used for team,
team-member, customer, and tag caches.

* fix(proxy): evict the cached key object after every reset_spend DB write

Greptile P1: eviction ran before the window-reset DB write committed, so a
request racing the reset could re-fetch and re-cache the pre-write row,
pinning that pod to the stale budget_limits for the rest of its own cache
TTL even after the write went through. Move the eviction to run last.

* test: pin real cache state and satisfy the test-quality gate

test_delete_cache_key_object_broadcasts_invalidation now asserts a real
UserApiKeyCache no longer holds the evicted entry, rather than only
inspecting a mock's call args. Suppress test-quality-ok on the
hash_token/_check_proxy_or_team_admin_for_key/_delete_cache_key_object/
publish_auth_cache_invalidation patches: none has an HTTP boundary to
fake, matching the pattern the file already uses for these same targets.

* fix(proxy): narrow budget_limits by the str branch, not the list branch

isinstance(x, list) in the else branch still leaves Sequence[object] | str
(a tuple satisfies Sequence without being a list), so json.loads() saw a
possible non-str argument. Check isinstance(x, str) instead, which narrows
each branch to exactly the type it needs.

* fix(proxy): persist advanced budget-window boundaries before zeroing counters

Greptile P1: publishing a zeroed window counter before the new reset_at
committed let a request racing the write compute window_start from the
stale boundary, re-sum the historical spend log rows the reset was
clearing, and put the counter right back above budget. Compute every
window's new boundary, persist all of them in one DB write, then zero
each window's Redis counter only once that write has landed.
2026-08-28 15:13:03 -07:00
tin-berri
fb80ba7c98
fix(spend): remove the proxy-wide autorouter savings baseline override (#38700)
Every complexity router now derives and records its savings baseline from
its hardest configured tier, and the spend writer always prices against the
decision-recorded baseline model and deployment id. A leftover
litellm_settings.autorouter_savings_baseline_model key is inert
2026-08-28 14:53:27 -07:00
tin-berri
e966369558
fix(shadow-eval): validate Anthropic SDK judge credentials (#38701)
* fix(shadow-eval): validate Anthropic SDK judge credentials

* test(shadow-eval): configure valid SDK judges
2026-08-28 14:24:46 -07:00