Commit graph

13884 commits

Author SHA1 Message Date
mateo-berri
c64318cfbd Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_post_call_policy_pipeline 2026-08-29 11:24:43 -07:00
tin-berri
2a5d09ee87
fix(policy): let the AI policy suggester drop sampling params its model refuses (#38594)
The suggester pins temperature=0.2 for tool-selection determinism and passed no
drop_params, so an operator-supplied reasoning model whose only accepted temperature is 1
made litellm raise UnsupportedParamsError and the whole suggestion fail. The default
gpt-4o-mini is unaffected; the failure needs the caller to name a model.

Every other internal LLM call the proxy makes on a user's behalf already opts in through
judge_acompletion, which sets drop_params=True on both dispatch paths. This was the one
caller outside that contract, so the sampling preference is now advisory here too and the
call degrades instead of dying.

Resolves LIT-6352
2026-08-29 10:58:27 -07:00
devin-ai-integration[bot]
30efcfd684
feat(mcp): support asymmetric (RS256) signing for MCP gateway session tokens (#38728)
* feat(mcp): support RS256 signing for MCP gateway session tokens

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: ruff format session token modules

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): enforce key strength on rotated public keys and unique kids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 10:18:59 -07:00
devin-ai-integration[bot]
0de1825450
fix(health): honor allow_requests_on_db_unavailable in readiness probe (#37640)
* fix(health): honor allow_requests_on_db_unavailable in readiness probe

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(health): bound readiness DB check and pass reconnect timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(health): bound whole readiness DB check with one deadline

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(health): keep readiness deadline fallback within lint budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(health): suppress TQ008 for proxy-global readiness patches

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(db): release reconnect lock when a waiting reconnect is cancelled

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: milan <milan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
2026-08-29 10:17:12 -07:00
Mateo Wang
cb7d41a5c6
Merge pull request #38739 from BerriAI/litellm_fix_tag_routing_reads_merged_metadata_tags
fix(proxy): tag routing misses proxy-merged tags when chat requests carry litellm_metadata
2026-08-29 10:14:05 -07:00
Mateo Wang
ed2f1836de
Merge pull request #38738 from BerriAI/litellm_fix_batch_list_page_past_unparseable_rows
fix(batches): fill a managed batch page past rows that will not parse
2026-08-29 10:13:29 -07:00
yucheng-berri
f0fadb7f99
test(e2e): add logging e2e coverage (s3_v2, gcs_bucket, team langfuse callback, datadog failure) (#38552)
* test: add logging e2e coverage (s3_v2, gcs_bucket, team langfuse callback, datadog failure)

Five new live e2e scenarios raising Logging & Guardrails registry coverage:
s3_v2 success and failure objects read back from the real S3 bucket,
gcs_bucket success record read back through the GCS JSON API (with
nextPageToken pagination and per-request bearer minting), team-scoped
Langfuse callback delivery with non-team isolation, and DataDog failure
event delivery queried by indexed model_group. datadog_reader gains
query-based variants of the marker search; the langfuse cell is a new
registry row. Bucket readers settle past a full flush interval so a
late duplicate cannot hide from the exactly-one assertions

* test: cover clock-skew day prefix in gcs read-back and retry team callback propagation

* test: key the s3 failure read-back on the provider error, not payload absence

* chore: rerun ci

* chore: rerun ci after config sync

* chore: rerun ci with pr lane env

* chore: rerun ci

* chore: rerun ci

* chore: rerun ci

* chore: rerun ci

* chore: rerun ci

* test: add guardrail e2e coverage (presidio masking, bedrock post and during call, moderation on messages) (#38553)

* test: add guardrail e2e coverage (presidio masking, bedrock post/during, moderation on messages)

* test: require the phone placeholder positively in the presidio masking predicate

* test: count only the 400 verdict body as a bedrock post_call block

* test(e2e): exempt the guardrail config echo from the post_call leak assertion

* test(e2e): pin the fail-closed contract for an unknown guardrail name (skipped, product gap)

* test(e2e): tolerate the readiness 503 from a transient db blip in the callback-config probes
2026-08-29 09:43:44 -07:00
Mateo Wang
c42ac262d3
Merge branch 'litellm_internal_staging' into fix_databricks_oauth_url 2026-08-29 07:00:58 -07:00
Mateo Wang
e48f8f016f
Merge pull request #38148 from mubashir1osmani/litellm_hosted_vllm_videos
feat(hosted_vllm): add vLLM-Omni videos API
2026-08-29 06:30:35 -07:00
Mateo Wang
6a3333d3c8
Merge pull request #38747 from BerriAI/litellm_aws_partition_helper
fix(aws): build every AWS endpoint and ARN from the region's partition (aws-cn, aws-us-gov)
2026-08-29 04:10:59 -07:00
Mateo Wang
39e4f1ae13
Merge pull request #38727 from BerriAI/litellm_aws_external_id_embed_sagemaker
fix(aws): forward aws_external_id in Bedrock embeddings and SageMaker credential loading
2026-08-29 03:30:32 -07:00
Mateo Wang
1fbf34e559
Merge pull request #38744 from BerriAI/litellm_fix_bedrock_batch_zero_count_retire
fix(bedrock): map real batch record counts and guard zero-count retire
2026-08-29 03:12:49 -07:00
mateo-berri
850b18d214 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_tag_routing_reads_merged_metadata_tags
# Conflicts:
#	tests/openai_endpoints_tests/test_e2e_openai_responses_api.py
2026-08-29 02:34:51 -07:00
mateo-berri
1947c65081 test(aws): type the new partition test parameters 2026-08-29 02:27:18 -07:00
mateo-berri
d5b5aca498 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_bedrock_batch_zero_count_retire 2026-08-29 02:21:28 -07:00
mateo-berri
2d11870793 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_batch_id_fallback_pin
# Conflicts:
#	tests/openai_endpoints_tests/test_e2e_openai_responses_api.py
2026-08-29 02:19:45 -07:00
mateo-berri
cd5cf84aac Merge branch 'litellm_internal_staging' into litellm_aws_partition_helper 2026-08-29 02:18:16 -07:00
mateo-berri
b0d485568b test(batches): pin the poller against retiring zero-count completed batches 2026-08-29 02:00:11 -07:00
mateo-berri
66947fcf7c test(responses): expect the bad-temperature 400 on a non-reasoning model 2026-08-29 01:58:15 -07:00
mateo-berri
8cf090b368 fix(router): arm the provider-scoped fallback pin only on resource-operating handlers 2026-08-29 01:46:33 -07:00
mateo-berri
7fbcd9c3ed fix(bedrock): treat partial record counts as unknown on batch retrieve 2026-08-29 01:40:07 -07:00
mateo-berri
81eb014e71 test(responses): use a reasoning-legal temperature in the gpt-5.5 extra_body merge test 2026-08-29 01:36:45 -07:00
mateo-berri
2968c246b9 test(responses): repair two tests broken by the gpt-5 temperature gate
PR #38593 stopped forwarding temperature to reasoning models, which left
test_extra_body_merges_with_request_data raising UnsupportedParamsError
and test_bad_request_bad_param_error no longer getting a rejection from
OpenAI because drop_params now eats the param. Both repairs are the same
hunks PR #38739 carries, so the branches merge clean in either order
2026-08-29 01:32:09 -07:00
mateo-berri
9da766a1f9 test: use a non-reasoning model for the responses bad-param e2e fixture 2026-08-29 01:22:49 -07:00
mateo-berri
ad8c1457d1 fix(aws): build every AWS endpoint and ARN from the region partition
Adds litellm/litellm_core_utils/aws_partition.py mapping a region to its
AWS partition (aws, aws-cn, aws-us-gov, and the iso partitions), its DNS
suffix, and its ARN prefix, and uses it at every AWS host and ARN build
site: bedrock (runtime, agent, agentcore, legacy client, batches, files,
realtime), sagemaker, polly, secrets manager, s3 log uploads, bedrock
passthrough routes, and rag ingestion. ARN detection now accepts
arn:aws-cn: and arn:aws-us-gov: prefixes.

STS region resolution now falls back to the configured aws_region_name
after the aws_sts_endpoint host and the AWS_REGION/AWS_DEFAULT_REGION env
vars, so cn and gov role assumption no longer silently signs against
us-west-2.

A partition sweep test walks every endpoint builder with cn regions and
asserts no amazonaws.com host or arn:aws: prefix comes out, plus an AST
guard that fails on any new f-string hardcoding either literal.
2026-08-29 01:21:59 -07:00
mateo-berri
6301a520a6 test: pin reasoning effort none in the responses extra_body fixture 2026-08-29 01:07:24 -07:00
mateo-berri
5d34fb20ff fix(bedrock): map real batch record counts and guard zero-count retire 2026-08-29 01:05:36 -07:00
mateo-berri
8fcfaf519c test: drop duplicated callback assertion 2026-08-29 00:53:24 -07:00
mateo-berri
0b14897a7e fix(router): pin batch, file, and fine-tuning job ids to their owning model group on fallback 2026-08-29 00:47:34 -07:00
mateo-berri
f58c3c0868 fix(proxy): fold litellm_metadata into metadata on chat routes so tag routing sees merged tags 2026-08-29 00:42:43 -07:00
mateo-berri
d7bd6ca614 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_aws_external_id_embed_sagemaker 2026-08-29 00:17:05 -07:00
mateo-berri
996019cd23 fix(policy_engine): keep post_call pipeline guardrail logging and reject background bypass
Post_call pipelines run step hooks against a copied request dict, so guardrail
writes into the metadata bucket (applied_guardrails for the response header,
standard_logging_guardrail_information for spend logs) were dropped when the
guardrail was the first writer. Merge those writes back onto the request on the
post_call allow path, keeping the request payload and the executor's per-step
guardrails activation flag out of it.

Background /v1/responses requests dodge the streaming 400: pre_call sees stream
unset, then the polling task forces stream=true with pre-call logic skipped and
the streaming branch returns before post_call_success_hook, silently bypassing
post_call pipelines. Reject background=true at pre_call the same way as
stream=true.

Also pin the run_in_parallel pipeline-managed exclusion in both hook loops with
regression tests.
2026-08-29 00:14:23 -07:00
Mateo Wang
ae7e50f096
Merge pull request #38607 from BerriAI/litellm_fix_files_pre_call_hook
fix(proxy): trigger async_pre_call_hook on POST /v1/files uploads
2026-08-28 23:58:05 -07:00
mateo-berri
44a5d7a47a fix(aws): forward aws_external_id in bedrock embeddings and sagemaker credential loading 2026-08-28 18:25:14 -07:00
mateo-berri
55569729b0 fix(policy_engine): scope pipeline-managed guardrail skips to the pipeline's mode 2026-08-28 18:16:35 -07:00
mateo-berri
aeac6a412c fix(policy_engine): skip pipeline-managed guardrails in the response-path guardrail loop 2026-08-28 17:59:15 -07:00
Mateo Wang
fef5d3d0f9
Merge pull request #38713 from BerriAI/litellm_fix_messages_stream_guardrail_logging_race
fix(guardrails): record post_call scans on native /v1/messages streams
2026-08-28 17:44:55 -07:00
mateo-berri
e6edd62f5d fix(policy_engine): propagate post_call pipeline replacement responses to the client 2026-08-28 17:40:48 -07:00
mateo-berri
dfcea2c186 fix(policy_engine): execute post_call guardrail pipelines on responses 2026-08-28 17:16:00 -07:00
devin-ai-integration[bot]
592518202c
feat(terraform): add litellm_jwt_key_mapping resource (#38714)
* feat(terraform): add litellm_jwt_key_mapping resource

Adds a Terraform resource for the proxy's JWT to virtual key mappings, so a
JWT client identified by a claim such as client_id, azp or sub maps to a
virtual key and inherits its models, budgets, rate limits and spend tracking.

Covers the four mapping endpoints: /jwt/key/mapping/new, /info, /update and
/delete. is_active is applied through a follow-up update because the create
endpoint always starts a mapping active, a dropped description is sent as an
empty string because the update endpoint ignores absent fields, changing the
mapped key rotates it in place, and changing the claim name or value forces
replacement since the update endpoint cannot change them.

* fix(terraform): revert key on failed jwt_key_mapping update

Classic SDKv2 persists a failed Update's diff-applied values to state
regardless of the error, so a rejected key rotation left the new key in
state while the proxy kept the old one and the next plan falsely converged.
Revert key via GetChange and resync description/is_active/computed fields
from a post-failure Read, since Read alone can't recover key (the proxy
never returns it).

Also drop the case-insensitive "mapping not found" body match: the proxy
raises 404 for all three not-found paths (info, update, delete), so
checking the status code alone is sufficient.

Clarify the docs: referencing a litellm_key resource's write-only key is
not a null-then-400 situation, it's a static "Missing required argument"
error at plan time, in every apply ordering.

* fix(terraform): stop leaving an active mapping behind on failed cleanup

Two issues flagged by review:

- Create has no way to ask the proxy for an inactive mapping, so an
  is_active=false mapping is briefly active while the follow-up
  deactivation runs. If that deactivation call itself fails, the mapping
  used to stay active and untracked. It's now deleted instead, closing
  the exposure rather than leaving it open indefinitely.
- On a failed update, only `key` was reverted before the recovery read.
  If that read also failed, description/is_active kept the rejected
  values, so a later plan could report false convergence. Now all three
  are reverted before the read runs.

Both come with regression tests, mutation-verified against the pre-fix
code.

* fix(deps): bump restrictedpython to 8.5 for GHSA-ffg3-p8fm-mjx2

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: retrigger ci

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): stub anthropic judge credentials in funnel seeding test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* revert(deps): keep uv.lock unchanged to keep the PR terraform-only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Fabrice Pont <fabrice.pont@doctolib.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-28 17:10:44 -07:00
yucheng-berri
470eb9620a
test(shadow_eval): configure the anthropic sdk judge in the funnel-seed test (#38717)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 16:59:22 -07:00
mateo-berri
1a5a856e3e fix(guardrails): defer native /v1/messages stream logging until post_call scans finish 2026-08-28 16:05:45 -07:00
Mateo Wang
4ebedf901a
Merge pull request #38696 from BerriAI/litellm_lit6376_dspy_streaming
fix(streaming): report response_cost and Anthropic citations from stream_chunk_builder
2026-08-28 15:31:23 -07:00
Mateo Wang
27c09248e4
Merge pull request #38593 from BerriAI/litellm_gpt5_default_reasoning_effort
fix(gpt-5): stop forwarding temperature and top_p to reasoning models that reject them
2026-08-28 15:20:41 -07:00
tin-berri
4e48d74455
feat(shadow_eval): measure both arms' cost so a job reports what the router would have saved (#38631)
The attempt row now prices the real arm (the payload's response_cost plus its own
routing classifier when it routed) beside the shadow arm (completion plus the
classifier cost the routing decision writes back), and flags turns litellm's
response cache served. A per-leg funnel table counts the eligible requests that
produced no row (lost the sampling dice, unjudgeable shape, concurrency shed),
so results can weigh judged rows against the traffic they stand for. Job results
gain per-slice and overall arm spends plus the coverage counts, the budget gates
charge the shadow arm's classifier spend against max_budget, and the dashboard
shows the measured cost comparison beside the win rate

Resolves LIT-6358

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 15:13:19 -07:00
Yassin Kortam
671f89d8bc
fix(proxy): reset a key's budget-window counters on spend reset (#38686)
* fix(proxy): reset a key's budget-window counters and broadcast the reset cross-pod

/key/{id}/reset_spend already reset the key's lifetime spend counter in
Redis, but a key with its own budget_limits (an extra time-windowed cap,
e.g. a daily budget layered on top of the lifetime max_budget) kept its
window counter untouched, so the key stayed 429'd on
"ExceededBudget: Key over <duration> budget" even after the admin action
reported spend back to $0.

Force-expire each window on reset: zero its Redis counter and restart the
window from now (window_start is derived as reset_at - budget_duration,
so reset_at must float to now + duration, not the next calendar boundary
get_budget_reset_time gives key creation - that boundary can still be
in the past relative to the spend that triggered the block).

Also close a second, narrower race: _delete_cache_key_object evicted the
cached key object only on the handling pod, so another pod could keep
serving the stale pre-reset object (and re-derive the pre-reset spend
counter via its own floor-marker cache) until its own TTL expired. It now
broadcasts the eviction, matching the pattern already used for team,
team-member, customer, and tag caches.

* fix(proxy): evict the cached key object after every reset_spend DB write

Greptile P1: eviction ran before the window-reset DB write committed, so a
request racing the reset could re-fetch and re-cache the pre-write row,
pinning that pod to the stale budget_limits for the rest of its own cache
TTL even after the write went through. Move the eviction to run last.

* test: pin real cache state and satisfy the test-quality gate

test_delete_cache_key_object_broadcasts_invalidation now asserts a real
UserApiKeyCache no longer holds the evicted entry, rather than only
inspecting a mock's call args. Suppress test-quality-ok on the
hash_token/_check_proxy_or_team_admin_for_key/_delete_cache_key_object/
publish_auth_cache_invalidation patches: none has an HTTP boundary to
fake, matching the pattern the file already uses for these same targets.

* fix(proxy): narrow budget_limits by the str branch, not the list branch

isinstance(x, list) in the else branch still leaves Sequence[object] | str
(a tuple satisfies Sequence without being a list), so json.loads() saw a
possible non-str argument. Check isinstance(x, str) instead, which narrows
each branch to exactly the type it needs.

* fix(proxy): persist advanced budget-window boundaries before zeroing counters

Greptile P1: publishing a zeroed window counter before the new reset_at
committed let a request racing the write compute window_start from the
stale boundary, re-sum the historical spend log rows the reset was
clearing, and put the counter right back above budget. Compute every
window's new boundary, persist all of them in one DB write, then zero
each window's Redis counter only once that write has landed.
2026-08-28 15:13:03 -07:00
tin-berri
42d278cfad
fix(router): tier-pinned reasoning_effort supersedes client effort carriers (#38698)
An auto-router tier pin injects reasoning_effort into request_kwargs, but
provider translations give a caller-supplied thinking, output_config.effort,
or reasoning carrier precedence over the reasoning_effort alias, so the pin
never reached the wire whenever the client expressed effort natively. Drop
the client's other encodings of the setting at the tier-param merge; a
client output_config keeps its non-effort fields
2026-08-28 14:53:33 -07:00
tin-berri
fb80ba7c98
fix(spend): remove the proxy-wide autorouter savings baseline override (#38700)
Every complexity router now derives and records its savings baseline from
its hardest configured tier, and the spend writer always prices against the
decision-recorded baseline model and deployment id. A leftover
litellm_settings.autorouter_savings_baseline_model key is inert
2026-08-28 14:53:27 -07:00
tin-berri
e966369558
fix(shadow-eval): validate Anthropic SDK judge credentials (#38701)
* fix(shadow-eval): validate Anthropic SDK judge credentials

* test(shadow-eval): configure valid SDK judges
2026-08-28 14:24:46 -07:00
yucheng-berri
3e280b1be9
fix(router): scrub fallback stamp keys in place and strip them at the proxy boundary (#38690)
PR #38586 changed the fallback-stamp scrub in async_function_with_fallbacks to
rebind kwargs[sibling] to a scrubbed copy instead of popping in place. Every
other router bucket write mutates the caller's dict in place, and everything
below the router resolves the metadata bucket by key presence, so on a proxy
request that carries litellm_metadata the copy becomes a detached object: the
proxy's post_call guardrail write-backs land in request_data while the spend
row is built from the router's copy. Result: guardrail_information and the
guardrail cost silently drop from the spend row on any request that planted a
reserved key, and an SDK caller aliasing one dict as both buckets loses the
router stamps entirely.

Scrub in place again, and move the anti-spoof to the proxy boundary: strip
attempted_fallbacks and original_model_group from client-supplied metadata and
litellm_metadata in add_litellm_data_to_request, next to the pricing-field
strip, so proxy traffic never carries a reserved key and the in-place pop only
ever fires for an SDK caller that planted one. Keep #38586's hop-stamp ordering
fix (caller keys first, stamps appended) untouched.
2026-08-28 14:18:56 -07:00