Commit graph

18680 commits

Author SHA1 Message Date
yassin
7ff8945da4 fix(router): sync both usage keys from Redis even when one increment is zero
A stream counted before its usage is known increments TPM by zero, so the
worker that served it never refreshed its local TPM value from Redis and
the first byte headers reported the token count another worker had already
consumed. Both pipeline operations now always run, matching the pre-change
callback, so the returned values refresh both worker local keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:45:48 +00:00
Yujong Lee
f9d423827f fix(rust_bridge): let LITELLM_RUST win over litellm.rust() for optional tiers
Parse the switch with pydantic TypeAdapter(bool) so 1/true/yes/on and 0/false/no/off all work, and treat an unparseable value as unset instead of off. PYTHON_ONLY and RUST_REQUIRED still ignore both switches

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:45:11 +00:00
Yuneng Jiang
2481146727
docs(e2e): say why the CLI-driving cells needed a driver fix, not a rule 2026-09-16 12:44:53 -07:00
yassin
0d605b7b45 Merge remote-tracking branch 'origin/main' into litellm_streaming_buffer_release_on_scan 2026-09-16 19:42:36 +00:00
mateo-berri
0dff64ce1a fix(bedrock): move a Nova invoke cache point behind an image or tool result to the last text block 2026-09-16 12:42:15 -07:00
Yuneng Jiang
39acea0754
feat(e2e): make Claude Code send the same bytes every build
The compat cells drove the CLI with a fresh HOME per invocation and the
pytest process's own working directory, and both reach the request body.
The system prompt names a memory directory built from
$CLAUDE_CONFIG_DIR/projects/<cwd slug>, so a per-invocation config
directory rewrote every body, and the CLI adds a git block for its
working directory, so inheriting the checkout rewrote every body once
per candidate. The device id churned for the same reason: the CLI mints
it once and persists it in .claude.json, which we threw away each call.

Nothing here was load-bearing. All three ride in metadata.user_id, whose
job is abuse detection, not quota, caching or continuity. So pin the
config directory and the working directory at fixed paths, seed the
device id, and pin the session id.

HOME stays fresh and empty per invocation, so the isolation is no weaker
than before, and the CLI's own state no longer outlives the pod either.
The working directory is deliberately not the checkout, so a
model-directed Read now sees an empty directory rather than the
repository.

A pinned session id needs --no-session-persistence beside it: the CLI
refuses a session id another live process holds, and the matrix runs its
cells across xdist workers. Without the flag, six of eight concurrent
invocations die on "Session ID is already in use".
2026-09-16 12:42:15 -07:00
kerry-berri
b6143b3711
Merge pull request #41457 from BerriAI/litellm-providers/price-sync
chore(prices): sync Google Gemini prices: 22 models
2026-09-16 12:37:16 -07:00
yassin
6faeb19c0c Merge remote-tracking branch 'origin/main' into litellm_streaming_buffer_release_on_scan 2026-09-16 19:33:16 +00:00
Yujong Lee
358d767c9e refactor(rust_bridge): route Bedrock transcription through the shared runtime
Replace the stateful transcription loader with NativeBinding pairs and call
runtime.run/arun from the Bedrock dispatch class so the RUST_REQUIRED catalog
row is load-bearing: missing native and admission declines are terminal, and
there is no Python replay. Cover the remaining runtime, OCR lifecycle and
configuration branches, and make decide() exhaustive.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:30:32 +00:00
mateo-berri
aa6d053a37 Merge remote-tracking branch 'origin/main' into litellm_otel_gen_ai_system_none 2026-09-16 12:30:20 -07:00
yassin
cb26454034 test(router): cover success callback racing the pre-header count and drop redundant docstrings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:28:50 +00:00
Mateo Wang
a1b1f1ef9d
Merge pull request #34455 from BerriAI/litellm_lit4767_empty_choices_streaming_guard
fix(responses): guard empty-choices chunks in the Responses API streaming bridge
2026-09-16 12:28:46 -07:00
yassin
8de51dfaab test(otel): describe which request metadata the pre-call seed reads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:25:40 +00:00
yassin
938b782b8a Merge remote-tracking branch 'origin/main' into litellm_customer_budget_prometheus_metrics 2026-09-16 19:25:23 +00:00
yassin
d5837ab97a test(guardrails): pin tool-call-only scan keys as non-empty
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:24:14 +00:00
yassin
c22916377f test(prometheus): cover customer budget series expiry by end_user ttl
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:22:30 +00:00
kerry
e46106e20b test: drop gemini-3.1-flash-lite-image capability pins
The per-route capability test hardcoded vendor facts, including function calling support on the gemini route, which the live model card says is not supported. Keep the backup-matches-main invariant and the routing tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:21:50 +00:00
Devin AI
8caf2cb61f fix(models): flag prompt caching on Vertex and Azure AI Grok rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:17:45 +00:00
ryan-crabbe-berri
7ba47a5b6e fix(budgets): page end-user cache invalidation after a budget reset
The budget-tier reset read every customer linked to an expiring tier into
one result set before the write, then invalidated their caches one key at
a time. Both of those scale with the customer count, so a large enough
deployment can OOM the proxy pod on the read, and the tail of the
population sits on a stale spend counter while the per-key invalidations
drain

PR #40639 moved the reset write itself to a link-based UPDATE, so that
pre-commit read no longer feeds the write. It only fed cache invalidation
and the service-logging counts, which means it can move after the commit.
This replaces it with a keyset walk over litellm_endusertable ordered by
user_id, taking RESET_BUDGET_JOB_BATCH_SIZE rows per page, the same shape
_reset_windows_for_source already uses, with no per-run page cap for the
same reason that walk has none: the cursor cannot survive the run, so a
cap would restart at the first customer on every tick and never reach the
tail

Each page's counter and cache keys now go out as one batched delete
through a new DualCache.async_delete_cache_keys, which drops the
in-memory entries and chunks the Redis DELETE at
DEFAULT_MAX_REDIS_BATCH_CACHE_SIZE

num_endusers_found and num_endusers_updated now report the customers
whose caches were invalidated after the commit rather than the rows read
before it, so both read 0 when the cascade write fails
2026-09-16 12:12:18 -07:00
yassin
30f02aa6da fix(otel): read only the caller's requester_metadata snapshot in the v2 pre-call hook
The pre-call hook passed the proxy's whole per-request metadata dict into the
request identity, so proxy-owned siblings such as requester_ip_address were
promoted alongside the caller's keys. Only the requester_metadata mapping is
read now, keyed under its wrapper, which keeps the default allowlist behaviour
unchanged

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:11:25 +00:00
Yujong Lee
7a1d433e7a refactor(rust_bridge): declarative route catalog and shared runtime selection
Replace the per-route enablement helpers (rust_enabled, rust_ocr_enabled, RUST_CHAT_COMPLETIONS_PROVIDERS, FallbackMode) with a single rule table in litellm/rust_bridge/catalog.py that maps a Context(route, provider, model, delivery) to one of four rollout tiers, and a pure decide() that turns tier plus process/env switches into a Decision. runtime.run/arun own the only fallback path: Python for PYTHON, native then Python on missing binding or admission decline for RUST_WITH_FALLBACK, raise for RUST_REQUIRED. OCR is the first route on the shared runtime; chat completions, Anthropic messages, and Responses websocket policy checks now read the catalog.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:10:42 +00:00
yassin
73a0c3bb8a Merge remote-tracking branch 'origin/main' into litellm_streaming_buffer_release_on_scan 2026-09-16 19:09:19 +00:00
yassin
2821ba9667 fix(prometheus): refresh default-budget customers and honor independent customer gauges
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:08:42 +00:00
Devin AI
03931180d7 merge main into litellm_registry_audit_2026_09_14 2026-09-16 19:03:36 +00:00
yassin
e105cde56a fix(router): sync worker-local usage cache to shared post-increment counters
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:02:15 +00:00
yucheng-berri
b04d530ecf
Merge pull request #41386 from BerriAI/litellm_otel_trace_correlation
fix(proxy): default litellm_trace_id to the OTel server span trace id
2026-09-16 11:58:45 -07:00
yassin
ad25fe1886 fix(guardrails): hold unscannable Responses windows and key terminal envelopes by output items
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:56:56 +00:00
yassin
a2108f02fb test(router): cover _increment_deployment_usage delta and unlimited deployment behavior
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:51:40 +00:00
mateo-berri
033aa8ba6d fix(bedrock): forward userContext in Knowledge Base Retrieve requests
The Bedrock vector store search only lifted retrievalConfiguration out of extra_body, so the caller's userContext (the Retrieve API's ACL identity) never reached Bedrock and ACL-enabled data sources answered with zero results. The transform now forwards userContext, taken from extra_body first and then from the top-level params where the OpenAI SDK's extra_body merge lands, as the caller sent it.
2026-09-16 11:50:36 -07:00
yassin
6e5cdd8f12 fix(router): count TPM/RPM usage before building rate-limit headers
Router.make_call now increments the deployment TPM/RPM counter before set_response_headers reads remaining usage, so the headers carry post-increment values directly and the in-flight subtraction workaround from LIT-2719 is removed. deployment_callback_on_success adds only the tokens not yet counted (streams) and never a second request. The counter key uses the resolved deployment name so wildcard routes are read back correctly, and the proxy strips the router-owned counted-tokens marker from client metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:45:01 +00:00
yassin
8c89cff0e0 feat(prometheus): add customer (end_user) budget gauges
Mirror the key, team, user and org budget gauges for customer objects with
litellm_remaining_customer_budget_metric, litellm_customer_max_budget_metric
and litellm_customer_budget_remaining_hours_metric. The gauges carry only the
end_user label, are emitted after each request and from the startup budget
refresh for every customer with a budget attached, and reuse the
enable_end_user_cost_tracking_prometheus_only opt-in and end_user series caps

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:39:56 +00:00
yassin
8a059cd4b4 fix(otel): promote nested metadata keys under the caller's dotted path
Strip only the proxy's requester_metadata. wrapper from an allowlisted key so
requester_metadata.trace_id lands as litellm.metadata.trace_id while other
dotted keys keep their full path and cannot collide on a shared leaf name

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:39:55 +00:00
yucheng
74d8328ad0 Merge remote-tracking branch 'origin/main' into litellm_lit7836_call_id_endpoint_logs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	tests/test_litellm/proxy/test_proxy_server.py
2026-09-16 18:39:55 +00:00
Devin AI
4f585d3931 fix(responses): drop top_p for gpt-5 reasoning models when drop_params is set
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:35:49 +00:00
mateo-berri
774fc6021b chore(proxy): drop restating docstrings on the passthrough header helpers and refresh the lazy OpenAPI snapshot 2026-09-16 11:35:43 -07:00
Mateo Wang
78848d01a3
Merge pull request #34427 from BerriAI/litellm_bedrock_rag_retrieval_filter
fix(rag): forward retrieval_filter from retrieval_config to vector store search
2026-09-16 11:34:44 -07:00
yassin
81799a131b Merge branch 'litellm_regenerate_no_body_secret_sync' into litellm_key_alias_update_secret_sync 2026-09-16 18:34:37 +00:00
yassin
bb380fec28 test(proxy): drop redundant docstrings from alias rename secret sync tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:34:37 +00:00
yassin
c438c3b2f8 test(proxy): drop redundant docstring from regenerate secret sync test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:33:50 +00:00
yassin
0143fe5583 fix(guardrails): hold tool-call windows until the final scan and expose Bedrock streaming flags to the UI
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:30:12 +00:00
yassin
be1664a485 fix(proxy): rename AWS Secrets Manager secret when key alias changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:29:29 +00:00
mateo-berri
938702a3a7 Merge branch 'main' of https://github.com/BerriAI/litellm into litellm_anthropic_passthrough_strip_virtual_key 2026-09-16 11:28:09 -07:00
yassin
4513f78df5 fix(migrations-check): read the table name past comments, ignore referential SET DEFAULT
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:14:36 +00:00
yassin
17059564a8 feat(otel): promote nested request metadata keys to litellm.metadata.* span attributes
baggage_metadata_keys entries such as requester_metadata.trace_id now resolve the caller's nested metadata.trace_id and stamp it on the LLM-call span as litellm.metadata.trace_id, in both the OTEL v2 logger and the legacy OpenTelemetry callback. Nested metadata mappings are flattened to dotted paths, only allowlisted leaves are promoted, and the requester_metadata blob itself is never promoted

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:08:14 +00:00
yucheng
1b5dacc717 fix(proxy): tolerate malformed auth spans and read the OTel span from request state for custom auth
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:05:38 +00:00
yassin
bf39aebcf1 ci(migrations): flag defaulted ADD COLUMN on request-log tables
Postgres 10 has no fast default path, so ADD COLUMN ... DEFAULT on
LiteLLM_SpendLogs rewrites the heap and every index under an ACCESS
EXCLUSIVE lock inside the boot-time migrate deploy. The checker now
reports it on LiteLLM_SpendLogs and LiteLLM_ErrorLogs; the two shipped
migrations that already do it are grandfathered

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:02:00 +00:00
Devin AI
732ac614cc test(passthrough): expect absent query params
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 17:58:50 +00:00
yassin
ed18626291 fix(proxy): sync AWS Secrets Manager on body-less key regenerate
POST /key/{key}/regenerate with no request body reaches async_key_rotated_hook with data=None, and the secret manager sync was gated on data being present, so the rotated key never reached AWS Secrets Manager and the revoked key stayed stored. Gate on response.token_id only and read the requested alias null-safely

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 17:56:39 +00:00
Zach Bernstein
242bff782f
fix(scim): clamp collection page size 2026-09-16 12:52:22 -05:00
kerry
686556fffc test(anthropic): build served-model stream chunks immutably
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 17:49:15 +00:00