Commit graph

49299 commits

Author SHA1 Message Date
mateo-berri
fe6f615c7a fix(proxy): log rejected unknown-model requests under a placeholder model name
A request whose model field matched no configured model was rejected with 400 but its failure row still persisted the raw client string as the model, so a client that concatenated its prompt into the model field wrote that prompt into LiteLLM_SpendLogs and the daily spend tables, where /user/daily/activity/aggregated returned it as a breakdown.models key. The spend log payload now records such rejections under the constant unknown-model, keeping the failed request counted without persisting client input as a model name.
2026-09-10 14:18:54 -07:00
mateo-berri
05f459d898 fix(redis): quiet every per-request Redis fallback while the breaker is open 2026-09-10 14:18:53 -07:00
mateo
1957bd388e fix(proxy): treat an open Redis breaker as a dropped spend counter update, not a cost tracking failure
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 21:15:02 +00:00
moe-berri
bc77aa05d2
Merge pull request #40608 from BerriAI/moe/lit-7493-zocdocauto-router-encrypted-codex-sub-agent-task-is
fix(router): classify encrypted delegated tasks with native Responses
2026-09-10 14:08:31 -07:00
ryan-crabbe-berri
2f114d44ed
Merge pull request #39190 from WolframRavenwolf/litellm_wandb_reasoning_effort
fix(wandb): preserve reasoning_effort in chat completions
2026-09-10 14:02:49 -07:00
mateo-berri
3b0ffaaa8d test(proxy): send valid organization requests so the licensed gate test proves the handler ran 2026-09-10 14:00:41 -07:00
mateo-berri
ad607516a2 fix(caching): log a refused async_increment as debug while the breaker is open 2026-09-10 13:59:15 -07:00
mateo
32b7daf691 fix(caching): make an open Redis circuit breaker a quiet cache miss
An open breaker raised a generic Exception on every skipped call, and DualCache caught
it and logged a full ERROR traceback each time. Under load that became hundreds of
traceback formats per second on every replica and pinned the proxies at 100% CPU.

Raise a typed RedisCircuitBreakerOpenError instead and have DualCache return its
in-memory result without logging for it. Classify redis-py pool exhaustion
(ConnectionError chained from TimeoutError) as a timeout so a latency blip goes
through the duration gate. Track a breaker generation so a call admitted before the
breaker opened cannot close it, leaving that to the HALF_OPEN probe. Rate limit the
LoggingWorker callback-error traceback to one per interval so a stalled logging
backend cannot start a second traceback storm.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 20:58:41 +00:00
mateo-berri
e004634396 docs(github): cover every regression class in the Affected release section 2026-09-10 13:58:31 -07:00
mateo-berri
4a3950cf67 refactor(realtime): type the Azure protocol picker with the streaming module's websocket protocol 2026-09-10 13:57:48 -07:00
devin-ai-integration[bot]
46a185d3cd
feat(rust_bridge): count budget-check input tokens in Rust on all LLM routes (#40381)
* feat(rust_bridge): count budget-check input tokens in Rust on all LLM routes

Rust counts input tokens from the raw JSON body with the GIL released inside the existing budget reservation, covering every LLM route the auth dependency guards. It only fires for models on the Anthropic tokenizer when a budget is set, and Python counts whenever Rust is off, missing, or declines a body shape.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(rust): count byte-level BPE tokens without the GPT-2 split regex (#40594)

The oniguruma run of the ByteLevel pre-tokenizer regex is about 90% of
encode_fast on a 100k token body (100 ms of the ~110 ms Rust admission
count in the gateway pod). A hand-written scanner that yields the same
pieces, then feeds the model directly, counts the same text in 10 ms.
It only engages for tokenizers with the Anthropic shape (optional NFKC,
ByteLevel without prefix space, no post-processor) and falls back to the
full encoder when the text contains an added token. Parity with
encode_fast is tested on random texts, the pieces are compared with the
real pre-tokenizer, and the \p{L}/\p{N}/\s tables are checked against
oniguruma for every code point.

NFKC runs through unicode-normalization-alignments, the crate and
Unicode tables NormalizedString::nfkc already uses, so the fast path
normalizes exactly what the full encoder would. Using the newer
unicode-normalization crate changed the count for 171 code points that
gained compatibility decompositions after Unicode 9 (U+32FF, U+A7F1..).
The fast normalizer is compared with the tokenizer's for every scalar
value and on random texts.

The scanner is built without mutable state: byte_char and mapped_len replace the const table builders and the reusable mapped buffer, and iter::successors replaces the stateful piece iterator. byte_chars_match_the_byte_level_alphabet checks the byte mapping against ByteLevel for every scalar value.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust_bridge): bound concurrent token-count encodes and share the Anthropic tokenizer predicate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 13:56:30 -07:00
mateo-berri
065ab11f0b Merge remote-tracking branch 'origin/litellm_internal_staging' into HEAD 2026-09-10 13:55:11 -07:00
devin-ai-integration[bot]
0e35c8fee9
fix(proxy): recreate the Prisma client when the writer session turns read-only (#40610)
The writer health probe only ran SELECT 1, which a read-only Postgres
session answers fine, so a pooled connection left pointing at a demoted
primary kept failing every write with SQLSTATE 25006 until the pod was
restarted. Probe transaction_read_only instead, treat a 25006 on the
request path as a signal to recreate the client, and back off
exponentially while the database as a whole stays read-only so a replica
or an in-progress failover does not get its engine killed every cycle.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 13:53:42 -07:00
mateo-berri
0e2402aa5d docs(github): add an Affected release section to the PR template
A fix for a perf, memory, or crash regression names the released or rc
version it regressed in and carries the `backport-stable` label, so the
stable release gate cherry-picks it onto the rc line before tagging.
2026-09-10 13:52:51 -07:00
devin-ai-integration[bot]
9cd1c4c29c
feat(terraform): prometheus metrics sidecar for the GCP Cloud Run gateway (#40614)
* feat(terraform): Prometheus metrics sidecar for the GCP Cloud Run gateway

gateway_metrics_port adds a metrics container running
litellm.proxy.prometheus_metrics_server next to the gateway, sharing the
PROMETHEUS_MULTIPROC_DIR over an in-memory volume, plus a Managed Service
for Prometheus collector sidecar that scrapes it over localhost and writes
to Cloud Monitoring. The gateway stays on port 4000 and the load balancer
routing is unchanged.

Resolves LIT-7502

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(terraform): reject fractional and collector health ports for gateway_metrics_port

Greptile review on #40614: 4000.5 fails integer port parsing and 13133 collides with the
gmp sidecar liveness listener. Also stop claiming the load balancer's own /metrics goes away

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 13:51:59 -07:00
mateo-berri
d3a0b0d45b fix(voyage): let caller params win for contextual auto-chunking and drop duplicate cost map entries
A flat list[str] sent to voyage-context-4 is treated as independent inputs and forwarded flat with
enable_auto_chunking=True, chunk_size=32000, and input_type=document unless the caller already set
input_type=query. Caller-supplied params now override the defaults instead of being clobbered.

The voyage-4 family and voyage-context-4 cost map entries already exist on litellm_internal_staging,
and voyage-4-nano is not served by the Voyage API, so those additions and their pricing test are dropped.
2026-09-10 13:51:45 -07:00
mateo-berri
9da3e63ae9 fix(redis): log an open circuit breaker once instead of a traceback per request and count sync timeouts as timeouts
While the Redis circuit breaker is open every guarded call was refused with a bare
Exception that each swallowing catch site logged as an ERROR traceback, so a
sub-second latency blip turned into thousands of tracebacks per minute and pinned
every replica's CPU. The sync guard also recorded socket timeouts as hard
connectivity failures, so with least-busy routing the breaker opened on the first
slow replies and the timeout-only min-duration guard never applied.

Refusals now raise RedisCircuitBreakerOpenError, and the catch sites route it
through log_redis_failure, which logs a refusal at DEBUG and everything else at
the caller's level. The sync guard passes is_timeout like the async one.
2026-09-10 13:50:34 -07:00
moe-berri
d266b76a8b fix(router): honor Codex reminders and map classifier failures 2026-09-10 13:48:21 -07:00
mateo-berri
8fe2094a55 refactor(realtime): move the Azure protocol picker into the Azure realtime handler 2026-09-10 13:46:13 -07:00
mateo-berri
4de441c181 test(ui): render TeamSSOSettings under a premium session so the organization dropdown tests fetch 2026-09-10 13:41:37 -07:00
devin-ai-integration[bot]
a9cec50960
feat(infra): scale gateway on per-pod RPS and TPS in Helm and Terraform (#40479)
* feat(infra): scale gateway on per-pod RPM and TPM in Helm and Terraform

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(helm): require the metrics server before rendering the gateway ServiceMonitor

The http port serves /metrics/ behind virtual-key auth, so a ServiceMonitor
pointed at it only collects 401s and the RPM/TPM HPA metrics never appear

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(infra): express gateway HPA, KEDA and ECS workload targets per second

Rename the per-pod request and token targets in both Helm charts and the
AWS module from per minute to per second, and shorten the recommended
Prometheus rate window to [1m] with no * 60 so the adapter and KEDA
signals are what the HPA compares against. ECS keeps CloudWatch's
60-second aggregation: the ALB target is 60x the per-second variable and
the token metric math divides the period Sum by 60 before dividing by
the running task count.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 13:36:23 -07:00
moe-berri
126dc37de4 fix(logging): reconstruct classifier audit redaction payloads 2026-09-10 13:33:36 -07:00
Joshua Valluru
93bd3d61b5 fix(mcp): preserve verified identity when warming OAuth cache 2026-09-10 13:20:49 -07:00
mateo-berri
8deb465346 fix(realtime): dial Azure's GA realtime upstream for GA clients
Azure realtime defaulted to the beta upstream whenever realtime_protocol
was not configured, so a GA client's session.update (session.type,
output_modalities, nested audio) was forwarded unchanged to
/openai/realtime and Azure rejected it with "Unknown parameter:
'session.type'" on gpt-realtime and gpt-realtime-1.5. The unset default
now follows the client the way the OpenAI handler already does: a client
that sends OpenAI-Beta: realtime=v1 keeps the beta upstream, any other
client gets /openai/v1/realtime. An explicit realtime_protocol in
litellm_params or LITELLM_AZURE_REALTIME_PROTOCOL still wins.
2026-09-10 13:18:52 -07:00
yujonglee
b0d66a15b8
feat(ocr): add core foundation and Mistral adapter (#40530)
* feat(ocr): add core foundation and transport primitives

* fix(ocr): decline missing Mistral credentials

* fix(rust): compile trace parity on Rust 1.98

* refactor(ocr): define native response capability

* refactor(auth): generalize missing API key errors

* refactor(core): keep URL helpers usage scoped

* refactor(ocr): support native responses across adapters

* refactor(ocr): preserve unmapped provider params

* refactor(ocr): distinguish request preparation from payload transforms

* refactor(ocr): trace payload transformation at codec boundary
2026-09-10 13:18:41 -07:00
mateo-berri
cc401c4041 fix(ui): skip organization fetches when the session is not premium 2026-09-10 13:16:01 -07:00
moe-berri
c2176c786e fix(router): redact cookies and preserve legacy classifier logs 2026-09-10 13:12:04 -07:00
Joshua Valluru
55ab5ee53d fix(mcp): preserve identity checks across cached OAuth credentials 2026-09-10 13:08:46 -07:00
ryan-crabbe-berri
81f1778f53 feat(proxy): gate organization endpoints on an enterprise license
The Admin UI and the docs already present Organizations as an enterprise
feature, but every /organization route served unlicensed proxies. A router
level dependency now enforces the license on all of them, and it resolves
the auth dependency first so a bad key still gets 401 rather than 403.

Claude-Session: https://claude.ai/code/session_01Se8ERtsqMQ3eVzWiLVMyNS
2026-09-10 13:05:58 -07:00
moe-berri
208d554c00 fix(router): validate encrypted classifiers after deployment selection 2026-09-10 13:04:36 -07:00
moe-berri
488bf6f596 Merge remote-tracking branch 'origin/litellm_internal_staging' into moe/lit-7493-zocdocauto-router-encrypted-codex-sub-agent-task-is
# Conflicts:
#	tests/test_litellm/router_strategy/test_complexity_router.py
2026-09-10 12:55:53 -07:00
moe-berri
83b1e9e0d7 test(router): keep classifier E2E verification local 2026-09-10 12:46:50 -07:00
moe-berri
692a311efb
Merge pull request #40599 from BerriAI/litellm_lit7492_codex_envelopes
fix(router): strip Codex harness envelopes before classification
2026-09-10 12:45:35 -07:00
moe-berri
33954d9cfe refactor(router): validate classifier snapshots without recursion 2026-09-10 12:35:59 -07:00
moe-berri
f4ebcef0a1 fix(router): classify encrypted delegated tasks with native Responses 2026-09-10 12:32:09 -07:00
mateo
ace6de88a4 fix(tests): load local model costs for DeepSeek flash regression
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 19:31:11 +00:00
moe-berri
181b3fd94a fix(router): classify new asks before reminder-only tails 2026-09-10 12:24:42 -07:00
moe-berri
41a147f2a0 fix(router): normalize Azure classifier audit payloads 2026-09-10 12:22:17 -07:00
mateo
896f0dff2e fix(model_prices): add provider-prefixed deepseek/deepseek-flash
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 19:14:28 +00:00
mateo
60c7dd8348 fix(model_prices): add deepseek-flash and gpt-live-1, bill DeepSeek legacy flash aliases at Flash rates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 19:13:49 +00:00
moe-berri
1697684b68 fix(router): scope Codex envelope defaults to Codex clients 2026-09-10 12:01:04 -07:00
moe-berri
153c962e15 feat(router): log classifier input and masked source request 2026-09-10 12:00:18 -07:00
Wolfram Ravenwolf
13ddd1ec64 fix(wandb): gate reasoning effort on model capabilities 2026-09-10 20:47:13 +02:00
moe-berri
5d6fec94b7 fix(router): strip Codex harness envelopes before classification 2026-09-10 11:33:21 -07:00
devin-ai-integration[bot]
6c69dd0f72
perf(proxy): reuse cached model group and deployment info in budget reservation (#40593)
Profiling the sidecar-enabled gateway at 700 rps showed ~2.4% of all samples
in get_model_group_info called per request from budget reservation, plus
get_deployment_model_info for tiered pricing tables. Both are read-only lookups
over the model list, so serve them from the Router's lru caches and clear the
deployment cache alongside the group cache when the model list changes.

The deployment-info cache is a per-router lru_cache built in __init__ rather
than a class-level decorated method, so it does not pin Router instances in a
process-wide cache and is dropped with the router.

A price data reload replaces litellm.model_cost without touching model_list, so
the reload replay also clears both caches; otherwise reservation would keep
pricing against the old catalog until an unrelated model-list change.

Co-authored-by: yassin <yassin@berri.ai>
2026-09-10 11:22:46 -07:00
devin-ai-integration[bot]
f84034f500
feat(mock): report admission-time input token count in mock_response usage (#40590)
* feat(mock): report admission-time input token count in mock_response usage

Mock completions always reported prompt_tokens=10, so spend tracking, TPM metrics, budgets and the tokens-per-minute autoscaling signal saw 10 tokens for a 100k-token request. Budget reservation now carries the admission-time input token count in the reservation record, and mock_completion reads it back so mock traffic exercises the same spend and TPM paths as real traffic without any extra tokenizer work.

* fix(mock): keep a zero admission input token count instead of falling back to 10

---------

Co-authored-by: yassin <yassin@berri.ai>
2026-09-10 11:22:26 -07:00
mrinal
10a5761bb9 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_health_check_db_storm 2026-09-10 18:14:25 +00:00
devin-ai-integration[bot]
2bf065f97d
fix(terraform): restore d.Partial(true) on a rejected /key/update (#40527)
The squash of #40512 onto a base that already carried #40514 left the Partial call on the metadata pre-read error path only, so a rejected /key/update again persisted the planned values into state and TestResourceKeyUpdateFailureKeepsPriorState fails on the default branch.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 11:09:18 -07:00
Mateo Wang
e907e5ee9b
Merge pull request #39237 from BerriAI/litellm_fix_dashscope_rerank_endpoint
fix(dashscope): remap chat-shaped api_base to the live rerank route
2026-09-10 10:31:48 -07:00
ryan-crabbe-berri
11ec6f7a36
Merge pull request #38509 from yinonkahta-p5/litellm_pointfive_logger
feat(pointfive): add the pointfive logging integration
2026-09-10 10:26:51 -07:00