Commit graph

48347 commits

Author SHA1 Message Date
mateo-berri
9da3e63ae9 fix(redis): log an open circuit breaker once instead of a traceback per request and count sync timeouts as timeouts
While the Redis circuit breaker is open every guarded call was refused with a bare
Exception that each swallowing catch site logged as an ERROR traceback, so a
sub-second latency blip turned into thousands of tracebacks per minute and pinned
every replica's CPU. The sync guard also recorded socket timeouts as hard
connectivity failures, so with least-busy routing the breaker opened on the first
slow replies and the timeout-only min-duration guard never applied.

Refusals now raise RedisCircuitBreakerOpenError, and the catch sites route it
through log_redis_failure, which logs a refusal at DEBUG and everything else at
the caller's level. The sync guard passes is_timeout like the async one.
2026-09-10 13:50:34 -07:00
moe-berri
d266b76a8b fix(router): honor Codex reminders and map classifier failures 2026-09-10 13:48:21 -07:00
mateo-berri
4de441c181 test(ui): render TeamSSOSettings under a premium session so the organization dropdown tests fetch 2026-09-10 13:41:37 -07:00
devin-ai-integration[bot]
a9cec50960
feat(infra): scale gateway on per-pod RPS and TPS in Helm and Terraform (#40479)
* feat(infra): scale gateway on per-pod RPM and TPM in Helm and Terraform

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(helm): require the metrics server before rendering the gateway ServiceMonitor

The http port serves /metrics/ behind virtual-key auth, so a ServiceMonitor
pointed at it only collects 401s and the RPM/TPM HPA metrics never appear

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(infra): express gateway HPA, KEDA and ECS workload targets per second

Rename the per-pod request and token targets in both Helm charts and the
AWS module from per minute to per second, and shorten the recommended
Prometheus rate window to [1m] with no * 60 so the adapter and KEDA
signals are what the HPA compares against. ECS keeps CloudWatch's
60-second aggregation: the ALB target is 60x the per-second variable and
the token metric math divides the period Sum by 60 before dividing by
the running task count.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 13:36:23 -07:00
moe-berri
126dc37de4 fix(logging): reconstruct classifier audit redaction payloads 2026-09-10 13:33:36 -07:00
Joshua Valluru
93bd3d61b5 fix(mcp): preserve verified identity when warming OAuth cache 2026-09-10 13:20:49 -07:00
yujonglee
b0d66a15b8
feat(ocr): add core foundation and Mistral adapter (#40530)
* feat(ocr): add core foundation and transport primitives

* fix(ocr): decline missing Mistral credentials

* fix(rust): compile trace parity on Rust 1.98

* refactor(ocr): define native response capability

* refactor(auth): generalize missing API key errors

* refactor(core): keep URL helpers usage scoped

* refactor(ocr): support native responses across adapters

* refactor(ocr): preserve unmapped provider params

* refactor(ocr): distinguish request preparation from payload transforms

* refactor(ocr): trace payload transformation at codec boundary
2026-09-10 13:18:41 -07:00
mateo-berri
cc401c4041 fix(ui): skip organization fetches when the session is not premium 2026-09-10 13:16:01 -07:00
moe-berri
c2176c786e fix(router): redact cookies and preserve legacy classifier logs 2026-09-10 13:12:04 -07:00
Joshua Valluru
55ab5ee53d fix(mcp): preserve identity checks across cached OAuth credentials 2026-09-10 13:08:46 -07:00
ryan-crabbe-berri
81f1778f53 feat(proxy): gate organization endpoints on an enterprise license
The Admin UI and the docs already present Organizations as an enterprise
feature, but every /organization route served unlicensed proxies. A router
level dependency now enforces the license on all of them, and it resolves
the auth dependency first so a bad key still gets 401 rather than 403.

Claude-Session: https://claude.ai/code/session_01Se8ERtsqMQ3eVzWiLVMyNS
2026-09-10 13:05:58 -07:00
moe-berri
208d554c00 fix(router): validate encrypted classifiers after deployment selection 2026-09-10 13:04:36 -07:00
moe-berri
488bf6f596 Merge remote-tracking branch 'origin/litellm_internal_staging' into moe/lit-7493-zocdocauto-router-encrypted-codex-sub-agent-task-is
# Conflicts:
#	tests/test_litellm/router_strategy/test_complexity_router.py
2026-09-10 12:55:53 -07:00
moe-berri
83b1e9e0d7 test(router): keep classifier E2E verification local 2026-09-10 12:46:50 -07:00
moe-berri
692a311efb
Merge pull request #40599 from BerriAI/litellm_lit7492_codex_envelopes
fix(router): strip Codex harness envelopes before classification
2026-09-10 12:45:35 -07:00
moe-berri
33954d9cfe refactor(router): validate classifier snapshots without recursion 2026-09-10 12:35:59 -07:00
moe-berri
f4ebcef0a1 fix(router): classify encrypted delegated tasks with native Responses 2026-09-10 12:32:09 -07:00
moe-berri
181b3fd94a fix(router): classify new asks before reminder-only tails 2026-09-10 12:24:42 -07:00
moe-berri
41a147f2a0 fix(router): normalize Azure classifier audit payloads 2026-09-10 12:22:17 -07:00
moe-berri
1697684b68 fix(router): scope Codex envelope defaults to Codex clients 2026-09-10 12:01:04 -07:00
moe-berri
153c962e15 feat(router): log classifier input and masked source request 2026-09-10 12:00:18 -07:00
Wolfram Ravenwolf
13ddd1ec64 fix(wandb): gate reasoning effort on model capabilities 2026-09-10 20:47:13 +02:00
moe-berri
5d6fec94b7 fix(router): strip Codex harness envelopes before classification 2026-09-10 11:33:21 -07:00
devin-ai-integration[bot]
6c69dd0f72
perf(proxy): reuse cached model group and deployment info in budget reservation (#40593)
Profiling the sidecar-enabled gateway at 700 rps showed ~2.4% of all samples
in get_model_group_info called per request from budget reservation, plus
get_deployment_model_info for tiered pricing tables. Both are read-only lookups
over the model list, so serve them from the Router's lru caches and clear the
deployment cache alongside the group cache when the model list changes.

The deployment-info cache is a per-router lru_cache built in __init__ rather
than a class-level decorated method, so it does not pin Router instances in a
process-wide cache and is dropped with the router.

A price data reload replaces litellm.model_cost without touching model_list, so
the reload replay also clears both caches; otherwise reservation would keep
pricing against the old catalog until an unrelated model-list change.

Co-authored-by: yassin <yassin@berri.ai>
2026-09-10 11:22:46 -07:00
devin-ai-integration[bot]
f84034f500
feat(mock): report admission-time input token count in mock_response usage (#40590)
* feat(mock): report admission-time input token count in mock_response usage

Mock completions always reported prompt_tokens=10, so spend tracking, TPM metrics, budgets and the tokens-per-minute autoscaling signal saw 10 tokens for a 100k-token request. Budget reservation now carries the admission-time input token count in the reservation record, and mock_completion reads it back so mock traffic exercises the same spend and TPM paths as real traffic without any extra tokenizer work.

* fix(mock): keep a zero admission input token count instead of falling back to 10

---------

Co-authored-by: yassin <yassin@berri.ai>
2026-09-10 11:22:26 -07:00
mrinal
10a5761bb9 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_health_check_db_storm 2026-09-10 18:14:25 +00:00
devin-ai-integration[bot]
2bf065f97d
fix(terraform): restore d.Partial(true) on a rejected /key/update (#40527)
The squash of #40512 onto a base that already carried #40514 left the Partial call on the metadata pre-read error path only, so a rejected /key/update again persisted the planned values into state and TestResourceKeyUpdateFailureKeepsPriorState fails on the default branch.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 11:09:18 -07:00
Mateo Wang
e907e5ee9b
Merge pull request #39237 from BerriAI/litellm_fix_dashscope_rerank_endpoint
fix(dashscope): remap chat-shaped api_base to the live rerank route
2026-09-10 10:31:48 -07:00
ryan-crabbe-berri
11ec6f7a36
Merge pull request #38509 from yinonkahta-p5/litellm_pointfive_logger
feat(pointfive): add the pointfive logging integration
2026-09-10 10:26:51 -07:00
Mateo Wang
218b3280d1
Merge pull request #39272 from BerriAI/litellm_fix_e2e_lint_pathspec
ci(lint): gate top-level tests/e2e and litellm files in the diff-scoped lint steps
2026-09-10 10:13:56 -07:00
ryan-crabbe-berri
c79c73f859
Merge pull request #40446 from BerriAI/litellm_bedrock_aws_session_tags
feat(bedrock): thread aws_session_tags into STS AssumeRole
2026-09-10 09:53:29 -07:00
ryan-crabbe-berri
3ddb920028 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_aws_session_tags 2026-09-10 09:42:22 -07:00
Mateo Wang
c1a83fc005
Merge pull request #40581 from BerriAI/litellm_registry_audit_2026_09_10
fix(model_prices): registry audit 2026-09-10, absorb open pricing PRs
2026-09-10 07:43:29 -07:00
eugene-yao-zocdoc
f0e3e031c3 test(redis): pin ElastiCache IAM signing and TLS coercion invariants
Strengthens the serverless test to assert ResourceType is signed rather
than merely present, ties _uses_tls to the redis-py kwarg coercion so the
two cannot drift, locks the stripped-kwarg name tuple to the test's
expectations, and adds "off" and "True" sentinel flag values.

Renames the provider builder's parameter to redis_settings.
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
c4fc20bcf9 fix(redis): accept every truthy flag and sign serverless ElastiCache caches
The ElastiCache IAM gate read `aws_iam_auth` and `ssl` with a helper that
only accepted the literal string "true", while the kwarg coercion that runs
later accepts "true", "1" and "yes". Type coercion happens after the gate, so
`REDIS_AWS_IAM_AUTH=1` silently skipped IAM auth and `REDIS_SSL=1` made the
"requires TLS" check fail closed on a connection that was in fact TLS. Both
helpers now share `_str_to_bool`.

AWS signs serverless cache tokens with an extra `ResourceType=ServerlessCache`
query parameter, so tokens minted for a serverless cache were rejected. Adds an
`aws_iam_serverless` setting (`REDIS_AWS_IAM_SERVERLESS`) that puts the
parameter into the signed URL, and lowercases the cache name because
ElastiCache lowercases it at creation time.
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
bf73c49d73 refactor(redis): drop unreachable frozen credentials guard 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
2b0711e4f5 style(redis): restore the Azure AD marker comment
The comment documents that the raw Azure client id, tenant id and secret
are deliberately kept off the connect function, so this branch should
never have dropped it
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
cd4073895f refactor(redis): type botocore credentials without Any
Deferred annotation evaluation keeps the type-checking-only botocore
import off the runtime path, so the alias only reintroduced typing.Any,
which the strict ruff budget now bans
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
aca81c263c fix(redis): validate signed ElastiCache IAM URL 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
f1cd2b03ca fix(redis): type signed ElastiCache IAM URL 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
0217a137fe fix(redis): require TLS for ElastiCache IAM auth 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
a7ad6023d6 ci: retry CodSpeed result upload
Generated with AI

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
9763bd80ae style(proxy): format coordination Redis validation
Generated with AI

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
921c263876 fix(proxy): validate coordination Redis mappings
Generated with AI

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
bb043fe29a test(redis): cover ElastiCache IAM failures
Generated with AI

Co-Authored-By: Claude Code
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
f66891d80f fix(redis): type ElastiCache IAM configuration
Generated with AI

Co-Authored-By: Claude Code
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
604b9edc53 test(redis): cover AWS IAM provider install on async cluster nodes
No existing test combined aws_iam_auth with startup_nodes, the shape
ElastiCache/Valkey Serverless deployments actually use.
2026-09-10 10:13:02 -04:00
eugene-yao-zocdoc
1273e46b8b feat(redis): add ElastiCache IAM authentication 2026-09-10 10:13:02 -04:00
mateo-berri
996ca7c5b2 test(cost): assert jina rerank spend at the registry rate 2026-09-10 07:04:19 -07:00
mateo
134d1f3899 fix(model_prices): registry audit 2026-09-10, absorb open pricing PRs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 13:17:14 +00:00