Commit graph

16834 commits

Author SHA1 Message Date
devin-ai-integration[bot]
0e35c8fee9
fix(proxy): recreate the Prisma client when the writer session turns read-only (#40610)
The writer health probe only ran SELECT 1, which a read-only Postgres
session answers fine, so a pooled connection left pointing at a demoted
primary kept failing every write with SQLSTATE 25006 until the pod was
restarted. Probe transaction_read_only instead, treat a 25006 on the
request path as a signal to recreate the client, and back off
exponentially while the database as a whole stays read-only so a replica
or an in-progress failover does not get its engine killed every cycle.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 13:53:42 -07:00
mateo-berri
d3a0b0d45b fix(voyage): let caller params win for contextual auto-chunking and drop duplicate cost map entries
A flat list[str] sent to voyage-context-4 is treated as independent inputs and forwarded flat with
enable_auto_chunking=True, chunk_size=32000, and input_type=document unless the caller already set
input_type=query. Caller-supplied params now override the defaults instead of being clobbered.

The voyage-4 family and voyage-context-4 cost map entries already exist on litellm_internal_staging,
and voyage-4-nano is not served by the Voyage API, so those additions and their pricing test are dropped.
2026-09-10 13:51:45 -07:00
mateo-berri
9da3e63ae9 fix(redis): log an open circuit breaker once instead of a traceback per request and count sync timeouts as timeouts
While the Redis circuit breaker is open every guarded call was refused with a bare
Exception that each swallowing catch site logged as an ERROR traceback, so a
sub-second latency blip turned into thousands of tracebacks per minute and pinned
every replica's CPU. The sync guard also recorded socket timeouts as hard
connectivity failures, so with least-busy routing the breaker opened on the first
slow replies and the timeout-only min-duration guard never applied.

Refusals now raise RedisCircuitBreakerOpenError, and the catch sites route it
through log_redis_failure, which logs a refusal at DEBUG and everything else at
the caller's level. The sync guard passes is_timeout like the async one.
2026-09-10 13:50:34 -07:00
moe-berri
d266b76a8b fix(router): honor Codex reminders and map classifier failures 2026-09-10 13:48:21 -07:00
moe-berri
126dc37de4 fix(logging): reconstruct classifier audit redaction payloads 2026-09-10 13:33:36 -07:00
Joshua Valluru
93bd3d61b5 fix(mcp): preserve verified identity when warming OAuth cache 2026-09-10 13:20:49 -07:00
mateo-berri
8deb465346 fix(realtime): dial Azure's GA realtime upstream for GA clients
Azure realtime defaulted to the beta upstream whenever realtime_protocol
was not configured, so a GA client's session.update (session.type,
output_modalities, nested audio) was forwarded unchanged to
/openai/realtime and Azure rejected it with "Unknown parameter:
'session.type'" on gpt-realtime and gpt-realtime-1.5. The unset default
now follows the client the way the OpenAI handler already does: a client
that sends OpenAI-Beta: realtime=v1 keeps the beta upstream, any other
client gets /openai/v1/realtime. An explicit realtime_protocol in
litellm_params or LITELLM_AZURE_REALTIME_PROTOCOL still wins.
2026-09-10 13:18:52 -07:00
moe-berri
c2176c786e fix(router): redact cookies and preserve legacy classifier logs 2026-09-10 13:12:04 -07:00
Joshua Valluru
55ab5ee53d fix(mcp): preserve identity checks across cached OAuth credentials 2026-09-10 13:08:46 -07:00
ryan-crabbe-berri
81f1778f53 feat(proxy): gate organization endpoints on an enterprise license
The Admin UI and the docs already present Organizations as an enterprise
feature, but every /organization route served unlicensed proxies. A router
level dependency now enforces the license on all of them, and it resolves
the auth dependency first so a bad key still gets 401 rather than 403.

Claude-Session: https://claude.ai/code/session_01Se8ERtsqMQ3eVzWiLVMyNS
2026-09-10 13:05:58 -07:00
moe-berri
208d554c00 fix(router): validate encrypted classifiers after deployment selection 2026-09-10 13:04:36 -07:00
moe-berri
488bf6f596 Merge remote-tracking branch 'origin/litellm_internal_staging' into moe/lit-7493-zocdocauto-router-encrypted-codex-sub-agent-task-is
# Conflicts:
#	tests/test_litellm/router_strategy/test_complexity_router.py
2026-09-10 12:55:53 -07:00
moe-berri
83b1e9e0d7 test(router): keep classifier E2E verification local 2026-09-10 12:46:50 -07:00
moe-berri
f4ebcef0a1 fix(router): classify encrypted delegated tasks with native Responses 2026-09-10 12:32:09 -07:00
mateo
ace6de88a4 fix(tests): load local model costs for DeepSeek flash regression
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 19:31:11 +00:00
moe-berri
181b3fd94a fix(router): classify new asks before reminder-only tails 2026-09-10 12:24:42 -07:00
moe-berri
41a147f2a0 fix(router): normalize Azure classifier audit payloads 2026-09-10 12:22:17 -07:00
mateo
60c7dd8348 fix(model_prices): add deepseek-flash and gpt-live-1, bill DeepSeek legacy flash aliases at Flash rates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 19:13:49 +00:00
moe-berri
1697684b68 fix(router): scope Codex envelope defaults to Codex clients 2026-09-10 12:01:04 -07:00
moe-berri
153c962e15 feat(router): log classifier input and masked source request 2026-09-10 12:00:18 -07:00
Wolfram Ravenwolf
13ddd1ec64 fix(wandb): gate reasoning effort on model capabilities 2026-09-10 20:47:13 +02:00
moe-berri
5d6fec94b7 fix(router): strip Codex harness envelopes before classification 2026-09-10 11:33:21 -07:00
devin-ai-integration[bot]
6c69dd0f72
perf(proxy): reuse cached model group and deployment info in budget reservation (#40593)
Profiling the sidecar-enabled gateway at 700 rps showed ~2.4% of all samples
in get_model_group_info called per request from budget reservation, plus
get_deployment_model_info for tiered pricing tables. Both are read-only lookups
over the model list, so serve them from the Router's lru caches and clear the
deployment cache alongside the group cache when the model list changes.

The deployment-info cache is a per-router lru_cache built in __init__ rather
than a class-level decorated method, so it does not pin Router instances in a
process-wide cache and is dropped with the router.

A price data reload replaces litellm.model_cost without touching model_list, so
the reload replay also clears both caches; otherwise reservation would keep
pricing against the old catalog until an unrelated model-list change.

Co-authored-by: yassin <yassin@berri.ai>
2026-09-10 11:22:46 -07:00
devin-ai-integration[bot]
f84034f500
feat(mock): report admission-time input token count in mock_response usage (#40590)
* feat(mock): report admission-time input token count in mock_response usage

Mock completions always reported prompt_tokens=10, so spend tracking, TPM metrics, budgets and the tokens-per-minute autoscaling signal saw 10 tokens for a 100k-token request. Budget reservation now carries the admission-time input token count in the reservation record, and mock_completion reads it back so mock traffic exercises the same spend and TPM paths as real traffic without any extra tokenizer work.

* fix(mock): keep a zero admission input token count instead of falling back to 10

---------

Co-authored-by: yassin <yassin@berri.ai>
2026-09-10 11:22:26 -07:00
mrinal
10a5761bb9 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_health_check_db_storm 2026-09-10 18:14:25 +00:00
Mateo Wang
e907e5ee9b
Merge pull request #39237 from BerriAI/litellm_fix_dashscope_rerank_endpoint
fix(dashscope): remap chat-shaped api_base to the live rerank route
2026-09-10 10:31:48 -07:00
ryan-crabbe-berri
11ec6f7a36
Merge pull request #38509 from yinonkahta-p5/litellm_pointfive_logger
feat(pointfive): add the pointfive logging integration
2026-09-10 10:26:51 -07:00
Mateo Wang
218b3280d1
Merge pull request #39272 from BerriAI/litellm_fix_e2e_lint_pathspec
ci(lint): gate top-level tests/e2e and litellm files in the diff-scoped lint steps
2026-09-10 10:13:56 -07:00
ryan-crabbe-berri
3ddb920028 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_aws_session_tags 2026-09-10 09:42:22 -07:00
mateo-berri
f6685b7858 fix(cost_calculator): keep base_model pricing off the regional row
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
A deployment with base_model set was priced from the region's own row once
completion_cost forwarded the response region into cost_per_token, which now
strips the provider prefix and finds bedrock/<region>/<base_model>. Explicit
pricing (base_model or custom pricing) suppresses the region for cost_per_token
the same way _select_model_name_for_cost_calc already does
2026-09-10 08:12:16 -07:00
eugene-yao-zocdoc
f0e3e031c3 test(redis): pin ElastiCache IAM signing and TLS coercion invariants
Strengthens the serverless test to assert ResourceType is signed rather
than merely present, ties _uses_tls to the redis-py kwarg coercion so the
two cannot drift, locks the stripped-kwarg name tuple to the test's
expectations, and adds "off" and "True" sentinel flag values.

Renames the provider builder's parameter to redis_settings.
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
c4fc20bcf9 fix(redis): accept every truthy flag and sign serverless ElastiCache caches
The ElastiCache IAM gate read `aws_iam_auth` and `ssl` with a helper that
only accepted the literal string "true", while the kwarg coercion that runs
later accepts "true", "1" and "yes". Type coercion happens after the gate, so
`REDIS_AWS_IAM_AUTH=1` silently skipped IAM auth and `REDIS_SSL=1` made the
"requires TLS" check fail closed on a connection that was in fact TLS. Both
helpers now share `_str_to_bool`.

AWS signs serverless cache tokens with an extra `ResourceType=ServerlessCache`
query parameter, so tokens minted for a serverless cache were rejected. Adds an
`aws_iam_serverless` setting (`REDIS_AWS_IAM_SERVERLESS`) that puts the
parameter into the signed URL, and lowercases the cache name because
ElastiCache lowercases it at creation time.
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
bf73c49d73 refactor(redis): drop unreachable frozen credentials guard 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
0217a137fe fix(redis): require TLS for ElastiCache IAM auth 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
bb043fe29a test(redis): cover ElastiCache IAM failures
Generated with AI

Co-Authored-By: Claude Code
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
f66891d80f fix(redis): type ElastiCache IAM configuration
Generated with AI

Co-Authored-By: Claude Code
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
604b9edc53 test(redis): cover AWS IAM provider install on async cluster nodes
No existing test combined aws_iam_auth with startup_nodes, the shape
ElastiCache/Valkey Serverless deployments actually use.
2026-09-10 10:13:02 -04:00
eugene-yao-zocdoc
1273e46b8b feat(redis): add ElastiCache IAM authentication 2026-09-10 10:13:02 -04:00
mateo-berri
0308b05c7a merge: bring litellm_internal_staging into litellm_bedrock_mantle_govcloud_cost_row 2026-09-10 07:08:58 -07:00
mateo-berri
996ca7c5b2 test(cost): assert jina rerank spend at the registry rate 2026-09-10 07:04:19 -07:00
mateo
134d1f3899 fix(model_prices): registry audit 2026-09-10, absorb open pricing PRs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 13:17:14 +00:00
dclark
a39c358785 fix(proxy): preserve local spend adjustments during reseeding 2026-09-10 13:44:08 +01:00
michelligabriele
93b15ed428
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_v1_models_alias_resolution 2026-09-10 13:26:18 +02:00
Yinon Kahta
123c01c439 feat(pointfive): answer the ui health check with a liveness ping 2026-09-10 14:02:35 +03:00
Yinon Kahta
29b02e9215 feat(pointfive): add pointfive to the dashboard logging integrations 2026-09-10 14:02:35 +03:00
Yinon Kahta
cb150846c2 feat(pointfive): list pointfive in the proxy callback registry 2026-09-10 14:02:35 +03:00
Yinon Kahta
6b27a16e67 feat(pointfive): add batching logger callback 2026-09-10 14:02:35 +03:00
Yinon Kahta
85d10c8558 feat(pointfive): add presigned upload client 2026-09-10 14:02:35 +03:00
Yinon Kahta
6d7b3a82ff feat(pointfive): add gzipped ndjson batch encoding 2026-09-10 14:02:35 +03:00
Yinon Kahta
0b898b47ca fix(http_handler): let put opt out of following redirects
get and post already take follow_redirects. put built the request and sent it
with the client default, so a caller uploading to a URL it did not choose had
no way to refuse a redirect. Same plumbing as the other two methods.
2026-09-10 14:02:35 +03:00