Commit graph

18968 commits

Author SHA1 Message Date
moe-berri
488bf6f596 Merge remote-tracking branch 'origin/litellm_internal_staging' into moe/lit-7493-zocdocauto-router-encrypted-codex-sub-agent-task-is
# Conflicts:
#	tests/test_litellm/router_strategy/test_complexity_router.py
2026-09-10 12:55:53 -07:00
moe-berri
83b1e9e0d7 test(router): keep classifier E2E verification local 2026-09-10 12:46:50 -07:00
moe-berri
f4ebcef0a1 fix(router): classify encrypted delegated tasks with native Responses 2026-09-10 12:32:09 -07:00
mateo
ace6de88a4 fix(tests): load local model costs for DeepSeek flash regression
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 19:31:11 +00:00
moe-berri
181b3fd94a fix(router): classify new asks before reminder-only tails 2026-09-10 12:24:42 -07:00
moe-berri
41a147f2a0 fix(router): normalize Azure classifier audit payloads 2026-09-10 12:22:17 -07:00
mateo
60c7dd8348 fix(model_prices): add deepseek-flash and gpt-live-1, bill DeepSeek legacy flash aliases at Flash rates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 19:13:49 +00:00
moe-berri
1697684b68 fix(router): scope Codex envelope defaults to Codex clients 2026-09-10 12:01:04 -07:00
moe-berri
153c962e15 feat(router): log classifier input and masked source request 2026-09-10 12:00:18 -07:00
Wolfram Ravenwolf
13ddd1ec64 fix(wandb): gate reasoning effort on model capabilities 2026-09-10 20:47:13 +02:00
moe-berri
5d6fec94b7 fix(router): strip Codex harness envelopes before classification 2026-09-10 11:33:21 -07:00
devin-ai-integration[bot]
6c69dd0f72
perf(proxy): reuse cached model group and deployment info in budget reservation (#40593)
Profiling the sidecar-enabled gateway at 700 rps showed ~2.4% of all samples
in get_model_group_info called per request from budget reservation, plus
get_deployment_model_info for tiered pricing tables. Both are read-only lookups
over the model list, so serve them from the Router's lru caches and clear the
deployment cache alongside the group cache when the model list changes.

The deployment-info cache is a per-router lru_cache built in __init__ rather
than a class-level decorated method, so it does not pin Router instances in a
process-wide cache and is dropped with the router.

A price data reload replaces litellm.model_cost without touching model_list, so
the reload replay also clears both caches; otherwise reservation would keep
pricing against the old catalog until an unrelated model-list change.

Co-authored-by: yassin <yassin@berri.ai>
2026-09-10 11:22:46 -07:00
devin-ai-integration[bot]
f84034f500
feat(mock): report admission-time input token count in mock_response usage (#40590)
* feat(mock): report admission-time input token count in mock_response usage

Mock completions always reported prompt_tokens=10, so spend tracking, TPM metrics, budgets and the tokens-per-minute autoscaling signal saw 10 tokens for a 100k-token request. Budget reservation now carries the admission-time input token count in the reservation record, and mock_completion reads it back so mock traffic exercises the same spend and TPM paths as real traffic without any extra tokenizer work.

* fix(mock): keep a zero admission input token count instead of falling back to 10

---------

Co-authored-by: yassin <yassin@berri.ai>
2026-09-10 11:22:26 -07:00
mrinal
10a5761bb9 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_health_check_db_storm 2026-09-10 18:14:25 +00:00
mubashir1osmani
1d911c766f Merge remote-tracking branch 'berri/litellm_internal_staging' into litellm_mistral_ocr_batches 2026-09-10 13:51:18 -04:00
Devin AI
431dcce6a7 test(proxy): give pre-call mocks real router_settings and fallbacks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 17:46:44 +00:00
mubashir1osmani
91e7df3d8f chore: regenerate cost-map schema and UI API types, drop test banner comments 2026-09-10 13:35:35 -04:00
Mateo Wang
e907e5ee9b
Merge pull request #39237 from BerriAI/litellm_fix_dashscope_rerank_endpoint
fix(dashscope): remap chat-shaped api_base to the live rerank route
2026-09-10 10:31:48 -07:00
Devin AI
c064e576ee fix(proxy): tolerate missing router_settings and non-list fallbacks in fallback resolution
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 17:29:20 +00:00
ryan-crabbe-berri
11ec6f7a36
Merge pull request #38509 from yinonkahta-p5/litellm_pointfive_logger
feat(pointfive): add the pointfive logging integration
2026-09-10 10:26:51 -07:00
Devin AI
dff08dcb55 fix(proxy): retry rate-limit fallbacks from a pristine request snapshot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 17:16:45 +00:00
Mateo Wang
218b3280d1
Merge pull request #39272 from BerriAI/litellm_fix_e2e_lint_pathspec
ci(lint): gate top-level tests/e2e and litellm files in the diff-scoped lint steps
2026-09-10 10:13:56 -07:00
ryan-crabbe-berri
3ddb920028 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_aws_session_tags 2026-09-10 09:42:22 -07:00
ryan-crabbe-berri
b40fc0ac22 refactor(bedrock): resolve AWS credentials from one typed auth struct
Every Bedrock and SageMaker call site hand-copied the same nine aws_* kwargs
into BaseAWSLLM.get_credentials, so each new auth param has to be threaded
into a dozen places and any site that misses one silently assumes the role
with the wrong parameters.

Introduce AwsAuthParams, a frozen pydantic model whose fields are exactly the
credential-shaped params get_credentials accepts, plus resolve_credentials on
BaseAWSLLM and pop_aws_auth_params for the call sites that must strip the keys
out of optional_params. Deriving AWS_AUTH_PARAM_KEYS from the model's fields
means the mirror list in common_utils can no longer drift from the struct.

Behavior is unchanged: the same values reach STS from the same call sites.
Dropping any one field from the resolver fails one of the new tests.

Claude-Session: https://claude.ai/code/session_01E6zsK1DBcXfbetkgX86fw2
2026-09-10 09:20:13 -07:00
mateo-berri
f6685b7858 fix(cost_calculator): keep base_model pricing off the regional row
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
A deployment with base_model set was priced from the region's own row once
completion_cost forwarded the response region into cost_per_token, which now
strips the provider prefix and finds bedrock/<region>/<base_model>. Explicit
pricing (base_model or custom pricing) suppresses the region for cost_per_token
the same way _select_model_name_for_cost_calc already does
2026-09-10 08:12:16 -07:00
eugene-yao-zocdoc
f0e3e031c3 test(redis): pin ElastiCache IAM signing and TLS coercion invariants
Strengthens the serverless test to assert ResourceType is signed rather
than merely present, ties _uses_tls to the redis-py kwarg coercion so the
two cannot drift, locks the stripped-kwarg name tuple to the test's
expectations, and adds "off" and "True" sentinel flag values.

Renames the provider builder's parameter to redis_settings.
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
c4fc20bcf9 fix(redis): accept every truthy flag and sign serverless ElastiCache caches
The ElastiCache IAM gate read `aws_iam_auth` and `ssl` with a helper that
only accepted the literal string "true", while the kwarg coercion that runs
later accepts "true", "1" and "yes". Type coercion happens after the gate, so
`REDIS_AWS_IAM_AUTH=1` silently skipped IAM auth and `REDIS_SSL=1` made the
"requires TLS" check fail closed on a connection that was in fact TLS. Both
helpers now share `_str_to_bool`.

AWS signs serverless cache tokens with an extra `ResourceType=ServerlessCache`
query parameter, so tokens minted for a serverless cache were rejected. Adds an
`aws_iam_serverless` setting (`REDIS_AWS_IAM_SERVERLESS`) that puts the
parameter into the signed URL, and lowercases the cache name because
ElastiCache lowercases it at creation time.
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
bf73c49d73 refactor(redis): drop unreachable frozen credentials guard 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
0217a137fe fix(redis): require TLS for ElastiCache IAM auth 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
bb043fe29a test(redis): cover ElastiCache IAM failures
Generated with AI

Co-Authored-By: Claude Code
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
f66891d80f fix(redis): type ElastiCache IAM configuration
Generated with AI

Co-Authored-By: Claude Code
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
604b9edc53 test(redis): cover AWS IAM provider install on async cluster nodes
No existing test combined aws_iam_auth with startup_nodes, the shape
ElastiCache/Valkey Serverless deployments actually use.
2026-09-10 10:13:02 -04:00
eugene-yao-zocdoc
1273e46b8b feat(redis): add ElastiCache IAM authentication 2026-09-10 10:13:02 -04:00
mateo-berri
0308b05c7a merge: bring litellm_internal_staging into litellm_bedrock_mantle_govcloud_cost_row 2026-09-10 07:08:58 -07:00
mateo-berri
996ca7c5b2 test(cost): assert jina rerank spend at the registry rate 2026-09-10 07:04:19 -07:00
mateo
134d1f3899 fix(model_prices): registry audit 2026-09-10, absorb open pricing PRs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 13:17:14 +00:00
dclark
a39c358785 fix(proxy): preserve local spend adjustments during reseeding 2026-09-10 13:44:08 +01:00
michelligabriele
93b15ed428
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_v1_models_alias_resolution 2026-09-10 13:26:18 +02:00
Yinon Kahta
123c01c439 feat(pointfive): answer the ui health check with a liveness ping 2026-09-10 14:02:35 +03:00
Yinon Kahta
29b02e9215 feat(pointfive): add pointfive to the dashboard logging integrations 2026-09-10 14:02:35 +03:00
Yinon Kahta
cb150846c2 feat(pointfive): list pointfive in the proxy callback registry 2026-09-10 14:02:35 +03:00
Yinon Kahta
6b27a16e67 feat(pointfive): add batching logger callback 2026-09-10 14:02:35 +03:00
Yinon Kahta
85d10c8558 feat(pointfive): add presigned upload client 2026-09-10 14:02:35 +03:00
Yinon Kahta
6d7b3a82ff feat(pointfive): add gzipped ndjson batch encoding 2026-09-10 14:02:35 +03:00
Yinon Kahta
0b898b47ca fix(http_handler): let put opt out of following redirects
get and post already take follow_redirects. put built the request and sent it
with the client default, so a caller uploading to a URL it did not choose had
no way to refuse a redirect. Same plumbing as the other two methods.
2026-09-10 14:02:35 +03:00
dclark
8b965880ef fix(proxy): prevent spend counter double counting 2026-09-10 11:29:37 +01:00
mateo
c956c24ade Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_deflake_20260905
Some checks are pending
Terraform Provider / gofmt, vet, build, test (push) Waiting to run
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Waiting to run
2026-09-10 09:19:52 +00:00
yucheng
239c126f29 refactor(guardrails): resolve the Agent 365 conversation id from the typed LiteLLM logging object
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
Terraform Modules / fmt, validate, test (aws) (push) Waiting to run
Terraform Modules / fmt, validate, test (gcp) (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Waiting to run
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Waiting to run
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 08:25:48 +00:00
yucheng
f4033ab84a fix(guardrails): keep Agent 365 missing-secret startup error readable after log redaction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 08:08:33 +00:00
yucheng
fcaf2d7d98 fix(otel): size the message ceiling so every vocabulary fits beside it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 08:00:51 +00:00