Commit graph

52158 commits

Author SHA1 Message Date
ryan-crabbe-berri
b40fc0ac22 refactor(bedrock): resolve AWS credentials from one typed auth struct
Every Bedrock and SageMaker call site hand-copied the same nine aws_* kwargs
into BaseAWSLLM.get_credentials, so each new auth param has to be threaded
into a dozen places and any site that misses one silently assumes the role
with the wrong parameters.

Introduce AwsAuthParams, a frozen pydantic model whose fields are exactly the
credential-shaped params get_credentials accepts, plus resolve_credentials on
BaseAWSLLM and pop_aws_auth_params for the call sites that must strip the keys
out of optional_params. Deriving AWS_AUTH_PARAM_KEYS from the model's fields
means the mirror list in common_utils can no longer drift from the struct.

Behavior is unchanged: the same values reach STS from the same call sites.
Dropping any one field from the resolver fails one of the new tests.

Claude-Session: https://claude.ai/code/session_01E6zsK1DBcXfbetkgX86fw2
2026-09-10 09:20:13 -07:00
mateo-berri
f6685b7858 fix(cost_calculator): keep base_model pricing off the regional row
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
A deployment with base_model set was priced from the region's own row once
completion_cost forwarded the response region into cost_per_token, which now
strips the provider prefix and finds bedrock/<region>/<base_model>. Explicit
pricing (base_model or custom pricing) suppresses the region for cost_per_token
the same way _select_model_name_for_cost_calc already does
2026-09-10 08:12:16 -07:00
Mateo Wang
c1a83fc005
Merge pull request #40581 from BerriAI/litellm_registry_audit_2026_09_10
fix(model_prices): registry audit 2026-09-10, absorb open pricing PRs
2026-09-10 07:43:29 -07:00
eugene-yao-zocdoc
f0e3e031c3 test(redis): pin ElastiCache IAM signing and TLS coercion invariants
Strengthens the serverless test to assert ResourceType is signed rather
than merely present, ties _uses_tls to the redis-py kwarg coercion so the
two cannot drift, locks the stripped-kwarg name tuple to the test's
expectations, and adds "off" and "True" sentinel flag values.

Renames the provider builder's parameter to redis_settings.
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
c4fc20bcf9 fix(redis): accept every truthy flag and sign serverless ElastiCache caches
The ElastiCache IAM gate read `aws_iam_auth` and `ssl` with a helper that
only accepted the literal string "true", while the kwarg coercion that runs
later accepts "true", "1" and "yes". Type coercion happens after the gate, so
`REDIS_AWS_IAM_AUTH=1` silently skipped IAM auth and `REDIS_SSL=1` made the
"requires TLS" check fail closed on a connection that was in fact TLS. Both
helpers now share `_str_to_bool`.

AWS signs serverless cache tokens with an extra `ResourceType=ServerlessCache`
query parameter, so tokens minted for a serverless cache were rejected. Adds an
`aws_iam_serverless` setting (`REDIS_AWS_IAM_SERVERLESS`) that puts the
parameter into the signed URL, and lowercases the cache name because
ElastiCache lowercases it at creation time.
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
bf73c49d73 refactor(redis): drop unreachable frozen credentials guard 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
2b0711e4f5 style(redis): restore the Azure AD marker comment
The comment documents that the raw Azure client id, tenant id and secret
are deliberately kept off the connect function, so this branch should
never have dropped it
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
cd4073895f refactor(redis): type botocore credentials without Any
Deferred annotation evaluation keeps the type-checking-only botocore
import off the runtime path, so the alias only reintroduced typing.Any,
which the strict ruff budget now bans
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
aca81c263c fix(redis): validate signed ElastiCache IAM URL 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
f1cd2b03ca fix(redis): type signed ElastiCache IAM URL 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
0217a137fe fix(redis): require TLS for ElastiCache IAM auth 2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
a7ad6023d6 ci: retry CodSpeed result upload
Generated with AI

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
9763bd80ae style(proxy): format coordination Redis validation
Generated with AI

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
921c263876 fix(proxy): validate coordination Redis mappings
Generated with AI

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
bb043fe29a test(redis): cover ElastiCache IAM failures
Generated with AI

Co-Authored-By: Claude Code
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
f66891d80f fix(redis): type ElastiCache IAM configuration
Generated with AI

Co-Authored-By: Claude Code
2026-09-10 10:13:03 -04:00
eugene-yao-zocdoc
604b9edc53 test(redis): cover AWS IAM provider install on async cluster nodes
No existing test combined aws_iam_auth with startup_nodes, the shape
ElastiCache/Valkey Serverless deployments actually use.
2026-09-10 10:13:02 -04:00
eugene-yao-zocdoc
1273e46b8b feat(redis): add ElastiCache IAM authentication 2026-09-10 10:13:02 -04:00
mateo-berri
0308b05c7a merge: bring litellm_internal_staging into litellm_bedrock_mantle_govcloud_cost_row 2026-09-10 07:08:58 -07:00
mateo-berri
996ca7c5b2 test(cost): assert jina rerank spend at the registry rate 2026-09-10 07:04:19 -07:00
mateo
134d1f3899 fix(model_prices): registry audit 2026-09-10, absorb open pricing PRs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 13:17:14 +00:00
dclark
a39c358785 fix(proxy): preserve local spend adjustments during reseeding 2026-09-10 13:44:08 +01:00
michelligabriele
93b15ed428
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_v1_models_alias_resolution 2026-09-10 13:26:18 +02:00
Yinon Kahta
123c01c439 feat(pointfive): answer the ui health check with a liveness ping 2026-09-10 14:02:35 +03:00
Yinon Kahta
29b02e9215 feat(pointfive): add pointfive to the dashboard logging integrations 2026-09-10 14:02:35 +03:00
Yinon Kahta
cb150846c2 feat(pointfive): list pointfive in the proxy callback registry 2026-09-10 14:02:35 +03:00
Yinon Kahta
224c7d6216 feat(pointfive): register the pointfive callback 2026-09-10 14:02:35 +03:00
Yinon Kahta
6b27a16e67 feat(pointfive): add batching logger callback 2026-09-10 14:02:35 +03:00
Yinon Kahta
85d10c8558 feat(pointfive): add presigned upload client 2026-09-10 14:02:35 +03:00
Yinon Kahta
6d7b3a82ff feat(pointfive): add gzipped ndjson batch encoding 2026-09-10 14:02:35 +03:00
Yinon Kahta
0b898b47ca fix(http_handler): let put opt out of following redirects
get and post already take follow_redirects. put built the request and sent it
with the client default, so a caller uploading to a URL it did not choose had
no way to refuse a redirect. Same plumbing as the other two methods.
2026-09-10 14:02:35 +03:00
dclark
8b965880ef fix(proxy): prevent spend counter double counting 2026-09-10 11:29:37 +01:00
mateo
c956c24ade Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_deflake_20260905
Some checks are pending
Terraform Provider / gofmt, vet, build, test (push) Waiting to run
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Waiting to run
2026-09-10 09:19:52 +00:00
yucheng
239c126f29 refactor(guardrails): resolve the Agent 365 conversation id from the typed LiteLLM logging object
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
Terraform Modules / fmt, validate, test (aws) (push) Waiting to run
Terraform Modules / fmt, validate, test (gcp) (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Waiting to run
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Waiting to run
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 08:25:48 +00:00
yucheng
5921e64446 refactor(guardrails): read key_alias from the typed UserAPIKeyAuth in Agent 365 payload
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 08:14:01 +00:00
yucheng
f4033ab84a fix(guardrails): keep Agent 365 missing-secret startup error readable after log redaction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 08:08:33 +00:00
yucheng
fcaf2d7d98 fix(otel): size the message ceiling so every vocabulary fits beside it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 08:00:51 +00:00
yucheng
492251c7bc fix(otel): cap per-index OpenInference message attributes span-wide
OpenInferenceMapper spelled every captured prompt and response message out as
two indexed attributes with no bound. A few dozen turns overran the OTel SDK's
128-attribute span limit, which evicts oldest first, so the gen_ai.* model,
provider, usage, cost and finish reason written before it were what got
dropped. Both directions now share one MAX_MESSAGE_ATTRS_PER_SPAN ceiling, the
response keeps at least half of it, and input.value / output.value still carry
the complete conversation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 07:32:54 +00:00
yucheng-berri
e5da59336d
feat(prometheus): bucket latency by input sequence length (#40059)
* feat(prometheus): bucket latency by input sequence length

* style: format startup resolver call

* fix(prometheus): handle unknown input lengths

* fix(prometheus): preserve disabled custom input length labels

* test(prometheus): seed startup snapshot in mocked logger test

* fix(proxy): preserve database setting types during startup

* fix(prometheus): distinguish missing usage and preserve config persistence

Keep quoted database-storage config values intact for legacy persistence readers while using a local boolean for early callback discovery. Distinguish absent provider usage from an explicitly reported zero when labeling latency metrics.

* fix(proxy): normalize input length flag from secret managers

* fix(prometheus): isolate input buckets and preserve missing usage

Keep built-in buckets on latency histograms, preserve unrelated custom labels, and classify raw incomplete usage and upstream total-only headers as unknown. Cover count conservation, failure callbacks, explicit zero, startup snapshots, and direct caller compatibility. Drop earlier branch budget changes.

* fix(proxy): defer Prometheus alerting until stored settings load

Reuse successful startup storage resolution and preserve callback deduplication across alerting reloads.

* test(prometheus): restore input length flag between tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Revert "test(prometheus): restore input length flag between tests"

This reverts commit f302e0b7bd.

* refactor(prometheus): make input length flag config/env only

Drop the Admin UI General Settings row, the safe DB override entry, the startup reorder that loaded DB litellm_settings before Prometheus callbacks, and the alerting-only Prometheus path. The flag now behaves like prometheus_emit_stream_label: litellm_settings in config.yaml or an os.environ reference, applied on restart.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 00:29:28 -07:00
yucheng
aee5b7a0d8 fix(agent_365): treat a verdict without a boolean allowed field as unavailable
A 200 body that is JSON but lacks a boolean allowed was recorded as a Defender
Block and rejected with 400. It is malformed, so it now routes through the
same unreachable_fallback handling as non-JSON and non-object bodies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 07:26:52 +00:00
Jon Walton
f610bdb54b
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_user_budget_webhook_alerts
# Conflicts:
#	tests/test_litellm/proxy/auth/test_auth_checks.py
2026-09-10 14:58:53 +08:00
yucheng
a1ab5d42cc chore(agent_365): regenerate lazy OpenAPI snapshot and dashboard API types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 06:58:39 +00:00
yucheng
a1e203f0a8 style(agent_365): sort test imports
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 06:44:40 +00:00
yucheng
e34c9206ae Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_agent365_mcp_guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/types/guardrails.py
#	ui/litellm-dashboard/src/app/(dashboard)/guardrails/_components/guardrail_garden_configs.ts
#	ui/litellm-dashboard/src/app/(dashboard)/guardrails/_components/guardrail_garden_data.test.ts
#	ui/litellm-dashboard/src/app/(dashboard)/guardrails/_components/guardrail_garden_data.ts
#	ui/litellm-dashboard/src/app/(dashboard)/guardrails/_components/guardrail_info_helpers.tsx
2026-09-10 06:43:02 +00:00
AaronHowell
f843e596e7 fix(responses): 同步上游亲和性改动
Co-authored-by: Bytechoreographer <Bytechoreographer@users.noreply.github.com>
2026-09-10 14:42:40 +08:00
yucheng
4227474d48 Revert "chore: update Next.js build artifacts (2026-09-10 05:25 UTC, node v24.21.0)"
This reverts commit 2ac03231d4.
2026-09-10 06:42:20 +00:00
devin-ai-integration[bot]
6bb60f34e3
fix(s3_v2): freeze refreshable credentials before signing and retry 403 uploads with a fresh signature (#40187)
RefreshableCredentials (IMDS roles) can refresh between the access key, secret and token reads SigV4 performs, producing a mixed-generation signature that S3 rejects with 403 and the log is dropped. Snapshot the credentials with get_frozen_credentials before signing, treat 403 like 500/503 in the upload retry loop, and fetch credentials plus sign again on every attempt in both the async and sync upload paths. Tests load a real botocore credential_process profile and fake only the HTTP boundary with httpx.MockTransport

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-09 23:14:16 -07:00
devin-ai-integration[bot]
dde19adde1
fix(proxy): bound concurrent key and spend-counter DB lookups to stop prisma pool thrash (#40387)
* fix(proxy): bound concurrent key and spend-counter DB lookups to stop prisma pool thrash

A cache-miss burst fanned every get_key_object DB fallback and SpendCounterReseed
point lookup into the prisma query-engine httpx pool at once. httpcore's request
assignment is O(queued x connections) per event, so the event loop spent most of its
time in pool bookkeeping and the logging worker's 20s wait_for tripped. Callers now
wait on a small shared semaphore (PROXY_DB_LOOKUP_MAX_CONCURRENCY, default 25)
instead of queueing inside httpcore

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): type the in-flight counting prisma fake

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): drop module docstring from db_lookup_gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): type the in-flight counting table fake

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 05:44:30 +00:00
yucheng
2ac03231d4 chore: update Next.js build artifacts (2026-09-10 05:25 UTC, node v24.21.0)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 05:25:49 +00:00
tin-berri
13837d319d
fix(openai): drop tool schema regex patterns OpenAI's validator cannot compile (#40485)
* fix(openai): drop tool schema regex patterns OpenAI's validator cannot compile

OpenAI validates function tool parameters with jsonschema's format checker,
which compiles every pattern with Python re. Claude Code's Artifact tool ships
an ECMA-262 pattern with \p{..} Unicode property escapes, so any OpenAI target
behind /v1/messages, /v1/responses or /v1/chat/completions 400s with
"Invalid schema for function 'Artifact': '...' is not a 'regex'" for every
model family. Drop only the patterns Python re rejects, keep the rest, at the
same seams that already flatten top-level combinators.

* fix(openai): walk only schema positions, iteratively, and drop regexes for every openai deployment

Review round: the regex sanitizer now walks JSON Schema applicator positions
only (properties, items, prefixItems, combinators, $defs, additionalProperties
and the rest), so a pattern key inside default, examples, const or a vendor
extension is data and stays. It also drops patternProperties keys Python re
cannot compile, which OpenAI checks the same way. The walk is level-order and
rebuilt deepest level first instead of recursive, so the code-quality recursion
gate passes and there is no depth cap below what a JSON parser admits. On the
chat wire an openai deployment with a custom api_base now drops such regexes
too, since that base is usually a proxy in front of the same validator, while
the lossier combinator flattening stays limited to api.openai.com hosts.
2026-09-09 22:08:30 -07:00