Commit graph

49299 commits

Author SHA1 Message Date
mateo-berri
b3f2de058b fix(cli): only guard the default settings.json with the lite up and autoroute backups on login --config-claude 2026-09-09 18:19:56 -07:00
Kerry Lu
3aa9267df2 fix(ui): drop blank streamed costs instead of reading them as zero
Number("") and Number(" ") both return 0, which passed the finite check, so a
provider reporting an empty cost got a fabricated $0.000000 metric instead of
having the unusable value omitted.

Both ingestion sites carried the same inline parsing, so this pulls it into one
parseUsageCost helper that keeps finite numbers and non-blank numeric strings and
drops everything else, including booleans, arrays and breakdown objects.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YNw8WvkvCSeTcE5qvergu3
2026-09-09 18:19:52 -07:00
Mateo Wang
b22ca7ac6d
Merge pull request #40464 from BerriAI/litellm_internal_copy_35771
fix(azure): respect DEFAULT_MAX_RETRIES in initialize_azure_sdk_client (internal copy of #35771)
2026-09-09 18:19:08 -07:00
mateo-berri
c1ca963d75 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit_5546_count_tokens_offload 2026-09-09 18:18:17 -07:00
mateo-berri
47eb7f4349 fix(token_counter): give each event loop its own count limiter
A single process-wide anyio.CapacityLimiter wakes waiters with the
asyncio.Event of whichever loop created it, so a fifth loop in another
thread (asyncio.run per request under Celery, gunicorn sync workers, or
run_async_function from function_with_fallbacks) waited forever once four
counts were in flight, and building it at import time raised
AsyncLibraryNotFoundError on anyio below 4.2. The limiter now lives in a
RunVar and is created on first use inside the running loop.

TOKEN_COUNTER_MAX_CONCURRENT_COUNTS reads from the environment like
TOKEN_COUNTER_MAX_EXACT_CHARS, and the proxy's tokenizer lookup runs on the
shared thread pool so a Hub download no longer holds a count slot.
2026-09-09 18:17:32 -07:00
mateo-berri
76c9342e9d fix(proxy): run migrations through python -m prisma when the prisma console script is not on PATH
The proxy probed for the Prisma CLI by spawning the bare `prisma` console
script, so a launcher whose PATH lacked the interpreter's bin directory
printed "prisma package not found", skipped every migration and served
traffic against an empty schema. Every Prisma command now falls back to
`python -m prisma` when the console script is not on PATH, and the boot
probe checks for the console script or the importable package instead of
spawning anything.
2026-09-09 18:17:12 -07:00
Mateo Wang
21ae29b759
Merge pull request #40284 from BerriAI/litellm_legacy_hook_streaming_pipeline_step
feat(guardrails): run legacy post-call hooks as streaming pipeline steps
2026-09-09 18:14:37 -07:00
Mateo Wang
53f3e70f02
Merge pull request #40461 from BerriAI/litellm_lit_7373_responses_custom_tool_call_guardrail
fix(guardrails): scan and rewrite Responses custom_tool_call output items
2026-09-09 18:13:00 -07:00
Joshua Valluru
493c98f9bf fix(mcp): normalize default ports in preview origins 2026-09-09 18:11:52 -07:00
Mateo Wang
0e088337a2
Merge pull request #31884 from BerriAI/litellm_add-claude-sonnet-5-pricing
fix(pricing): rolling model registry update: Bedrock gpt-6-astra, gpt-image-2.5, Cohere rerank 4, Vertex Grok 4.3/4.6/4.20, Gemini 3.5 audio, OpenAI web search fee, xAI Imagine video, Lyria 3.5, Voyage, ChatGPT GPT-5.5/5.6, Bedrock Mantle, Scaleway dates
2026-09-09 18:11:37 -07:00
tin-berri
e34c4c8edc
fix(ui): scope shadow eval models to configured chat groups (#40488)
Co-authored-by: Claude Code <noreply@anthropic.com>
2026-09-09 18:09:28 -07:00
Joshua Valluru
7c1d060b7a fix(mcp): enforce OAuth identity binding across credential lifetime 2026-09-09 18:04:36 -07:00
Joshua Valluru
3523f3731c fix(mcp): bind inherited preview credentials to the saved origin 2026-09-09 18:04:36 -07:00
Joshua Valluru
415538056e chore: sync MCP OAuth binding with current default branch 2026-09-09 18:03:46 -07:00
tin-berri
317b29e69d
feat(cli): add lite configure claude and lite unconfigure claude (#40319)
Persistently route Claude Code through a LiteLLM proxy with a long-lived virtual key or the stored lite login, turn on gateway model discovery so /model lists the proxy's models, optionally pick the model Claude Code starts on, and record what changed so unconfigure restores only the keys the user has not touched since. lite login --config-claude writes through the same receipt and is undoable too. The two settings merges (lite up / --config-claude and lite autoroute) collapse into one credential-aware merge
2026-09-09 18:02:23 -07:00
mateo-berri
b554d79914 fix(cli): resolve Claude Code's settings file through CLAUDE_CONFIG_DIR
Claude Code reads settings.json from CLAUDE_CONFIG_DIR when it is set,
while lite login --config-claude wrote to ~/.claude/settings.json and
lite claude checked that same file for its apiKeyHelper. With the
override set, lite could drop ANTHROPIC_AUTH_TOKEN because the helper
lives in a file Claude Code never reads, leaving it with no key at all.

claude_settings_path(environ) now picks the file the way Claude Code
does, and both commands go through it. The CLI test directory gets an
autouse conftest that isolates HOME, USERPROFILE, and CLAUDE_CONFIG_DIR
per test and fails any test that writes the developer's real
settings.json, restoring it first.
2026-09-09 18:01:11 -07:00
Joshua Valluru
245369764e fix(mcp): preserve edited settings in static connection previews 2026-09-09 17:49:52 -07:00
mrinal
3258b732c7 chore(ui): bump smol-toml to fix GHSA-7w5x-hrqm-74c2 osv-scan failure
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 00:46:54 +00:00
kerry-berri
b8d7f68aeb
Apply suggestion from @greptile-apps[bot]
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-09-09 17:46:12 -07:00
joshua-berri
a8bb2f93b9
Merge pull request #40494 from BerriAI/litellm_fix_mcp_dotted_properties_5777
fix(ui): preserve dotted MCP tool argument names
2026-09-09 17:45:00 -07:00
Joshua Valluru
892d20d86f fix(mcp): refresh tools when editing server connection settings 2026-09-09 17:34:24 -07:00
Kerry Lu
1213d7d399 test(e2e): cover /embeddings and assert memory and fallbacks in the Redis timeout test
Add an embeddings case with its own closed-port primary and mock backup (the fallback map in
the gateway config gains the pair; LiteLLMParamsBody.mock_response accepts the list an embedding
mock needs). Assert from /metrics that the proxy's resident memory grows by no more than 200 MB
across each case where the process collector reports it (Linux), that the router counted a
successful fallback for every request, and that every spend row is a success.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZDULyJPp17ZFiJenRxs2T
2026-09-09 17:17:33 -07:00
tin-berri
6b721de3e5
fix(databricks): route Unity model services through AI Gateway (#40492)
Co-authored-by: Claude Code <noreply@anthropic.com>
2026-09-09 17:15:48 -07:00
Joshua Valluru
0b21b99ebd fix(ui): type MCP argument resolver results explicitly 2026-09-09 16:57:15 -07:00
ryan-crabbe-berri
2085a37b82 test(e2e): issue the JWT suite's tokens from a real Keycloak realm
The suite used to mint its own RS256 tokens from a stand-in issuer, which
could only ever prove the proxy agreed with the tests: the claims were
whatever the tests chose to sign. Every JWT bug worth catching lives in the
shape of what an identity provider really emits, so the suite now runs
against Keycloak (realm in idp_realm.json), provisions a group and a user per
test through its admin API, and signs in through the direct-access grant.

That changes what the tokens look like: sub is Keycloak's opaque user uuid
rather than a friendly name, groups arrives from a protocol mapper, aud is
the IdP's own audience, and the JWKS carries an encryption key beside the
signing key so the proxy has to select on kid. The expiry case now takes a
one-second token from a second client in the realm and waits for it to lapse
instead of forging a stale exp.

The proxy config the suite needs is unchanged. CI runs it against a Keycloak
deployed beside the ephemeral stack, which lives in the releaser repo.

Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
2026-09-09 16:56:14 -07:00
Joshua Valluru
fb9bb60e80 test(mcp): cover auth diagnostics through HTTP handler 2026-09-09 16:55:52 -07:00
Joshua Valluru
70c238dde0 fix(ui): preserve dotted MCP tool argument names 2026-09-09 16:54:43 -07:00
ryan-crabbe-berri
a650178ebe
Merge pull request #33703 from BerriAI/litellm_fix_jwt_key_mapping_cascade_delete
fix(jwt): cascade-delete JWT key mappings when their virtual key is deleted
2026-09-09 16:52:20 -07:00
mrinal-berri
3bc2aea4cd
Merge branch 'litellm_internal_staging' into litellm_fix_health_check_db_storm 2026-09-09 16:51:02 -07:00
tin-berri
2000642592
fix(router): resolve route candidate ids through the router's own resolver (#40491)
get_candidate_model_ids_for_route (added in #40280 for the encrypted-content affinity
check) reconstructed the candidate pool by unioning the model_name and team indexes with
pattern_router.route. That diverged from how the router actually resolves a route: it took
a union instead of the first matching path, and pattern_router.route only matches the
literal name, so a provider-qualified pattern (matched by get_deployments_by_pattern, which
retries the {provider}/{model} form) was missed and the default deployment was ignored.

For an affinity follow-up on a wildcard or team-public route, that mismatch could strip
encrypted reasoning on a same-group cooldown, or return a 503 on a real cross-path switch.

Delegate the non-model_name case to _try_early_resolve_deployments_for_model_not_in_names,
the same resolver _common_checks_available_deployment uses, so candidate membership follows
the router's real precedence. With include_team_models left off it stays read-only and does
not raise. Behavior for concrete model groups and routing groups is unchanged.


Claude-Session: https://claude.ai/code/session_01KAumQbhzk6jdWWHFLA8Jar

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-09 16:43:50 -07:00
ryan-crabbe-berri
969d152c4b fix(jwt): evict jwt_key_mapping cache when a virtual key is deleted
The FK cascade drops the LiteLLM_JWTKeyMapping row, but the cached
jwt_key_mapping:{claim}:{value} entry still resolved to the deleted token
hash, so every JWT call from that identity failed until
virtual_key_mapping_cache_ttl expired instead of auto-registering against a
recreated key. delete_verification_tokens now snapshots the mapping cache
keys before the delete and evicts them across replicas afterwards, the same
way /key/regenerate already does.

Claude-Session: https://claude.ai/code/session_011Tn3657NkV6ojLqewL64Kb
2026-09-09 16:40:25 -07:00
Mateo Wang
b7dad8b44e
Merge pull request #40294 from shotsan/fix/empty-choices-handling
fix(convert_dict_to_response): handle empty choices list without raising 500 APIError
2026-09-09 16:34:47 -07:00
Kerry Lu
2e4461ec45 test(e2e): register the Redis timeout deployments through /model/new
The e2e directive has every test create its deployments through the management API and delete
them on teardown. Drop the static model_list from the gateway config; the test now registers the
closed-port primary and the mock backup itself, and the fallback map stays in router_settings
where proxy-level config belongs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZDULyJPp17ZFiJenRxs2T
2026-09-09 16:33:18 -07:00
mateo-berri
a68fe4e4d9 fix(responses): keep the client's usage shape when the logging copy cannot re-validate the response 2026-09-09 16:31:48 -07:00
mateo-berri
7563d94e65 test(responses): assert the streamed response.completed event echoes the named tool_choice 2026-09-09 16:29:55 -07:00
mateo-berri
e27a018a6c fix(router): pin bridge-replayed encrypted reasoning to the deployment that minted it
The encrypted_content_affinity check only read the pin from the Responses
input, which /v1/messages builds after the router has picked a deployment,
so a model group spread across OpenAI orgs sent follow-up turns to the
wrong org and got invalid_encrypted_content back. The check now also
decodes the pin from bridge-tagged thinking and redacted_thinking blocks
in the Anthropic messages. The bridge also keeps a deployment's own
include list next to reasoning.encrypted_content instead of replacing it.
2026-09-09 16:29:52 -07:00
Kerry Lu
2d7e5999b7 test(e2e): pause Redis writes for the run and prove the timeouts from /metrics
A loopback Redis answers many commands inside the 1 ms socket timeout, so nothing guaranteed
the failure path ran. The test now holds the proxy's Redis in CLIENT PAUSE WRITE for its
duration, so every write the proxy sends, the spend counter increment included, outlives the
timeout, and lifts the pause in teardown. Reads stay live so the control connection can do that.

Enable the prometheus callback in the gateway config and assert from /metrics that the proxy
counted at least the breaker's five timeouts and that, during each case, it saw fresh timeouts,
a breaker transition, or an open breaker rejecting every call. The open breaker is the state a
customer's worker sits in, and cost tracking fails on every request either way.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZDULyJPp17ZFiJenRxs2T
2026-09-09 16:26:03 -07:00
mateo-berri
5db99ca7ba test(e2e): register the poller batch cost join as its own uncovered cell
The retrieve writer and the CheckBatchCost poller are two different spend
writers, and the in-run failed batch only proves the first. Split the
key_attribution batch cell in two: retrieve_batch_cost_joins_retrieving_key,
which test_terminal_batch_cost_row_joins_the_retrieving_key claims, and
poller_batch_cost_joins_creating_key, which no test claims yet and so shows
up on the coverage dashboard as a P1 gap instead of hiding behind the
retrieve leg. The rationale records why one run cannot hand the poller a
completed batch on a stack that boots a fresh Postgres per build.
2026-09-09 16:23:45 -07:00
Joshua Valluru
c972bbe80f fix(mcp): send debug headers immediately for GET streams 2026-09-09 16:22:41 -07:00
mateo-berri
a248ff2c6e fix(guardrails): fail open on Responses tool-call rewrites in another shape and keep item names
Guardrails that hand back tool_calls in their own shape (vendor JSON, user code output) raised a KeyError on the non-stream Responses write-back. Returned tool calls are now validated before comparison; a shape or count that does not line up leaves every tool-call item unchanged and logs a warning naming the guardrail. A tool call's name is written back only when the guardrail changed it, so a nameless custom_tool_call no longer picks up the custom_tool placeholder.
2026-09-09 16:14:50 -07:00
mateo-berri
654afb5477 fix(e2e): assert the batch cost join on an in-run failed batch instead of a cross-run baton
The batch list is served from LiteLLM_ManagedObjectTable whenever the managed
files hook is loaded, and the Buildkite e2e stacks bundle a fresh Postgres per
build, so a prior run's marker batch is never listed and the baton could only
ever pass vacuously. Each run now creates a batch OpenAI fails at validation
within seconds, retrieves it by its raw provider id with the same key until it
is failed, and asserts the {provider_batch_id}_batch_cost row that retrieve
writes joins the key's token hash and alias
2026-09-09 16:14:07 -07:00
mateo-berri
e27a549aa9 test(bedrock): type the probe credential overrides 2026-09-09 16:14:00 -07:00
mateo-berri
7e137958e9 fix(cli): let the apiKeyHelper supply Claude Code's key under lite claude
`lite login --config-claude` writes an apiKeyHelper into ~/.claude/settings.json,
and `lite claude` then also exported ANTHROPIC_AUTH_TOKEN, so Claude Code opened
with its "Both ANTHROPIC_AUTH_TOKEN and apiKeyHelper set" banner and, because the
env token wins, never ran the helper that was meant to refresh the key.

When the launch key is the stored login key and settings.json carries exactly
the helper lite wrote for this base URL, `lite claude` now leaves the token out
of the env (dropping an inherited one) so Claude Code asks the helper. An
explicit --api-key or LITELLM_PROXY_API_KEY, a helper for another proxy, a
hand-written helper, or no settings file keep the previous behavior.
2026-09-09 16:12:40 -07:00
ryan-crabbe-berri
b1fb478e37 feat(proxy): restrict aws_session_tags on model management to proxy admins
Team admins can create and edit team models through /model/new, PUT /model/update
and PATCH /model/{id}/update. The proxy forwards aws_session_tags to STS under its
own identity, so a team admin could pick tags that unlock aws:PrincipalTag gated
resources. Only proxy admins may now set or change aws_session_tags there; an
unchanged tag set still passes so team admins can edit other fields
2026-09-09 16:10:04 -07:00
mateo-berri
cc0c6087e3 chore: merge litellm_internal_staging into litellm_lit_7022_azure_ai_passthrough_config 2026-09-09 16:09:57 -07:00
Mateo Wang
8b2983bd90
Merge pull request #40449 from BerriAI/litellm_databricks_reasoning_content
fix(databricks): keep top-level reasoning_content from OpenAI-compatible gateway models
2026-09-09 16:07:15 -07:00
mateo-berri
96e46ba8ff fix(responses): stamp streamed usage cost when the provider usage arrives as a dict
Perplexity's Responses payload fails ResponsesAPIResponse validation on
truncation "" and is kept as an unvalidated model, so its usage stays a
plain dict and _stamp_responses_usage_cost raised AttributeError on every
streamed completion once reasoning made the cost non-zero. Validate the
dict into ResponseAPIUsage before stamping, keeping a provider-reported
cost when it carries one.

Resolves LIT-7391
2026-09-09 16:06:55 -07:00
mateo-berri
fbc6fb56ae fix(bedrock): route gpt-6-astra reasoning_effort to reasoning.effort and mark Nova 2 tool_choice
The converse reasoning gate only matched openai.gpt-5, so gpt-6-astra fell through to
Anthropic's thinking block and Bedrock rejected the call with 400 Unknown parameter:
'thinking'. Match any openai.gpt-<digit> model at the three gate sites instead.

Nova 2 lite and pro accept forced tool_choice on Converse (verified live on
us.amazon.nova-2-lite-v1:0), so the nine Nova 2 registry keys now advertise
supports_tool_choice. The invoke dispatcher also forwards json_mode to Nova like it
already does for Anthropic and TwelveLabs.
2026-09-09 16:06:19 -07:00
Kerry Lu
939ac04a93 test(e2e): cover /v1/responses in the Redis timeout test and fail the primary for real
The Responses path returns a mock for any mock_response string, so the InternalServerError
sentinel never failed there. Point the primary deployment's api_base at a closed port instead,
which fails every endpoint the same way, then parametrize the test over /chat/completions and
/v1/responses and register both on the coverage cell.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZDULyJPp17ZFiJenRxs2T
2026-09-09 16:03:36 -07:00
mateo-berri
21e6c6e2f3 test(convert_dict_to_response): expect the type-naming error for a null choices value 2026-09-09 15:56:41 -07:00