Commit graph

54050 commits

Author SHA1 Message Date
berriai-litellm-provider-info-sync[bot]
aa31c0b962
feat(pricing): add github_copilot/claude-haiku-5.5 (#45656)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-09 13:22:35 -07:00
Mateo Wang
3c56cd31f5
chore(greptile): load nested AGENTS.md files as review context scoped to their directories (#45636)
* chore(greptile): load nested AGENTS.md files as review context scoped to their directories

* refactor(greptile): use a stdlib dataclass so the files.json check runs without uv

* test(greptile): write fixture AGENTS.md files through a helper instead of rebinding a loop variable
2026-10-09 13:22:29 -07:00
berriai-litellm-provider-info-sync[bot]
bbc8d66323
chore(cost-map): add openai gpt-rosalind-discovery from the pricing page (#45653)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-09 12:51:58 -07:00
devin-ai-integration[bot]
07e315ebfb
fix(ui): show team alias in key edit Team dropdown and allow clearing it (#45629)
* fix(ui): show team alias in key edit Team dropdown and allow clearing it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): show alias-less teams by ID once in the key edit Team dropdown

* fix(ui): keep the key organization when its team is cleared

* fix(ui): show a dash for teams without an alias, keeping the ID underneath

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 19:47:03 +00:00
devin-ai-integration[bot]
dafa9cbd4c
ci: run every build_and_test job in every CircleCI pipeline (#43152)
* ci: run the main branch CircleCI job set on rc/* branches

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: let litellm_rc_* branches run the main-only CircleCI jobs

The 26 build_and_test jobs filtered to main now share one anchor that also admits branches named litellm_rc_*, so a run-ci pipeline on such a branch at an rc SHA runs the full suite. Every other PR keeps the integration-only subset, and the migration schedule stays on main

* ci: run every build_and_test job in every CircleCI pipeline

Drop the main-only branch filter from the 26 build_and_test jobs, so a run-ci PR pipeline into any base, an rc branch pipeline and main's scheduled pipeline all run the full suite. The migration cron stays on main

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 12:33:14 -07:00
berriai-litellm-provider-info-sync[bot]
de6a0717b7
feat(azure): add azure_ai/Microsoft-Decision-1 (#45639)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-09 12:01:13 -07:00
devin-ai-integration[bot]
fff86d6376
fix(proxy): classify /v1/messages pass-through streams as Anthropic on any host (#45602)
Co-authored-by: Sunny-Soni00 <sunny.s@atriauniversity.edu.in>
2026-10-09 11:57:03 -07:00
devin-ai-integration[bot]
f2f8df2858
feat(ui): show the public JWKS for LiteLLM-signed Anthropic credentials (#45527)
* feat(ui): show the public JWKS for LiteLLM-signed Anthropic credentials

The edit modal of a saved Anthropic credential whose identity source is the
LiteLLM internal issuer now fetches GET /credentials/{name}/jwks and shows the
document with a copy button, so an admin can register it in the Claude Console
without curl. Picking the internal issuer before the credential is saved shows
a hint to reopen it once saved

* fix(ui): wait for a fresh JWKS before offering one cached from an earlier open

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-09 11:53:10 -07:00
yuneng-jiang
cddc7cde97
test: fix stale and flaky tests across CircleCI, GHA and Buildkite (#45530)
* test(integration): poll the partition lock witness off the event loop

The witness poll ran a blocking psycopg connect and query on the test's event loop, so the partition DDL task only progressed during the 20ms sleeps and missed the 3s deadline on loaded runners. Polling through asyncio.to_thread lets the DDL reach the held lock concurrently, and the deadline is 10s since it only bounds how long the DDL takes to start waiting

* test(e2e): retry the reasoning turn until the model emits a reasoning item

Whether gpt-5.4-mini emits a reasoning item next to a forced function call is up to the model, and it skipped it on four scheduled runs while a same-SHA rerun passed. The first turn now retries up to three times before the unchanged reasoning-item assertion

* test(e2e): retry a Bedrock Converse model once when it reports no cache tokens

Opus on Bedrock Converse reported zero cache tokens on half of the scheduled runs while a same-SHA rerun passed. A model that comes back uncached is asked once more inside the 5 minute window, and the non-zero cache token assertion is unchanged

* test(e2e): poll until a deleted stored response stops being retrievable

Azure kept serving a just-deleted streamed response for a moment, so a single retrieve did not raise. Both delete tests now poll retrieve until it returns an error and still require a 4xx

* test(e2e): give the Together prompt cache five primed attempts

Together documents its prefix cache as best-effort, and three primed attempts came back with zero cached tokens once while a same-SHA rerun passed

* test(e2e): allow 30s for the idle RSS reading at session start

A loaded router replica took longer than 10s to answer the session-start memory read, which pytest reruns cannot recover. The RSS budget assertion is unchanged

* test(e2e): scope the MCP submission check to its own card

The register call timed out at 15s on a loaded stack and leaked a submitted server, so the retries matched two "3 passing, 1 failing" cards. Registration gets 60s and the check reads only the submitted server's card

* test(router): wait for the primary success record before counting it

The shadow fan-out test waited only for the two shadow success events and then asserted exactly one primary success, which the logging worker sometimes had not delivered yet. It needed a rerun on 5 of 59 main runs

* test(proxy): give the spend-log cleanup run a 1s budget

A 0.25s budget could expire before the first batch on a loaded xdist worker, leaving rows_deleted at 0. One second still stops the 50-batch, 5-second loop on the deadline, which is what the test proves

* test(autorouter): release the slow token count at 2.05s instead of 2.3s

The count only has to finish after a quarter of the 8s worker budget, and the planner gives it 3s, so releasing at 2.3s left 0.7s of slack that a loaded runner used up. 2.05s is still past the quarter mark with almost 1s of slack

* test(xai): run the reasoning-effort tests on grok-4.6 with a prompt that needs reasoning

grok-4.7 now returns reasoning tokens without any reasoning text, so reasoning_content was never set and both tests failed on every scheduled run. Called directly, grok-4.6 returned reasoning text 6 of 6 times on a step-by-step arithmetic prompt but only some of the time on a bare greeting

* test(integration): give the two worker-kill chaos tests the owned-proxy time budget

Each test spends about 50s on the burst and then up to graceful_stop_seconds() stopping its owned proxy, which overran the 90s default pytest timeout on loaded runners. They now use the same 2 * graceful_stop_seconds() + 120 budget as other owned-proxy tests

* test(integration): compare regional image responses without their created timestamp

Two identical generations that straddle a second boundary differ only in created, which failed the byte-for-byte comparison. Every other field and both spend rows are still compared

* test(integration): count only the S3 logger's flush task

The set of asyncio tasks created while the logger starts also picked up client close finalizers left by earlier tests, so the count ranged from 1 to 8. The test now counts the periodic_flush task it owns and cancels

* test(integration): send guardrail timeout probes eight at a time

All 44 probes went at once to a two-worker proxy, so on a loaded runner some guardrails used their whole 1s timeout before their request left the proxy, and the sink never saw them. A different set of providers failed on each run. Eight in flight still overlaps the waits

* test(integration): count a killed Arize chaos worker as gone once it is a zombie

psutil reports an unreaped zombie as still running, so the 10s death check failed whenever the supervisor was slow to reap the SIGKILLed worker, and the stop then overran the 90s default timeout. The test now accepts a zombie or a missing process and has the owned-proxy time budget

* test(integration): send the Typesafe connection test its real model and endpoint

The test passed the proxy alias as litellm_params.model, and /health/test_connection lays the request over the stored deployment, so the alias replaced the real typesafe/ model and the check failed with "LLM Provider NOT provided" on every run since it landed in #45481. It now sends the provider model, api_base and api_key directly

* test(integration): keep DB-stored models out of the usage-routing Redis read test

The owned proxy inherited store_model_in_db and the job's shared database, so a model another test left behind joined the router and added its cooldown key to the MGET the test compares exactly. Reproduced locally with one /model/new model present (3 failed), and green with model loading from the database turned off

* test(vertex_ai): prove the batch upload streams by laziness instead of peak memory

The two tracemalloc ratio tests flaked on unrelated PRs because peak memory on a shared xdist worker includes other threads' allocations and garbage from earlier tests. They are replaced by deterministic checks of the same property: the upload stream parses and maps a row only when it is pulled, so a body whose tail is not JSON yields its valid rows first, and a Path source reflects a row rewritten on disk after the upload started. Making the parse, the output, or the file read eager fails these tests

* test(integration): keep the alias test_connection call as a known bug

* test: type the counting mapper and the Bedrock rerun helper

* test: annotate the new test locals as Final and build them as tuples

* test(vertex_ai): prove pull-driven transforms with the garbage tail alone
2026-10-09 11:40:14 -07:00
nate-berri
0519882fef
chore: ignore .mypy_cache (#45628)
Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
2026-10-09 18:26:19 +00:00
tin-berri
1dbed8e1e2
fix(liteadmin): detach keys from teams and default to Sonnet 5.5 (#45625) 2026-10-09 11:17:53 -07:00
berriai-litellm-provider-info-sync[bot]
fe955f5e42
fix(bedrock): add the bare Pegasus 1.5 row and take GPT-5.x context windows from the model cards (#45621)
* fix(bedrock): add the bare Pegasus 1.5 row and take GPT-5.x context windows from the model cards

Price-Sync: litellm-providers

* fix(bedrock): whitelist the bare Pegasus 1.5 id for the converse routing check

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: Kerry Lu <klu@berri.ai>
2026-10-09 11:06:54 -07:00
yuneng-jiang
6afdf482de
test: move offline anthropic prompt caching tests to tests/unit (#45616)
The two offline tests in tests/local_testing/test_anthropic_prompt_caching.py failed on main because
#24071 started passing logging_obj to client.post, so assert_called_with on the patched
AsyncHTTPHandler.post no longer matched. The request bodies themselves were still correct

Both now live in tests/unit/llms/anthropic/chat and assert the body and headers that reach the
wire through respx instead of patching our own HTTP handler. The coverage allowlist entry keeps the
five tests that still need provider credentials
2026-10-09 10:53:08 -07:00
devin-ai-integration[bot]
61e5f2dd3c
feat(decisions): add hosted_vllm provider (#45501)
* feat(decisions): add hosted_vllm provider

Adds HostedVLLMDecisionsConfig so hosted_vllm/<served model> works on
POST /v1/systemone and POST /v1/decisions against a vLLM server that
serves the System One body (vllm-project/vllm#59299, after v0.31.0).
vLLM only supports choice questions and rejects others, so the base
decisions config gains a per-provider health_check_questions attribute
and the evaluation-mode health probe sends a choice question for
hosted_vllm while every other provider keeps the noul probe.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): move the hosted_vllm placeholder key to constants

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 10:41:51 -07:00
Krrish Dholakia
18e87d3039
fix(ui): title usage overview chart "Daily usage" and label the model table "Top models" (#45604)
* fix(ui): title usage overview chart "Daily usage" and label the model table "Top models"

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): clarify usage chart subtitle as top 8 models plus Other

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 09:56:59 -07:00
yuneng-jiang
5e1c5c0bd1
test: move offline proxy auth, hook, spend and pass-through tests to tests/unit (#45568)
* test: move offline proxy auth, guardrail hook, spend, budget reset, GCS payload and pass-through tests into tests/unit

Relocate 53 legacy tests that pass with no network or keys into the tests/unit
files that mirror the code they exercise, drop 2 llm_guard tests already covered
by test_llm_guard_call_type_aliases, and delete the emptied legacy files.

* test: freeze the reset clock and run the real reset methods in the moved budget reset tests

Pin datetime in reset_budget_job and timezone_utils to a fixed instant and await
the service-hook tasks the job spawns instead of sleeping. Drive the per-user
failure with a malformed budget_duration row from the prisma mock so the real
_reset_budget_for_key/user/team run instead of patched replacements.

* test: use fixed timestamps in the moved pass-through and spend log tests

* test: mock the moderation HTTP call and type the moved moderation and spend log tests

Serve the OpenAI moderation response through an httpx MockTransport so the real
async_make_request runs, annotate the moved tests' locals with Final, and pass
get_logging_payload typed kwargs instead of an untyped input_args dict.

* test: drive the moved proxy tests through boundaries and annotate their locals with Final

Serve GCS downloads and the Router moderation call through httpx mocks with fake
Google credentials, let the real pass-through success handler fail on the mocked
upstream response, and fold the Vertex live route endpoint check into the existing
route test. Annotate moved locals with Final, build the budget reset rows without
rebinding, and type the remaining fakes without **kwargs
2026-10-09 09:49:13 -07:00
yuneng-jiang
3250ccc802
test: move offline logging, secret manager and provider tests from legacy dirs into tests/unit (#45554)
Moves 58 legacy nodes (cyberark, logging callback manager, langfuse handler and helpers, SQS, app crypto, bedrock nova and invoke, cloudflare, litellm_proxy, nvidia nim, replicate, triton, convert_dict_to_response, get_model_info, vcr live-call probe) into the tests/unit files that mirror the modules they exercise. Calls that reached huggingface.co, langfuse or 0.0.0.0 are refused at the HTTP boundary with respx. Tests that need credentials, a license, a spawned server, sleeps, a subprocess, or assert nothing stay in their legacy files
2026-10-09 09:48:10 -07:00
yuneng-jiang
3eeca75ca0
test: move offline router, retry and latency tests from local_testing to tests/unit (#45551)
* test: move offline router, latency-routing, retry and exception tests from local_testing into tests/unit

Moves 29 offline nodes into the tests/unit files that mirror the code they exercise.
Hosted-provider calls are mocked at the HTTP boundary with respx, and sleeps that only
built a start/end gap are replaced by explicit timestamps. Legacy files left empty are deleted.

* ci: record the live prompt caching cases left in tests/local_testing as unrun

Moving the router prompt caching test to tests/unit leaves only live Anthropic and Vertex cases in tests/local_testing/test_anthropic_prompt_caching.py, which every CircleCI -k already deselects

* test: freeze the clock in the moved lowest-latency routing tests

The latency logger buckets usage by the current minute, so a run that crosses a minute boundary could grow the cache and fail the memory check. The moved tests now use fixed timestamps and a frozen clock in the logger module

* ci: describe the mixed prompt caching file accurately in the coverage allowlist
2026-10-09 09:47:04 -07:00
devin-ai-integration[bot]
6d4fa56ac4
fix(model_prices): registry audit 2026-10-09, together qwen3.7-max price, azure kimi-k2.7-code retirement (#45595)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 09:02:58 -07:00
berriai-litellm-provider-info-sync[bot]
63a4f3f2f3
fix(gemini): sync gemini deprecation dates with the changelog and deprecations page (#45491)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-09 04:56:18 -07:00
devin-ai-integration[bot]
bbe5595a99
fix(responses): keep a flagged hosted deployment's prompt cache breakpoint on the chat bridge (#45500)
* fix(responses): keep a flagged hosted deployment's prompt cache breakpoint on the chat bridge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): read the serving provider's row for prompt cache breakpoints on the chat bridge and base_model

* fix(anthropic_cache_control_hook): rename the module-level provider resolver so the recursion check passes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): follow the credential dialog's new name field and auth method id in the federation e2e spec

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-09 04:00:39 -07:00
devin-ai-integration[bot]
dd31339e05
test: wait on conditions instead of wall-clock in guardrail parallelism and scope-option tests (#45571)
* test: wait on conditions instead of wall-clock in guardrail parallelism and scope-option tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: annotate the guardrail barrier locals as Final

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 03:41:48 -07:00
devin-ai-integration[bot]
b005af81e6
refactor(types): replace Any with proven types in 3 files (#45565)
* refactor(types): replace Any with proven types in 5 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): drop Azure AI Search and OpenRouter image edit seams that reject previously accepted payloads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 03:24:16 -07:00
joshua-berri
a4fd58501a
fix(mcp): never forward the caller's LiteLLM key to MCP servers (#45401)
* fix(mcp): never forward the caller's LiteLLM key to MCP servers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(mcp): tidy caller-key scrub

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): avoid master-key scanner false positive

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): match admission when scrubbing the caller key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(mcp): drop header copy churn

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): normalize bearer variants in caller-key match

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): inject admission header name and cover provider keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): exercise real REST header extraction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): cover caller-key scrub on the real REST path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): exclude gateway admission header from upstream forwarding

* fix(mcp): skip empty admission headers when scrubbing keys

* fix(mcp): use scrubbed server credentials for auth probes

* fix(mcp): match probe credential precedence to upstream calls

* test(mcp): simplify probe header assertion

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-09 03:21:07 -07:00
Mateo Wang
a3236ba947
feat(ui): configure OpenAI workload identity federation from the LLM Credentials and Add Model forms (#45528)
* feat(ui): configure OpenAI workload identity federation from the LLM Credentials and Add Model forms

* fix(ui): let a federated OpenAI credential omit the service account when the proxy env provides it

* fix(ui): clear the federation API Base error when the admin switches the credential to an API key
2026-10-09 00:30:44 -07:00
devin-ai-integration[bot]
cdba36a1f0
chore: bump litellm-enterprise 0.1.75 -> 0.1.76, litellm-proxy-extras 0.4.107 -> 0.4.108 (#45534)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-09 07:29:44 +00:00
devin-ai-integration[bot]
ee47da334d
test(proxy): inject the clock into the v1 parallel request limiter and pin it in its tests (#45522)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-09 00:16:03 -07:00
devin-ai-integration[bot]
488594e03f
refactor(proxy): inject one UTC clock read per operation into gateway tracking, PTU rollup and Mavvrik export (#45520)
* fix(proxy): read the UTC clock once per operation in gateway tracking, PTU rollup and Mavvrik export

* refactor(proxy): default the injected clocks to get_utc_datetime instead of three private copies

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-09 00:15:38 -07:00
yuneng-jiang
a7ff709f3d
fix(ui): let the model picker remove selections that are no longer available (#43944)
* fix(ui): let ModelSelect deselect selections that are no longer offered

A selected model that is no longer served (removed from config, deleted,
or outside the org ceiling) never appeared in the dropdown, so once it sat
past the fifth chip nothing could remove it. List such selections in an
Unavailable group ahead of the offered options so they can be found and
unchecked.

* fix(ui): skip the Unavailable group when no live models were loaded

A failed or empty model list made every selection look unavailable. Only
flag selections as unavailable when there is a live model list to compare
against, and cover ordering with several unavailable selections.

* fix(ui): only suppress the Unavailable group when the model list failed to load

Keying the guard on the context-filtered list hid the group when an
organization ceiling excluded every selection, which brought the dead end
back for team forms. Key it on the proxy model list having loaded.

* fix(ui): keep other selections when removing one beside a special option

Removing an unavailable model while a special option stayed selected
collapsed the whole selection to the special option, dropping the other
saved values. Collapse only when a special option is newly picked. Also
skip the Unavailable group while an org team's model ceiling is unknown,
since the offered list is empty for that reason alone.
2026-10-09 00:00:26 -07:00
devin-ai-integration[bot]
0d17f954c0
feat(credentials): add display_name and make credential_name immutable on PATCH (#43148)
* feat(credentials): add credential_alias and make credential_name immutable on PATCH

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): query credential rows by role to stay under the no-node-access budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): expect credential_alias in load_credential_list dump

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): drive the credential SearchSelect by placeholder and option roles

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): pass the decrypted CredentialItem to update_db_credential during master key rotation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): trim credential_alias in CredentialModal so whitespace-only input clears the alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(credentials): replace credential_alias with display_name and reject edits to config credentials

credential_name stays the immutable reference key. display_name is a nullable, trimmed, max 255
character label set on POST or PATCH (omit keeps it, null clears it, blank is a 400). Reads return
display_name plus an in-memory source tag (db or config), and PATCH or DELETE on a config-defined
credential answers 400 instead of a misleading 404

* feat(ui): show credential display names everywhere and lock config credentials

The credentials table shows the display name with the credential name beneath it, badges config
credentials and disables their edit and delete actions. The edit modal adds an editable display
name next to the read-only credential name and only sends it when it changed. Every credential
picker and reference (add model, model info, models table, vector store form and info) shows the
label while still submitting credential_name

* feat(cli): set display names on credentials and show them in the list

* refactor(credentials): keep the new credential types inside the lint gate budgets

* refactor(cli): print credential command JSON through one helper

* fix(credentials): treat an empty credential_name on PATCH as omitted and pin the 404 for vanished rows

A blank credential_name never renamed anything, so PATCH accepts it again instead of answering 400. New tests pin a trimmed display_name on PATCH, the repository carrying display_name, and a 404 (not the config-owned 400) for a DB credential another worker already deleted. The hydration helper no longer copies display_name, since none of its callers read it, and the CLI update passes display_name straight through because the exactly-one check already makes it None when clearing.

* fix(ui): keep a model's credential when its picker text is emptied and skip the admin-only list for other roles

Clearing the search text in the model edit form's credential picker used to submit null, silently detaching the model's credential on save. None is now the only way to clear it. The models table fetches /credentials only for proxy admins, since everyone else got a 403 on each page load and falls back to the raw name anyway. Also fixes a type error in the re-use credential dialog and adds tests for the display name surviving a provider switch, the 255 character limit, the request payload, and the vector store None choice.

* fix(client): percent-encode the credential name when updating its display name

A name with ? or # was cut short in the URL, so the PATCH landed on a different credential or 404'd.

* fix(credentials): answer 405 with Allow: GET for PATCH and DELETE on config credentials

A config-defined credential exists (GET returns it) but is read-only through the API, which is what 405 Method Not Allowed means. 400 described a malformed request, and 404 would claim the credential does not exist.

* fix(proxy): ignore a display_name set on config credential_list entries so non-string values cannot fail boot

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 23:49:56 -07:00
devin-ai-integration[bot]
5c6ea040b0
test: update stale MCP call_tool and usage card assertions, add timeout headroom to request-log index boot test (#45523)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 23:44:46 -07:00
devin-ai-integration[bot]
145d382089
ci: run only integration jobs on PR CircleCI pipelines and move router and guardrails suites to GHA (#45509)
Restrict 26 build_and_test jobs to main with job-level branch filters, delete litellm_router_unit_testing in favor of a router-unit-tests GHA shard, and narrow guardrails_testing to the license-dependent test while the rest of tests/guardrails_tests runs in a guardrails-tests GHA shard

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 22:58:55 -07:00
devin-ai-integration[bot]
d34ad35281
fix(anthropic): count a leading system run through count_tokens' system parameter (#45463)
* fix(anthropic): count a leading system run through count_tokens' system parameter

Anthropic's count_tokens rejects role "system" at the head of messages, so
/v1/responses/input_tokens with instructions, and /v1/messages/count_tokens
with a system-role message, fell back to the local tokenizer. The shared
Anthropic count_tokens transformation now lifts the leading run of system
messages into the top-level system parameter, the way the chat path sends
it, after any system the caller set. Anthropic direct, Azure AI Anthropic,
and Bedrock Mantle share that transformation.

* test(integration): cover the count_tokens leading-system lift across Anthropic, Azure AI and Bedrock Mantle

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 22:34:25 -07:00
yuneng-jiang
72ab736863
test(routing): wait on deployment registration instead of boot model-info traffic (#45267) 2026-10-08 21:31:53 -07:00
devin-ai-integration[bot]
4b975e6f49
ci: render lint, unit and smoke checks as <tier> / <job> with one collector per tier (#45480)
* ci: restructure lint, unit and smoke workflows

* ci: preserve source formatting in tier workflows

* ci: retire per-shard coverage flags and tighten the tier workflow guards

* test(ci): freeze the workflow startup safety models

* ci: keep the current required check names running until the ruleset moves to the tier collectors

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-09 04:28:41 +00:00
devin-ai-integration[bot]
50ab9a3211
fix(token_counter): price a base64 PDF document per page instead of as one image (#45301)
* fix(token_counter): price a base64 PDF document per page instead of as one image

A `document` or `file` block carrying inline PDF bytes was priced like a single
image (85 tokens), so when a provider's count-tokens endpoint rejected the
model (Bedrock Opus) the local fallback answered 116 for a 12-page PDF the
provider then billed at 35941 input tokens. The counter now reads the PDF with
pypdf and prices each page as its extracted text plus the image Anthropic
renders it to (1568 px long edge, 1.15 MP, 750 pixels per token), falling back
to the old image pricing when pypdf is missing or the bytes are not a readable PDF.

* refactor(token_counter): count PDF pages as they are read

Sum each page's text and rendered-image tokens straight from the pypdf
reader instead of materializing a page list first, keep the fallback
to image pricing atomic when a page cannot be read, and annotate the
new tests' locals as Final

* test(integration): audit cells for page-priced PDF documents in count_tokens, pre-call checks and spend

* test(integration): release the held peer when the concurrent budget wait times out

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 21:20:46 -07:00
berriai-litellm-provider-info-sync[bot]
4df5006f29
fix(bedrock): add the claude-sonnet-4-5 EOL date from the Bedrock model card (#45496)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 20:53:52 -07:00
joshua-berri
04f0ea8c33
feat(mcp): translate input requests and bind continuations (#45464)
Some checks failed
Unit Tests / misc (push) Blocked by required conditions
Unit Tests / misc-dirs (push) Blocked by required conditions
Unit Tests / proxy-auth (push) Blocked by required conditions
Unit Tests / proxy-hooks-client (push) Blocked by required conditions
Unit Tests / proxy-infra (push) Blocked by required conditions
Unit Tests / proxy-infra-root (push) Blocked by required conditions
Unit Tests / proxy-server (push) Blocked by required conditions
Unit Tests / unit (push) Blocked by required conditions
Unit Tests / responses-caching-types (push) Blocked by required conditions
Unit Tests / Lens Python 3.10 (push) Waiting to run
Unit Tests / assert-shard-coverage (push) Waiting to run
Unit Tests / auth-checks (push) Blocked by required conditions
Unit Tests / budgets (push) Blocked by required conditions
Unit Tests / custom-logging (push) Blocked by required conditions
Unit Tests / db-and-spend (push) Blocked by required conditions
Unit Tests / endpoints-and-responses (push) Blocked by required conditions
Unit Tests / guardrails-hooks (push) Blocked by required conditions
Unit Tests / jwt-and-keys (push) Blocked by required conditions
Unit Tests / key-generation (push) Blocked by required conditions
Unit Tests / logging-misc (push) Blocked by required conditions
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Code Quality Checks / code-quality (push) Has been cancelled
Unit Tests: Documentation Validation / documentation (push) Has been cancelled
Code Quality Checks / python-310-import-smoke (push) Has been cancelled
CI Coverage / assert-ci-coverage (push) Has been cancelled
Helm unit test / unit-test (push) Has been cancelled
Lens Worker Image / lens-worker-image (amd64, ubuntu-latest) (push) Has been cancelled
Lens Worker Image / lens-worker-image (arm64, ubuntu-24.04-arm) (push) Has been cancelled
UI Unit Tests / ui-unit-tests (push) Has been cancelled
Lens Worker Image / Publish Lens development index (push) Has been cancelled
* feat(mcp): translate input requests and bind resumable continuations

* fix(mcp): keep continuations bound to their original upstream

* fix(mcp): sync advertised protocol API schema

* test(mcp): run interaction regressions in GitHub CI

* test(mcp): set source paths for interaction proxy processes

* test(mcp): verify salt guidance across interaction carriers

* fix(mcp): preserve interaction authentication and elicitation policy

* fix(mcp): bound preflight bodies and preserve challenge routes

* test(mcp): type interaction regression fixtures

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-08 20:10:10 -07:00
devin-ai-integration[bot]
0ae62c5d02
feat(ui): add the evaluation mode to the Add Model form (#45481)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 19:53:27 -07:00
berriai-litellm-provider-info-sync[bot]
ea43e27485
fix(bedrock): take gpt-6.1-sol context window from the Bedrock model card (#45488)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 19:50:47 -07:00
devin-ai-integration[bot]
deaf88e1a4
fix(ci): import seed_tracing_fixtures from the pytest scripts path in rust trace tests (#45485)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 19:41:22 -07:00
berriai-litellm-provider-info-sync[bot]
2c29e9c360
fix(bedrock): add gpt-6.1-sol ultrafast tier prices from the Bedrock model card (#45482)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 19:28:25 -07:00
devin-ai-integration[bot]
fe1e8d1182
fix(router): match deployment pricing ids against the cost map only within the deployment's provider (#45472)
* fix(router): match deployment pricing ids against the cost map only within the deployment's provider

A deployment model_info.id that equals another provider's catalog key
(e.g. baseten/zai-org/glm-5.2 on an openai-compatible api_base) was merged
into that provider's built-in row by register_model, keeping
litellm_provider=baseten on the entry. _check_provider_match then rejected
the row at request time and the deployment billed $0. register_model now
takes a keyword-only custom_llm_provider used to scope
_get_builtin_model_info_for_registration and
_resolve_builtin_model_cost_entry, and the router passes the deployment's
provider through. The stored entry stays provider-less for
non-colliding ids.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(utils): pass custom_llm_provider straight to the registration lookup

Drop the per-entry lookup-provider expression in register_model per review:
the kwarg alone scopes _get_builtin_model_info_for_registration, and
_resolve_builtin_model_cost_entry keeps its main signature and caller.
_register_custom_pricing_for_request passes the provider through so
router-originated per-request registrations get the same scoped lookup.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(utils): drop docstring and simplify set restore in per-request collision test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(router): type the colliding-id deployment fixture

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 19:20:21 -07:00
joshua-berri
911aff2752
refactor(sdk): separate core AWS and tokenizer dependencies (#44447)
* refactor(sdk): separate core AWS and tokenizer dependencies

* fix(sdk): preserve runtime tokenizer alias compatibility

* test(sdk): compare tokenizer installs through existing entry point

* fix(sdk): preserve bearer headers and optional dependency interfaces

* fix(sdk): retain safe tokenizer fallback diagnostics

* test(sdk): inject tokenizer dependency for fallback diagnostics

* fix(core): preserve AWS dependency errors with retries

* fix(core): centralize optional AWS dependency handling

* refactor(core): reuse optional import helper for Invoke streams

* test(core): compare complete shared dependency requirements

* fix(core): keep tokenizer logging import compatible with main

* fix(core): preserve native token decoding and Responses dependency errors

* test(core): isolate optional dependency import failures

* fix(sdk): retain typed exception message access

* test(sdk): scope HTTP verification environment changes

* fix(sdk): preserve bearer request typing across Bedrock callers

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-08 19:07:57 -07:00
devin-ai-integration[bot]
79c3de46cf
fix(bedrock): remove the bare openai.gpt-6.1-sol cost-map row that AWS cannot invoke (#45479)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 18:52:02 -07:00
devin-ai-integration[bot]
390bea6a53
fix(proxy): save file details for every batch output file so they list and retrieve (#41761)
* fix(proxy): register bedrock batch output files with a file object so they list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): guard output file size when content is missing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover managed batch output file listings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): persist metadata for batch output files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover batch output file listing and retrieval end to end

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve provider metadata for managed batch files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): update managed batch output registration mocks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): inject fake prisma client in managed batch output file tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): retry provider file details and refresh fallback batch output entries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover fallback batch output file details refresh after provider recovers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover model-name managed batch file retrieval

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(managed_files): bound the batch output file lookup so a slow provider cannot stall GET /v1/batches

* fix(managed_files): skip the provider lookup for a fallback entry written moments ago

* fix(proxy): safely refresh managed batch file details

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep batch listing provider-free

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): return metadata for empty S3 objects

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): strengthen batch fallback regression assertions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(proxy): describe provider reads and DB writes for batch output files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): sanitize caller-derived ids in managed file logs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(proxy): simplify batch and file endpoint docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate API types from proxy OpenAPI spec

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(test): resolve managed files lint violations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): save refreshed managed file details through ManagedFileRepository

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): classify retrieved file purpose by its own bucket prefix

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(proxy): move batch and file endpoint behavior notes to litellm-docs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): resolve strict lint violations in file retrieval

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): mark retrieval request metadata handoffs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve managed batch type-check diagnostics

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): avoid duplicate final response bindings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Yucheng He <yucheng@berri.ai>
2026-10-08 18:51:45 -07:00
joshua-berri
f48d837cd2
feat(sdk): build core from an independent packaging manifest (#44340)
* feat(sdk): build core from an independent packaging manifest

* test(sdk): compare rebuilt core payload without generated SBOM identity

* fix(sdk): retain native build configuration in the core sdist

* test(sdk): collect coverage from the core build entry point

* fix(ci): isolate core packaging coverage configuration

* fix(test): identify installed core metadata on Python 3.10

* fix(packaging): preserve Git ignore rules in core staging

* fix(packaging): support source-only core builds

* fix(packaging): reject overlapping SDK distributions

* fix(packaging): align core requirements with current main

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-08 18:40:14 -07:00
berriai-litellm-provider-info-sync[bot]
70fd2c19cf
feat(bedrock): add twelvelabs pegasus 1.5 inference profiles (#45477)
* feat(bedrock): add twelvelabs pegasus 1.5 inference profiles

Price-Sync: litellm-providers

* test(bedrock): whitelist twelvelabs pegasus 1.5 invoke profiles

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: Kerry Lu <klu@berri.ai>
2026-10-08 18:32:50 -07:00
devin-ai-integration[bot]
c5c5154c4b
fix(cli): pin pi compat flags so lite pi stops sending store to Anthropic models (#40739)
* fix(cli): pin pi compat flags for gateway models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cli): annotate pi compatibility field

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cli): clarify pi compatibility suppression

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cli): drop the stale mutable-ok suppression on the pi compat block

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 18:30:27 -07:00
berriai-litellm-provider-info-sync[bot]
591637d475
chore(cost-map): sync openrouter prices from the models API (#45469)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 17:54:22 -07:00