Commit graph

45221 commits

Author SHA1 Message Date
ryan-crabbe-berri
7e7ac69258
test: gate the test tree on fifteen assertion and handler rules it already satisfies (#38361) 2026-08-26 16:05:34 -07:00
Mateo Wang
9e6d9e5964
Merge pull request #38417 from BerriAI/litellm_image_edit_health_probe_moderation_safe
fix(health): make the image_edit health probe moderation-safe
2026-08-26 16:03:14 -07:00
Mateo Wang
855a8bc764
Merge pull request #36397 from ousamabenyounes/litellm_fix_gemini_web_search_unique_queries_36377
fix(vertex_ai): bill Gemini grounding per unique web search query
2026-08-26 16:02:07 -07:00
mateo-berri
ab1b7bf3b6 fix(cost): price gemini-live-2.5-flash-native-audio realtime sessions
The GA vertex model had no cost map entry, and the realtime cost handler
accepted the router's price-less auto-registered deployment entry for the
session.created model at zero-defaulted rates, so sessions billed 0.0 even
when base_model pointed at the priced preview key. Adds the GA entry at its
published rates and makes the handler fall through zero-defaulted candidates
unless their cost map entry explicitly declares pricing.
2026-08-26 15:53:37 -07:00
mateo-berri
2f796530d0 test(health): parse the probe PNG without mutation 2026-08-26 15:50:51 -07:00
mateo-berri
31f0d82f00 fix(ptu): zero the Maps grounding rate on PTU deployments 2026-08-26 15:50:36 -07:00
mateo-berri
6df307fef8 fix(prompts): validate a prompt replacement before swapping and isolate per-row sync failures 2026-08-26 15:49:31 -07:00
mateo-berri
9003b02c3c Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_gemini_web_search_unique_queries_36377
# Conflicts:
#	tests/test_litellm/llms/vertex_ai/gemini/test_vertex_and_google_ai_studio_gemini.py
2026-08-26 15:39:17 -07:00
Mateo Wang
f7220556e1
Merge pull request #38395 from BerriAI/litellm_gemini_live_voice
fix(gemini-realtime): keep the client's voice on Vertex AI native-audio Live
2026-08-26 15:34:53 -07:00
mateo-berri
54b57575d7 test: restore model cost map via monkeypatch 2026-08-26 15:34:25 -07:00
Hamza Shah
5d4f8b36a6
fix(fireworks_ai): stop using the trace id as the session affinity key (#35754)
get_fireworks_session_id fell back to litellm_trace_id when no session id was
given. That id is generated per request (uuid4 when absent), so x-session-affinity
carried a different value every time and Fireworks prompt caching never hit;
cached_tokens stayed 0 across identical prompts.

The None path the original change described was effectively unreachable because
of it. Drop the fallback so affinity comes only from an id the caller actually
supplied: litellm_session_id, session_id, or metadata.session_id.

Callers who were relying on a trace id for affinity can pass litellm_session_id
instead, which is stable across the requests they want grouped.

Co-authored-by: mubashir1osmani <mubashir.osmani777@gmail.com>
2026-08-26 18:33:56 -04:00
mateo-berri
241daa4cb7 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_gemini_maps_grounding_cost
# Conflicts:
#	type-discipline-budget.json
2026-08-26 15:33:28 -07:00
Mateo Wang
a8f3a74360
Merge pull request #38397 from BerriAI/litellm_deepseek_vision_forwarding
fix: forward image content lists to DeepSeek vision models
2026-08-26 15:32:25 -07:00
mateo-berri
ce8d6f7d25 fix(mcp): token refresh and M2M egress honor the admin-entered token URL 2026-08-26 15:31:45 -07:00
mateo-berri
b9a790899a fix(gemini): bill Google Maps grounding as its own SKU
Gemini API Maps-grounded prompts were billed as web search and Vertex AI Maps-grounded prompts were not billed at all. Classify grounding metadata per candidate into web search vs Maps requests, carry a distinct google_maps_grounding_requests usage counter through non-streaming and streaming paths, and price it via the new google_maps_grounding_cost_per_query cost map key with per-query and per-prompt defaults keyed off web_search_billing_unit. Fixes #35906
2026-08-26 15:31:27 -07:00
mateo-berri
4d1d7b446f test(speech): type the bridge spend regression test helpers 2026-08-26 15:27:38 -07:00
mateo-berri
416984a5e0 fix(health): make the image_edit health probe moderation-safe
The image_edit health probe sent a 512x512 solid-gray PNG with the generic
chat prompt "test from litellm", an ambiguous pair OpenAI's gpt-image-1
output moderation sometimes rejects as moderation_blocked, which reported a
working deployment as unhealthy. The probe now sends a blue circle on a
white background with a descriptive edit prompt, and a provider moderation
verdict (ContentPolicyViolationError or a moderation_blocked error body) is
treated as proof the endpoint works rather than as an unhealthy deployment.
2026-08-26 15:23:08 -07:00
mateo-berri
fb13b47ee5 test: type the pricing test helpers 2026-08-26 15:21:24 -07:00
mateo-berri
fd751a5023 fix(proxy): key lazy openapi stubs off registered features, not sys.modules 2026-08-26 15:18:55 -07:00
mateo-berri
5130bafda8 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_prompt_patch_sync 2026-08-26 15:18:11 -07:00
mateo-berri
d565860f60 fix(cost-map): correct gemini-live native-audio text input rate 2026-08-26 15:15:05 -07:00
ryan-crabbe-berri
52b7bea6f3
Merge pull request #37708 from BerriAI/litellm_team_member_budget_no_reset
fix(team): allow no-reset default budgets for team members
2026-08-26 15:14:30 -07:00
mateo-berri
6918214266 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_dotprompt_model_swap
# Conflicts:
#	tests/test_litellm/integrations/dotprompt/test_prompt_manager.py
2026-08-26 15:13:57 -07:00
mateo-berri
3418d7baf9 fix(speech): keep proxy metadata and completion cost through the TTS completion bridge 2026-08-26 15:12:41 -07:00
mateo-berri
f824ca7433 fix(responses): run prompt hook before provider credential resolution in sync responses() 2026-08-26 15:08:23 -07:00
Mateo Wang
632a007967
Merge pull request #38406 from BerriAI/litellm_fix_db_router_settings_overwrite
fix(proxy): stop empty DB router_settings lists from clobbering yaml fallbacks
2026-08-26 15:06:26 -07:00
mateo-berri
aabbc3204b fix(gemini-realtime): drop the native-audio speechConfig strip on Google AI Studio too
Live probes against every gemini_native_audio model on both providers show
setup accepts a valid prebuilt voice and 1007s only unknown voice names, so
the strip predicate rested on a false premise and silently discarded the
client's voice on AI Studio native-audio sessions
2026-08-26 15:06:23 -07:00
mateo-berri
cd9dcb55b4 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_regenerate_lazy_openapi_snapshot 2026-08-26 15:04:13 -07:00
mateo-berri
b687fe2b50 fix(prompts): propagate PATCHed prompt templates to every worker and pod 2026-08-26 15:03:32 -07:00
Yuneng Jiang
3576c773eb
fix(proxy): serialize the search tool router refresh
reload_search_tools_from_db is a read-modify-write of the shared llm_router
global: it reads the whole table, merges the config tools in, and replaces
router.search_tools wholesale. Two of those interleaving lets the older
snapshot's assignment land last and put back a tool the newer one deleted, so a
revoked tool keeps serving on the provider key it carried until the next reload.

Take MODEL_RECONCILE_LOCK, which add_deployment already uses to serialize the
same shape of work on the same global. It has to go on this entry point rather
than in _init_search_tools_in_db, because _init_non_llm_objects_in_db calls that
while already holding the lock and asyncio.Lock is not reentrant.

A separate search-tools-only lock would not close the race: the periodic
reconcile reaches _init_search_tools_in_db under MODEL_RECONCILE_LOCK, so only
that same lock orders an endpoint refresh against a cron tick.

Ordering across workers is unchanged and still reconciles on the next tick.
2026-08-26 15:02:19 -07:00
tin-berri
f0a122d35f
refactor(ui): read the auto-router tier set through one row list (#38408)
The dashboard resolved the complexity-router tier set three different ways: a
private TIER_KEYS in build_complexity_router_config.ts, TIER_ORDER in
complexity_router_tiers.ts, and TIER_KEYS in ComplexityRouterConfig.tsx. The
edit modal went further and re-implemented the whole create payload builder,
kept in sync only by a comment reading "Mirrors buildComplexityRouterConfig".

tier_rows.ts now owns the tier set. Every consumer reads activeTierRows(value)
and a row carries its own id, so the plan-mode floor and per-model params point
at a row rather than at a position, and the leaves that already wanted entries
(buildAutoRouterTestTargets, getRequiredModels, model_info_view) take them.
buildUpdatedComplexityRouterConfig becomes preserve-unmanaged-keys around the
shared builder instead of a second copy of it.

Also drops the literal ", ]" that renders as visible text in two DialogFooter
blocks on the auto-router routing-test and connection-test dialogs, left over
from a JSX array-to-fragment conversion.

No behaviour change: all 566 tests over the touched modules pass with fixture
shape changes only, no assertion edited.
2026-08-26 14:53:06 -07:00
yuneng-jiang
4148cf7d7d
Merge pull request #38309 from BerriAI/litellm_azure_ai_unprocessable_retry_tests
test(azure-ai): pin the 422 retry that drops the field the provider rejected
2026-08-26 14:52:41 -07:00
Mateo Wang
77bce45100
Merge pull request #38405 from BerriAI/litellm_lit6243_prompt_cache_min_tokens
fix(cost-map): correct prompt_cache_min_tokens for Claude Fable 5 and backfill Anthropic re-export entries
2026-08-26 14:51:56 -07:00
Mateo Wang
9888830207
Merge pull request #38344 from ksk2023/fix-cost-alias-double-prefix
fix(cost_calculator): resolve real cost key when model_name alias contains '/'
2026-08-26 14:51:53 -07:00
Mateo Wang
16e9efccaf
Merge pull request #38404 from BerriAI/litellm_fix_prompt_data_double_nest
fix(prompts): reject keyed prompt_data with prompt_id and populate prompt version
2026-08-26 14:47:59 -07:00
mateo-berri
e9f3963869 refactor(proxy): build snapshot fragments immutably to satisfy the type-discipline gate 2026-08-26 14:44:31 -07:00
mateo-berri
2548e960f1 fix(cost-map): correct Gemini TTS and native-audio rates
Gemini 2.5 Flash Preview TTS, Gemini 2.5 Pro Preview TTS, and the three
gemini-2.5-flash-native-audio entries carried rates copied from the text
models, so audio output was billed 2x to 6x under Google's published
prices. Set the published per-token rates on all ten keys, add
output_cost_per_audio_token to the native-audio entries, and drop the
long-context tier rates Google does not publish for Pro TTS.
2026-08-26 14:43:21 -07:00
mateo-berri
898ff74673 refactor(proxy): type the snapshot fragments and wrap a long test line 2026-08-26 14:39:51 -07:00
mateo-berri
3b9f6ee2aa fix(proxy): apply empty DB router_settings lists only where yaml sets no value 2026-08-26 14:37:29 -07:00
mateo-berri
afe5a240e5 fix(proxy): regenerate lazy OpenAPI snapshot and guard it in CI
The committed snapshot behind /openapi.json for unloaded lazy features had drifted on 30 of 31 fragments and never had one for a2a_registration or gemini_agents, so those routes showed as placeholder GET stubs or old docstrings until traffic loaded them. Regenerate the snapshot and schema.d.ts, make the check-ui-api-types job and make check regenerate the snapshot and fail on drift, and make the generator refuse to write a snapshot when any feature fails to import so a broken import cannot silently drop fragments.
2026-08-26 14:32:04 -07:00
mateo-berri
f772cad959 fix(cost-map): backfill prompt_cache_min_tokens for the remaining Claude 4.x re-export entries 2026-08-26 14:31:45 -07:00
mateo-berri
43b8ed0fd9 fix(prompts): validate only the litellm_params a PATCH sends 2026-08-26 14:20:26 -07:00
mateo-berri
dbc819dc77 fix(prompts): apply prompt templates before routing on /v1/responses and honor ignore_prompt_manager_model
On /v1/responses the prompt template ran inside litellm.aresponses, after the
router had already resolved a deployment and injected its api_key/api_base, so a
prompt whose metadata.model pointed at another provider sent the old
deployment's credentials cross-provider (401). The proxy now runs the prompt
template for aresponses in the pre-call hook, before routing, so the router
picks the deployment that matches the swapped model. As a backstop, the SDK
refuses a cross-provider swap when explicit credentials are already present
instead of forwarding them.

ignore_prompt_manager_model and ignore_prompt_manager_optional_params saved on
a prompt were only read by the generic manager, so dotprompt prompts ignored
them on every endpoint. PromptManagementBase now merges the prompt spec's flags
with the per-request ones for every manager, and the generic manager no longer
drops caller flags when no spec is present.
2026-08-26 14:12:28 -07:00
mateo-berri
f334108f33 docs(prompts): sync lazy openapi snapshot and dashboard schema with the fixed create_prompt example 2026-08-26 14:11:21 -07:00
mateo-berri
951cef1e98 fix(cost_calculator): strip duplicated region segment from alias cost keys 2026-08-26 14:10:05 -07:00
mateo-berri
465ebb1bdd fix(mcp): join discovery for a clientless DCR bridge still missing its registration endpoint 2026-08-26 14:08:09 -07:00
Mateo Wang
f57e4b812c
Merge pull request #38403 from BerriAI/litellm_lit3373_valkey_acl
fix(caching): require the namespace delimiter when checking already-namespaced redis keys
2026-08-26 14:04:51 -07:00
ryan-crabbe-berri
870328f8cc test(team): fake the budget table instead of patching new_budget and update_budget
The three member-budget tests patched litellm internals and asserted only on
the mock, which tripped the TQ002 and TQ008 test-quality ratchet. Fake the
prisma budget table on the shared client and assert on the row that reaches
the database plus the returned team payload.
2026-08-26 14:00:45 -07:00
mateo-berri
caa97eea22 fix(cost-map): correct prompt_cache_min_tokens for Claude Fable 5 and backfill Anthropic re-export entries 2026-08-26 14:00:39 -07:00
mateo-berri
6d1b295dae fix(proxy): stop empty DB router_settings lists from clobbering yaml fallbacks 2026-08-26 13:58:00 -07:00