Commit graph

46981 commits

Author SHA1 Message Date
mateo-berri
d3ff59fd8a fix(cost-map-sync): keep the reconcile step green when arming auto-merge fails
A failed `gh pr merge --auto` in the fallback arm aborted the step under
`set -e` before the warning and before `sync` was written, so every later
tick went red on the same PR. The failure now prints a warning naming the
PR for a human to merge and the tick carries on.
2026-09-05 00:47:33 -07:00
mateo-berri
b2402f8abf fix(guard-main-branch): accept a cost map sync branch only when the sync bot opened the PR
A `litellm_cost_map_sync_*` head now also needs a Bot author to pass the
main guard, so a person cannot borrow the prefix to route a change past
`litellm_internal_staging`. The error names the author type it saw.
2026-09-05 00:22:33 -07:00
mateo-berri
340569b09a fix(cost-map-sync): keep the bot merging without a ruleset bypass and stop pinning the prices it owns
The reconcile only looks at bot PRs against the branch the run is on, so a
dispatch from another branch never counts as the open sync PR. When a green
bot PR cannot be merged because the app is not a bypass actor yet, the run
arms auto-merge and warns instead of failing every tick. A base with no
required checks falls back to every check so a dispatch there can still
merge, and a tick whose catalogs and base match the last no-op sync skips
the install and the script.

guard-main-branch accepts litellm_cost_map_sync_* heads so the bot keeps
working once main is the default branch again.

The tests that pinned exact OpenRouter prices and limits now check that the
entries exist and are priced: the sync owns those values, and a pin would turn
every legitimate reprice into a red bot PR that pauses syncing.
2026-09-05 00:11:12 -07:00
mateo-berri
26f74e662d fix(cost-map-sync): write every Vercel tier list at the row's shared breakpoints
The cost calculator walks the input_cost_per_token_above_* thresholds and reads
the output and cache prices at the same threshold, so an output or cache tier
that broke at a breakpoint no input tier had was stored and never billed. Each
list is now written at the union of the row's breakpoints, priced from the tier
that covers that breakpoint. Today's 41 tiered catalog rows are aligned, so the
synced map is byte-identical; the regression test bills a mismatched row through
generic_cost_per_token
2026-09-04 23:36:41 -07:00
mateo-berri
9478a32d22 fix(cost-map-sync): reconcile only sync PRs the App opened from this repo
The reconciler picked the open sync PR by title and branch prefix alone,
so a fork PR carrying the same title and a litellm_cost_map_sync_ branch
could pass the guard with its own repricing and be merged with the App
token. It now lists PRs authored by the App (--author app/<slug>) and
drops cross-repository heads, and without the App it never selects a PR,
which also removes the unreachable no-App warning branch.
2026-09-04 23:13:18 -07:00
mateo-berri
58b3037e4b fix(sync-cost-map): map vercel tiers, inherit family traits, hold out-of-bounds changes, and reconcile the open bot PR
Vercel long-context tiers become *_above_<N>k_tokens keys when contiguous on a whole thousand, and a row whose tiers do not fit is skipped with a warning. Image and audio output are priced per token, and a row with an unpriced non-text output is skipped instead of billed as text. A new entry inherits the traits no catalog expresses (cache minimum, adaptive thinking, sampling params, system messages, thinking always on) from its same-mode root, found by the bare name or its longest dash prefix. The max_tokens / max_output_tokens pair moves as a unit. Shrinking limits, prices crossing zero or moving more than 10x, and every price on an already-priced varies_by_provider row are held back and listed as warnings for a human commit. Updated entries keep their curated key order with new keys appended sorted.

The workflow's own token is read-only and every write uses the GitHub App token; without the App a scheduled run explains why it cannot open a PR. Each tick first reconciles the open bot PR: a conflicting one is closed and re-synced, a green one is merged, a red one is left for a human, and a sync only runs when none is open. The sync step runs with --no-dev and only when it will be used.

The hardcoded map schema in test_utils.py gains the 32k tier keys the synced map now carries.
2026-09-04 23:00:08 -07:00
mateo-berri
a13daf7c5c fix(ci): print the uncapped sync report to the workflow log so capped warnings stay findable 2026-09-04 19:26:09 -07:00
mateo-berri
f41b19a120 fix(ci): cap each sync PR body section so a large first sync stays under GitHub's body limit 2026-09-04 19:17:58 -07:00
mateo-berri
5cc9690cb5 fix(sync-cost-map): reject non-finite catalog prices and rename the guard in the reasoning-effort docstring 2026-09-04 18:56:15 -07:00
mateo-berri
d9e7448938 feat(ci): add the cost map sync bot for openrouter and vercel_ai_gateway 2026-09-04 18:34:37 -07:00
mateo-berri
61bed79566 feat(ci): add the cost map guard check
Replace test-model-map.yml with a pull_request_target guard that validates the
cost map, its backup, and its generated schema on every PR, and additionally
enforces the sync bot contract on litellm_cost_map_sync_* branches: only the
three cost map files may change, no model or field is removed, and the special
root keys stay untouched.
2026-09-04 17:18:17 -07:00
tin-berri
b3c867c7b2
fix(auto_router): derive tier definitions in prompt editor (#39688) 2026-09-04 16:48:33 -07:00
ryan-crabbe-berri
d23bec84c4
Merge pull request #39196 from BerriAI/litellm_guardrail_usage_cost_rollup
feat(guardrails): roll up Bedrock guardrail cost per usage counter
2026-09-04 15:21:16 -07:00
tin-berri
6dff3a5f72
fix(complexity_router): fall back to a live peer when the decided tier model is fully cooled down (#39675)
A complexity tier can name several model groups, but the pool pick and the session-pin
replay both returned a group without consulting deployment health, so a group whose every
deployment was in cooldown was still routed to and the request died at the router's
zero-deployment check while a healthy peer sat in the same tier.

Gate the decided response at the pre-routing hook's exits, the seam the modality gate
already occupies, so every arm that can place a request is covered by one owner: a fresh
classification, a replayed or escalated pin, a plan-mode floor, a context-window
escalation, an adaptive pick, and whatever arm is added next.

Peers come from the decided tier only. Climbing to a higher tier costs more than the
classifier asked for and is left to a follow-up. The gate fails open on every uncertainty:
an unreadable cooldown view, a decision carrying no tier, a group the router knows no
deployments for, or a tier whose peers are all cooling.
2026-09-04 22:11:50 +00:00
ryan-crabbe-berri
df68edca76
Merge pull request #39808 from BerriAI/litellm_lit_5379_jwt_mapping_cache_invalidation
fix(jwt): invalidate JWT key mapping cache on /key/regenerate
2026-09-04 14:53:19 -07:00
Mateo Wang
922659fb15
Merge pull request #39388 from BerriAI/litellm_registry_audit_2026_09_02
fix(model_prices): verified registry audit, Databricks Sep-2026 catalog, realtime image pricing, deprecation dates
2026-09-04 14:51:40 -07:00
ryan-crabbe-berri
9bd34adb6d Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_guardrail_usage_cost_rollup
# Conflicts:
#	type-discipline-budget.json
2026-09-04 14:41:55 -07:00
ryan-crabbe-berri
f59021f0e0
Merge pull request #39807 from BerriAI/litellm_lit_6690_relax_routing_group_name
fix(ui): accept any routing group name the backend accepts
2026-09-04 14:39:37 -07:00
ryan-crabbe-berri
6c81a5c423 feat(guardrails): store untracked units on the rollup row instead of nulling cost
A row that received both priced and unpriced increments used to collapse
to cost NULL, throwing away the priced subtotal and making every unit on
it read as untracked. The rollup now carries a second column,
untracked_units, that the aggregator increments for units with no known
price while cost keeps accruing for the rest, so cost covers exactly
units - untracked_units. Rows written before the migration keep cost
NULL and still read as untracked in full

The endpoints read untracked units off the column (or the whole row for
a legacy NULL) rather than from a NULL filter, and the policies overview
now fills totalUntrackedUsageUnits, which the previous commit missed

Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
2026-09-04 14:38:08 -07:00
moe-berri
336891bbc7
Merge pull request #39696 from BerriAI/litellm_bound_classifier_timeout
fix(router): bound auto-router classifier latency
2026-09-04 14:37:52 -07:00
ryan-crabbe-berri
49cac6fc5b fix(jwt): evict mapping cache after DB write in /jwt/key/mapping update and delete
Evicting before the mutation commits left a race: a concurrent JWT
request could re-cache the old mapping between the eviction and the
commit, keeping a deleted or renamed claim authorized until the cache
TTL expired. Flagged by review on PR #39808.
2026-09-04 14:35:42 -07:00
yuneng-jiang
2849aee57d
fix(health): probe test_connection with the credential the request names (#39801)
* fix(health): probe test_connection with the credential the request names

/health/test_connection matches the request's model string against the
configured deployments and merges the match's litellm_params underneath the
request. A request that named a stored credential but no key of its own
still satisfied the "request sets no connection fields" test, so it inherited
the matched deployment's api_key and api_base, and load_credentials_from_list
then skipped the named credential because api_key was already set.

A wildcard route covering the model is enough to match, so the Add Model
page's Test Connect probed with an unrelated deployment's key while echoing
back the credential that was selected.

Naming a credential the configuration does not name now withholds the
configuration's credential fields, the same set already withheld from a
request that supplies its own endpoint. Naming no credential still inherits
them, as documented.

* test(health): drop test docstrings that restate their own names

* test(health): assert the credential probe on the wire, not on the call args

The connection-test regressions patched litellm.ahealth_check and read the
params handed to it. Driving the endpoint through the app with respx faking
the upstream instead lets the real credential resolution run, so the tests
assert the key and host that actually go out, which is what the bug was about.

It also drops three of the five patched proxy internals; the two that are left
are proxy-global wiring with no injection seam, the same ones the image_edit
connection test already has to reach for.

* chore(ui): regenerate schema.d.ts for the test_connection docs change
2026-09-04 14:18:41 -07:00
ryan-crabbe-berri
52b746e8ea test(key): annotate regenerate JWT mapping test patches for TQ008 2026-09-04 14:18:09 -07:00
ryan-crabbe-berri
2f7ee39545 fix(jwt): invalidate JWT key mapping cache on /key/regenerate
/key/regenerate carries the JWT-to-key mapping to the new token via FK
cascade, but the jwt_key_mapping cache entry kept resolving the old
(now invalid) token for up to virtual_key_mapping_cache_ttl. Snapshot
the key's mapping cache keys before the token update and evict them
with evict_and_broadcast so every worker drops the stale entry.

Also share the cache-key format through jwt_key_mapping_cache_key and
upgrade the /jwt/key/mapping CRUD endpoints from local-only deletes to
evict_and_broadcast, closing the same cross-worker staleness there.
2026-09-04 14:11:31 -07:00
ryan-crabbe-berri
1548be8235 feat(guardrails): report the usage units a guardrail's cost leaves out
A row's cost sums only the daily rows that carry a tracked cost, so it
silently under-reports whenever some rows are NULL (pre-migration days,
old pods mid-rollout, an unpriced counter). Both usage endpoints now
return the per-counter units behind those NULL rows next to the cost
(untrackedUsageUnits / totalUntrackedUsageUnits on the overview,
untracked_usage_units on the detail), so a partial cost is never mistaken
for a complete one and the reader can see exactly what it excludes

Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
2026-09-04 14:09:59 -07:00
ryan-crabbe-berri
6f18a4d81e fix(ui): accept any routing group name the backend accepts
The create form rejected names with slashes or spaces even though the
proxy stores and routes any non-empty string. Drop the client-only
character pattern and trim the name before the required check so a
whitespace-only name is still refused

Claude-Session: https://claude.ai/code/session_01HkaXiD6gssHnx3kqu1rR8C
2026-09-04 14:06:34 -07:00
Mateo Wang
300d335255
Merge pull request #39361 from BerriAI/litellm_fix_mantle_host_re_anchor
fix(bedrock_mantle): anchor MANTLE_HOST_RE so custom Mantle hosts are honored
2026-09-04 13:40:02 -07:00
tin-berri
8beca1d58d
fix(auto-router): route 1M complex tier to GPT Sol (#39797)
* feat(ui): add 1M context auto-router preset

* feat(ui): use heuristic v2 for 1M preset

* fix(ui): keep 1M preset test within lint budget

* fix(auto-router): route 1M complex tier to GPT Sol

* test(auto-router): update 1M complex tier expectation
2026-09-04 13:33:06 -07:00
ryan-crabbe-berri
189bd857f4
Merge pull request #39794 from BerriAI/litellm_v2_org_update_public
feat(organization): expose PATCH /v2/organization/{organization_id} in the OpenAPI schema
2026-09-04 13:32:52 -07:00
ryan-crabbe-berri
38a1b44993
Merge pull request #39793 from BerriAI/litellm_v2_org_update_validation
fix(organization): reject negative limits and unparseable budget_duration on PATCH /v2/organization
2026-09-04 13:32:07 -07:00
ryan-crabbe-berri
a502c728ac
Merge pull request #39670 from BerriAI/litellm_fix_org_update_null_budget_limits
fix(organization): clear org budget limits when PATCH /organization/update sends null
2026-09-04 13:31:56 -07:00
Mateo Wang
338a37d8cd
Merge pull request #39632 from BerriAI/litellm_lit6874_fireworks_perplexity_off_peak_pricing
fix(cost): honor off_peak_pricing in the fireworks_ai and perplexity cost calculators
2026-09-04 13:20:52 -07:00
Mateo Wang
44b1cc7b0f
Merge pull request #39589 from BerriAI/litellm_fix_v1_messages_midstream_timeout_failure_logging
fix(proxy): log mid-stream /v1/messages failures as failures with partial usage
2026-09-04 13:20:30 -07:00
moe-berri
6234399f9e fix(router): keep circuit-open fallbacks out of session pins
An open classifier circuit routed through the ordinary heuristic or
classifier_fallback path, and both causes are pin-worthy, so a session
whose turn landed on the cooldown fallback held that model for the whole
session_affinity TTL and never reclassified after the breaker closed.

The circuit-open signal now blocks the pin, and _classifier_failure_outcome
tags its outcomes through one helper instead of reassigning a Final.
2026-09-04 13:16:34 -07:00
ryan-crabbe-berri
7a717740dd test(organization): assert rejected values write nothing to the DB 2026-09-04 13:11:31 -07:00
mateo
50d6b26a86 fix(registry): mark baseten GLM-5.3 as vision-capable per Baseten vision docs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 20:10:58 +00:00
ryan-crabbe-berri
0eb2363074 test(organization): assert route publicity through the production OpenAPI generator 2026-09-04 13:10:37 -07:00
ryan-crabbe-berri
3e4a884b25 feat(organization): expose PATCH /v2/organization/{organization_id} in the OpenAPI schema 2026-09-04 13:03:18 -07:00
ryan-crabbe-berri
976ff0a785 fix(organization): 422 on negative limits and unparseable budget_duration in v2 update 2026-09-04 12:58:33 -07:00
mateo
f5157a63eb test: allow 128k and 256k tiered cache fields in registry schema test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 19:56:30 +00:00
ryan-crabbe-berri
d71fe43c1a
Merge pull request #39682 from BerriAI/litellm_lit_4738_per_user_usage_pagination
fix(ui): paginate per-user usage with the shared server-side DataTable footer
2026-09-04 12:53:12 -07:00
ryan-crabbe-berri
3bdc5ecd0e refactor(ui): drop the per-user usage page clamp now handled by the shared DataTable
The shared DataTable clamps a server-mode page index whenever rowCount no
longer reaches it (#39776), including the empty-dataset case this table's
own clamp skipped because it required total_pages > 0. Remove the local
clamp and cover the empty case through the component so the wiring into
the shared behavior is what the tests prove
2026-09-04 12:47:37 -07:00
ryan
7fde31fe08 fix(ui): fall back to the last page when per-user usage shrinks under the current page
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 12:45:26 -07:00
ryan
fd42bddee6 fix(ui): reset per-user usage page in the same render as the tag filter change
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 12:45:26 -07:00
ryan
11f272e08b fix(ui): paginate per-user usage with the shared server-side DataTable footer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 12:45:26 -07:00
devin-ai-integration[bot]
205a5e9d6c
feat(mcp): use x-mcp-<access_group>-* headers as default upstream credentials for group members (#39717)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 12:45:24 -07:00
ryan-crabbe-berri
62087c5d5f
Merge pull request #39776 from BerriAI/litellm_datatable_server_page_clamp
fix(ui): clamp server-paginated DataTable page index when rowCount shrinks
2026-09-04 12:43:38 -07:00
yuneng-jiang
4ad4db22f1
Merge pull request #39773 from BerriAI/litellm_/pr-39770-test-failure-c4aacd
test(caching): drive the redis stall burst off the clock, not asyncio.wait_for
2026-09-04 12:33:01 -07:00
yuneng-jiang
150182d76a
Merge pull request #39772 from BerriAI/litellm_ci_remove_pytest_x_flag
ci: report every failing test in a job instead of stopping at the first
2026-09-04 12:32:50 -07:00
mateo
1f0611a8b9 fix(registry): drop Together MiniMax M2.7 and revert Qwen2.5 7B Turbo pricing, both non-serverless
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 19:25:10 +00:00