Commit graph

42521 commits

Author SHA1 Message Date
mubashir1osmani
8bf7dfed24 Merge remote-tracking branch 'berri/litellm_internal_staging' into litellm_playground_realtime
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
# Conflicts:
#	ui/litellm-dashboard/src/components/llm_calls/fetch_models.tsx
2026-08-08 12:32:15 -07:00
devin-ai-integration[bot]
cfd64d45a8
fix(ui): show team BYOK models in team fallback settings (#36241)
* fix(ui): show team BYOK models in team fallback settings

Team router settings loaded fallback options from /model_group/info, which resolves models without a team, so a team's own BYOK deployments were never selectable in its own fallback config. Load the team-scoped listing when a team id is present.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): ignore stale team model responses in router settings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): use react-query for fallback model listing in router settings accordion

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-08-08 19:28:57 +00:00
mubashir1osmani
0cca947803 fix(ui): restore minimal realtime playground client
Drop the custom cancel/queue/dual-buffer state machine and the duplicate
OPEN_AI_REALTIME_VOICES list. Reuse OPEN_AI_VOICE_SELECT_OPTIONS and the
original thin WebSocket event handler (append deltas, response.done
fallback) with the shared ChatComposer shell
2026-08-08 12:25:42 -07:00
mubashir1osmani
af04d3a6e1 fix(ui): stop realtime responses from stalling or splitting bubbles
Text turns use text-only modalities with VAD off, and voice re-enables
audio+VAD only while recording so residual VAD cannot race a second
response. Stream text and transcript into one draft, always update the
last assistant bubble (even past status lines), and unlock with a cancel
timeout so the UI cannot stick mid-response forever
2026-08-08 12:25:42 -07:00
mubashir1osmani
037bbe2e9f fix(ui): queue realtime pending messages and drop new comments
Keep every text submission while a response is cancelled, not only the
latest. Remove newly added source comments called out in review
2026-08-08 12:25:42 -07:00
mubashir1osmani
d9281b48c9 fix(ui): prevent concurrent realtime responses and garbled transcripts
Track in-progress responses so text/mic turns cannot race response.create,
prefer a single transcript stream per turn, and queue sends until the active
response finishes or is cancelled
2026-08-08 12:25:42 -07:00
mubashir1osmani
9056ff27b7 feat(ui): restyle realtime playground with shared chat composer
Move RealtimePlayground off Ant Design onto shadcn controls and the
shared ChatComposer, use the realtime-safe voice list, and reuse the
same composer for Compare message input
2026-08-08 12:25:42 -07:00
mubashir1osmani
b6475bd343 fix(ui): surface model load failures and align request key with models
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Throw when both model endpoints fail so ChatUI can show an error instead of
an empty list. Use the same resolved key for model listing and chat requests,
and drop the fetch_models JSDoc
2026-08-08 12:25:42 -07:00
mubashir1osmani
f40ef53894 fix(ui): show human labels on playground SDK type select
Match the SelectValue children pattern so OpenAI SDK / Azure SDK
render instead of the raw openai/azure values
2026-08-08 12:25:42 -07:00
mubashir1osmani
e72add6882 fix(ui): show human labels on playground virtual key source select
Base UI SelectValue renders the raw value unless children supply the
label, so map session/custom to Current UI Session and Virtual Key
2026-08-08 12:25:41 -07:00
mubashir1osmani
7785796121 fix(ui): load playground models for virtual keys via /v1/models
Prefer the key-scoped OpenAI models list when the Virtual Key source is
selected, enrich with mode from model_group/info, and debounce custom key
input so models appear for the key's access set
2026-08-08 12:25:41 -07:00
mubashir1osmani
170d68f32d fix(ui): size chat composer textarea with CSS field-sizing
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Drop direct el.style.height mutation in favor of field-sizing:content
2026-08-08 12:25:41 -07:00
mubashir1osmani
404cb97d96 style(ui): strengthen playground chat composer border and shadow
Make the shared chat input stand out with a fuller border, layered
shadow, and a slightly stronger focus ring
2026-08-08 12:25:41 -07:00
mubashir1osmani
3a6b2578b7 feat(ui): adopt vercel-style chat composer for playground
Replace the compact single-line input with a PromptInput-style composer:
taller auto-growing textarea, rounded card shell, footer tools, and
stop button while a request is in flight
2026-08-08 12:25:41 -07:00
mubashir1osmani
5e5f2306a7 fix(ui): exclude unknown model modes from playground endpoint filters
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Modes outside ModelMode (batch, rerank, ocr, etc.) must not collapse to
chat-compatible, or conversational endpoints surface unusable models
2026-08-08 12:25:41 -07:00
mubashir1osmani
dad2dc1158 fix(ui): restore playground model filtering by endpoint
Bring back the prior Chat model dropdown filter (including chat models
on responses/anthropic/interactions and image models on image_edits), and
map mode realtime so the realtime endpoint only lists compatible models
2026-08-08 12:25:41 -07:00
mubashir1osmani
f2e3b6a568 fix(ui): drop null guardrail names from MultiSelect options 2026-08-08 12:25:41 -07:00
mubashir1osmani
757d5a4dcd fix(ui): apply coy theme after code props on ReasoningContent 2026-08-08 12:19:17 -07:00
mubashir1osmani
4aff515a8e fix(ui): cast syntax highlighter theme through unknown 2026-08-08 12:15:09 -07:00
mubashir1osmani
38e1fe8ffd fix(ui): extract chat handlers and type MCP tool pool 2026-08-08 12:11:33 -07:00
Mateo Wang
8b16ee1dc2
Merge pull request #36277 from BerriAI/litellm_make_check_fallback
build(lint): rename make pre-commit to make check with a working-tree fallback
2026-08-08 12:08:34 -07:00
mubashir1osmani
01d55b9d0f fix(ui): guard null voice selection from shadcn Select 2026-08-08 12:07:36 -07:00
mubashir1osmani
d78a016bc5 fix(ui): green playground shadcn CI after staging merge
Normalize MCP description nullability for MultiSelect, preserve model
selection via functional setState, and update endpoint/vector-store
selector tests for the shadcn combobox API
2026-08-08 12:01:44 -07:00
mubashir1osmani
0302283eba chore: merge litellm_internal_staging into playground shadcn stack 2026-08-08 11:56:47 -07:00
yuneng-jiang
b0fd3e1e30
Merge pull request #36288 from BerriAI/litellm_sync_main_into_internal_staging
chore(ci): sync main into internal staging
2026-08-08 11:25:43 -07:00
Mateo Wang
f6df762b25
test: roll back live router replay membership between tests (#36278)
Since #35491, every Router joins the module-global _live_routers weak set at
construction, and every model cost map swap replays the deployments of every
member on top of the freshly adopted map. #36039 isolated the register_model
ledger half of that replay but not this half: under pytest-xdist, a Router
created by an earlier test in the same worker that was still referenced (or
simply not yet garbage collected) re-registered its deployments during
TestPriceDataReloadIntegration::test_distributed_reload_check_function, and
register_model hydrated the sparse mocked gpt-3.5-turbo entry into a full
ModelInfo dict, failing the exact-equality assert (reruns cannot help since
the polluting router survives in the worker process)

The autouse isolate_litellm_state fixture now snapshots _live_routers before
each test and restores its membership on teardown, so a test's routers stop
contributing to cost map rebuilds once the test ends. A canary pair in
test_conftest_isolation.py asserts the rollback
2026-08-08 10:45:43 -07:00
Yuneng Jiang
09323fcc4a
chore(ci): sync main into internal staging 2026-08-08 10:42:07 -07:00
mateo-berri
fb7861fbfd build(lint): count deleted files toward check triggers 2026-08-08 10:41:06 -07:00
Mateo Wang
4d9defd573
Merge pull request #36282 from BerriAI/litellm_decrease_anys_fable3
chore(typing): clear 1.4k basedpyright Any errors across 21 hotspot files
2026-08-08 10:36:46 -07:00
Mateo Wang
1d0cba7f7c
Merge pull request #35551 from BerriAI/devin_ai_require_managed_files_read_paths_35530 2026-08-08 10:22:18 -07:00
Mateo Wang
8fdb1c1cf2
Merge pull request #36161 from BerriAI/litellm_ruff_external_strict_rules 2026-08-08 09:21:13 -07:00
mateo-berri
20eb7bb437 chore(typing): clear 1.4k basedpyright Any errors across 21 hotspot files
Typing-only pass over the 21 files with the highest reportAny and
reportExplicitAny density among self-contained modules: management
endpoints, guardrails, streaming internals, response transformations,
MCP server, enterprise managed files, and vector store management.

Whole-tree basedpyright drops from 148,648 to 146,984 errors (-1,664),
with reportAny -1,111 and reportExplicitAny -296. No rule increased
repo-wide and no file regressed on any rule. No cast(), type: ignore,
noqa, suppression comments, or new Any annotations anywhere in the diff,
and no runtime behavior changes.

Budgets ratcheted by make lint-budget-update: basedpyright -1,663 across
48 rules, ruff-strict -86, type-discipline -110.
2026-08-08 08:14:29 -07:00
mateo-berri
f038be22db build(lint): rename make pre-commit to make check with a working-tree fallback 2026-08-08 03:25:35 -07:00
mateo-berri
8c0556abf6 fix(proxy): authenticate managed ids before routing 2026-08-08 02:42:34 -07:00
mateo-berri
b01eacd67c ci: run the new fine-tuning and vector store file test dirs 2026-08-08 01:36:30 -07:00
mateo-berri
5f7a663005 fix(proxy): enforce require_managed_files on every raw provider id route
require_managed_files was only checked on upload, so raw provider ids still
reached the batch, fine-tuning and vector store file routes. Ownership rows
exist only for managed ids, so those requests were forwarded under shared
credentials with no tenant check: knowing another tenant's id was enough to
read, run against, cancel or delete their object.

Generalise the file-id guard to validate_managed_id_requirement(resource_id,
resource_kind) and call it on batch create/retrieve/cancel, fine-tuning
create/retrieve/cancel (training_file and validation_file both) and the shared
vector store file id resolver. Behaviour is unchanged when the setting is off.
2026-08-08 01:33:32 -07:00
Mateo Wang
e24a9146e3
Merge pull request #36252 from BerriAI/litellm_ci_concurrency_guards
ci: give the remaining pull_request workflows a concurrency group
2026-08-08 00:28:01 -07:00
Mateo Wang
c28cbb804c
Merge pull request #35141 from BerriAI/litellm_vertex_batch_create_error_propagation
fix(vertex_ai): surface real error/status on vertex batch create instead of IndexError 500
2026-08-07 23:28:36 -07:00
mateo-berri
10209a8f91 Merge branch 'litellm_internal_staging' into devin_ai_require_managed_files_read_paths_35530 2026-08-07 23:17:28 -07:00
mateo-berri
5cd027cbbc fix(lint): let the ratchet guard recognise a graduated rule
A budget rule that graduates into a config's hard-fail select list rightly
leaves the budget file, but the ratchet guard read any disappearance as a
silently raised ceiling. Teach it the pairing between ruff-strict-budget.json
and ruff.toml: a dropped rule is excused only when the paired config's
lint.extend-select (minus lint.ignore) now hard-fails it, so deleting a rule
without graduating it still trips the guard.
2026-08-07 23:11:23 -07:00
mateo-berri
f304b7b19f refactor(lint): graduate the 35 zero-violation strict rules into ruff.toml
Every strict-gate rule whose budget ceiling was already 0 moves into the base
config's lint.extend-select, so editors and ruff check --fix surface the
diagnostics directly and the budget file shrinks to rules with real debt.
Graduates stay in ruff-strict.toml's select so the strict RUF100 pass keeps
policing their stale noqa directives, and base external entries they made
redundant (FURB, I001, RUF010, RUF022, RUF023, RUF051) are dropped so base
RUF100 polices those directly. UP037 had two violations hidden behind a star
import; importing Literal explicitly fixes them so UP037 can graduate too.
New drift tests pin the invariants: every strict-selected rule is budgeted or
hard-failed by base, every base-owned rule stays visible to exactly one
RUF100 pass, and graduated rules fail the normal ruff run.
2026-08-07 23:10:33 -07:00
mateo-berri
7bffbbd1f2 refactor(vertex_ai): drop unreachable post-path status checks in batches handler
HTTPHandler.post and AsyncHTTPHandler.post call raise_for_status before returning, so the status_code != 200 branches after the create and cancel POSTs could never run. Non-2xx already surfaces as httpx.HTTPStatusError from inside the client. The checks after GETs stay: the get helpers return without raising. Tests that faked a non-raising POST response are replaced by HTTPStatusError propagation coverage.
2026-08-07 23:00:50 -07:00
Devin AI
21df36ed09 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_vertex_batch_create_error_propagation 2026-08-08 05:42:05 +00:00
Mateo Wang
dee667edb7
Merge pull request #36257 from BerriAI/litellm_ruff_external_split
fix(lint): make strict-gate noqas survive base ruff and flag stale ones
2026-08-07 22:39:50 -07:00
mateo-berri
c3536c29a0 fix(lint): cover every base-owned ruff rule in the strict gate's external list 2026-08-07 21:38:53 -07:00
mateo-berri
557d14cc71 fix(lint): make strict-gate noqas survive base ruff and flag stale ones 2026-08-07 21:27:00 -07:00
Mateo Wang
d0758a291c
Merge pull request #36215 from BerriAI/litellm_declare_model_info_pricing
refactor(types): declare mirrored pricing fields on ModelInfo
2026-08-07 21:06:14 -07:00
mateo
ede84eee15 ci: give the remaining pull_request workflows a concurrency group
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-08 03:25:38 +00:00
mateo
f668c10609 chore(ui): regenerate dashboard api types for ModelInfo pricing fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-08 03:25:09 +00:00
mateo
24ac999cf7 fix(types): drop Final on SPECIAL_MODEL_INFO_PARAMS for star-import rebinding
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-08 03:25:09 +00:00