* refactor(litellm-rust): move provider transforms into litellm-core + crate allowlist test
* feat(litellm-rust): ai-gateway absorbs route I/O (io/) with lib+server feature split
* refactor(litellm-rust): point python-bridge at litellm-ai-gateway
* build(litellm-rust): macOS pyo3 dynamic_lookup linker flag for cdylib builds
* docs(litellm-rust): 3-crate map in README/AGENTS + refresh CLAUDE boundary
* refactor(litellm-rust): update workspace members to the three crates
* feat(ui): track frontend lint counts in a committed snapshot
Persist the eslint budget-rule counts (no-explicit-any, complexity,
max-depth) to eslint-metrics.json so the trend is queryable straight
from git history and can later feed a dashboard. A CI drift check
regenerated from the same lint report keeps the snapshot honest, so a
PR that shifts a count has to run npm run lint:metrics and commit it
* fix(ui): harden lint-metrics drift check and eslint failure handling
Make the drift comparison symmetric over the union of committed and
actual keys so a phantom rule left in eslint-metrics.json (for example
after a rule is dropped from eslint-budgets.json) is caught instead of
silently passing. Only swallow eslint's lint-errors exit code in the
generator and rethrow anything else, so a fatal eslint failure surfaces
its real output rather than a confusing ENOENT on the missing report
Deleting every budget window from a virtual key looked like it saved but
reverted on reload, while editing a window persisted. The key edit form set
budget_limits to undefined once the window list was emptied, and
JSON.stringify drops undefined keys, so /key/update received no budget_limits
field at all and model_dump(exclude_unset=True) skipped the existing
clear-on-empty branch. Sending [] instead lets the backend store JSON null and
clear the stored windows, matching how it already treats an explicit empty list
Resolves LIT-3742
A user provisioned with "All Proxy Models" stores the literal
"all-proxy-models" sentinel in user.models. get_direct_access_models looked
that string up as a real model_name via get_model_list, which matched no
deployment, so /v2/model/info marked every model direct_access=false and the
Models + Endpoints page rendered empty for such users when they have no teams.
The model dropdown / Playground worked because get_key_models already expands
the sentinel to the full proxy model list, hence the inconsistency in the
report.
Expand the sentinel to all non-team deployment ids via
get_model_ids(exclude_team_models=True), the same call the PROXY_ADMIN branch
in the caller already uses. This fixes both /v1/model/info and /v2/model/info
since they share _populate_team_access_on_models. Empty user.models stays "no
direct access" to match get_key_models semantics.
Fixes#22791
* fix(anthropic): drop unsupported speed param with drop_params
Anthropic fast mode (speed) is Opus 4.6/4.7/4.8 on the direct API only.
Strip speed when the model map lacks supports_speed and drop_params is set,
for both chat completions and /v1/messages passthrough.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(ci): allow supports_speed in model map schema
The new supports_speed flag on Opus entries must pass JSON schema
validation in test_aaamodel_prices_and_context_window_json_is_valid.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(review): raise on unsupported speed without drop_params
Passthrough /v1/messages now raises UnsupportedParamsError when speed
is unsupported and drop_params is false. Emit drop warning from
map_openai_params when speed is silently skipped.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(anthropic): gate speed param by routed provider, not just model id
Vertex, Azure, and Bedrock reuse the shared Anthropic transform and strip
their provider prefix first, so a bare `claude-opus-4-8` resolved to the
direct-API model-map entry (`supports_speed: true`) and forwarded `speed`
upstream, producing the same 400 that drop_params is meant to prevent.
Gate fast mode on `custom_llm_provider == "anthropic"` so it stays on the
direct Anthropic API across both the chat completions and `/v1/messages`
passthrough paths, and collapse the duplicated drop/raise logic in
map_openai_params into the shared `_maybe_drop_speed_param` helper.
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* docs: realtime pre-warmed connection pool design + raw-passthrough follow-up
* docs: realtime pool benchmark, repro steps, and deploy guidance
* feat: expose Router::deployments() for host-side upstream enumeration
* refactor: split realtime dial/splice and add warm-handoff entry point
* feat: pre-warmed upstream realtime connection pool with fresh-dial fallback
* feat: add realtime pool handle to gateway AppState
* feat: try warm pooled upstream before fresh-dial in realtime service
* feat: thread realtime pool through the realtime route bridge
* feat: build and pre-warm the realtime pool at gateway startup
* build: lean Dockerfile for the realtime gateway (default features, env stand-in)
Minimal multi-stage image for load-testing the realtime pool: builds the
gateway with default features (no python-config, no libpython), runs on a
debian-slim base (~157MB), and reads model_list from the OPENAI_REALTIME_MODEL
env stand-in. No config.yaml or pip install needed. Build context is the repo
root; only litellm-rust/ is included via the sidecar .dockerignore.
* docs: add generic ai-gateway benchmarking skill
Teaches an agent how to benchmark any ai-gateway endpoint: deploy the gateway,
run one load generator against both provider-direct and the gateway with the
same protocol, phase-decompose latency (dial/session/first-token/total),
compare at scale, and report success% + p50/p95. Documents the
benchmarks/<endpoint>/ layout and the hard no-committed-keys rule.
* docs: drop per-endpoint realtime benchmark README
The measured results table lives in the PR description (numbers go stale in a
committed README). The benchmarks/realtime/ dir now holds only the sanitized
load-gen harness; the generic method is in benchmarks/SKILL.md.
* test: add sanitized realtime WS load-gen harness for gateway benchmarks
Copies the ws-bench Go load generator (main.go, go.mod, go.sum, Dockerfile,
run.sh) into benchmarks/realtime/. Measures dial/session/first-audio/total per
WebSocket connection against both OpenAI-direct and the gateway. No keys are
hardcoded — the bearer token comes from -key / $OPENAI_API_KEY; default host
is api.openai.com.
* test: add hosted-runner serve.sh wrapper for the realtime harness
Render one-off jobs don't surface stdout via the Logs API, so on a hosted
runner the long-lived web service runs the leg and PUBLISHES the result: serve.sh
decodes the base64 flag list, runs wsbench teeing output to /tmp/web/result.txt,
then serves it over HTTP so the result is fetchable at /result.txt. No secrets
are written to the served file (the -key is only in wsbench's argv). The
Dockerfile now copies both run.sh and serve.sh.
* test: make serve.sh publish result atomically and clear stale output
Remove any prior result.txt at startup and write the new run to a .partial file
that's atomically moved into place only once complete. Prevents a fetcher from
reading a previous run's numbers while the current run is still in flight.
* perf: refill the realtime pool concurrently so warm supply keeps up at scale
The replenisher dialed missing warm sockets sequentially, so a full refill cost
needed x handshake (~needed x 350ms). Under high connect rates the pool drained
faster than it refilled and ~85% of connects missed (measured: only ~15% pool
hits at 5000/500). Firing the dials together with join_all refills in ~one
handshake window, keeping warm supply close to peak concurrent connects so the
sub-millisecond warm handoff becomes the median rather than the lucky-hit tail.
Each warm_one dial is independent (no shared state until the final push under the
lock), so concurrent refill is safe. Pool unit tests unchanged and passing.
* build: drop Dockerfile.lean
The lean load-test image isn't worth carrying in the repo; deploy the gateway
however you normally do and set the pool env vars.
* docs: slim benchmarks/realtime to a README (harness moved to its own repo)
Drop the Go load-gen, Dockerfile, run.sh, serve.sh, and the generic SKILL.md from
the repo. The harness now lives at github.com/ishaan-berri/litellm-realtime-bench;
benchmarks/realtime/README.md carries the results table and links there for repro.
* docs: add realtime route README with pooling design + diagram
Pooling is now documented as a section in src/routes/realtime/README.md next to
the code (handoff diagram, sizing rule, config, notes) instead of the standalone
REALTIME_POOL_DESIGN.md RFC. Update the realtime_pool.rs doc pointer to it.
* fix: satisfy clippy manual_flatten on concurrent pool refill
Use .into_iter().flatten() instead of an if-let-Ok in the for loop over the
join_all results, and let rustfmt wrap it. Clears the CI clippy -D warnings
failure; fmt + clippy + cargo test all green locally.
---------
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
The per-user BYOK, OAuth (OBO), and env-var management endpoints resolved the
target MCP server through a DB-only lookup (get_mcp_server / get_all_mcp_servers_for_user).
A server defined in config.yaml lives only in the in-memory registry and never
gets a row in LiteLLM_MCPServerTable, so those endpoints raised 404 "MCP Server
<id> not found" (or 403 for non-admins) before any credential could be stored,
leaving config-server users unable to connect and forced to re-authorize forever.
Route all three through a single registry-aware resolver: DB first, then the
in-memory registry (built into LiteLLM_MCPServerTable via _build_mcp_server_table,
the same fallback fetch_mcp_server already uses), then the canonical
get_allowed_mcp_servers authorization the MCP gateway enforces on tool calls.
Admins get a 404 for an unknown id; non-admins get 403 for a missing-or-forbidden
server so server ids stay non-enumerable. This also closes a gap where the two
store endpoints performed no per-server authorization at all.
Re-pins LITELLM_BUILD_IMAGE and LITELLM_RUNTIME_IMAGE across all 6 Dockerfiles
from the prior digests (openssl 3.6.2-r3) to the current chainguard wolfi-base
digest c61ac691 (openssl 3.6.3-r2, >= the fixed 3.6.3-r0). The runtime stage is
the shipped image, so the runtime digest is what actually resolves the
customer-facing CVE; the build image is bumped too for hygiene. Two Dockerfiles
tracked a second equally-stale digest; both are unified onto the patched one.
The bedrock/mantle chat and messages paths hardcoded the public Mantle host and ignored api_base, so private VPC/VPCE/GovCloud endpoints could not be used. Route URL construction through a shared helper that prefers api_base and aws_bedrock_runtime_endpoint before falling back to the regional public host.
Co-authored-by: Shivam Rawat <shivamrawat@Shivams-MacBook-Pro.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(model_prices): correct regional processing uplift assignment
gpt-4.1, gpt-4o, gpt-5, and their variants were incorrectly carrying
the 10% EU/US regional processing uplift multiplier. Per OpenAI's
pricing docs, the uplift applies only to models released on or after
2026-03-05 (gpt-5.4 series and gpt-5.5 series).
Removes the uplift from: gpt-4.1, gpt-4.1-mini, gpt-4.1-nano,
gpt-4o, gpt-4o-2024-08-06, gpt-4o-2024-11-20, gpt-4o-mini, gpt-5,
gpt-5-pro, gpt-5-mini, gpt-5-nano.
Adds the uplift to: gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.4-pro,
gpt-5.5, gpt-5.5-pro.
* fix(model_prices): apply same regional uplift correction to backup file
* fix(model_prices): add regional uplift to date-versioned gpt-5.4/5.5 siblings
* test(model_prices): update data residency tests to use gpt-5.4 as the uplift model
The tests were using gpt-5 which no longer carries the regional processing
uplift after correcting which models have it. Switch to gpt-5.4 (released
2026-03-05, the cutoff date) and add a regression parametrize covering
all pre-cutoff models to pin that they stay uplift-free.
* test(batches): use gpt-5.4 for data residency uplift assertion
batch_cost_calculator's data residency uplift test still pinned gpt-5,
which no longer carries the regional processing uplift after this change.
Switch it to gpt-5.4 (the canonical post-cutoff uplift model), matching
the llm_cost_calc test update.
---------
Co-authored-by: mgalbato <37748295+mgalbato@users.noreply.github.com>
Bumps the 12 packages osv-scanner flags on litellm_internal_staging, taking
the scan from 24 known vulnerabilities to zero. vcrpy goes to 8.2.1 first so
aiohttp can move to 3.14.1 (vcrpy <= 8.1.1 cannot import aiohttp 3.14), then
the two aiohttp ignore entries are dropped from osv-scanner.toml. The
langchain stack moves together since langchain 1.3.9 requires langgraph 1.2.x.
Runtime deps cryptography (48.0.1), starlette (1.3.1), python-multipart
(0.0.32), pydantic-settings (2.14.2) and pypdf (6.13.3) are bumped via relock,
and the dashboard's js-yaml, ws and form-data overrides are bumped too.
Also removes the paths filter on the OSV workflow so it runs on every PR
rather than only when a lockfile changes, which is why it never showed up on
recent code-only PRs
* add RealtimeTransformResult type for realtime transforms
* add RealtimeProviderConfig pure trait in litellm-core
* add core realtime module
* register realtime module in litellm-core lib
* add OpenAI realtime transform + complete_url parity in providers
* add openai realtime module
* add openai provider module
* register openai provider module in providers lib
* add realtime() fn that invokes OpenAI GA realtime API end to end
* register realtime route module in providers lib
* wire tokio/tokio-tungstenite/futures-util into providers crate
* add tokio, tokio-tungstenite, futures-util to rust workspace deps
* update Cargo.lock for realtime websocket deps
* docs: add litellm-rust provider/route contributor guide
* add typed RealtimeEvent; make RealtimeTransformResult hold typed events
* type RealtimeProviderConfig trait on RealtimeEvent instead of raw strings
* type OpenAI realtime passthrough transforms on RealtimeEvent
* type realtime() fn on RealtimeEvent end to end (parse/serialize at host edge)
* docs: add typed-contracts core rule to core CLAUDE.md
* harden complete_url: default bare host / unknown scheme to wss://
---------
Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
* fix: redact config and MCP secrets in read-only admin views
GET /config/field/info and the MCP server list/detail endpoints returned
secret-bearing fields to any caller with an admin view, including
read-only admins. They now return those fields in full only to a full
PROXY_ADMIN; every other caller gets the reduced, non-admin view, while
non-sensitive fields remain readable. Regression tests cover the
role-based visibility on both endpoints, including that a full admin
still sees everything needed to populate the edit form.
* fix: redact nested secrets in config field info for non-admins
/config/field/info returned structured general_settings fields verbatim to
any admin-view caller, so a view-only admin reading database_args received
the nested aws_web_identity_token (a DynamoDB role-assumption credential) in
plaintext. Recurse into dict/list field values and redact secret leaves for
non-PROXY_ADMIN callers, leaving non-secret siblings and full-admin reads
unchanged
* fix: redact secret config values in /config/list for non-admins
/config/list shared the same _user_has_admin_view gate as /config/field/info
but returned each field value unredacted, so a view-only admin reading the
list received pass_through_endpoints upstream Authorization headers verbatim.
Route every general_settings value through a shared role-aware redactor
(extracted from /config/field/info) covering the top-level and nested field
paths, so non-PROXY_ADMIN callers get secret-bearing fields redacted while
full-admin reads stay unchanged
* chore(ci): allowlist _redact_secret_values_in_obj in recursive_detector
The config secret redactor recurses over JsonValue, which is acyclic, and
its depth is bounded by the operator-authored general_settings schema. Add
it to the recursive_detector ignore list alongside the other bounded
nested-redaction helpers (mask_dict, _redact_sensitive_litellm_params)
* proxy: cap recursive secret redaction depth at 10
Match the cap on _redact_sensitive_litellm_params (the closest analog
in the proxy, also recursive, key-name driven, returns a sentinel).
The previous justification — bounded by operator-authored schema depth,
JsonValue acyclic — is true today but is a property of the threat model,
not an enforced invariant of the function. If a code path is ever added
that pipes external input into general_settings (config import,
migration tooling, JWT-driven settings, …) the assumption silently
breaks. A local cap makes the invariant local.
The cap branch fails closed: at _REDACT_SECRET_MAX_DEPTH the whole
subtree is replaced with 'REDACTED' rather than returned verbatim. A
future refactor that flips this to fail-open would let a deeply nested
credential leak; the new regression test test_redact_secret_values_in_obj_fails_closed_at_max_depth
guards against that.
Updates the recursive_detector ignore-list rationale to point at the
numeric cap rather than the structural argument.
* test: actually exercise the depth cap in fails-closed test
The previous fixture stored the leaf under the secret-named key
'aws_web_identity_token', which the recursor's key-name short-circuit
redacts regardless of the cap — so the test passed both with and
without the cap in place. Empirically confirmed: under an uncapped
mutant the old fixture still hides the secret (key-name catches it),
the new fixture leaks it (only the cap can stop it). Swap the leaf
key to a non-secret name so the cap is the only redaction path
exercised, making the test fail on mutation as advertised.