Commit graph

39873 commits

Author SHA1 Message Date
Ishaan Jaff
e2916e8e31
fix(ai-gateway): close realtime auth review gaps 2026-06-24 15:35:44 -07:00
Ishaan Jaff
b7f37f3ff0
fix(proxy): move gateway auth under rust control plane 2026-06-24 15:27:25 -07:00
Ishaan Jaff
acbd5991c8
test(proxy): 5xx (http + proxy) propagation 2026-06-24 13:40:27 -07:00
Ishaan Jaff
f13272b216
docs(ai-gateway): strict-by-default auth (per-connection re-verify); cache opt-in rationale 2026-06-24 13:40:27 -07:00
Ishaan Jaff
bfdfee2b7c
fix(ai-gateway): auth cache off by default — re-verify every connection (closes budget/rate-limit bypass; cache opt-in for high-RPS routes) 2026-06-24 13:40:27 -07:00
Ishaan Jaff
95e41e94f5
fix(proxy): propagate 5xx (HTTPException or ProxyException) instead of masking as 401 (Greptile) 2026-06-24 13:40:27 -07:00
Ishaan Jaff
1faedc5f4d
docs(ai-gateway): explain auth-cache perf default + strict-mode knob 2026-06-24 13:27:13 -07:00
Ishaan Jaff
56ed05418e
perf(ai-gateway): key cache on by default (TTL 60s); set LITELLM_AUTH_CACHE_TTL_SECS=0 for strict per-connection verification 2026-06-24 13:27:13 -07:00
Ishaan Jaff
4c7bd5bf17
fix(proxy): restore verify_key body mangled by suggestion merge; propagate 5xx instead of masking as 401 2026-06-24 13:27:13 -07:00
Ishaan Jaff
0a977b8d1a
docs(ai-gateway): add data-plane auth env vars to render blueprint 2026-06-24 13:20:06 -07:00
Ishaan Jaff
df0faa86e6
docs(ai-gateway): document delegating virtual-key auth to the LiteLLM proxy 2026-06-24 13:20:06 -07:00
ishaan-berri
44e84c95bd
Update litellm/proxy/auth/internal_auth_endpoints.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-06-24 13:15:28 -07:00
Ishaan Jaff
1b0120dc0b
fix(proxy): builtin dict generics (UP006 strict budget) + hide internal auth route from OpenAPI/UI schema 2026-06-24 13:09:52 -07:00
Ishaan Jaff
2561850482
test(proxy): cover model forwarding in verify 2026-06-24 13:03:28 -07:00
Ishaan Jaff
a99822ebc0
fix(proxy): forward model into verify so can_key_call_model enforces model access (veria High) 2026-06-24 13:03:28 -07:00
Ishaan Jaff
c547badbb1
perf(ai-gateway): keep cache get() O(1); reconcile order on insert (Greptile P2) 2026-06-24 13:03:28 -07:00
Ishaan Jaff
36fdb698ea
fix(ai-gateway): don't leak auth-backend detail to client; thread model + key cache by (route,model,key) 2026-06-24 13:03:28 -07:00
Ishaan Jaff
d767ae7750
feat(ai-gateway): KeyAuthenticator.verify takes model for model-access enforcement 2026-06-24 13:03:28 -07:00
Ishaan Jaff
d7059c6c58
fix(ai-gateway): redact upstream URL from auth errors (without_url); forward model to verify 2026-06-24 13:03:28 -07:00
Ishaan Jaff
1584b703da
feat(ai-gateway): extractor forwards request path; cache key includes route 2026-06-24 12:46:22 -07:00
Ishaan Jaff
1d04e56e6e
feat(ai-gateway): send the serving route to the verify endpoint 2026-06-24 12:46:22 -07:00
Ishaan Jaff
a5203ed000
feat(ai-gateway): thread route through KeyAuthenticator + plain-English auth/client docs 2026-06-24 12:46:22 -07:00
Ishaan Jaff
391dd3321c
test(proxy): cover synthetic-request route + Bearer normalization 2026-06-24 12:46:22 -07:00
Ishaan Jaff
ea2c7ccf14
fix(proxy): verify against a synthetic request on the gateway's route; require route; Bearer-prefix the key 2026-06-24 12:46:22 -07:00
Ishaan Jaff
97f5194a40
test(proxy): cover data-plane gate + verify_key 2026-06-24 12:33:40 -07:00
Ishaan Jaff
68b1b6d341
feat(proxy): register internal auth router 2026-06-24 12:33:40 -07:00
Ishaan Jaff
1d76bf0c9b
feat(proxy): /internal/v1/auth/verify gated by the data-plane key 2026-06-24 12:33:40 -07:00
Ishaan Jaff
f2119c2eb6
build(ai-gateway): add sha2 dep for key hashing 2026-06-24 12:33:40 -07:00
Ishaan Jaff
7ef78966b3
feat(ai-gateway): realtime route uses UserApiKeyAuth extractor 2026-06-24 12:33:40 -07:00
Ishaan Jaff
6465b9e7dc
feat(ai-gateway): construct PythonAuthClient from data-plane env at startup 2026-06-24 12:33:40 -07:00
Ishaan Jaff
5c04ec2ba1
feat(ai-gateway): AppState carries authenticator + key cache 2026-06-24 12:33:39 -07:00
Ishaan Jaff
9c6ca78da9
feat(ai-gateway): wire auth submodules + re-exports 2026-06-24 12:33:39 -07:00
Ishaan Jaff
8170e0ed18
feat(ai-gateway): bounded (200) TTL key cache 2026-06-24 12:33:39 -07:00
Ishaan Jaff
7097ae530a
feat(ai-gateway): UserApiKeyAuth struct + extractor (the route 'decorator') 2026-06-24 12:33:39 -07:00
Ishaan Jaff
2538c6e940
feat(ai-gateway): PythonAuthClient — verify keys via the Python control plane 2026-06-24 12:33:39 -07:00
Ishaan Jaff
f98a1d12cc
feat(ai-gateway): KeyAuthenticator trait — the auth swap seam 2026-06-24 12:33:39 -07:00
ishaan-berri
bd759182ca
refactor(litellm-rust): dissolve providers into core + ai-gateway (strict 3-crate layers) (#31218)
* refactor(litellm-rust): move provider transforms into litellm-core + crate allowlist test

* feat(litellm-rust): ai-gateway absorbs route I/O (io/) with lib+server feature split

* refactor(litellm-rust): point python-bridge at litellm-ai-gateway

* build(litellm-rust): macOS pyo3 dynamic_lookup linker flag for cdylib builds

* docs(litellm-rust): 3-crate map in README/AGENTS + refresh CLAUDE boundary

* refactor(litellm-rust): update workspace members to the three crates
2026-06-24 12:22:43 -07:00
ryan-crabbe-berri
f2f6cacb19
feat(ui): track frontend lint counts in a committed snapshot (#31157)
Some checks are pending
LiteLLM Rust / rustfmt, clippy, test (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* feat(ui): track frontend lint counts in a committed snapshot

Persist the eslint budget-rule counts (no-explicit-any, complexity,
max-depth) to eslint-metrics.json so the trend is queryable straight
from git history and can later feed a dashboard. A CI drift check
regenerated from the same lint report keeps the snapshot honest, so a
PR that shifts a count has to run npm run lint:metrics and commit it

* fix(ui): harden lint-metrics drift check and eslint failure handling

Make the drift comparison symmetric over the union of committed and
actual keys so a phantom rule left in eslint-metrics.json (for example
after a rule is dropped from eslint-budgets.json) is caught instead of
silently passing. Only swallow eslint's lint-errors exit code in the
generator and rethrow anything else, so a fatal eslint failure surfaces
its real output rather than a confusing ENOENT on the missing report
2026-06-24 11:35:32 -07:00
ryan-crabbe-berri
8f4389246d
fix(ui): persist budget window deletion on virtual keys (#31107)
Deleting every budget window from a virtual key looked like it saved but
reverted on reload, while editing a window persisted. The key edit form set
budget_limits to undefined once the window list was emptied, and
JSON.stringify drops undefined keys, so /key/update received no budget_limits
field at all and model_dump(exclude_unset=True) skipped the existing
clear-on-empty branch. Sending [] instead lets the backend store JSON null and
clear the stored windows, matching how it already treats an explicit empty list

Resolves LIT-3742
2026-06-24 09:19:53 -07:00
Sameer Kankute
8bca05d311
fix(anthropic): sanitize tool_use ids on native /v1/messages path (#31094) 2026-06-24 07:57:46 -07:00
mubashir1osmani
e0c8a6b483
fix(proxy): expand all-proxy-models sentinel in direct access lookup (#31153)
A user provisioned with "All Proxy Models" stores the literal
"all-proxy-models" sentinel in user.models. get_direct_access_models looked
that string up as a real model_name via get_model_list, which matched no
deployment, so /v2/model/info marked every model direct_access=false and the
Models + Endpoints page rendered empty for such users when they have no teams.
The model dropdown / Playground worked because get_key_models already expands
the sentinel to the full proxy model list, hence the inconsistency in the
report.

Expand the sentinel to all non-team deployment ids via
get_model_ids(exclude_team_models=True), the same call the PROXY_ADMIN branch
in the caller already uses. This fixes both /v1/model/info and /v2/model/info
since they share _populate_team_access_on_models. Empty user.models stays "no
direct access" to match get_key_models semantics.

Fixes #22791
2026-06-23 22:23:14 -07:00
Krrish Dholakia
d0706c17fe
fix(anthropic): drop unsupported speed param with drop_params (#31152)
* fix(anthropic): drop unsupported speed param with drop_params

Anthropic fast mode (speed) is Opus 4.6/4.7/4.8 on the direct API only.
Strip speed when the model map lacks supports_speed and drop_params is set,
for both chat completions and /v1/messages passthrough.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): allow supports_speed in model map schema

The new supports_speed flag on Opus entries must pass JSON schema
validation in test_aaamodel_prices_and_context_window_json_is_valid.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(review): raise on unsupported speed without drop_params

Passthrough /v1/messages now raises UnsupportedParamsError when speed
is unsupported and drop_params is false. Emit drop warning from
map_openai_params when speed is silently skipped.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(anthropic): gate speed param by routed provider, not just model id

Vertex, Azure, and Bedrock reuse the shared Anthropic transform and strip
their provider prefix first, so a bare `claude-opus-4-8` resolved to the
direct-API model-map entry (`supports_speed: true`) and forwarded `speed`
upstream, producing the same 400 that drop_params is meant to prevent.

Gate fast mode on `custom_llm_provider == "anthropic"` so it stays on the
direct Anthropic API across both the chat completions and `/v1/messages`
passthrough paths, and collapse the duplicated drop/raise logic in
map_openai_params into the shared `_maybe_drop_speed_param` helper.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-06-23 22:22:49 -07:00
Mateo Wang
c0a146929c
chore: clarify rule about trailing periods (#31175)
I notice Claude applies the rule irregularly. This aims to fix that
2026-06-23 22:05:59 -07:00
ishaan-berri
6072019b0d
perf: pre-warm upstream realtime connection pool to cut session-establishment latency (#31163)
* docs: realtime pre-warmed connection pool design + raw-passthrough follow-up

* docs: realtime pool benchmark, repro steps, and deploy guidance

* feat: expose Router::deployments() for host-side upstream enumeration

* refactor: split realtime dial/splice and add warm-handoff entry point

* feat: pre-warmed upstream realtime connection pool with fresh-dial fallback

* feat: add realtime pool handle to gateway AppState

* feat: try warm pooled upstream before fresh-dial in realtime service

* feat: thread realtime pool through the realtime route bridge

* feat: build and pre-warm the realtime pool at gateway startup

* build: lean Dockerfile for the realtime gateway (default features, env stand-in)

Minimal multi-stage image for load-testing the realtime pool: builds the
gateway with default features (no python-config, no libpython), runs on a
debian-slim base (~157MB), and reads model_list from the OPENAI_REALTIME_MODEL
env stand-in. No config.yaml or pip install needed. Build context is the repo
root; only litellm-rust/ is included via the sidecar .dockerignore.

* docs: add generic ai-gateway benchmarking skill

Teaches an agent how to benchmark any ai-gateway endpoint: deploy the gateway,
run one load generator against both provider-direct and the gateway with the
same protocol, phase-decompose latency (dial/session/first-token/total),
compare at scale, and report success% + p50/p95. Documents the
benchmarks/<endpoint>/ layout and the hard no-committed-keys rule.

* docs: drop per-endpoint realtime benchmark README

The measured results table lives in the PR description (numbers go stale in a
committed README). The benchmarks/realtime/ dir now holds only the sanitized
load-gen harness; the generic method is in benchmarks/SKILL.md.

* test: add sanitized realtime WS load-gen harness for gateway benchmarks

Copies the ws-bench Go load generator (main.go, go.mod, go.sum, Dockerfile,
run.sh) into benchmarks/realtime/. Measures dial/session/first-audio/total per
WebSocket connection against both OpenAI-direct and the gateway. No keys are
hardcoded — the bearer token comes from -key / $OPENAI_API_KEY; default host
is api.openai.com.

* test: add hosted-runner serve.sh wrapper for the realtime harness

Render one-off jobs don't surface stdout via the Logs API, so on a hosted
runner the long-lived web service runs the leg and PUBLISHES the result: serve.sh
decodes the base64 flag list, runs wsbench teeing output to /tmp/web/result.txt,
then serves it over HTTP so the result is fetchable at /result.txt. No secrets
are written to the served file (the -key is only in wsbench's argv). The
Dockerfile now copies both run.sh and serve.sh.

* test: make serve.sh publish result atomically and clear stale output

Remove any prior result.txt at startup and write the new run to a .partial file
that's atomically moved into place only once complete. Prevents a fetcher from
reading a previous run's numbers while the current run is still in flight.

* perf: refill the realtime pool concurrently so warm supply keeps up at scale

The replenisher dialed missing warm sockets sequentially, so a full refill cost
needed x handshake (~needed x 350ms). Under high connect rates the pool drained
faster than it refilled and ~85% of connects missed (measured: only ~15% pool
hits at 5000/500). Firing the dials together with join_all refills in ~one
handshake window, keeping warm supply close to peak concurrent connects so the
sub-millisecond warm handoff becomes the median rather than the lucky-hit tail.
Each warm_one dial is independent (no shared state until the final push under the
lock), so concurrent refill is safe. Pool unit tests unchanged and passing.

* build: drop Dockerfile.lean

The lean load-test image isn't worth carrying in the repo; deploy the gateway
however you normally do and set the pool env vars.

* docs: slim benchmarks/realtime to a README (harness moved to its own repo)

Drop the Go load-gen, Dockerfile, run.sh, serve.sh, and the generic SKILL.md from
the repo. The harness now lives at github.com/ishaan-berri/litellm-realtime-bench;
benchmarks/realtime/README.md carries the results table and links there for repro.

* docs: add realtime route README with pooling design + diagram

Pooling is now documented as a section in src/routes/realtime/README.md next to
the code (handoff diagram, sizing rule, config, notes) instead of the standalone
REALTIME_POOL_DESIGN.md RFC. Update the realtime_pool.rs doc pointer to it.

* fix: satisfy clippy manual_flatten on concurrent pool refill

Use .into_iter().flatten() instead of an if-let-Ok in the for loop over the
join_all results, and let rustfmt wrap it. Clears the CI clippy -D warnings
failure; fmt + clippy + cargo test all green locally.

---------

Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
2026-06-23 21:49:56 -07:00
tin-berri
360adbe765
fix(mcp): resolve config-defined servers in per-user credential and env-var endpoints (#31171)
The per-user BYOK, OAuth (OBO), and env-var management endpoints resolved the
target MCP server through a DB-only lookup (get_mcp_server / get_all_mcp_servers_for_user).
A server defined in config.yaml lives only in the in-memory registry and never
gets a row in LiteLLM_MCPServerTable, so those endpoints raised 404 "MCP Server
<id> not found" (or 403 for non-admins) before any credential could be stored,
leaving config-server users unable to connect and forced to re-authorize forever.

Route all three through a single registry-aware resolver: DB first, then the
in-memory registry (built into LiteLLM_MCPServerTable via _build_mcp_server_table,
the same fallback fetch_mcp_server already uses), then the canonical
get_allowed_mcp_servers authorization the MCP gateway enforces on tool calls.
Admins get a 404 for an unknown id; non-admins get 403 for a missing-or-forbidden
server so server ids stay non-enumerable. This also closes a gap where the two
store endpoints performed no per-server authorization at all.
2026-06-23 21:38:10 -07:00
ishaan-berri
a6b7dcc7d6
build: add Dockerfile + render blueprint for rust ai-gateway (#31154)
* build: make rust ai-gateway Dockerfile config.yaml-based (repo-root context)

* build: add dockerignore to shrink repo-root context for ai-gateway

* build: add sample realtime config.yaml for rust ai-gateway

* build: point ai-gateway render blueprint at config.yaml + repo-root context

* docs: document config.yaml as primary path for rust ai-gateway

* build: run rust ai-gateway container as non-root user

* ci: re-trigger flaky otel fake-openai-endpoint cooldown
2026-06-23 20:41:50 -07:00
ishaan-berri
1d5ab42e14
feat: add minimal rust router + axum ai-gateway calling router.realtime (2/2) (#31135)
* add CoreError::Routing variant for deployment selection failures

* add minimal Rust Router (simple-shuffle) mirroring router.py spec

* add litellm-router crate manifest

* add ai-gateway POST /v1/realtime handler calling router.realtime

* add ai-gateway health routes

* wire ai-gateway routes into the axum app

* add ai-gateway AppState holding the shared router

* add ai-gateway axum server entrypoint

* add litellm-ai-gateway binary crate manifest

* docs: add ai-gateway folder-architecture AGENTS.md

* register router + ai-gateway crates and axum/rand deps in workspace

* update Cargo.lock for router + ai-gateway crates

* split router: extract model_list types into deployment module

* split router: extract routing policy into strategy module

* split router: move Router orchestration into router module

* router lib: wire submodules and re-export public API

* add read_model_list helper reusing ProxyConfig env/secret resolution

* add GIL-activity tracker (records acquisitions, 30s window)

* add GET /health/gil endpoint for polling GIL activity

* add pyo3 load_router_from_config bridge (feature-gated, load-time only)

* register /health/gil route in ai-gateway

* wire build_router: load from python config when feature enabled

* add optional pyo3 dep + python-config feature to ai-gateway

* update Cargo.lock for optional pyo3 dependency

* fix: satisfy strict ruff budget (FA100) in read_model_list

* test: cover read_model_list env resolution + empty config

* ai-gateway: bind localhost by default, warn on bad PORT/missing keys, wire gateway key

* ai-gateway: add gateway_key to AppState for realtime auth

* ai-gateway: require bearer auth + map unknown model to 404 on /v1/realtime

* ai-gateway: move python interop into python/ with load-time-only AGENTS.md

* ai-gateway: document auth, gil, and python folder in AGENTS.md

* core: add router module (model_list types + simple-shuffle selection)

* ai-gateway: dispatch realtime via core router + providers (drop router crate dep)

* update Cargo.lock: fold router into core

* workspace: drop crates/router member and litellm-router dep

* read_model_list: reuse ProxyConfig.get_config (includes + os.environ + DB) instead of thin yaml read

* ai-gateway: constant-time bearer compare + 500 (not 503) for unconfigured key

* ai-gateway: trim stored gateway key to match trimmed bearer token

* ai-gateway: add subtle dep for constant-time comparison

* workspace: add subtle dependency

* update Cargo.lock for subtle

* core router: make strategy a folder (one module per strategy, simple_shuffle)

* providers: make realtime() a streaming splice (client stream <-> OpenAI) instead of collect

* providers: add futures-channel dev-dep for the streaming live test

* ai-gateway: make /v1/realtime a WebSocket (auth before upgrade, splice typed events)

* ai-gateway: dispatch realtime as a stream splice

* ai-gateway: route /v1/realtime via GET (WebSocket), drop POST

* ai-gateway: enable axum ws feature + futures-util

* update Cargo.lock for ws feature + futures-channel

* core router: add has_deployment() for pre-flight model checks

* ai-gateway: extract auth into auth/ module (single master key, LITELLM_MASTER_KEY)

* ai-gateway routes: adopt router()-per-module template + merge in app()

* ai-gateway: document auth/ + routes template in AGENTS.md

* ai-gateway: realtime route as thin handler + service + transport

* ai-gateway: auth as a RequireMasterKey extractor (idiomatic axum FromRequestParts)

* ai-gateway: docs for auth extractor + simplified route template

* ai-gateway: collapse realtime route to mod.rs + service.rs; docs for extractor/template

* providers realtime: enforce idle timeout around the splice (reap stalled sessions)

* ai-gateway: rename realtime service timeout param to idle_timeout

---------

Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
2026-06-23 19:16:34 -07:00
yucheng-berri
fda08dd727
fix(docker): bump wolfi-base digest to patch openssl CVE-2026-34182 (#31133)
Re-pins LITELLM_BUILD_IMAGE and LITELLM_RUNTIME_IMAGE across all 6 Dockerfiles
from the prior digests (openssl 3.6.2-r3) to the current chainguard wolfi-base
digest c61ac691 (openssl 3.6.3-r2, >= the fixed 3.6.3-r0). The runtime stage is
the shipped image, so the runtime digest is what actually resolves the
customer-facing CVE; the build image is bumped too for hygiene. Two Dockerfiles
tracked a second equally-stale digest; both are unified onto the patched one.
2026-06-23 17:51:25 -07:00
Shivam Rawat
8dfe702e5b
fix(bedrock-mantle): honor api_base for VPC endpoint routing on bedrock/mantle/... (#31141)
The bedrock/mantle chat and messages paths hardcoded the public Mantle host and ignored api_base, so private VPC/VPCE/GovCloud endpoints could not be used. Route URL construction through a shared helper that prefers api_base and aws_bedrock_runtime_endpoint before falling back to the regional public host.

Co-authored-by: Shivam Rawat <shivamrawat@Shivams-MacBook-Pro.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-23 17:31:59 -07:00
yuneng-jiang
aab732d874
chore(ci): bump litellm version (#31139)
* bump: version 1.90.0 → 1.91.0

* adding uv lock
2026-06-23 23:25:42 +00:00