mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-16 23:41:43 +00:00
39680 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e989016c2d |
fix(auth_v2): match request paths with keyMatch so obj patterns span segments
Use keyMatch instead of keyMatch2 in the Casbin matcher so a "/*" or "/scim/v2/*" obj pattern unambiguously spans path separators - a require_permission check on a multi-level route like /api/v1/models now matches the granting policy rather than risking a 403. keyMatch is the canonical trailing-wildcard route matcher; the anchored act matcher is unchanged, so a "GET" policy still cannot grant "GETX". |
||
|
|
a7f9d5d6b9 |
test(auth_v2): pin the mTLS forwarded-DN gate against XFF spoofing
The forwarded subject-DN trust gate keys on the raw socket peer and prefers a verified TLS-layer cert: - an untrusted peer cannot smuggle a forged DN by claiming a trusted address via X-Forwarded-For (the gate ignores XFF) - a verified client cert from the ASGI TLS extension wins over a proxy-forwarded DN header Full auth_v2 suite: 175 passing. |
||
|
|
c9e7fd829c |
fix(auth_v2): prefer verified mTLS cert and confirm forwarded-DN gates on the socket peer
The forwarded subject-DN trust gate already compared the raw transport peer (request.client.host / ASGI scope client) against trusted_proxy_cidrs, never the XFF-resolved IP, so an attacker spoofing X-Forwarded-For cannot defeat it. Make that ordering explicit and stronger: a genuinely verified client certificate from the ASGI TLS extension is now preferred when present, and the spoofable forwarded-header path is only consulted as a fallback, still gated on the direct socket peer. |
||
|
|
f158a6c9e0 |
ci(auth_v2): run auth_v2 tests in a dedicated shard to avoid test_models basename clash
Adding tests/test_litellm/proxy/auth_v2 to the proxy-auth shard collided with proxy/client/test_models.py (both bare test_models, no package __init__), failing collection with an import file mismatch. Run the auth_v2 suite in its own job with the xmlsec1 apt packages instead; the base apt-packages input added earlier is reused. |
||
|
|
ac971d7e8b |
test(auth_v2): pin the role allowlist on the OIDC-login and SAML-SSO paths
The veria review-response fix gates IdP-asserted roles through the same per-provider allowlist on every role-bearing path, not just JWT bearer tokens. - rbac: filter_claim_roles (the shared gate) denies a self-asserted role by default, filters to the allowlist, and only admits platform roles behind the explicit allow_platform_roles flag - saml: a signed SSO assertion asserting platform_admin yields a session whose Principal has no roles by default, and is filtered to the allowlist when set - oidc: the login callback's identity build (map userinfo -> gate roles -> session) denies platform_admin by default and filters to the allowlist Full auth_v2 suite: 173 passing. |
||
|
|
44ac50493e |
ci(auth_v2): run the proxy auth_v2 suite with xmlsec1 in the proxy-auth shard
The proxy/auth_v2 tests were not picked up by any shard because the proxy-auth path list matched the literal proxy/auth directory, not its sibling. Add the path and install xmlsec1 so the pysaml2 signing tests run rather than skip; the install is gated behind a new optional apt-packages input so the other shards are unchanged. |
||
|
|
200f674b94 |
test(auth_v2): pin the token-claim hardening and provisioning security fixes
Regression coverage for the security review fixes (H1/M1/M3/S7): - resolver: a deactivated SCIM user (active=False) is rejected 403; claims whose keys start with "_" never surface on the Principal - authenticators: H1 privilege escalation - a self-asserted token role grants nothing without a per-provider allowlist, the allowlist filters roles, and platform-level roles need an explicit allow_platform_roles gate - rbac: the Casbin act matcher is anchored, so a "GET" policy does not grant "GETX" - scim: PATCH that targets the read-only id (replace, remove, no-path replace) is rejected 400 with the record's id unchanged; unauthenticated/under-scoped requests render a SCIM Error body (401/403); /Schemas is a ListResponse envelope - saml: a replayed signed assertion is rejected 401 (single-use), and an unsolicited IdP-initiated response is rejected 401 when allow_unsolicited is off (default secure) Full auth_v2 suite: 165 passing. |
||
|
|
99efd314f7 |
fix(auth_v2): enforce the role allowlist on the SAML SSO path too
The per-provider role allowlist + platform-role gate only covered the bearer-JWT path, so a SAML IdP could still mint platform_admin (or any Role) through its attribute->roles mapping. Extract the filter into a shared rbac.filter_claim_roles, add allowed_roles/allow_platform_roles to SAMLConfig (default empty = no roles from the assertion), and apply it in the ACS before the attributes become claims, so SSO paths enforce the same role policy as token paths. The JWT path now reuses the same helper. |
||
|
|
fc6d51cfc0 |
fix(auth_v2): gate OIDC login roles through the provider allowlist
The browser OIDC login/callback path accepted IdP-asserted roles straight into the session, so a malicious or misconfigured IdP could assert platform_admin over SSO and have it land in the Principal - the same escalation the bearer token path already closes. The callback now runs the mapped claims through the shared _apply_role_policy with the matched provider config before minting the session, so roles outside allowed_roles are dropped and platform roles require allow_platform_roles. With the defaults (empty allowlist, platform off) no IdP-asserted role survives. |
||
|
|
341e75ab45 |
test(auth_v2): repoint to frozen sub-package layout and pin H1 group provisioning
Follow the sub-package split (oidc/saml/scim/*) and the token-claim hardening: - import config models (OIDCProviderConfig, SAMLConfig) from the top-level package and the moved helpers from their concrete sub-modules (saml.router, saml.config, oidc.router) so tests are stable against __init__ re-export churn - resolver: a token group claim is no longer authoritative on its own; it becomes a TeamIdentity only when it resolves to a provisioned SCIM Group in the store (split into provisioned vs not-provisioned cases) Full auth_v2 suite green (153) and stable across repeated runs. |
||
|
|
a9be3d23e0 |
refactor(auth_v2): drop duplicate SCIM error helper and private re-exports
S7 is already handled self-contained by the SCIM router's route_class (renders 401/403 as a SCIM Error while preserving WWW-Authenticate), so remove the redundant errors.scim_error_response and its AuthSecurity docstring note. Also stop re-exporting private underscore helpers from the oidc/saml sub-package __init__s; the public names (config + build_*_router) remain re-exported and test code references the concrete modules for internals. |
||
|
|
fec8e0a039 |
fix(auth_v2): complete the freeze batch (re-exports, SCIM error helper, bounded SAML login state)
- Re-export the test/public helpers from the sub-package __init__s so existing imports resolve: oidc exposes _provider_key/_user_from_userinfo, saml exposes _map_attributes/_metadata_source/_user_from_mapped. - S7: add errors.scim_error_response(exc) rendering the RFC 7644 SCIM Error schema; a host registers it as the exception handler for the SCIM routes (noted on the AuthSecurity docstring). - veria-ai MEDIUM: bound the SAML outstanding-request map with a TTL (300s) and max-size eviction, the same treatment as the session store, so unauthenticated /auth/saml/login traffic can no longer accumulate login state unbounded. - Silence a fastapi/starlette generic-Request override quirk on the SCIM route's get_route_handler so mypy is clean at the freeze sha. |
||
|
|
75509686f7 |
fix(auth_v2): reject SCIM id mutation and SCIM-shape auth errors
PATCH no longer lets a scim:write principal reassign the read-only id
attribute (RFC 7643): any operation whose path targets id, or a no-path value
object carrying id, is rejected with 400 instead of rewriting the record
identifier and clobbering another resource. Authentication failures on the
guarded SCIM routes now render the SCIM Error schema (RFC 7644) for 401 and
403 via a router-scoped route class, preserving the WWW-Authenticate
challenge, rather than the generic {detail} body.
|
||
|
|
422c4df16e |
fix(auth_v2): close token-claim privilege escalation and related hardening
Token-flows and credential-flows security review fixes: - H1 (privilege escalation): a validly-signed token could self-assert platform_admin and arbitrary teams via the roles/groups claims. Roles from a token are now filtered to a per-provider allowlist (OIDCProviderConfig.allowed_roles, default empty = none) with platform-level roles gated behind an explicit allow_platform_roles flag; groups become authoritative TeamIdentity only when they resolve to a provisioned SCIM Group in the store. Introspection responses carry no roles (no per-provider policy applies). - M1: enforce iss on introspection (OAuth2IntrospectionConfig.issuer) in addition to aud. - M2: bound JWKS refetch with cache_jwk_set + a 300s lifespan and add a 10s PyJWKClient timeout, so an unknown-kid stream can't amplify into per-request network fetches. - M3: require https for issuer/jwks_uri/introspection_endpoint (loopback excepted for dev). - rbac: anchor the Casbin act matcher (^(...)$) so a "GET" policy can't grant "GETX". - basic auth: switch the reference store to pbkdf2_hmac-sha256 (600k iterations); the verifier protocol still lets deployments plug argon2/bcrypt. - LOW: generic invalid_token description instead of echoing PyJWT internals; guard the introspection response.json() and require active to be boolean true. |
||
|
|
9b9cc60994 |
refactor(auth_v2): split oidc/saml/scim into sub-packages
Per the revised design §2, each protocol that carries its own config/routes becomes a sub-package while the shared core stays flat. oidc.py -> oidc/router.py with oidc/config.py (OIDCProviderConfig); saml.py -> saml/router.py with saml/config.py (SAMLConfig + the attribute-map default); scim.py -> scim/router.py. Each sub-package __init__ re-exports its public names so call sites read from litellm.proxy.auth_v2.saml import SAMLConfig, build_saml_router. SessionConfig moves next to the SessionStore it configures in session.py. git mv preserves history; AuthConfig now composes the protocol configs from their sub-packages. |
||
|
|
0ee4397a59 |
test(auth_v2): adapt the suite to the AuthSecurity refactor and renames
Repoint the whole test surface off install_auth/app.state onto the AuthSecurity composition root: build AuthSecurity(config, store, ...) and declare routes with Security(auth.principal[, scopes]), auth.require_roles, auth.require_permission; mount routers via build_*_router(auth). Apply the PEP8 renames (JWTVerifier, APIKeyAuthenticator, OIDCAuthenticator, MutualTLSAuthenticator, OIDCProviderConfig, SAMLConfig, MutualTLSConfig, RBACEngine.has_any_role). SAML moves to the shared SessionStore + "litellm_session" cookie and session.safe_relay_state; the build_authenticators tests assert concrete types now that the scheme attribute is gone. No coverage lost; 151 tests pass. Note: routes use the default-value Security() style instead of Annotated[...] because this module runs under future annotations, where an Annotated marker is stringified and FastAPI re-evaluates it in module globals, which cannot see the closure-local auth instance. |
||
|
|
5f02c88369 |
fix(auth_v2): authenticate OIDC login sessions and adapt SCIM to AuthSecurity
The OIDC callback now mints a server-side session and sets the shared session cookie (httponly, secure, samesite=lax) before redirecting, so login yields an authenticated session that the shared SessionAuthenticator resolves; previously it returned the user record as JSON and left the caller unauthenticated. The login flow owns its CSRF protection without Starlette SessionMiddleware: it stores state, nonce and the PKCE (S256) verifier in the short-lived oauth_txn_store keyed by a temporary cookie, then on callback consumes the transaction one-time, checks the returned state, exchanges the code with the verifier and validates the nonce against the id_token. A replayed or expired state finds no transaction and returns 400. SCIM moves onto the AuthSecurity DI surface: build_scim_router(auth) reads auth.resolver and guards writes with Security(auth.principal, scopes= ["scim:write"]) instead of app.state and get_current_principal. |
||
|
|
4ec48302a5 | docs(auth_v2): one-line docstrings on AuthSecurity Security() entrypoints (D1) | ||
|
|
8302f55995 |
refactor(auth_v2): replace install_auth with AuthSecurity DI
Per user direction, drop install_auth/AuthContext/app.state entirely; the enforcement layer is now an AuthSecurity object whose bound methods are the FastAPI Security() dependencies. The app constructs AuthSecurity(config, resolver) once and passes auth.principal / auth.require_roles / auth.require_permission to Security(); routers take the instance explicitly via build_*_router(auth) and read auth.resolver/auth.config rather than request.app.state. Browser sessions are unified behind a shared SessionStore + SessionAuthenticator (session.py): one cookie, keyed on identity["method"], so SAML and the upcoming OIDC login flow share one store. AuthSecurity owns the post-login session_store and a short-TTL oauth_txn_store for OIDC state/nonce/PKCE; SessionConfig moves the cookie/TTL/redirect settings off SAMLConfig. SAML keeps its protocol-specific outstanding/replay state local. Fold in the standing judge findings: collapse the triplicated http/oauth2/oidc bearer paths into one _authenticate_bearer_jwt helper, inject the introspection async-client via a factory instead of importing litellm inline, drop the dead scheme attribute from the authenticators, and rename to PEP 8 acronym casing (JWTVerifier, OIDCAuthenticator, OIDCProviderConfig, SAMLConfig, RBACEngine, APIKeyAuthenticator, MutualTLSAuthenticator). __all__ now exports AuthSecurity, Role and the resolver protocols. scim.py and oidc.py move to build_*_router(auth) separately. |
||
|
|
c512a49fa5 |
test(auth_v2): pin the hardened auth behaviors from the security fixes
Cover the security fixes landed in |
||
|
|
450349965c |
fix(auth_v2): honor nested SCIM patch paths, align /Schemas, unshadow filter
PATCH now applies dotted attribute paths like name.givenName instead of silently dropping them, and rejects unsupported value-filter paths (emails[type eq "work"].value) with a 400 SCIM Error so behavior matches the advertised patch support. /Schemas now uses the ListResponse envelope like the other discovery endpoints, and the list route's query parameter no longer shadows the builtin while keeping the RFC 7644 ?filter= wire contract. |
||
|
|
6f3fc5eba4 |
fix(auth_v2): close mTLS spoofing, introspection audience, SAML replay, and deactivated-user gaps
Address the SSO and credential-flow security review findings: - mTLS (HIGH): the forwarded subject-DN header was trusted unconditionally, so any caller could send it and mint a service-account principal. Trust it only when the immediate peer is inside trusted_proxy_cidrs (same model as XFF) and fail closed otherwise; the ASGI-TLS-extension path already fails closed when no verified cert is present. - OAuth2 introspection (HIGH): RFC 7662 responses were accepted regardless of audience. Enforce the response aud against OAuth2IntrospectionConfig.audience and reject active tokens whose audience does not match. - SAML (HIGH): default allow_unsolicited to False so IdP-initiated/login-CSRF responses are rejected, add a single-use assertion-id replay cache, and bind the post-login redirect to the RelayState stored against the matched InResponseTo request rather than trusting the echoed form field. - Deactivated users (M1): the resolver now rejects a credential that resolves to a SCIM user with active=False, so deactivation actually blocks authentication. - Stop carrying underscore-prefixed carrier keys (raw api key, basic password) into Principal.claims, which is documented for audit logging. |
||
|
|
ca896ac073 |
fix(auth_v2): make SCIM discovery public and return 404 on missing DELETE
RFC 7644 requires /ServiceProviderConfig, /ResourceTypes and /Schemas to be publicly readable; split them onto an unguarded router while Users and Groups stay behind scim:write. DELETE on a missing User or Group now returns a 404 SCIM Error instead of a misleading 204. |
||
|
|
71a189bf65 |
fix(auth_v2): harden HTTP basic, SAML sessions, and JWKS fetch
Address Greptile security findings in the authenticator, SAML and config layers: - HTTP Basic accepted any password and copied the cleartext password into Principal.claims. Verify the password against an injected BasicAuthVerifier (InMemoryBasicAuthStore holds username -> salted sha256, constant-time compared with hmac.compare_digest) and stop putting the password in the credential; basic with no configured verifier now rejects rather than trusting the caller. - SAML session cookie gains the Secure flag (httponly and samesite=lax already set), gated by SamlConfig.cookie_secure. - SAML session store gains TTL expiry and max-size eviction (SamlConfig.session_ttl_seconds / session_max_size) so it can no longer grow unbounded or hand out stale sessions. - JWKS signing-key lookup ran synchronously inside the async request path and blocked the event loop on a cache miss; run the JWT verify off-loop via starlette run_in_threadpool on the http-bearer, oauth2 at+jwt and oidc paths. |
||
|
|
2dd336a795 |
test(auth_v2): move tests under proxy/ to mirror the module relocation
The module moved from litellm/auth_v2 to litellm/proxy/auth_v2 (commit
|
||
|
|
104a5e1443 |
refactor(auth_v2): move module under litellm/proxy/auth_v2
The module imports FastAPI and is proxy-only, so it belongs under litellm/proxy beside the legacy litellm/proxy/auth rather than at the top level. git mv preserves history; the package is self-contained so the relative imports are unchanged, and the scim2-models mypy override is path-independent. |
||
|
|
55a332bb91 |
test(auth_v2): cover Casbin-backed RBAC hierarchy and permissions
RBAC moved to an embedded Casbin enforcer: require_roles now honors the role hierarchy and require_permission gates object/action against the policy. - rbac: RbacEngine.has_role inherits down the g-rules (platform_admin satisfies an org_admin/team_member gate, org_admin satisfies org_viewer, team_admin satisfies team_member) and never climbs (team_member fails an org_admin gate); enforce honors the default policy (platform_admin any obj/act incl keyMatch2 on /scim/v2/*, platform_viewer read-only, org_viewer no write) and an operator CSV fully replaces the in-code defaults - security: require_roles passes a higher role through a lower-role gate via the hierarchy; require_permission allows platform_admin, denies a viewer on write with detail "Forbidden", and 401s when unauthenticated; an RbacEngine injected onto the AuthContext overrides the default policy (operator CSV path) Replaces the removed has_any_role coverage. Mutation-checked: dropping the hierarchy lookup or short-circuiting enforce fails these. |
||
|
|
3dc660a135 | test(auth_v2): cover OAuth2 token introspection over the cached async client | ||
|
|
e309003c84 |
feat(auth_v2): back RBAC with Casbin
Replace the hand-rolled has_any_role set check with a Casbin-backed RbacEngine (per design 03 §4). The engine wraps casbin.Enforcer over an embedded RBAC model (request sub/obj/act, g role hierarchy, keyMatch2 on obj, regexMatch on act) and a default in-code policy: platform_admin inherits org_admin/team_admin/platform_viewer, org_admin inherits org_viewer, team_admin inherits team_member; grants platform_admin /* .*, platform_admin /scim/v2/* .*, platform_viewer /* GET. Operators can replace the whole policy with a CSV via AuthConfig.casbin_policy_path (FileAdapter); no DB adapter yet. require_roles now honors the hierarchy through the enforcer's grouping (get_implicit_roles_for_user) instead of exact-match membership, so a platform_admin passes a require_roles(ORG_ADMIN) gate; signature and 403 semantics are unchanged. New require_permission(obj, act) dependency runs get_current_principal then RbacEngine.enforce and 403s on deny. The engine is built in install_auth and injectable for tests via a new rbac kwarg. Scope checks stay plain SecurityScopes (a token property, not policy). Adds casbin to the proxy extra (pure python, no native deps). |
||
|
|
7cf35ccc0a |
test(auth_v2): cover SCIM scim:write guard and SAML RelayState redirect
Follow the auth module updates: SCIM routes now require the scim:write scope,
and the SAML ACS/login flow redirects to a validated RelayState instead of
returning JSON.
- scim: authenticate every request with a scoped key, and pin the guard
directly: no credential -> 401 with WWW-Authenticate, an authenticated
principal without scim:write -> 403 insufficient_scope
- saml: assert ACS returns 303 to a safe RelayState ("/dashboard") and falls
back to default_redirect_path for an absolute/"//host" RelayState; assert
GET /login threads ?next= through as a validated RelayState; add a garbage
SAMLResponse -> 401 case
A mutation spot-check confirmed the open-redirect tests fail when the
_safe_relay_state guard is bypassed.
|
||
|
|
d3878ee9a9 |
fix(auth_v2): require scim:write auth on all SCIM routes
The SCIM router mounted /scim/v2/* with no security dependency, so any caller could create or delete users and groups unauthenticated. Guard the whole router with Security(get_current_principal, scopes=["scim:write"]) per design 03 §11, so provisioning callers authenticate with the same bearer token or API key as every other route and the scope gates them: unauthenticated requests now 401, an authenticated principal without scim:write gets 403 insufficient_scope. Also document two deployment facts uncovered alongside this: uvicorn's --proxy-headers rewrites request.client from X-Forwarded-For before this module's trusted_proxy_cidrs check runs and silently bypasses it (install_auth docstring), and the scheme_order precedence where HTTP precedes openIdConnect so a bearer JWT is labeled bearer_jwt rather than oidc (both verify identically). |
||
|
|
03faccea17 |
test(auth_v2): add test suite for the standards-based auth module
Cover every layer of litellm/auth_v2 with tests that fail when the behavior regresses, not just for coverage. Highlights: - authenticators: JwtVerifier enforces signature, aud, iss, exp, required claims, and at+jwt typ via an injected jwks_client (real RS256 against an in-test RSA keypair, no monkeypatching); per-scheme apiKey/http-bearer/ http-basic/oauth2/oidc/mTLS extraction and fail-fast on present-but-invalid - security: OR precedence first-match-wins, a present-but-invalid api key does not fall through to a valid bearer, scope -> 403 insufficient_scope, role -> 403, missing credential -> 401 with WWW-Authenticate, network wired onto the principal - resolver: sha256 api-key lookup (wrong key never resolves), claims-driven principal build (groups -> teams, roles filtered to the Role enum), mTLS -> service account - network: trusted-proxy XFF honored only from a trusted peer, right-to-left parse skips chained proxies, spoofed XFF from an untrusted peer ignored - scim: Users/Groups create/get/patch/list/delete round-trip plus malformed body -> SCIM 400 Error and discovery endpoints - oidc: userinfo -> scim2_models.User mapping and the upsert seam - saml: a real pysaml2 IdP mints a signed assertion; ACS provisions the user, sets a session cookie, and authenticates with method=saml, while tampered and unsigned assertions are rejected (skipped when xmlsec1 is absent) - models/rbac/config: frozen Credential, Role validation, scope/role helpers, SamlConfig metadata validation A mutation spot-check confirmed the suite fails when JWT verification or the api-key hash lookup is broken. |
||
|
|
da1a088a4f |
feat(auth_v2): redirect after SAML ACS with validated RelayState
Replace the test-convenience JSON body from /acs with the standard SP flow: set the session cookie, then 303 redirect to the RelayState the IdP echoes back, or to SamlConfig.default_redirect_path (default "/") when it is absent. RelayState is validated to block open redirects - only relative paths are honored (must start with a single "/", reject "//", any scheme, and backslashes), and anything else falls back to the default. GET /login threads a ?next= query param through as RelayState with the same validation so the post-login landing page survives the round trip. |
||
|
|
677762bf60 |
refactor(auth_v2): align SamlConfig with the revised design doc
Match the updated 03-design.md SAML spec: rename sp_entity_id to entity_id, collapse the split idp_metadata_path/idp_metadata_inline into one idp_metadata field accepting inline XML, a local path, or a remote URL, and default the attribute_map to the common Okta/Entra claims (email, givenName, surname, groups). Make assertion signing mandatory by hardcoding want_assertions_signed rather than exposing it as a togglable field. Map givenName/surname into the SCIM User's Name (given/family/formatted) and email into emails, and fail the ACS closed with 401 on any parse or signature-verification error. |
||
|
|
0b74ffa9c6 |
feat(auth_v2): implement full SAML 2.0 SP via pysaml2
Replace the deferred SAML thin-adapter stub with a working Service Provider built on pysaml2: an SP metadata endpoint, an SP-initiated /login that redirects to the IdP, and an ACS handling the HTTP-POST binding that verifies the signed assertion, maps NameID and attribute statements into a scim2_models.User, and upserts it through the same ProvisioningStore seam SCIM and OIDC use. A SamlAuthenticator reads the post-ACS session cookie and resolves to the one normalized Principal like every other scheme; AuthMethod gains a SAML member. IdP metadata loads from a file path or inline XML via SamlConfig, and install_auth mounts the router and authenticator when SAML is enabled. pysaml2 pulls pyOpenSSL transitively without pinning it, and older pyOpenSSL caps cryptography below 46 and breaks at import against the version this proxy already requires; pin pyOpenSSL>=26 so the resolver stays on a cryptography-46-compatible release. pysaml2 also needs the system xmlsec1 binary at runtime (brew install libxmlsec1 on macOS, apt-get install xmlsec1 libxmlsec1-dev on Debian); SamlConfig.xmlsec_binary can point at it when it is not on PATH. |
||
|
|
a0a59a2197 |
feat(auth_v2): add standards-based auth and identity module
New additive litellm/auth_v2 package: a thin orchestration layer over PyJWT, Authlib and scim2-models behind FastAPI's native Security() primitives that normalizes every credential into one standards-shaped Principal carrying org/team/user and network identity. Authentication, identity resolution, authorization and enforcement are kept as separate layers. Five authenticators cover the OpenAPI scheme types (apiKey, http bearer-JWT/basic, oauth2 at+jwt + introspection, openIdConnect, mutualTLS); a shared JwtVerifier enforces signature, issuer, audience and exp on every JWT path via a cached PyJWKClient. RBAC is a flat Role enum plus scope/role checks wired through SecurityScopes. Missing or invalid credentials return 401 with an RFC 9110/6750 WWW-Authenticate challenge, scope failures return 403 insufficient_scope. SCIM 2.0 Users/Groups/PATCH/discovery and an Authlib OIDC login flow share one ProvisioningStore seam; the SAML SP is a documented thin adapter pending pysaml2. The module is unimported by the proxy app and depends on nothing in litellm/proxy/auth. |
||
|
|
2bbf688613 |
build(auth_v2): add Authlib and scim2-models for the auth_v2 module
Pull in the OSS libraries the standards-based auth module orchestrates: Authlib for the OIDC login flow and scim2-models for SCIM 2.0, and switch PyJWT to the [crypto] extra so JWKS-backed RS256 verification is explicit (cryptography was already a proxy dependency). scim2-models ships py.typed but its generic, alias-driven models trip mypy's call-arg check though they work at runtime, so treat the library as untyped at the boundary in both litellm/mypy.ini (used by CI) and the root pyproject mypy config. |
||
|
|
dff25fef44
|
feat(proxy): add option to disable server-side prepared statements for DB lookups (#29984) | ||
|
|
3bd3951e37
|
fix(proxy): recover from cached-plan errors by reconnecting the Prisma client (#29983) | ||
|
|
1436ee9092
|
fix(mcp): drop orphaned per-user credential rows when an MCP server is deleted (#30141) | ||
|
|
7899463c6a
|
fix(callbacks): forward callback_settings to callback initializers and guard consumers against non-dict values (#30161)
* fix(datadog): pass callback_specific_params so DatadogCostManagementLogger receives cost_tag_keys (#29590) * fix(datadog): pass callback_specific_params so DatadogCostManagementLogger receives cost_tag_keys * test(proxy): regression test that load_config forwards callback_specific_params * fix(proxy): guard lakera_prompt_injection callback_specific_params against non-dict Addresses review feedback: forwarding callback_settings as callback_specific_params (so DatadogCostManagementLogger receives cost_tag_keys) exposed the lakera_prompt_injection branch, which did lakeraAI_Moderation(**callback_specific_params ["lakera_prompt_injection"]) with no type guard. A config like `callback_settings: {lakera_prompt_injection: "any-string"}` then hit `**"any-string"` -> TypeError: argument after ** must be a mapping, not str. Guard the lakera branch with isinstance(dict), matching the existing presidio and datadog_cost_management branches (non-dict values fall back to {}). Add a regression test asserting initialize_callbacks_on_proxy ignores a non-dict value instead of crashing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: inject fake lakera_ai module to avoid importing the real one CI fix for the lakera regression test: it stubbed litellm.proxy.proxy_server with a SimpleNamespace and then monkeypatch.setattr'd the real lakera_ai module, which forces importing it — and lakera_ai does `from litellm.proxy.proxy_server import LiteLLM_TeamTable`, absent on the stub -> ImportError under proxy-infra tests. Inject a fake lakera_ai module into sys.modules instead, so the callbacks branch's `from ...lakera_ai import lakeraAI_Moderation` resolves to the stub without loading the real module. The guard under test (isinstance(dict) in the lakera branch) is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(callbacks): guard compression/websearch interceptors against non-dict callback_settings (#30153) #29590 forwards the full callback_settings dict into initialize_callbacks_on_proxy, which activates the compression_interception and websearch_interception consumers. Their initialize_from_proxy_config read the callback_settings subkey without an isinstance(dict) guard, so a non-dict value such as `compression_interception: true` reached from_config_yaml(...).get(...) and aborted proxy startup with AttributeError. #29590 added that guard for lakera_prompt_injection but not for these two Mirror the isinstance(dict) guard already used by the lakera, presidio, and datadog branches so a non-dict value is ignored and the callback initializes with defaults. A parametrized test feeds every callback_settings consumer a non-dict value through initialize_callbacks_on_proxy to catch a future consumer that forgets the guard * fix(callbacks): normalize non-dict callback_specific_params to empty dict A blank callback_settings: key in YAML loads as None, and config.get('callback_settings', {}) returns None because dict.get only falls back to the default when the key is absent. Forwarding that value verbatim to initialize_callbacks_on_proxy made the first '<name>' in callback_specific_params membership test raise TypeError: argument of type 'NoneType' is not iterable, aborting proxy startup. Same failure for any non-dict root such as callback_settings: true. Normalize the value at the function boundary so both callsites (and any future ones) initialize callbacks with their defaults instead of crashing. --------- Co-authored-by: Hedi Daoud <150018939+hdaoud23@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
20e453f698
|
feat(cli): per-agent lite claude / codex / opencode commands that wrap coding agents through the proxy (#29850)
* feat(cli): add `litellm-proxy run -- <agent>` to wrap coding agents through the proxy Wraps Claude Code, Codex, OpenCode, and any other coding agent so all of its LLM traffic routes through a LiteLLM proxy, with the agent-vault style of "just works" DX: one `run -- <agent>` command, auto SSO login when interactive, env-key "agent mode" for containers/CI, and a fail-fast key check against the proxy so bad credentials error immediately instead of deep inside the agent. The wrapped binary is detected by name to pick the right variables. Claude Code gets ANTHROPIC_BASE_URL (the bare proxy root, so it appends /v1/messages) and ANTHROPIC_AUTH_TOKEN, with any stray ANTHROPIC_API_KEY cleared so the proxy token wins. Codex and OpenCode get OPENAI_BASE_URL (proxy + /v1) and OPENAI_API_KEY. Unrecognized commands get both sets so they work either way. `litellm-proxy claude-code` remains as a shortcut for `run -- claude`. The core logic is split into dependency-injected helpers (agent_profile, build_agent_env, verify_proxy_key, run_agent) so env wiring, the preflight, and the launch handoff are unit-tested without monkeypatching, alongside CliRunner tests for auth resolution, agent mode, and auto-login. Mutation-tested the env profiles, preflight, and agent-mode branch to confirm the tests fail when the behavior is broken. https://claude.ai/code/session_0154VpLXW7mMvk5wfbgPRJa6 * Make each coding agent its own litellm-proxy command Replace the `run -- <agent>` interface and the `claude-code` shortcut with top-level commands generated per known agent, so launching is just `litellm-proxy claude`, `litellm-proxy codex`, or `litellm-proxy opencode`, with everything after the agent name forwarded straight to it. This drops the ceremony of `run --` and cuts typing. The `--model`/`--small-fast-model` wrapper flags are gone; pass the agent's own model flag instead, or export the model env vars (the wrapper preserves what you already have set), which keeps the surface minimal and avoids intercepting flags the agent owns. Rename the module to agents.py to match. * fix(cli): route `litellm-proxy codex` through the proxy via a custom provider Codex ignores OPENAI_BASE_URL (it always dials api.openai.com over the Responses WebSocket transport), so the OpenAI env profile alone left `litellm-proxy codex` talking to OpenAI directly instead of the proxy. Point Codex at the proxy with a custom provider passed as `-c` config overrides, and force the HTTP/SSE Responses transport with supports_websockets=false since the proxy does not speak the Responses WebSocket protocol. The provider reads its key from OPENAI_API_KEY, which the agent env already exports. The overrides are injected ahead of the user's args so they precede Codex's subcommand. Claude Code and OpenCode are unaffected; they honor the exported env vars. Adds regression tests for the per-agent launch args and the injection ordering. Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com> * Rename litellm-proxy CLI command to lite The proxy management CLI was invoked as litellm-proxy, which is a lot to type for an everyday command. Rename the console script entry point to lite and update the in-CLI usage examples, help text, error messages and docs to match. * fix(sso): stop CLI auth success page from hanging on "Closing..." The CLI opens the SSO success page with webbrowser.open, so the tab is not script-opened and the browser refuses window.close(). The countdown would end on "Closing..." and the tab would sit there forever. Drop the countdown and just show "You can now close this window and return to your terminal." from the start, while still attempting window.close() once so the tab auto-closes in the rare case the browser allows it. Add a regression test asserting the manual-close instruction is always present and the misleading countdown/"Closing..." text is gone. * fix(cli): reattach controlling terminal after SSO login, keep litellm-proxy alias When the first `lite claude` has to log in via browser SSO, completing the login could leave stdin detached from the terminal, so a TUI agent like Claude Code would start in non-interactive mode and exit with "Input must be provided". The wrapper now reopens the controlling terminal onto stdin just before handoff when the session started interactively; piped or redirected input is detected up front and left alone, so agent-mode and non-interactive use are unchanged. Also keep the `litellm-proxy` console script as an alias for `lite` so existing scripts and CI that invoke `litellm-proxy` keep working; both names map to the same CLI. * feat(install): make the curl installer need only curl, not a pre-existing Python The installer now lets uv provision a managed Python 3.13 when no suitable interpreter is found, instead of aborting. The minimum is also bumped from 3.9 to 3.10 to match the package's requires-python (>=3.10), so a system Python 3.9 is no longer selected only for uv tool install to reject it. * feat(cli): add thin litellm[cli] install path (install-cli.sh + brew) for the lite CLI On a developer laptop the `lite` CLI only needs `lite login` and running coding agents through a proxy, but the sole install path was `litellm[proxy]`, which drags in the whole server tree (fastapi, uvicorn, boto3, polars, cryptography, litellm-enterprise). The CLI's heavy imports are all guarded, so it runs on the base SDK plus just rich, pyyaml and requests. Add a `cli` extra carrying exactly those three, a `scripts/install-cli.sh` curl one-liner that installs `litellm[cli]`, and a `BerriAI/homebrew-litellm` tap formula with a release runbook under `packaging/homebrew/`. The installer passes no `--python`, so uv honours litellm's requires-python and provisions a managed interpreter, skipping a too-old (3.9) or too-new (3.14+) system Python instead of failing to resolve. A pyproject thin-contract test asserts the `cli` extra keeps the deps the CLI imports and never leaks a server-only dependency from `proxy`, so the laptop install cannot silently re-bloat * fix(install): let uv pick the Python via --python-preference system Both installers detected a system Python with a floor-only check and forced it with `uv tool install --python <interp>`. On a host whose only Python is outside litellm's requires-python (a too-old 3.9 or, increasingly, a too-new 3.14) that forced an incompatible interpreter and the resolve failed. Drop the detection and pass `--python-preference system`: uv reuses a compatible system Python when present and downloads a managed one otherwise, always honouring requires-python * test(router): filter aiohttp unclosed-session gc noise in test_async_fallbacks test_async_fallbacks asserts the last three captured log records are the router's fallback messages. Under the litellm_router_testing job (pytest -k router -n 4) many router tests share the module-level in_memory_llm_clients_cache (max 200, ttl 3600s). Older cached OpenAI/Azure clients get evicted while their aiohttp ClientSession is still open, and when the gc reclaims them aiohttp emits "Unclosed client session"/"Unclosed connector" through the asyncio logger. Those records land in caplog mid-test and push the expected router logs out of the last-three window, so the assertion flips to failing non-deterministically. These warnings are async cleanup noise, not router debug logs, so filter them out exactly like the existing leaked-task warnings before asserting order. The assertion on the three router fallback messages is unchanged. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
a4a3348801
|
[internal copy of #28007] Fix/gcp model garden streaming (#28363)
* fix(vertex): stream Model Garden Gemma/Qwen responses correctly through /v1/messages * test(vertex): cover _CombinedChunkSplitter defensive branches * test(databricks): rename test file to avoid duplicate basename collision * fix(databricks,anthropic): defensive token defaults; document single-mode splitter Address greptile P2 concerns: - databricks: default usage token fields to 0 when constructing ChatCompletionUsageBlock from a partially populated usage block — matches the defensive pattern used in ollama/vertex_ai/cohere/bedrock. - _CombinedChunkSplitter: clarify in the docstring that an instance is single-mode (sync or async, not both), since the two iteration paths hold independent upstream iterator references. Co-authored-by: Claude <claude@anthropic.com> --------- Co-authored-by: Steven Kessler <9701252+stvnksslr@users.noreply.github.com> Co-authored-by: Claude <claude@anthropic.com> |
||
|
|
410b892f77
|
fix(register_model): preserve built-in cache pricing when registering custom overrides under unmapped keys (#30044)
* fix(spend-tracking): fall back to direct spend-counter increment when reservation reconcile fails When the reservation-reconcile path in `_reconcile_budget_reservation_for_counter_update` hits a Redis error, it now correctly returns an empty set so that `increment_spend_counters` re-runs the direct increment for the affected counters. Previously, the function logged the failure, invalidated the reserved counters, and still returned the reserved counter keys, which caused the caller to skip the direct increment. With the increment skipped and the counter deleted, the next request reseeded the counter from `LiteLLM_VerificationToken.spend`, a column the batched flusher only updates every few seconds, so the enforced cross-pod spend value collapsed to a stale snapshot and budget gating stopped firing for affected keys. Adds a regression test that exercises the failure path with a flaky redis backend and asserts the actual response cost lands in the shared counter. * fix(register_model): preserve built-in cache pricing when registering custom overrides under unmapped keys When a custom-priced model is registered under a key shape that get_model_info cannot resolve (e.g. litellm_params.model set to bedrock/bedrock/us.anthropic.claude-sonnet-4-6 or another non-canonical alias), register_model previously fell back to an empty existing_model. The merged entry then carried only the fields the user set explicitly (input/output cost, provider) and dropped cache pricing. Downstream the cost calculator defaulted cache_creation_input_token_cost and cache_read_input_token_cost to 0, silently dropping the bulk of the bill for cache-heavy Anthropic traffic. register_model now attempts to resolve a canonical built-in entry by stripping provider prefixes, region prefixes, and provider-specific suffixes before giving up. When a variant resolves, its defaults (notably cache pricing) are inherited while the user's explicit overrides still win. When nothing resolves and the user supplied no cache pricing, it logs a warning instead of silently under-billing. * fix(router): inherit built-in cache pricing on deployments with partial custom pricing A deployment configured with only input_cost_per_token and output_cost_per_token under model_info was being registered under its model_info.id with no cache cost fields. The cost calculator then defaulted cache_creation_input_token_cost and cache_read_input_token_cost to 0, silently billing cache_read and cache_creation tokens at zero. For cache-heavy Anthropic traffic this drops the bulk of the bill. When the deployment's litellm_params.model resolves to a built-in cost-map entry, pull the cache pricing fields from there before registering. User-specified cache fields still win on merge; only missing fields are inherited. Pairs with the register_model fallback added earlier in this branch: that handles unmapped key shapes like bedrock/bedrock/x, this handles deploy-id keys whose backend model is mapped. * fix(register_model): inherit only cache pricing on unmapped-key fallback, not provider The unmapped-key fallback in register_model copied the entire resolved built-in entry, so registering openai/command-r-plus inherited the cohere built-in's litellm_provider and get_model_info(custom_llm_provider=openai) could no longer resolve it. Restrict the fallback to the cache-pricing fields, matching the router-side _inherit_builtin_cache_pricing, so the cache-cost dropout stays fixed without clobbering the registered provider. Add a direct unit test for Router._inherit_builtin_cache_pricing so the router coverage check sees it, and pin the fixed spend-counter contract: when reservation reconcile fails the counter must hold the directly incremented cost rather than being left at None. |
||
|
|
a75ed0079c
|
chore(ui): make knip recognize .mjs scripts and openapi-typescript (#30052)
The knip entry/project globs only matched scripts/**/*.ts, so the two .mjs scripts went unanalyzed and produced "no matches" config hints. openapi-typescript was also reported as unused because gen-api-types.mjs invokes its binary through a dynamic execFileSync path that knip cannot trace statically; ignoreDependencies records that it is genuinely used. |
||
|
|
f9293d40c4
|
fix(proxy): self-heal startup/reload prisma reads on engine disconnect (#28803) | ||
|
|
3b40ac987f
|
Litellm oss 090626 (#30021)
* fix(mcp): report scoped server name during initialize (#29865) * fix mcp scoped server name * Update litellm/proxy/_experimental/mcp_server/mcp_context.py Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * test(mcp): cover scoped server name in the SSE initialize handler --------- Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * fix(ui): show all session logs in the drawer, not just the first 50 (#29795) * fix(ui): show newest session logs first * test(ui): keep session log pagination coverage * fix(ui): show all session logs in the drawer, not just the first page The session detail drawer fetched session logs via sessionSpendLogsCall without page/page_size, so it only ever received the backend default of one page (50 rows). Sessions with more than 50 calls had the rest unreachable in the UI (#29153). sessionSpendLogsCall now takes page/page_size, and the drawer fetches the first page, reads total_pages, then fetches the remaining pages and accumulates them before the existing client-side sort. This keeps the single continuous list (and the selected-log lookup and keyboard navigation, which all assume the full session) correct. Fetching is bounded by a page cap, and the sidebar shows a "showing most recent N" note if a session exceeds it. The rows are lightweight metadata (the endpoint excludes messages/response), so the full set is small; request/response bodies are still loaded per log on demand. * fix(ui): default session drawer to most recent log, newest first Open a session with its most recent log selected, and order the sidebar newest-first to match the all-sessions logs overview. MCP calls stay grouped last. The latest log by time is computed explicitly, since the MCP grouping means it is not always the first row. * Apply fetching pages in batches suggestion from @greptile-apps[bot] Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * fix(ui): derive session total from accumulated rows when backend omits it Compute the session total after all pages are fetched, falling back to the accumulated row count rather than the first page's. Guards the truncation note against a backend response that omits total but spans multiple pages. --------- Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * fix(proxy): handle Mistral multipart passthrough (#29927) * fix(proxy): handle Mistral multipart passthrough * chore: satisfy passthrough ci formatting * test(proxy): cover Mistral passthrough in CI shard * fix(vertex_ai): use REP host for context caching on eu/us multi-region endpoints (#29573) Context caching built the cachedContents URL as https://{location}-aiplatform.googleapis.com, which is an invalid host for the eu/us multi-region endpoints and returns 404. The inference path already resolves these to the REP host (https://aiplatform.{geo}.rep.googleapis.com) via get_vertex_base_url(); reuse that helper in _get_token_and_url_context_caching so caching uses the same host as inference. Adds tests covering the eu/us multi-region cachedContents URLs (v1 and v1beta1). Fixes #29571 * Support per-model encrypted content affinity config (#29760) Co-authored-by: shin-berri <shin-laptop@berri.ai> Co-authored-by: yuneng-jiang <yuneng@berri.ai> * fix: propagate upstream status code in proxy API exception handler (#29402) * fix: propagate upstream status code in proxy API exception handler When Google GenAI / Vertex returns a 404 for deprecated or missing models via streamGenerateContent, the exception was falling through to a generic handler that defaulted to 500. Now provider exceptions carrying a valid HTTP status_code correctly propagate it through to the ProxyException. * fix: apply black formatting to common_request_processing.py * fix: tighten status code range to 400-599 and deduplicate ProxyException raise * fix(tests): use valid vertex_location in context caching tests Replace "test_location" (contains underscore) with "us-central1" so tests pass the regex validation added in get_vertex_base_url(). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(sdk): add xAI OAuth provider (#29866) * Add xAI OAuth provider * Update oauth.py Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * Fix xAI OAuth CI failures * Add xAI OAuth coverage tests * Move xAI OAuth coverage tests to core utils * Address xAI OAuth review comments * Prevent xAI OAuth api_base token exfiltration * Treat blank xAI OAuth api keys as absent * Wrap invalid xAI OAuth JSON responses * Use xAI OAuth behind explicit flag --------- Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * fix(proxy) #27734 allow clearing budget_duration and team_member fields by sending null on /key/update and /team/update (#27751) * fix(proxy): allow clearing budget_duration and team_member fields by sending null on /key/update and /team/update Fixes #27734 Sending null for budget_duration, team_member_budget, team_member_budget_duration, team_member_rpm_limit, or team_member_tpm_limit via /key/update or /team/update returned 200 OK but silently ignored the null value. The fields remained unchanged in the database. Root causes: - /key/update: prepare_key_update_data() popped budget_duration from the update dict but never re-added it (or budget_reset_at) when the value was None. - /team/update: _set_budget_reset_at() only acted when budget_duration was non-None, leaving a stale budget_reset_at in the DB. - /team/update: team_member_* null values bypassed the budget table update entirely because should_create_budget() requires at least one non-None field. * test(proxy): cover no-budget-row path in clear_team_member_budget_fields * fix(presidio): unmask PII tokens in Anthropic native SSE streaming bytes (#30028) * fix(presidio): unmask PII tokens in Anthropic native SSE streaming bytes When output_parse_pii=true on the Anthropic native path (anthropic/claude-*), response chunks arrive as raw bytes in SSE format. _stream_pii_unmasking was yielding those bytes unchanged, so <PERSON_1> tokens were never replaced with the original values before reaching the caller. Add _unmask_sse_bytes_chunk to parse each data: line, find content_block_delta / text_delta events, and apply _unmask_pii_text before re-encoding. Wire it into _stream_pii_unmasking so bytes chunks are unmasked when pii_tokens exist. * fix(presidio): handle CRLF line endings and non-ASCII PII in SSE unmask Strip trailing \r before the [DONE] guard so CRLF-terminated SSE chunks don't bypass it and silently swallow a JSONDecodeError. Add ensure_ascii=False to json.dumps so non-ASCII replacement values like accented names are preserved as UTF-8 on the wire rather than being \uXXXX-escaped. Add regression tests for both cases. * feat(bedrock_mantle): path-aware Responses routing (/v1/responses vs /openai/v1/responses) (#29925) * feat(bedrock_mantle): path-aware Responses routing (/v1/responses vs /openai/v1/responses) Bedrock Mantle serves the Responses API on two upstream paths: - gpt frontier models (gpt-5.5 / gpt-5.4) on /openai/v1/responses - every other Responses-capable model (e.g. gpt-oss) on the standard /v1/responses BedrockMantleResponsesAPIConfig gains a `use_openai_path` flag; the provider gate in utils.py picks the path per model: openai.gpt-* (non gpt-oss) -> /openai/v1/responses; any model declared mode=responses (price-map entry or user model_info) -> /v1/responses; everything else returns None and keeps the existing chat-completions emulation. Adds gpt-5.5 / gpt-5.4 price-map entries, registry wiring, and the routing-matrix tests. * feat(bedrock_mantle): data-driven frontier routing via use_openai_responses_path Addresses the Greptile review point that frontier detection should be a price-map field rather than a hardcoded name match. The gate now routes a model to /openai/v1/responses when its price-map entry declares use_openai_responses_path, so a frontier model whose name does not follow the openai.gpt- convention can be onboarded by JSON alone. The name-convention check is kept as a fallback that needs no price-map entry, which preserves zero-change routing for a future gpt-6 before its entry loads. gpt-5.5 / gpt-5.4 get the flag in both price maps. Adds tests for the data-driven flag path and for the flag presence on the gpt-5.x entries; both branches are mutation-tested. * test(model_prices): allow use_openai_responses_path in price-map schema The model_prices_and_context_window.json schema validator (test_aaamodel_prices_and_context_window_json_is_valid) enforces additionalProperties: false, so the new use_openai_responses_path flag on the gpt-5.5 / gpt-5.4 entries failed validation. Add it to the schema as a boolean, alongside the other supports_* / capability flags. * Add Tensormesh serverless models to the model cost map (#30037) * Add Tensormesh serverless models to the model cost map * Flag reasoning support on the Tensormesh models that expose thinking mode * fix(proxy): invalidate stale key spend counter after budget reset or manual spend update (#30001) * fix(proxy): reconcile stale key spend counter after budget reset * fix(proxy): invalidate stale key spend counter after budget reset or manual spend update * fix(proxy): remove read-time stale counter reconciliation to prevent budget bypass * revert: undo unrelated formatting changes in enterprise directory * test(proxy): add unit test for key spend update invalidating counter * test(proxy): fix mocked update_data and hash token expectations in unit test * fix(proxy): use Responses-API transformer in pass-through cost tracking (#29728) The `elif is_responses:` branch of `openai_passthrough_handler` was calling the chat-completions `transform_response` on a Responses API payload. The chat-completions transformer expects `choices: [...]` in the raw response; the Responses API uses `output: [...]` and `usage.input_tokens` / `usage.output_tokens` (not `prompt_tokens` / `completion_tokens`). The result was a KeyError 'choices' deep inside `convert_to_model_response_object`, swallowed by the surrounding `except Exception` in the handler, and the SpendLogs row was written by the fallback path with zeroed-out tokens, spend, and model. This bug silently undercounts cost for every successful pass-through call to either OpenAI's `/v1/responses` or Azure's `/openai/v1/responses` (deployments configured for the Responses API). Reproduced 2026-06-04 against a real Azure OpenAI Responses API deployment proxied through LiteLLM v1.88.0. Fix: use the dedicated `OpenAIResponsesAPIConfig.transform_response_api_response` for the Responses branch. This transformer already exists in LiteLLM (`litellm/llms/openai/responses/transformation.py`) and knows the Responses-API on-the-wire shape. `litellm.completion_cost` already handles `ResponsesAPIResponse` natively with `call_type="responses"`, so no downstream changes are needed. Tests: test_responses_api_uses_responses_transformer_not_chat_completions NEW. Real regression test — exercises the openai_passthrough_handler with a real-shaped Responses payload (no `choices`, has `output` and Responses-API `usage` keys) and NO mocked `get_provider_config`. Pre-fix: raises KeyError 'choices' inside the chat-completions transformer (the bug). Post-fix: returns a ResponsesAPIResponse, completion_cost is called with call_type="responses" and a ResponsesAPIResponse instance (asserted). Verified to fail on un-fixed handler + pass on fixed handler before commit. test_responses_api_cost_tracking UPDATED. Old test mocked `get_provider_config` (no longer called in the responses branch post-fix). Now mocks the Responses transformer directly (`OpenAIResponsesAPIConfig.transform_response_api_response`) to test the downstream cost-calc contract. Out of scope for this PR (separate followup): - Recognizing *.cognitiveservices.azure.com (the newer Azure OpenAI hostname) in the is_openai_*_route checks. Separate PR. Co-authored-by: shin-berri <shin-laptop@berri.ai> Co-authored-by: yuneng-jiang <yuneng@berri.ai> * fix(skills): execute DB skills by matching the litellm_skill_ tool name prefix (#30116) Skill IDs are generated as litellm_skill_<uuid> and the model-facing tool name is the sanitized skill ID, but the post-call execution gates in SkillsInjectionHook only ran tools whose name starts with "skill_", so DB skills were silently returned to the client as raw tool calls. Fixes #28122. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(anthropic): synthesize content_block_start when Responses stream omits output_item.added (#30115) * fix(team): reserve team budget raises for proxy admins on /team/update (#30030) The caller's PERSONAL max_budget was the wrong yardstick for /team/update: a team's spend ceiling has nothing to do with the admin's own key budget. That comparison was an unintended side effect of reusing _check_user_team_limits() (which exists for the /team/new path) and broke the UI, which re-sends the unchanged budget on every save. New behavior on /team/update for standalone teams: - A team admin (already authorized via _verify_team_access) may freely KEEP or LOWER the team budget, and change models/tpm/rpm, without being gated by their personal limits. - GROWING a team's spend ceiling is a budget-authority action reserved for proxy admins -> 403 for team admins. "Growing" covers both raising max_budget above the team's current finite value and removing the cap entirely (max_budget=null, detected via model_fields_set so an explicit null is distinguished from an omitted field). For a team that currently has no cap, setting a finite value is a restriction and is allowed. - Org-scoped teams remain governed by _check_org_team_limits() (capped by the org budget). Also reverts the #29525 existing_team_max_budget workaround in _check_user_team_limits() back to the create-only form; /team/new still enforces the creator's personal caps. docs(access_control): resolve the contradiction in the team-admin section — team admins can keep/lower the budget and manage rate limits/models, but cannot raise the team budget (proxy-admin only). tests: unit + behavior coverage for raise-blocked, cap-removal-blocked (team admin), raise/removal allowed (proxy admin), uncapped-team restriction allowed, keep/lower/resend allowed, and unchanged create-path guards. Co-authored-by: Cursor <cursoragent@cursor.com> * test(ui): data-driven App Router migration E2E smoke (default + server-root-path) (#29974) * test(ui): add a data-driven App Router migration E2E smoke Add a growing Playwright smoke for migrated pages: for each segment it deep-links to the path route, asserts the URL and that the dashboard shell rendered, then clicks off to a legacy page and asserts navigation still works. Driven by e2e_tests/fixtures/migratedPages.ts, so adding a page is one line. Runs in two situations against the same proxy: the default mount (npm run e2e:migration) and a non-root SERVER_ROOT_PATH mount (npm run e2e:migration:root). globalSetup now logs in at `${SERVER_ROOT_PATH}/ui/login` so the admin storage state is valid under a prefix. Seeded with api-reference; append the rest as their migrations merge. * test(ui): support headed slow-motion + watch pauses in the migration smoke Honor SLOWMO in the server-root-path config (the default config already did), and add an env-gated E2E_WATCH_MS pause so a headed run lingers on each state. Both are no-ops by default, so CI behavior is unchanged. * test(ui): make the migration smoke a sidebar-click user journey Rework the smoke from deep-linking to a real navigation journey: start at the landing page, click the migrated page in the sidebar (expanding submenus for nested items), assert the path route rendered, reload it (the check a wrong server_root_path breaks), bounce to a legacy page and back, and — once two pages are migrated — navigate directly between two migrated pages. Verifies via URL + shell render, driven by the same fixture list. * test(ui): address review on the migration smoke Escape ROOT and segment before interpolating them into RegExp URL matchers so a future segment containing regex metacharacters can't silently widen the match. Make the server-root-path config fail fast when SERVER_ROOT_PATH is unset instead of silently re-running the default mount and passing without exercising the prefix. * test(ui): drop unused watch helper and fix stale smoke README * test(ui): run the migration smoke under a server root path in CI * test(ui): harden + instrument the server-root-path proxy reboot in CI * test(ui): run the server-root-path migration smoke as its own CI job Replace the in-place proxy reboot in e2e_ui_testing with a dedicated e2e_ui_testing_server_root_path job that boots the proxy once with SERVER_ROOT_PATH=/litellm, matching how every other proxy variant in the config gets its own job rather than killing and relaunching the live proxy. The reboot was failing deterministically: after pkill -9 and relaunch the prefixed proxy never came back up on :4000 (connection refused), so the smoke never ran. The readiness step that was supposed to surface the cause could never reach its boot-log tail because CircleCI runs steps under bash -eo pipefail and the preceding `curl -sv ... | tail` aborted the step with curl's exit 7. Booting the proxy as the job's own background step lets any boot crash land in that step's log instead of being swallowed. The default e2e_ui_testing job is unchanged aside from dropping the reboot, prefixed-readiness, and prefixed-smoke steps; the migration smoke still runs at the root mount there via the default Playwright config. * fix(proxy): extend response headers hook to streaming, TTS, image gen, and pass-through (#24232) * fix(proxy): extend response headers hook to streaming, TTS, image gen, and pass-through * test: mock post_call_response_headers_hook in audio speech route tests * chore(ui): remove dead App Router route stubs under (dashboard) (#30045) models-and-endpoints, organizations, and virtual-keys each had a page.tsx route under (dashboard)/ that is not in MIGRATED_PAGES, so the sidebar and deep links never resolve to it and the route is unreachable. Each was a thin wrapper that handed the shared view empty or no-op props (empty modelData with a no-op setModelData, hardcoded empty organizations, no-op setUserRole/setUserEmail), so reaching one would render a degraded page in any case. The real wrapper belongs in the PR that flips each page into MIGRATED_PAGES, written with eyes on it and a test This continues the dead-scaffolding cleanup from #28891. The shared components these wrappers rendered (ModelsAndEndpointsView, OrganizationFilters) stay, since the legacy ?page= switch in app/page.tsx and src/components still import them * fix(ui/mcp): reset OAuth state on create-server modal close so a prior server's token no longer leaks into the next add-server session (#30000) * fix(ui/mcp): reset OAuth hook state on modal close so a prior server's token no longer leaks into the next add-server session * fix(ui/mcp): clear in-flight OAuth guard on reset and reset form/tools on modal close so nothing leaks on a parent-driven dismiss * fix(mcp): allow team access-group grants in OAuth authorize/token access check (#30041) * fix(mcp): honor team access-group grants in OAuth authorize/token access check * test(mcp): mock build_effective_auth_contexts in non-admin authorize tests for isolation * docs(security): require a reproduction video for vulnerability reports (#30048) (#30063) With AI models capable of automated vulnerability discovery now publicly available, we expect a large increase in report volume, much of it unverified. Requiring a video of the exploit running against a live instance raises the bar for submissions and keeps triage focused on reproducible issues. Reports without a video will be closed and reopened if one is added later. Co-authored-by: stuxf <70670632+stuxf@users.noreply.github.com> * feat(ui): add admin flag to disable in-product UI nudges for everyone (#29796) * feat(ui): add admin flag to disable in-product UI nudges for everyone Admins can now suppress the survey and Claude Code feedback popups for all users via a single disable_ui_nudges UI setting, instead of relying on each user dismissing them individually. * fix(ui): suppress nudges while ui settings are loading Gate nudgesDisabled on the ui-settings loading state so an admin with disable_ui_nudges on doesn't see the survey prompt flash, and the getInProductNudgesCall fetch doesn't fire, on a cold page load before the flag resolves. Falls back to showing nudges if the fetch errors. * test(ui): wrap CreateKeyPage test in QueryClientProvider page.tsx now calls useUISettings (react-query), which needs a QueryClient that layout.tsx supplies in production but the test did not. Add the provider and mock getUiSettings so the query resolves. * chore(ui): remove dead dashboard files and unused dependencies (#30047) * chore(ui): remove dead dashboard files and unused dependencies knip flagged seven orphaned source/config files with no importers and five declared dependencies that nothing in the tree uses. Removing them shrinks the dashboard bundle's source surface and keeps the manifest honest; vite stays installed transitively via vitest, so test tooling is unaffected. * fix(ci): restore serverRootPath.config.ts referenced by SERVER_ROOT_PATH workflow The dead-code sweep removed e2e_tests/serverRootPath.config.ts, but its spec (tests/login/serverRootPathRedirect.spec.ts) and the test_server_root_path.yml workflow step still depend on it, so the redirect e2e job failed to load a config that no longer existed. * fix(proxy): authorize batch files using upload target_model_names (LIT-3593) (#30009) * fix(proxy): authorize batch files using upload target_model_names (LIT-3593) After replace_model_in_jsonl, body.model is a stripped provider id. Reverse-mapping it via resolve_model_name_from_model_id is first-match on model_list and caused false 403s when multiple deployments share the same stripped name. Use target_model_names from the unified file id instead. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(proxy): restore resolve_model_name_from_model_id for JSONL fallback path (LIT-3593) Restores the reverse-lookup for the JSONL body.model fallback path so that legacy/pre-target_model_names managed files still map stripped provider IDs back to proxy aliases before auth. Also cleans up redundant `or None`. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * Revert "fix(proxy): restore resolve_model_name_from_model_id for JSONL fallback path (LIT-3593)" This reverts commit |
||
|
|
2fe9feda71
|
fix(caching): restore stored prompt_tokens on embedding cache hits instead of recomputing (#30046) | ||
|
|
e15b37a18e
|
Add Claude Fable 5 across Anthropic, Bedrock, Vertex AI, and Azure AI (#30064)
* Add Claude Fable 5 across Anthropic, Bedrock, Vertex AI, and Azure AI
Adds cost map entries for claude-fable-5 ($10/$50 per MTok, 1M context,
128K output, adaptive thinking only) on the Anthropic API, Bedrock
converse (base, global, and us/eu geo inference profiles at the 10%
regional premium), Vertex AI, and Azure AI (Microsoft Foundry, which
serves Fable 5 with the full 1M context window unlike Opus 4.8).
Registers anthropic.claude-fable-5 in BEDROCK_CONVERSE_MODELS, lists the
model in the setup wizard, and extends the reasoning effort e2e grid.
The Bedrock, Vertex, and Azure grid cells carry fail_reason markers
until the CI accounts are provisioned: Bedrock needs the provider data
sharing opt-in Fable 5 requires, and the Foundry resource needs a
claude-fable-5 deployment.
The first-party entry carries provider_specific_entry {us: 1.1} for the
inference_geo premium and deliberately no fast multiplier since Fable 5
has no fast mode.
https://claude.ai/code/session_01MZarYYT3aS7DxaNjoax6Gm
* Drop removed sampling params for Claude 4.7+ when drop_params is set
Fable 5, Opus 4.7, and Opus 4.8 removed sampling params: the API rejects
top_p, top_k, and any temperature other than 1 with a 400. LiteLLM was
forwarding them even with drop_params enabled because the Anthropic and
Bedrock converse transformations passed temperature/top_p through
unconditionally.
Mirror the GPT-5/o-series handling: temperature=1 still passes through,
other values and any top_p are dropped when drop_params is set, and
without drop_params a clean client-side UnsupportedParamsError tells the
caller how to opt in, instead of surfacing the raw provider error.
https://claude.ai/code/session_01MZarYYT3aS7DxaNjoax6Gm
* Drive sampling param gating from the cost map and cover top_k
Greptile review follow-ups on the sampling param fix: the restriction for
Fable 5 / Opus 4.7 / 4.8 is now declared as supports_sampling_params: false
on every affected cost map entry (perplexity excluded; that route is
OpenAI-compatible and maps sampling params upstream) and read back through
a tri-state map lookup, keeping the name check only as a fallback for
provider-routed ids whose hosted map entries predate the flag, the same
layering supports_adaptive_thinking uses. top_k bypasses map_openai_params
as a provider-specific kwarg, so it is gated at the shared
AnthropicConfig.transform_request boundary (direct, Bedrock invoke, Vertex,
Azure) and in the Bedrock converse _handle_top_k_value path, with
drop_params threaded through the converse transform helpers.
Also updates the reasoning effort grid cell count assertion for the four
Fable 5 rows added on this branch (29 x 11 cells).
https://claude.ai/code/session_01MZarYYT3aS7DxaNjoax6Gm
* Declare supports_sampling_params in the cost map schema
The model map validation schema uses additionalProperties: false, so the
new flag must be declared for the 28 entries that carry it; this was the
one failing job (misc / Run tests) on the previous commit.
https://claude.ai/code/session_01MZarYYT3aS7DxaNjoax6Gm
* fix(bedrock): gate top_k=0 on converse to match Anthropic boundary
Truthiness check let top_k=0 silently disappear on models that removed
sampling params, while AnthropicConfig.transform_request treats 0 as
present and raises UnsupportedParamsError (or drops when drop_params is
set). Switch to 'is not None' so converse, direct Anthropic, invoke,
Vertex, and Azure all behave the same for top_k=0.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
|
||
|
|
2cd7e87485
|
fix(proxy): authorize batch files using upload target_model_names (LIT-3593) (#30009)
* fix(proxy): authorize batch files using upload target_model_names (LIT-3593)
After replace_model_in_jsonl, body.model is a stripped provider id. Reverse-mapping it via resolve_model_name_from_model_id is first-match on model_list and caused false 403s when multiple deployments share the same stripped name. Use target_model_names from the unified file id instead.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(proxy): restore resolve_model_name_from_model_id for JSONL fallback path (LIT-3593)
Restores the reverse-lookup for the JSONL body.model fallback path so that
legacy/pre-target_model_names managed files still map stripped provider IDs
back to proxy aliases before auth. Also cleans up redundant `or None`.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* Revert "fix(proxy): restore resolve_model_name_from_model_id for JSONL fallback path (LIT-3593)"
This reverts commit
|