Commit graph

14 commits

Author SHA1 Message Date
devin-ai-integration[bot]
797353f13a
fix(otel): name postgres service spans by operation and table (#44240)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
devin-ai-integration[bot]
53d2ab6b05
fix(proxy): emit postgres service spans only on real DB reads in auth cache helpers (#44148)
log_db_metrics wrapped whole cache-first auth helpers and always emitted a ServiceTypes.DB success event, so in-memory cache hits showed up as postgres <fn> spans and DB service metrics. The decorator now installs a ContextVar witness that _TrackedPrismaEngine marks on every Prisma query and transaction call, and the DB event is emitted only when the witness was marked. Real reads keep their existing call_type names, the failure path and the PROXY batch-write branch are unchanged, and Redis instrumentation is untouched.

A decorated helper that reaches Prisma only through another decorated helper (get_key_object -> get_object_permission, get_team_object_by_alias -> get_object_permission, get_tag_object -> get_tag_objects_batch) used to emit two events for one query. The inner wrapper now marks its witness as reported when it emits a success or DB failure event, and only unreported activity is handed up to the enclosing witness, so the inner event is the one that survives. An outer helper that also queries Prisma directly or through undecorated callees still gets its own event.

Tests: get_user_object and get_org_object cache hits emit no DB event; a get_user_object miss through the generated Prisma client emits exactly one postgres get_user_object event; decorator-level tests cover nested calls emitting only the inner event, outer calls with their own query, inner non-DB failures, bounded lookups and sibling-request isolation.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:27 -07:00
tin-berri
f63d989ff9
feat: add Bespoke Nimble gateway and OSS classifier support (#44246)
* feat: add Bespoke Nimble gateway and OSS classifier support

* feat: accept Ollama's nimble model name for the Bespoke provider

* test: exempt the POST-only bespoke decisions route from the all-methods check

test_pass_through_routes_support_all_methods requires every built-in
pass-through route to accept every HTTP method unless it is listed in
PROTOCOL_CONSTRAINED_PASS_THROUGH_ROUTES. /bespoke/v1/systemone is
POST-only like /laya/v1/systemone, so the test failed at this branch
and passed at the merge base. List it alongside Laya.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-02 19:37:08 -07:00
yujonglee
677205b3f5
fix(proxy): share ownership permissions for spend logs and traces (#44239)
* refactor(proxy): extract shared spend log read policy

* test(proxy): use named bindings for spend scope regression

* test(proxy): reuse existing spend log query harness

* test(proxy): cover spend log permission lookup adoption

* chore(proxy): relocate existing spend query baseline

* refactor(proxy): make scope query returns explicit

* refactor(proxy): inject deferred log permission lookup

* test(proxy): cover teamless management compatibility lookup

* refactor(proxy): compose user and team log grants

* refactor(proxy): share generic authorization composition

* refactor(proxy): compose trace read permissions

* refactor(proxy): centralize spend and trace authorization

* refactor(proxy): strengthen spend and trace scope types

* refactor(proxy): flatten log read scope into owned logs

Replace the AnyOf grant tree with a flat OwnedLogs(user_id, team_ids) scope,
and OwnedTraces(logs, api_key_hash) for traces, since every consumer flattened
the tree back into that shape.

A caller with no user id now gets an empty scope instead of matching ownerless
rows through Prisma's IS NULL. The dead request_id guard in ui_view_spend_logs
is removed, and the management facets inject the log team lookup and reuse
read_scope_sql instead of the list shim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(proxy): run spend scope tests through one SQLite emulator

Replace the string-matching payload emulator and the hand-rolled Prisma where
interpreter with one SQLite helper that runs the real scope SQL. Session scope
tests now go through the endpoint, including the no-user caller that must not
match ownerless rows. Drop duplicated lookup-failure and trace mapping cases.

load_permitted_log_team_ids returns no teams without a database instead of
relying on the resolver's broad except.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(proxy): unify log and trace ownership permissions

* test(tracing): align fixtures with ownership read scopes

* refactor(tracing): align query scopes with row ownership

* refactor(spend): make ownership SQL predicates explicit

* test(spend): validate ownership SQL against PostgreSQL

* docs(traces): drop key-row visibility from query help guide

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): reach the empty-memberships branch in team lookup test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate dashboard API types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 01:38:54 +00:00
yujonglee
688d791fa0
feat(traces): type queries and align read access with log visibility (#44228)
* wip

* wip

* test(traces): separate root status from diagnostic error counts

* test(traces): cover normalization precedence and fallbacks

* chore(cache): remove stray comments from trace PR

* test(traces): name lens test for shared query path

* fix(traces): place query implementation before test module

* test(traces): use unified read scope in migration tests

* ci(rust): allow feature checks to finish

* ci(mcp): allow dependency resolution to finish

* fix(traces): preserve key visibility and safe spend attribution
2026-10-02 21:55:42 +00:00
devin-ai-integration[bot]
b21e44cbf9
feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key (#42375)
* test(e2e): jwt auto_register map-existing-key repro

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): exclude blocked keys from auto_register_map_existing_key reuse

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(jwt): route existing-key lookup through VerificationTokenRepository

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): stop requiring LITELLM_SALT_KEY for the owned JWT gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): gate the owned JWT gateway tests behind E2E_OWNED_GATEWAY

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): only reuse keys that can call LLM routes in auto_register_map_existing_key

Skip Admin UI session keys and keys whose allowed_routes restrict them to
anything other than llm_api_routes (management, read_only, password-reset
sessions). Mapping a JWT to one of those left the user with 401s or 403s on
every LLM call, since the mapping persists.

* fix(jwt): scope auto_register_map_existing_key reuse to the JWT-resolved team

Only reuse a key whose team_id matches the team auth_builder resolved for
the JWT (no team matches no team), so a personal key can no longer bypass
the resolved team's model and budget limits.

With the flag on, the first JWT request now falls through to the same
virtual-key checks later mapped requests get, instead of returning early,
so a reused key's own limits apply from request one rather than 200 then
403. Flag off keeps the early return unchanged.

* fix(jwt): keep the early return when no master key is set

Without a master key the generic virtual-key path returns a bare
INTERNAL_USER object, so falling through on the first auto-registered
request dropped the key's team, models and budgets. Only fall through when
a master key is configured.

Tests now assert the reused key per team rather than the query shape, and
cover the flag-off early return and the no-master-key case.

* test(jwt): assert on race-loser's returned key, not only mocks (TQ002)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(jwt): close the auto_register_map_existing_key race, shared-claim and expiry holes

A key auto_register just minted is never adopted by a concurrent request, so the race loser's cleanup can no longer delete a key another request mapped and cascade its mapping away (503, user left with no key)

Reuse only happens when the claim value is the JWT-resolved user_id. A shared claim such as azp or client_id falls back to minting, so one user can no longer land on another user's personal key and budget

Only keys that never expire are reused, so an expiring key can no longer pin the claim to a permanent 401

Integration tests on a real proxy and Postgres cover all three. The race test holds the first mapping insert in a Postgres relay, so the interleaving is forced rather than timed. The where-clause shape unit tests are replaced by these, since only a real database proves the filter

* test(e2e): create the reused key in the team the JWT resolves to

The flag only reuses a key in the JWT-resolved team, and this identity's groups claim resolves to its team, so a teamless key was never eligible and the test could not pass

* test(integration): match the held statement across TCP reads

The relay looked for the trigger inside one read, so an insert split across two reads was never held and the race test would fail waiting for it. It now matches one exact trigger over a window that keeps the end of the previous read

* fix(jwt): gate key reuse on the claim field, not on the claim value

Requiring the claim value to equal the resolved user_id skipped reuse for users matched through the sso_user_id or case-insensitive email fallback, whose stored user_id differs from the JWT sub. That is the lookup LIT-5378 asks for. Reuse is now allowed when the virtual key claim is the user_id or user_email JWT field, globally or for the token's issuer, which still keeps shared claims such as azp or client_id on the mint path

* fix(jwt): let an issuer's own user field replace the global one when gating key reuse

An issuer that identifies users by uid no longer treats the global sub field as a user identity claim, so a shared sub under that issuer mints instead of reusing a personal key

* test(jwt): make the flag-off test fail when the flag no longer gates key reuse

The flag-off test used a config where sub was not a user identity claim, so deleting the flag check still passed. Configure user_id_jwt_field=sub so only the flag keeps the lookup off, and drop test docstrings

* chore(lint): drop mutable-ok suppressions that LIT013 flags as no-ops

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Mrinal Chanshetty <mchanshetty@Mrinals-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 12:11:33 -07:00
tin-berri
0dc23406eb
feat: add Laya gateway and OSS classifier providers (#43626)
* feat: add Laya gateway and classifier backend

* test: cover the pass-through model_group pin and repair the shard fakes

MockRequest in tests/pass_through_unit_tests gains an httpx.URL and an ASGI
scope, which get_request_route now reads inside
_init_kwargs_for_pass_through_endpoint, and the POST-only /laya/v1/systemone
route joins the protocol-constrained exemptions. A built-in pass-through pins
metadata.model_group to the resolved model so a client cannot choose its own
per-model budget key; test_pass_through_endpoints now proves that on a
non-Laya route and drops a duplicated assertion.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-02 11:29:36 -07:00
devin-ai-integration[bot]
54a51c80df
feat(proxy): bounded daily activity routes for all usage entities (#43408)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 16:57:23 -07:00
devin-ai-integration[bot]
a38fff6560
fix(proxy): enforce key/team vector_stores allowlist on /v1/rag/query (#43953)
* add test case for /rag/query and stronger auth check

* style(proxy): ruff format auth_checks.py

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): build rag query vector store ids immutably and test the no-registry path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): integration coverage for /v1/rag/query vector store allowlist

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): audit cells for /v1/rag/query vector store allowlist

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): consolidate vector store allowlist audit coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): type RAG vector store request body

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 13:48:41 -07:00
yujonglee
ec605826d4
feat: improve trace ingestion and trace details (#43975)
* refactor: separate OTLP HTTP decoding from trace codec

* feat: complete trace ingestion and read paths

* fix: encode OTLP protobuf errors in Rust

* fix: raise OTLP body limit to 16 MiB

* test: cover OTLP auth body parsing boundary

* refactor: parse OTLP media type into enum

* fix: enforce OTLP body size at HTTP boundary

* perf: preserve shared OTLP metadata across ingestion

* bench: compare owned and shared trace resource fanout

* refactor: extract shared storage and Python conversion caches

* refactor: keep shared storage owned by traces

* test: keep trace loopback coverage in Rust

* test(proxy): adapt trace coverage to injected access context

* fix(tracing): satisfy stacked branch lint checks

* refactor(tracing): use immutable ingestion payloads

* fix(tracing): declare native error encoder export

* test(proxy): resolve trace access through dependency

* fix(tracing): align merged normalizer types and bridge tests

* fix(tracing): address ingestion and diagnostic review findings

* fix(proxy): preserve body parsing for partial request scopes

* test(proxy): use valid HTTP scopes in request fixtures

* test(proxy): complete auth request flow scopes
2026-10-01 13:45:33 -07:00
devin-ai-integration[bot]
39e31958f8
test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy (#43998)
* test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): stub HIBP through respx by disabling the aiohttp transport

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): share the httpx transport fixture across proxy unit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): restore proxy globals without a missing-value sentinel

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): package moved dirs and stub the login breach check at the HTTP boundary

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): isolate the mcp server manager per test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 10:11:45 -07:00
moe-berri
d9f73245be
feat(lens): track worker spend through virtual keys (#43989)
* feat(lens): bill worker analysis through virtual keys

* fix(lens): pin the verified worker image and add setup proof

* fix(lens): preserve network checks and redact billed analysis logs

* test(lens): preserve legacy worker result submission during upgrade

* fix(lens): enforce trusted worker IPs and restore coverage uploads

* docs(lens): explain trusted proxy requirements for worker allowlists

* fix(lens): yield to worker disconnects after the synthetic body
2026-10-01 09:47:30 -07:00
devin-ai-integration[bot]
2d034bb35b
perf(proxy): hold one spend counter batch across admission and across post-call accounting (#43369)
Auth's spend counter MGET scope spans common checks, model budget check and reservation;
reservation increments go out as one pipeline; post-call reconcile adjustments ride the ordinary
increment pipeline and update_cache uses one batched read. Over-budget reservation counters are
charged one at a time so a rejection never touches the counters after it; post-call counter keys are
derived from ids without validating a UserAPIKeyAuth.

Resolves LIT-8881

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 15:05:33 -07:00
devin-ai-integration[bot]
248f0eb159
ci: move tests/proxy_unit_tests to tests/unit/proxy and run the proxy-db shards from litellm-tests (#42903)
* ci: fix the litellm-tests unit job with sysmon coverage, an env allowlist and coverage upload on failure

* test: replace key-dependent proxy, enterprise and mcp unit tests with synthetic values and integration and e2e coverage

* test: drop key reads at the legacy proxy, enterprise and mcp paths and wire the gemini pass-through split

* ci: move caching, proxy-extras, gateway and enterprise tests into tests/unit and run them from litellm-tests under their legacy flags

* ci: move caching, proxy-extras, gateway and enterprise tests into tests/unit and run them from litellm-tests under their legacy flags

* ci: move tests/proxy_unit_tests to tests/unit/proxy and run the proxy-db shards from litellm-tests

* ci: fail the unit shard when circleci tests split errors

* test: drop restating comments from the gemini pass-through split

* build: point the local proxy unit targets at the nested tests/unit/proxy tree

* ci: exit the unit shard cleanly when circleci tests split assigns it no files

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-09-24 22:59:11 +00:00