Commit graph

15 commits

Author SHA1 Message Date
devin-ai-integration[bot]
1238cfe90f
feat(proxy): add LITELLM_FIPS_MODE startup gate with provider assertion and loud password migration failure (#42700)
* feat(proxy): LITELLM_FIPS_MODE startup gate with provider assertion, TLS verify guard and loud password migration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): match ssl_verify off detection to runtime str_to_bool semantics

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): drop tautological fips probe test and satisfy CodeQL return checks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): inject the fake Prisma client through the module boundary instead of patching _setup_prisma_client

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 01:11:52 -07:00
devin-ai-integration[bot]
0ed1c08f02
feat(anthropic): workload identity federation and pluggable identity sources (#44448)
* feat(anthropic): workload identity federation and pluggable identity sources

Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one
commit on top of litellm_internal_staging without the dashboard changes.

Deployments on anthropic/ without a static api_key can exchange an OIDC
workload assertion for a short-lived sk-ant-oat01 token through a shared
RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file,
an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment,
per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation
fields are server-owned: refused inline in request bodies and on
POST /model/new, proxy-admin only on credentials, and the token exchange
is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS
adds a host. GET /credentials/{name}/jwks exports the public key set of a
LiteLLM-signed credential for the Claude Console.

The OpenAI federation trio from #39613 rides along on the backend side with
the same server-owned handling.

Fixes #28607
Resolves LIT-6107

Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>

* fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries

The files handler enabled workload identity on batch-result downloads but never received the
deployment's litellm_params, so a deployment authenticating through a named credential could only
mint from process-wide env vars. It now threads litellm_params through to the auth header the way
the batch retrieve path already does.

LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme,
so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the
gateway. Entries are now parsed as network locations whether or not they carry a scheme.

* fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle

* test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries

* fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly

* fix(proxy): decrypt stored litellm_params before the WIF write gate

* fix(proxy): hide WIF secret references from /health output

* fix(proxy): keep the proxy error shape on credential endpoint refusals

* fix(proxy): hide identity token file paths from /health output

* fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working

The Bedrock Claude Platform route already reads anthropic_workspace_id from
optional_params, so banning that spelling as a server-owned federation
parameter broke a pre-existing client capability. The federation field is now
anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID),
which restores the base branch's behavior for Bedrock callers, drops the
Bedrock-specific hint from the refusal message, and deletes the unconditional
ban constant that no longer had a reader

* fix(auth): share one exchanged token across workers reading the same assertion

Anthropic accepts each identity assertion exactly once, so two uvicorn
workers reading the same token file both minting from it means the second
exchange is denied with jti_reused. Minted tokens now land in a per-user
0700 cache directory guarded by a file lock, so workers on the same host
reuse one exchange until the token expires or the assertion rotates. A 401
is only retried when the re-read assertion actually differs, and the denial
hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache
and an empty value disables it

* fix: keep anthropic federation from being shadowed or leaked

An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated
deployment sent an empty x-api-key on every call instead of minting a token.
Blank values now read as unset, and a real static key on a federated deployment
logs once that it outranks federation and nothing is being federated.

The exchange-host allowlist matched hostnames only, so a second process on
another port of an allowed host was trusted with the workload's identity token.
An entry that names a port now trusts that port alone, while a bare host still
trusts every port.

The shared token store exists so the workers reading one projected token file do
not each spend its single-use jti. A source that mints its own assertion per
exchange shares nothing with another worker, so it no longer writes a live token
to disk for a lookup that can never hit.

* fix: unlink a staged token file a failed write leaves behind

The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads.

* refactor: move anthropic jwks derivation behind a provider-owned tagged union

* fix: unlink the staged token file when its write fails at close

A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token.

* fix(anthropic): close the staging descriptor before writing the shared token file

* fix(wif): judge federation writes by what they set, not what is stored

The admin gate read the stored deployment, so a team admin lost edit, delete
and Test Connection on any deployment carrying federation params. It now
returns early unless the submitted fields touch the federation surface, and a
Test Connection probe that points the deployment at its own api_base is still
refused, with the 403 no longer wrapped into a 500

The rest of the same review pass: POST /model/new refuses only a blocking
value of `blocked`, so a client that always sends `blocked: false` is not
turned away; a request body can no longer pick which federated identity to
mint as by naming a stored credential; an advisory refresh the executor
refuses disarms the entry instead of wedging the identity until the follower
timeout; the static-key shadow warning resolves its env fallback inside the
cache instead of once per request; credential writes drop nulls before
storing them; the token exchange validates the endpoint URL before reading an
assertion and keeps refusing redirects across a client heal; /health hides
every server-owned federation field from non-admins; and the async create_file
and create_batch paths say which setting is missing when the provider resolves
no URL

* fix(proxy): let a deployment write name a federated credential

reject_federated_credential_reference runs from is_request_body_safe, which
pre_db_read_auth_checks calls on every route, so it also fired on POST
/model/new, /model/update, /model/{id}/update and /health/test_connection. A
proxy admin could no longer attach a federated credential to a deployment over
the API or the Admin UI, leaving a static config.yaml entry as the only way to
configure the feature the rejection told the caller to go configure, and
_reject_non_admin_wif_write never got to make the call it exists to make.

is_request_body_safe now takes the route and skips only the credential-reference
check on the routes that reach can_user_make_model_call. Federation fields typed
inline into a body stay refused everywhere, and a call naming a federated
credential still cannot pick the identity it mints as.

* refactor(proxy): derive health display policy from the federation key sets

The health check module hand-copied the five workload identity fields whose
value is a credential, so a shared proxy surface named provider-specific
parameters and a newly added secret-bearing field would have gone on being
displayed until someone remembered both places

WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of,
types/utils derives secret_bearing_wif_litellm_params from it, and the health
layer splats that tuple the same way it already splats the admin-only one

* fix(anthropic_wif): treat blank identity-source fields as unset

* test(proxy): classify the federation params in the credential slot registry

main's registry test (#43298) now fails the build for any credential-named
deployment param without a classification. The five federation fields that
carry a token, a token file path, or a signing or client secret reference are
Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak
settings name a URL, a client id, an auth method, or a scope and are NotSecret

* fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics

Register the 18 Anthropic and 3 OpenAI federation params as frozen
ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel
and the request-body ban list read one declaration. Pass the deployment
api_base through to the count-tokens handler instead of a pre-suffixed
URL, which doubled the /count_tokens path on main's prompt-cache
predictor. Add the five litellm_anthropic_wif_* families to the
all-metrics Grafana dashboard.

* fix(credentials): gate PATCH on WIF fields resolved from model_id

The credential PATCH handler checked server-owned workload identity
federation fields only on the values the caller sent, while a body that
named a deployment through model_id had its credential values resolved
after that check. A non-admin could therefore copy a federated
deployment's WIF fields onto an ordinary credential. Resolve the incoming
values first and run the non-admin gate on them, matching the POST path

* fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header

Count-tokens walked its own credential ladder: a static key, else skip minting when
ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it
forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the
same deployment authenticated with that token. The handler now takes the auth header that
AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills
use, and merges the oauth beta a minted or consumer token carries with the token-counting beta

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
Co-authored-by: mateo-berri <happymvw@gmail.com>
2026-10-03 17:08:30 -07:00
devin-ai-integration[bot]
fe683ea139
feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models (#44136)
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models

GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.

* fix(proxy): offer a Codex service tier only when every deployment of the model lists it

* fix(codex-catalog): an invalid service_tiers value offers no tier for the model

* fix(codex-catalog): read service tiers off the deployments the key's team can route to

A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them

The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default

* test(codex-catalog): drop the redundant module docstring and sort the imports

* test(integration): add the Codex catalog audit cells and the multi-worker convergence note

* test(integration): clean up every catalog test model and answer the refresh GET

* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers

Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.

* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut

The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns

The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 20:42:55 +00:00
devin-ai-integration[bot]
4c6c84afb7
perf(proxy): stop prompt-cache eligibility from tokenizing the whole conversation (#44221)
is_prompt_caching_valid_prompt ran the full Python token_counter over every message to compare against the deployment's prompt cache minimum, 500 to 1000 ms at 440k to 740k tokens on every request that reaches the prompt_caching pre-call check, Rust on or off. messages_reach_token_count does the same arithmetic as token_counter(...) >= threshold and stops at the first message that reaches the threshold. Groups with one healthy deployment skip the prefix hash and pin lookup, which cannot change the result for them

Four fixed name span events make the pre-LLM phases measurable with OTel v2: litellm.request.body_received (with body_bytes) once per body read before parsing, on the JSON, binary and form branches, body_parsed, pre_call_completed, and deployment_selected emitted once per pick inside Router.async_get_available_deployment and get_available_deployment with attempt, reason and model group, so every router surface, retry and fallback is covered. Measured locally on /v1/chat/completions, /v1/messages and /v1/responses at 440k tokens with Rust on and off against a fake upstream

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
devin-ai-integration[bot]
797353f13a
fix(otel): name postgres service spans by operation and table (#44240)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
devin-ai-integration[bot]
e340e546e2
feat(traces): tracing development seed (#44363)
* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* wip

* wip

* wip

* chore(trace): checkpoint ongoing Rust migration

* refactor(trace): group Python bridge under trace package

* refactor(traces): read span conventions through a Convention trait

Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.

The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.

* feat(trace): export Rust-owned wire schemas and enforce contract bounds

* fix(trace): bound quoted counts in ClickHouse wire schemas

* feat(trace): generate Python wire contracts with datamodel-code-generator

* test(trace): validate migrated callers and generated contracts at the native boundary

* refactor(traces): rename normalization convention to format

* fix(traces): reconcile spend evidence and preserve unknown costs

* feat(traces): normalize additional telemetry formats

* test(traces): cover captured normalization fixtures

* refactor(traces): isolate SDK normalization rules

* feat(tracing): seed all trace exports for local dashboard

* fix(clickhouse): preserve custom LiteLLM request metadata

* docs(traces): define normalization module boundaries

* docs(traces): define resolution and OTLP boundaries

* fix(ui): normalize nullable trace message names

* refactor(traces): split resolver modules and cover resolution behavior

* test(traces): replace normalization snapshots with behavior assertions

* fix(ui): align dashboard API contracts with generated types

* refactor(traces): type normalization and storage boundaries

* fix(traces): seed captured SDK spend and preserve provider identities

* wip

* test(traces): verify guide discovery and content ordering

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:59:51 +00:00
yucheng-berri
5d42cb7cfa
fix(guardrails): encrypt guardrail litellm_params secrets at rest (#43627)
* fix(guardrails): encrypt guardrail litellm_params secrets at rest

* fix(guardrails): keep salt-key encryption on master key rotation and retry rows edited mid-rotation

- rotate guardrail params under LITELLM_SALT_KEY when set, matching the key reads decrypt with
- re-read and retry a row whose updated_at moved during rotation, up to GUARDRAIL_ROTATION_ATTEMPTS
- build decrypted Guardrail rows and the rotation count without mutating locals

* refactor(guardrails): retry guardrail rotation by bounded recursion instead of a rebound cursor

- each attempt re-reads the row and recurses with attempts_left - 1, so no loop variable is rebound
- cover the give-up path after GUARDRAIL_ROTATION_ATTEMPTS writes

* test(guardrails): drive the real guardrail rotator from the master key rotation test

- inject an encrypted guardrail row through the prisma client instead of replacing the GuardrailRegistry method
- assert the written params decrypt under the new master key

* Annotate guardrail param encryption collections for type-discipline gate

* Type guardrail param recursion through validated JSON containers

* Type guardrail registry test helpers and drop section comment

* Reject client-supplied encrypted values in guardrail litellm_params

* Allow depth-bounded contains_encrypted_marker in the recursion detector

* Keep a loaded guardrail when its DB params do not decrypt with the current key

* Apply other DB edits while keeping loaded values that do not decrypt, including PATCH models

* Keep the loaded guardrail when an undecryptable param has no loaded value

* Drop suppressions the type discipline gate on main now reports as unused

* Assert what the reinitialized guardrail holds after an edit to an undecryptable one

* Drive the rotation sync tests through a registered guardrail instead of patching reinitialize

* Type the rotation test helpers and drop the new test docstrings

* fix(guardrails): refuse to approve a submission whose params do not decrypt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 22:49:41 -07:00
yujonglee
677205b3f5
fix(proxy): share ownership permissions for spend logs and traces (#44239)
* refactor(proxy): extract shared spend log read policy

* test(proxy): use named bindings for spend scope regression

* test(proxy): reuse existing spend log query harness

* test(proxy): cover spend log permission lookup adoption

* chore(proxy): relocate existing spend query baseline

* refactor(proxy): make scope query returns explicit

* refactor(proxy): inject deferred log permission lookup

* test(proxy): cover teamless management compatibility lookup

* refactor(proxy): compose user and team log grants

* refactor(proxy): share generic authorization composition

* refactor(proxy): compose trace read permissions

* refactor(proxy): centralize spend and trace authorization

* refactor(proxy): strengthen spend and trace scope types

* refactor(proxy): flatten log read scope into owned logs

Replace the AnyOf grant tree with a flat OwnedLogs(user_id, team_ids) scope,
and OwnedTraces(logs, api_key_hash) for traces, since every consumer flattened
the tree back into that shape.

A caller with no user id now gets an empty scope instead of matching ownerless
rows through Prisma's IS NULL. The dead request_id guard in ui_view_spend_logs
is removed, and the management facets inject the log team lookup and reuse
read_scope_sql instead of the list shim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(proxy): run spend scope tests through one SQLite emulator

Replace the string-matching payload emulator and the hand-rolled Prisma where
interpreter with one SQLite helper that runs the real scope SQL. Session scope
tests now go through the endpoint, including the no-user caller that must not
match ownerless rows. Drop duplicated lookup-failure and trace mapping cases.

load_permitted_log_team_ids returns no teams without a database instead of
relying on the resolver's broad except.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(proxy): unify log and trace ownership permissions

* test(tracing): align fixtures with ownership read scopes

* refactor(tracing): align query scopes with row ownership

* refactor(spend): make ownership SQL predicates explicit

* test(spend): validate ownership SQL against PostgreSQL

* docs(traces): drop key-row visibility from query help guide

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): reach the empty-memberships branch in team lookup test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate dashboard API types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 01:38:54 +00:00
yuneng-jiang
b877a38e5f
fix(proxy): stop queued registry read-throughs spending the resync budget (#44277)
* fix(proxy): stop queued registry read-throughs spending the resync budget

RegistryReadThrough.attempt serializes misses behind one lock, but every
request that was queued behind the first one still spent a unit of the
20-per-5s resync budget and re-ran the DB resync, even though the first
request had already loaded the object. A burst of more than 20 requests
for a model created on another worker therefore exhausted the budget and
the rest got 400 "Invalid model name".

attempt now checks whether the key is already loaded once it holds the
lock and returns early without touching the budget. Models check the
router's model names and deployment ids, guardrails and agents reuse
their existing registry lookups.

* test(proxy): gate the queued read-through test on events and cover each registry's loaded check

The queued-requests test now holds the first resync on an asyncio.Event instead of
a timed sleep and records calls in a recorder with tuple and frozenset state. New
tests show the model, guardrail and agent read-throughs each answer an object that
is already loaded without reading the database, so rewiring any registry's loaded
check now fails a test

* test(proxy): keep the queued read-through recorder inside its test and type the agent registry fixture
2026-10-02 18:12:56 -07:00
yujonglee
ec605826d4
feat: improve trace ingestion and trace details (#43975)
* refactor: separate OTLP HTTP decoding from trace codec

* feat: complete trace ingestion and read paths

* fix: encode OTLP protobuf errors in Rust

* fix: raise OTLP body limit to 16 MiB

* test: cover OTLP auth body parsing boundary

* refactor: parse OTLP media type into enum

* fix: enforce OTLP body size at HTTP boundary

* perf: preserve shared OTLP metadata across ingestion

* bench: compare owned and shared trace resource fanout

* refactor: extract shared storage and Python conversion caches

* refactor: keep shared storage owned by traces

* test: keep trace loopback coverage in Rust

* test(proxy): adapt trace coverage to injected access context

* fix(tracing): satisfy stacked branch lint checks

* refactor(tracing): use immutable ingestion payloads

* fix(tracing): declare native error encoder export

* test(proxy): resolve trace access through dependency

* fix(tracing): align merged normalizer types and bridge tests

* fix(tracing): address ingestion and diagnostic review findings

* fix(proxy): preserve body parsing for partial request scopes

* test(proxy): use valid HTTP scopes in request fixtures

* test(proxy): complete auth request flow scopes
2026-10-01 13:45:33 -07:00
devin-ai-integration[bot]
a76b59db9f
test(proxy): move middleware, spend_tracking, pass_through, common_utils and root proxy tests into tests/unit/proxy (#44015)
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 18:23:31 +00:00
tin-berri
de06c93767
feat(router): opt in to prompt-cache cost routing (#43232)
* feat(router): opt in to prompt-cache cost routing

* fix(router): address prompt-cache routing review

* ci: include cache-routing regressions in coverage
2026-09-26 18:13:13 -07:00
devin-ai-integration[bot]
e0fb89bc82
fix(proxy): keep the submitted body out of 422 validation errors (#43231)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-25 17:05:53 -07:00
Mateo Wang
25fb7810c2
fix(spend): attribute CLI session spend to the per-user cli-session alias instead of the hashed session token (#40541)
* fix(spend): attribute CLI session spend to the per-user cli-session alias instead of the hashed session token

A CLI session token is a fresh random secret on every login, so since v1.99 each
login's spend rows carried a different sha256 hash as api_key and the usage APIs
could resolve neither key_alias nor user_email for them. Spend rows and logging
callbacks now attribute a session request to its stable alias,
cli-session-<user_id>, and the usage endpoints derive that alias and owner from
the key itself instead of scanning for a matching digest

* fix(spend): resolve the CLI session team from the user's first team in usage metadata

A cli-session key carries no team of its own in the DB, so the usage
breakdown showed team_id None for it and the export grouped it as
Unassigned. The login attaches the user's first team to the session, so
the recovery mirrors that rule for cli-session keys only.

* fix(spend): claim the session team only for a single-team user

The CLI login attaches a team on its own only when the user has exactly
one; a user in several teams picks one per login, so usage metadata for
the alias would otherwise name a team the login may not have used.

* test(pass_through): mark the mocked auth object as a plain key

The logged key follows the alias only for a session token; a bare
MagicMock reads as one, so the test names the field it relies on.

* fix(spend): attribute CLI session pass-through, queue, and managed batch spend to the cli-session alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): only treat the exact cli-session-<created_by> value as a batch key alias

A managed object row written by an older build can still carry the raw per-login
session token, which shares the cli-session- prefix. Matching on the prefix alone
would have surfaced that token as a trusted alias and persisted it verbatim in the
batch cost spend log, so the alias check now requires the exact per-user value and
every other prefixed value keeps going through redaction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): log proxy executed batch rows under the cli-session alias instead of the session token

_row_metadata set user_api_key from the raw bearer token while user_api_key_hash carried the alias, so the spend log redaction rejected the alias as untrusted and hashed the random session token instead

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): attribute semantic search embedding spend to the cli-session alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): scope /key/spend/report for a CLI session to the cli-session alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): use the cli-session alias for websearch spend, prometheus failure labels and the parallel limiter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(spend): drop explanatory docstrings on get_logged_api_key and attach_user_details

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend): only recover cli-session usage keys whose suffix is a known user

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 18:21:47 -07:00
devin-ai-integration[bot]
248f0eb159
ci: move tests/proxy_unit_tests to tests/unit/proxy and run the proxy-db shards from litellm-tests (#42903)
* ci: fix the litellm-tests unit job with sysmon coverage, an env allowlist and coverage upload on failure

* test: replace key-dependent proxy, enterprise and mcp unit tests with synthetic values and integration and e2e coverage

* test: drop key reads at the legacy proxy, enterprise and mcp paths and wire the gemini pass-through split

* ci: move caching, proxy-extras, gateway and enterprise tests into tests/unit and run them from litellm-tests under their legacy flags

* ci: move caching, proxy-extras, gateway and enterprise tests into tests/unit and run them from litellm-tests under their legacy flags

* ci: move tests/proxy_unit_tests to tests/unit/proxy and run the proxy-db shards from litellm-tests

* ci: fail the unit shard when circleci tests split errors

* test: drop restating comments from the gemini pass-through split

* build: point the local proxy unit targets at the nested tests/unit/proxy tree

* ci: exit the unit shard cleanly when circleci tests split assigns it no files

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-09-24 22:59:11 +00:00