Commit graph

48867 commits

Author SHA1 Message Date
Taranum Wasu
fcaf4b0cc9
Merge 89d7a3890f into 252c71c0b2 2026-09-15 20:28:07 -04:00
Yassin Kortam
252c71c0b2
Merge pull request #34939 from BerriAI/litellm_dd_span_user_email
Some checks failed
Postgres Tests / proxy-security (push) Has been cancelled
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Has been cancelled
Postgres Tests / proxy-behavior (push) Has been cancelled
Unit Tests: Documentation Validation / documentation (push) Has been cancelled
Unit Tests / caching-local (push) Has been cancelled
Unit Tests / core-utils (push) Has been cancelled
Unit Tests / enterprise-package (push) Has been cancelled
Unit Tests / enterprise-routing (push) Has been cancelled
Unit Tests / integrations (push) Has been cancelled
Unit Tests / All Other Providers (push) Has been cancelled
Unit Tests / Vertex AI (push) Has been cancelled
Unit Tests / misc (push) Has been cancelled
Unit Tests / proxy-auth (push) Has been cancelled
Unit Tests / proxy-endpoints (push) Has been cancelled
Unit Tests / proxy-extras (push) Has been cancelled
Unit Tests / proxy-infra (push) Has been cancelled
Unit Tests / proxy-server (push) Has been cancelled
Unit Tests / responses-caching-types (push) Has been cancelled
Unit Tests: Proxy DB Operations / auth-checks (push) Has been cancelled
Unit Tests: Proxy DB Operations / budgets (push) Has been cancelled
Unit Tests: Proxy DB Operations / custom-logging (push) Has been cancelled
Unit Tests: Proxy DB Operations / db-and-spend (push) Has been cancelled
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Has been cancelled
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Has been cancelled
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Has been cancelled
Unit Tests: Proxy DB Operations / key-generation (push) Has been cancelled
Unit Tests: Proxy DB Operations / logging-misc (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-runtime (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-server-core (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-utils (push) Has been cancelled
feat(dd_span_tagger): emit litellm.user_email span tag for JWT-authenticated requests
2026-09-15 12:40:19 -07:00
Yassin Kortam
af6dc1db08
Merge pull request #40993 from BerriAI/litellm_health_check_skip_save_on_failed_read
Some checks failed
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
fix(health): skip background health check DB writes when the latest-row read fails
2026-09-14 15:58:05 -07:00
ryan-crabbe-berri
3f32a19d55
Merge pull request #37762 from BerriAI/litellm_model_hub_description
feat(model_hub): surface model_info.description in Model Hub
2026-09-14 15:03:21 -07:00
Yassin Kortam
1f8bae7eab
Merge pull request #41058 from BerriAI/litellm_fix_wrapper_async_double_sync_success_handler
fix(utils): stop wrapper_async submitting the sync success handler twice
2026-09-14 12:43:32 -07:00
Yassin Kortam
8d5c165547
Merge pull request #41061 from BerriAI/litellm_model_info_skip_jsonable_encoder
perf(proxy): serialize /model/info listing once with orjson
2026-09-14 12:43:01 -07:00
yassin
15ff8d18e7 test(proxy): cover the CLI single-model branch of /model/info JSON serialization
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 10:07:10 +00:00
yassin
b86de179dc fix(utils): stop wrapper_async submitting the sync success handler twice
_client_async_logging_helper re-submitted logging_obj.success_handler to the
executor after _dispatch_success_logging had already done so, running the same
success pipeline twice per async request and racing on shared logging state.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 10:02:51 +00:00
yassin
3cc6a70466 docs(proxy): describe the /model/info JSON response and regenerate schema.d.ts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 09:48:26 +00:00
yassin
93e6770d68 perf(proxy): serialize /model/info listing once with orjson
FastAPI walked the multi-megabyte /model/info payload through jsonable_encoder
before json.dumps on every request. Return a prebuilt orjson Response instead,
keeping jsonable_encoder as the fallback for datetimes and other non-native values

Resolves LIT-5724

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 09:37:00 +00:00
Oliver Jensen
b3882d8e43
Merge pull request #40107 from BerriAI/litellm_forced_password_reset
feat(auth): breached password detection and forced change

BREAKING CHANGE: users can no longer change their password by issuing a request with a password parameter to /user/update; this has been replaced with /user/password/change dedicated to secure password change.
2026-09-14 10:17:21 +02:00
Oliver Jensen
f9da8a19b6
test(models): stop the password serialization test matching field-name substrings 2026-09-14 09:59:45 +02:00
Oliver Jensen
40118bd158
test(auth): annotate the session-minting patch for the test-quality gate 2026-09-14 09:59:45 +02:00
Oliver Jensen
38c04638ef
fix(lint): clear the one-over LIT002 and inline-object budget hits 2026-09-14 09:59:45 +02:00
Oliver Jensen
e77d11d8d7
refactor(auth): type the breach-screen DB dicts and flatten the session-key kwargs
Annotate screen_login_password_for_breach's update/where dicts with
prisma input TypedDicts and replace authenticate_user's conditional
dict splat with plain keyword arguments, clearing the LIT002 lines
this branch added in login_utils.py. No behavior change: an unflagged
login now passes allowed_routes=None and metadata={} explicitly, which
are the parameter defaults
2026-09-14 09:59:45 +02:00
Oliver Jensen
4f2836bc60
feat(auth): screen the login password inline and restrict the session on a fresh breach hit
A breach found during a login previously only flagged the account for the
NEXT login, handing out one free unrestricted 24h session. The HIBP screen
is now awaited before the session key is minted (worst case one 5s window
per user per 24h, fail-open unchanged), so a fresh hit restricts the
current session and the dashboard routes straight to change-password.

Also repairs two casualties of merge f5e47974db that the layout tests
caught: the lost usePathname import and a call to migratedHref, which
staging renamed to uiHref.
2026-09-14 09:59:45 +02:00
Oliver Jensen
671d032b20
feat(auth): force password reset for breached or admin-set passwords 2026-09-14 09:59:45 +02:00
Oliver Jensen
faf755345a
fix(auth): clear the CI gates on the change-password PR
The Terraform endpoint audit wanted POST /user/password/change covered
or allowlisted; it is a caller-scoped one-shot action, so allowlist it
next to /user/bulk_update. leftnav.test.tsx mocked next/navigation
without useRouter, which SidebarAccountMenu now calls, so every render
in that file threw. The two unannotated audit-log patches in
test_password_endpoints.py get their test-quality-ok reasons.

Also removes the LIT002 violations the PR added: prisma input TypedDicts
annotate the where/data dicts, a shared HTTPExceptionErrorDetail
TypedDict covers the HTTPException detail dicts, and the route decorator
takes a tags tuple.
2026-09-14 09:59:44 +02:00
Oliver Jensen
d79a893e37
feat(auth): add self-service change-password endpoint
Admin password sets on /user/update and per-user /user/bulk_update stay
supported and policy-enforced. The request model hides the password from
repr so management alerts never format the plaintext, and the all_users
bulk path rejects passwords instead of writing one plaintext value to
every row.
2026-09-14 09:59:44 +02:00
Oliver Jensen
5bb2c9e76f
fix(auth): annotate the strict-rule suppressions the merged gates now count
The staging merge brought BLE001 into the strict ruff set and lowered the
LIT002 ceiling, so the HIBP fail-open except and the params/headers dicts
in password_policy.py now need their noqa and mutable-ok reasons. The
headers dict moves to an annotated Final so the suppression fits the line
limit.
2026-09-14 09:59:43 +02:00
Oliver Jensen
1d18d11fcf
Apply suggestion from @greptile-apps[bot]
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-09-14 09:59:43 +02:00
Oliver Jensen
0bb0218d0b
fix(auth): screen bulk-update passwords concurrently before any db write
/user/bulk_update awaited a separate HIBP lookup for each user in the
batch, so a degraded-slow HIBP (5s timeout per lookup) could stretch a
500-user batch to ~2500s and time out the request after some updates
had already persisted.

validate_passwords_bulk dedupes the batch's passwords, strength-checks
first, then fires every needed HIBP lookup concurrently, bounding the
worst case at one 5s timeout window. bulk_update_processed_users now
screens the whole batch before the serial update loop, so a rejected
password fails only its own entry and validation failures precede any
persistence.
2026-09-14 09:59:43 +02:00
Oliver Jensen
a9a0bcb9f8
move hibp url to constants 2026-09-14 09:59:43 +02:00
Oliver Jensen
fcf7cb6e0c
fix(ui): regenerate schema.d.ts for the new_user password docstring 2026-09-14 09:59:43 +02:00
Oliver Jensen
1f0ab3d176
fix(auth): drop general_settings import left unused in new_user 2026-09-14 09:59:43 +02:00
Oliver Jensen
e2ea7e97a5
fix(auth): document /user/new password rejection and format password_policy 2026-09-14 09:59:43 +02:00
Oliver Jensen
bf8df3ab02
hibp support in password policy 2026-09-14 09:59:43 +02:00
devin-ai-integration[bot]
daa665e578
build(deps): re-suppress GHSA-h7x2-h6g9-p789 in osv-scan, mlflow still has no fixed release (#41036)
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-13 21:46:25 -07:00
Taranum01
89d7a3890f test(proxy): annotate the test-quality suppressions in the budget test
The test-quality budget gate has no headroom (every TQ rule is seeded
at exactly its current count on the gate's base, and the only legal
direction is down), so the new TQ005 / TQ008 violations introduced by
the budget test would otherwise block the PR.

Annotate each offending line with a `# test-quality-ok: <reason>`
comment that names the actual reason:
- TQ005 (process-wide global mutation) — the test deliberately flips
  ``litellm.enforce_end_user_model_max_budget_on_master_key`` behind a
  feature flag, which is the public behaviour the test asserts.
- TQ008 (SDK internal patching) — the proxy internals touched here have
  no public seam yet; rewriting the tests around a public seam is a
  follow-up. Until then the patches are the only signal we have.

The lint gate now reports `OK: every TQ rule is within its test-suite
ceiling (base c2c2a623c0)`.
2026-09-13 21:34:59 +05:30
Taranum01
ba42ccb1a7 fix(ui): regenerate schema.d.ts after enabling audit search param
The audit-logging endpoint in enterprise/litellm_enterprise added a
``search`` query parameter; the regenerated types hadn't picked it up.

Run ``pnpm run gen:api`` from ui/litellm-dashboard (which loads the
local litellm_enterprise edit-install so the audit endpoint is in the
spec) and commit the result. ``Check UI API Types Sync`` is happy again.
2026-09-13 21:26:17 +05:30
Taranum01
b939bb0b72 fix(proxy): restore AST-visible end-user budget check in custom auth
Two structural-test invariants broke during the upstream merge:

1. `test_master_key_auth_sets_via_virtual_key_marker` expected
   `_user_api_key_obj.via_virtual_key = True` after
   `update_valid_token_with_end_user_params`. The conflict resolution
   ate that line.

2. `test_custom_auth_also_skips_budget_checks_for_zero_cost_models`
   walks the AST of `_run_post_custom_auth_checks` and asserts that
   `is_end_user_within_model_budget` is called directly inside it,
   guarded by `skip_budget_checks`. The helper-extraction refactor
   moved that call into `_enforce_end_user_model_max_budget_checks`,
   which hides it from this function's AST.

Keep `_enforce_end_user_model_max_budget_checks` for the main auth
path and master-key path (where it is the only caller and the helper
is the right factoring). Inline the check back into
`_run_post_custom_auth_checks` to match upstream's structure and
satisfy the structural invariant without rewriting tests.
2026-09-13 21:13:47 +05:30
Taranum01
4e433bc9a1 fix(ui): regenerate schema.d.ts after upstream soft_budget removal
Upstream removed the ``soft_budget`` field (and one audit-log search
parameter) from the proxy's OpenAPI spec. The dashboard's generated
types hadn't been regenerated, so ``Check UI API Types Sync`` flagged
the diff.

Regenerated with ``npm run gen:api`` per ui/litellm-dashboard/CLAUDE.md
("schema.d.ts is generated ... never hand-edit it").

Also includes the bench GC-disable fix from the previous commit.
2026-09-13 16:28:11 +05:30
Taranum01
26b35080ba test(benchmarks): disable CPython GC during the benchmark session
Stabilises CodSpeed measurements of the LLM-completion benchmarks by
removing GC-induced noise from the per-iteration instruction count.

CPython's cyclic collector fires on its own clock and, because the
multi-turn benchmark only allocates a few KB per iteration, a collection
that lands mid-iteration inflates the per-iteration count by tens of
percent — exactly the magnitude of the flake that caused #32136's
test_completion_multi_turn to be flagged as a -25% regression.

The existing ``inline_logging_executor`` fixture already proved the
pattern works: deferring asynchronous executor work to a per-iteration
inline call removes background-thread scheduling noise. GC is the same
class of artefact — non-deterministic, runs orthogonally to the code
under test — and gets the same treatment.

The deferred collection runs once at session teardown; ``mock_response``
keeps the benchmarks on synthetic allocations so nothing escapes into
real tracing.

Verified locally: the multi-turn benchmark's standard deviation drops
from ~0.37 ms to ~0.001 ms across 20 × 1000-iteration runs, i.e. the
GC-attributable variance is now ~370× smaller.
2026-09-13 16:21:53 +05:30
Taranum01
88d3650d67 Merge upstream/litellm_internal_staging into fix/31842-end-user-model-max-budget
Resolves the merge conflicts blocking PR #32136.

Conflicts resolved:

1. litellm/__init__.py — both branches add an independent module-level
   bool flag. Kept both:
   - `enforce_end_user_model_max_budget_on_master_key` (this PR)
   - `block_requests_for_models_without_pricing` (upstream)

2. litellm/proxy/auth/user_api_key_auth.py — four conflict regions in
   the master-key auth path, the virtual-key auth path, the DB lookup
   auth path, and the custom-auth post-checks path. The upstream rewrite
   inlined some end-user parameter population; this PR keeps its
   `_enforce_end_user_model_max_budget_checks()` /
   `_maybe_enforce_master_key_end_user_model_max_budget()` helper paths
   so the master-key enforcement behaviour introduced by #32136 is
   preserved.

Verified with `python3 -c "import ast; ast.parse(...)"` on both files.
2026-09-13 16:06:17 +05:30
yassin
6d2c4899b0 fix(health): skip background health check DB writes when the latest-row read fails
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-13 09:44:22 +00:00
Mateo Wang
c2c2a623c0
Merge pull request #39846 from BerriAI/litellm_bedrock_mantle_govcloud_cost_row
Some checks are pending
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
fix(bedrock_mantle): price GovCloud regions from the regional cost row and accept region-prefixed model names
2026-09-12 21:13:58 -07:00
devin-ai-integration[bot]
62b3a93219
build(deps): bump smol-toml to 1.8.0 to clear GHSA-7w5x-hrqm-74c2 in osv-scan (#40478)
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 21:13:54 -07:00
Mateo Wang
b1a61f510c
Merge pull request #35918 from Lee-Si-Yoon/feat/friendli-model-metadata-sync
feat(friendli): auto-sync Friendli model metadata into price registry
2026-09-12 21:13:52 -07:00
Shivam Rawat
e8d671c94a
Merge pull request #36585 from BerriAI/litellm_remove_user_soft_budget_docstring
docs(user endpoints): remove unsupported soft_budget param from user docstrings
2026-09-12 21:13:46 -07:00
devin-ai-integration[bot]
8851148330
fix(router): preserve Azure Entra ID params in reusable credentials (#40889)
CredentialLiteLLMParams omitted tenant_id, client_id, client_secret,
azure_scope, azure_username and azure_password, so the strict dump used
by credential reuse and Azure client init dropped them and the reused
credential ended with no auth at all

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 21:13:45 -07:00
Yassin Kortam
036bfc08fc
docs(e2e): ban unit tests under tests/e2e (#33852)
The e2e harness exists to prove product features end to end against a live
proxy. The prior Hard Rule carved out an exception for "tests that cover the
harness itself" and pointed at coverage_registry/test_collector.py, which in
practice invited unit tests of harness helpers to be staged alongside e2e
work. That is the wrong tool: harness logic that is worth locking down does
not need a mock-driven unit test living under tests/e2e.

Drop the carve-out. The Hard Rule now reads that no unit tests of any kind
belong under tests/e2e, and the passing mention of unmarked harness coverage
in the transport section is removed so the doc no longer contradicts itself.

coverage_registry/test_collector.py still exists on disk and is left in place
for now; whether to relocate or remove it is a separate decision.
2026-09-12 21:13:43 -07:00
Mateo Wang
939d320246
Merge pull request #40618 from BerriAI/litellm_pr_template_affected_release
docs(github): add an Affected release section to the PR template
2026-09-12 21:13:38 -07:00
devin-ai-integration[bot]
77dc1a6c03
fix(anthropic-adapter): surface mid-stream provider errors as Anthropic error events (#33352)
* fix(anthropic-adapter): surface mid-stream provider errors as Anthropic error events

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* style(anthropic-adapter): drop added comments per repo convention

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-09-12 21:13:35 -07:00
Mateo Wang
386d29ee67
Merge pull request #38867 from BerriAI/litellm_hide_admin_tabs_view_only
fix(ui): hide admin write-form tabs on the models page from view-only admins
2026-09-12 21:13:34 -07:00
ryan-crabbe-berri
760119681c
Merge pull request #40814 from BerriAI/litellm_gate_health_services_alert_tests
fix(proxy): gate the webhook test alert on proxy admins
2026-09-12 21:13:30 -07:00
Mateo Wang
70e3f5a02e
Merge pull request #39836 from BerriAI/litellm_lit_6975_bedrock_files_delete_list
feat(bedrock): support file delete and list for S3-backed managed files
2026-09-12 21:13:27 -07:00
ryan-crabbe-berri
1ce3690257
Merge pull request #40657 from BerriAI/litellm_lit_7358_session_token_grant_resolver
fix(auth): refresh lite login session token grants from the live user and team rows
2026-09-12 21:13:25 -07:00
Mateo Wang
a978ad2227
Merge pull request #39068 from BerriAI/litellm_spend_log_request_id_call_id
fix(spend_logs): store litellm_call_id and match it in request_id lookups
2026-09-12 21:12:57 -07:00
yuneng-jiang
daa2b0248a
Merge pull request #40172 from BerriAI/litellm_remove_main_guard
ci: remove main branch source guard
2026-09-12 21:10:34 -07:00
mateo-berri
8608a03bd8 Merge origin/litellm_internal_staging into litellm_spend_log_request_id_call_id
Keeps the base's rule that a non-admin id lookup matching no spend-log row answers 403, so the detail route never consults cold storage without an owner row
2026-09-12 21:04:25 -07:00