Commit graph

49317 commits

Author SHA1 Message Date
michelligabriele
eddb29d90e
fix(proxy): dedup latest health checks in SQL and gate the DB save per window 2026-09-03 11:33:44 +02:00
mateo-berri
2c08e7abf8 test(ai-gateway): pin the crypto provider to ring and prove the dial installs it
The two tests that shipped with the fix both called ensure_crypto_provider
themselves, so deleting the call from connect_upstream left the whole suite
green, and swapping ring for aws-lc-rs did too.

Adds an integration test, which gets its own process, that dials wss:// at a
local plain-TCP listener through the public Responses WebSocket entrypoint and
asserts an Err plus an installed provider. Without the install in the dial it
panics with the original CryptoProvider message. A unit test now compares the
installed provider's cipher suites and key-exchange groups against ring's, so
the choice of backend is pinned rather than assumed.

Also names tls12 in the workspace rustls features: it already arrives through
reqwest and tokio-rustls, so the graph is unchanged, but a direct dependency
should say it needs TLS 1.2 rather than inherit it.
2026-09-03 02:27:00 -07:00
mateo
d4a480fb6e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_deflake_20260902 2026-09-03 09:18:54 +00:00
mateo-berri
c79d1d12ae fix(ai-gateway): install a rustls crypto provider before dialing upstream WebSockets
The gateway's dependency graph turns on two rustls crypto backends at once:
reqwest's rustls-tls pulls in ring, and litellm-core's bedrock-auth pulls in
aws-lc-rs through aws-config. rustls 0.23 refuses to guess between them, so
ClientConfig::builder panics, and that is exactly how tokio-tungstenite builds
its TLS config. Every outbound WebSocket dial killed its tokio worker and the
client saw the socket vanish with no close frame.

reqwest and the AWS SDK both pick a provider explicitly, so only the tungstenite
path was affected. Route all three dial sites through one helper that installs
ring once per process before connecting.
2026-09-03 02:15:57 -07:00
mateo-berri
0a62195db2 fix(utils): redact credentials nested inside lists and tuples
redact_credentials_in_payload only recursed into mappings, so a
credential-named key one level inside a list or tuple, the shape
extra_body and metadata routinely carry, still reached stdout under
set_verbose. Rebuild sequences element by element too, keeping the
container's own type so the printed repr is unchanged apart from the
secret.
2026-09-03 02:11:32 -07:00
mateo-berri
64601fd7ae fix(utils): redact credential kwargs from the set_verbose request line
`litellm.set_verbose = True` printed the caller's kwargs verbatim to stdout, so
`api_key` and its siblings landed in terminals and container log drains in
plaintext while the same statement's logger emission was already redacted.

Mask the kwargs at the source with a shared helper in
`litellm_core_utils/sensitive_data_masker.py`, reusing the existing
`SensitiveDataMasker` key classification and the `REDACTED` marker
`secret_redaction.py` already owns, so both debug surfaces agree.
2026-09-03 01:55:56 -07:00
mateo-berri
dc28f0eb4b fix(vector-stores): report search failures on the responses API surface too 2026-09-03 01:54:09 -07:00
mateo-berri
1d1eb4264f build(ai-gateway): keep the committed enterprise wheels out of the build context 2026-09-03 01:49:40 -07:00
Devin AI
e1b2d9de3c fix(images): forward gpt-image supported params like background to OpenAI and Azure
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 08:45:23 +00:00
mateo-berri
fbb799e240 ci: drop the ai-gateway image from the coverage allowlist now that a job builds it 2026-09-03 01:44:20 -07:00
mateo-berri
b55833b1f9 fix(ai-gateway): build the release image again and cover it in CI
The image had two independent breaks. The Dockerfile pinned rust 1.90 while
the repo pins 1.98 in rust-toolchain.toml and never copied it in, so the first
cargo call died on crates needing a newer rustc. The runtime stage then ran
pip install on the root pyproject, which builds with maturin against the
python-bridge crate, so metadata generation failed with no Cargo manifest and
no Rust toolchain in that stage.

Copy rust-toolchain.toml into the builder so every cargo call uses the pinned
channel, build the wheel in the builder stage where cargo and python3-dev
already live, and have the runtime stage install that artifact instead of
compiling anything. Add the ai-gateway image job to the rust workflow so a
broken build fails a PR instead of surfacing on a release.
2026-09-03 01:40:00 -07:00
mateo
4e9c6b5dd4 refactor(model_armor): type the buffered stream chunks as object
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 08:26:43 +00:00
mateo-berri
2647c2890f fix(oci): keep the stream's own finish reason instead of a synthetic stop chunk
The OCI wrapper overrides chunk_creator wholesale, so it never recorded the finish reason or marked the terminal chunk as sent. The shared end-of-stream finalizer then appended a synthetic chunk whose finish_reason was always stop, which downgraded a tool_calls completion for any client that reads the finish reason off the last chunk.
2026-09-03 01:25:32 -07:00
mateo-berri
1b4d2e25db fix(proxy): stop putting the literal string "None" in error payloads
A blocked guardrail (and any other HTTP error the proxy converts) came back
with "type": "None" and "param": "None", because the converters passed the
string "None" as the getattr default instead of None. OpenAI types error.type
as a required string and error.param as nullable, so type now falls back to
the type its status code stands for and param serializes as JSON null.

Covers the non-streaming body, the SSE error frame, the client-disconnect
frame, and the unclassified-exception path, so every unified LLM endpoint and
the anthropic endpoints return the same shape.
2026-09-03 01:23:48 -07:00
mateo-berri
e464e749e0 fix(vector-stores): fall back to annotate when the failure mode is unrecognized
litellm_settings keys are set on the litellm module with no allowlist, so a typo
in vector_store_search_failure_mode reached assert_never and turned every
vector-store request into a 500. Validate the configured value and fall back to
the permissive default with a warning naming the supported modes.
2026-09-03 01:23:28 -07:00
mateo-berri
fa5a90e08e fix(bedrock): stop the Moonshot invoke transform from resolving AWS credentials
AmazonMoonshotConfig.transform_request called
_get_boto_credentials_from_optional_params purely for its side effect of
popping the aws_* keys off optional_params, then threw the result away. On
a box whose default AWS profile uses login_session without botocore[crt],
that call raises, so a bearer-token bedrock/invoke/moonshot.* deployment
still 500s with MissingDependencyException even after the rest of this
branch skips the chain.

It now filters the aws_* keys into a local dict the way the Qwen, OpenAI
and Claude 3 invoke transformations already do, so no credentials are
resolved and the caller's optional_params keeps the keys sign_request
reads afterwards.
2026-09-03 01:20:27 -07:00
mateo-berri
461a3a3ea5 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_spend_logs_provider_response_id 2026-09-03 01:16:28 -07:00
mateo-berri
c85da0a75f fix(logging): key bridged /v1/messages rows on the id the caller received
/v1/messages against a non-Anthropic model answers with the Responses id,
but the spend row was built from a fresh ModelResponse, so it landed on a
chatcmpl- uuid nobody can look up. Carry that id through the same way the
Anthropic branch now does, and make the passthrough spend assertions fail
on an empty lookup instead of skipping past it.
2026-09-03 01:15:06 -07:00
mateo-berri
6499ca349f fix(vector-stores): import assert_never from typing_extensions for Python 3.10 2026-09-03 01:11:22 -07:00
mateo
3190f42abf refactor: clear fresh tech debt from the last 24 hours (2026-09-03)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 08:09:51 +00:00
mateo-berri
f823d9ae0d Merge remote-tracking branch 'origin/litellm_vector_store_hook_router_injection' into litellm_vector_store_surface_retrieval_failure 2026-09-03 01:09:34 -07:00
mateo-berri
445ccc14b2 fix(vector-stores): surface retrieval failures to the API caller
A vector store search that fails is swallowed by the pre-call hook, so the
request goes to the model with an un-augmented prompt and the caller gets a
200 answering from the model's own knowledge with no way to tell the
knowledge base was skipped.

Failed searches now ride the same channel their successes already use: a
vector_store_search_failures entry on provider_specific_fields naming the
store id, provider, and error. That is additive and always on. For callers
who would rather fail than answer ungrounded, litellm_settings
vector_store_search_failure_mode: error raises VectorStoreSearchError (400)
instead; the default stays annotate, today's permissive behavior.

The hook's outer catch-all also now names the requested vector store ids in
its log line, and only wraps the augmentation itself, so the fail-closed
raise is not swallowed by it.
2026-09-03 00:53:55 -07:00
mateo-berri
3b814179c8 refactor(logging): name the response-id helper for what it reads, not the provider 2026-09-03 00:50:11 -07:00
mateo-berri
6edb72f79c Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_azure_ai_reclassify
# Conflicts:
#	tests/test_litellm/test_main.py
2026-09-03 00:47:14 -07:00
mateo-berri
58575c77a5 test(e2e): correlate anthropic passthrough spend rows by the served message id 2026-09-03 00:46:07 -07:00
mateo-berri
54f4fa2e1b test(passthrough): look up anthropic spend rows by the message id the caller received 2026-09-03 00:42:37 -07:00
mateo-berri
ef453c943a Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_ui_build_check_image_boundary 2026-09-03 00:38:51 -07:00
yuneng-jiang
658f50663d
fix(ui): keep Virtual Keys list state in the URL so it survives leaving the page (#39481)
* fix(ui): keep Virtual Keys list state in the URL so it survives leaving the page

The search term, sort, pagination and drawer filters lived in component
state, so navigating away from Virtual Keys and back reset the table to an
unfiltered first page. Move them into query state alongside the existing
?key= deep link, which also makes a filtered view shareable.

* fix(ui): namespace the Virtual Keys filter params and bound page inputs

The unprefixed team_id filter hijacked the /api-keys create-key deep link,
which already takes team_id as a prefill, so ?create=true&team_id=X silently
filtered the list underneath the modal. Prefix the four drawer filters.

Now that page and page_size come from the address bar, clamp them to what
/key/list accepts instead of forwarding 0, negatives or an int64-overflowing
page straight through, and trim filter values arriving from a URL the same
way the drawer already trims them.

* fix(ui): fall back to a sortable column when the URL names an unknown one

A hand-edited or stale sort_by reached /key/list, which 400s it, leaving the
Virtual Keys page on its loading skeleton with no error. Validate it against
the fields the table's own headers can produce, and clear sort_by rather than
blanking it when a sort is reset so the URL stays clean.

Also replaces a default-state URL assertion that ran before any query-state
write could land, so it could not fail for the regression it named.

* fix(ui): use TanStack's functionalUpdate instead of a hand-rolled updater resolver

The local helper narrowed typeof updater === "function" against an
unconstrained T, which TypeScript cannot do because T itself may be a function
type, so next build failed to type check. table-core already exports the same
helper.
2026-09-03 00:35:32 -07:00
mateo-berri
6582dbe1c0 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_coerce_multipart_numeric_fields 2026-09-03 00:33:18 -07:00
mateo-berri
0e537d212a Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_vector_store_hook_router_injection 2026-09-03 00:32:33 -07:00
mateo-berri
20a55146cc Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_mistral_voxtral_tts_speech 2026-09-03 00:31:24 -07:00
mateo-berri
0a2581c14c fix(mistral): keep the deployment voice default and drop the unreachable api base fallback
Review turned up two real problems in the TTS path.

Router.aspeech forwarded voice=None whenever the caller omitted it, which overwrote a
voice set in the deployment's litellm_params, so a configured fallback voice was
ignored on voice-less requests. It now leaves the key alone when no voice is passed.

get_complete_url also fell back to MISTRAL_API_BASE, but speech() always receives a
non-null api_base from get_llm_provider, whose mistral branch only reads
MISTRAL_AZURE_API_BASE and otherwise hardcodes the public host. That branch could
never run, and its unit test asserted a behavior the real path does not have. The
working override is api_base on the deployment, now pinned by an end-to-end test
2026-09-03 00:30:07 -07:00
mateo-berri
1cfc402763 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_oci_streaming_chunk_ids 2026-09-03 00:28:02 -07:00
mateo-berri
fd2104a61e refactor(oci): build the identity-pinned stream chunk in one shot
Address the two Greptile P2 notes: construct the chunk through the shared
creator instead of mutating its choices afterwards, and drop the decorative
divider comment from the new tests.
2026-09-03 00:27:54 -07:00
mateo-berri
46d7e92845 fix(spend_tracking): key /v1/messages spend rows on the msg_ id the client received
POST /v1/messages returns an Anthropic-shaped body whose `id` is the only
request id the caller ever sees, but the spend row was written with a
`chatcmpl-<uuid>` (non-streaming) or the bare `litellm_call_id` (streaming and
the /anthropic/v1/messages passthrough), so
GET /spend/logs?request_id=msg_... returned [].

The logging conversion now carries the provider's response id through:
_handle_anthropic_messages_response_logging seeds the ModelResponse it builds
with the Anthropic id, and the passthrough logging handler prefers the id it
read off the response body or the message_start chunk over litellm_call_id.
get_spend_logs_id already prefers response_obj["id"], so the spend row and
standard_logging_object["id"] now both carry the id the client holds.
2026-09-03 00:26:18 -07:00
Mateo Wang
066d5f0694
Merge pull request #39502 from BerriAI/litellm_/triage-slack-message-4bd4e6
fix(test): drop the duplicate embedding_executor arg in the Bedrock KB fake handler
2026-09-03 00:26:08 -07:00
mateo-berri
62c7e84448 fix(proxy): parse numeric multipart fields on /v1/images/edits back into numbers
Every field of a multipart form arrives as a string, so `n` reached the
provider as "2" and Bedrock Nova Canvas rejected the request with
"expected type: Number, found: String". Restore the type the request
schema declares at the boundary where the form is parsed, driven by the
schema's own type hints so the helper covers any int- or float-typed
field on any multipart endpoint.
2026-09-03 00:22:54 -07:00
mateo-berri
c6bd682a5c Merge branch 'litellm_internal_staging' into litellm_fix_failing_request_slowdown
Resolve type-discipline-budget.json by taking the lower limit per rule so no
ceiling ratchets back up.

Staging's 66a3d24b3f left a duplicate embedding_executor parameter in the
Bedrock KB fake handler, which makes ruff fail on the whole tests tree. Drop
the duplicate here so this branch compiles; #39502 makes the same change on
staging.
2026-09-03 00:20:51 -07:00
mateo-berri
8572544b44 refactor(proxy-extras): pull the migrate deploy recovery branches into a budget helper
_setup_database_v2 decided the next attempt budget inline in eight
branches, each rebinding budget before continuing. The branches now live
in _budget_after_deploy_failure, which returns the budget the next pass
runs under, and the two identical idempotent-recovery blocks share
_mark_migration_applied. The loop backs off whenever a pass spent an
attempt, which is the same set of paths that slept before.
2026-09-03 00:20:46 -07:00
mateo-berri
71f5b89499 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_python_version_ci
Resolves two conflicts:

- tests/test_litellm/vector_stores/test_main.py: staging moved search() to a
  RouterVectorStoreEmbeddingExecutor while this branch parametrized the same
  test over query; keep both the executor assertions and the parametrize.
- tests/logging_callback_tests/test_bedrock_knowledgebase_hook.py: staging
  carries a duplicate embedding_executor kwarg that makes the file a
  SyntaxError; drop the trailing duplicate.
2026-09-03 00:16:19 -07:00
mateo-berri
9fb403a80f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bedrock_bearer_skip_sigv4_chain 2026-09-03 00:12:00 -07:00
mateo-berri
b503bcabea test(vector-stores): cover the hook's default proxy runtime wiring 2026-09-03 00:09:27 -07:00
mateo-berri
69efad0384 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_oci_streaming_chunk_ids 2026-09-03 00:09:23 -07:00
mateo-berri
75e7f4c4a5 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_mistral_voxtral_tts_speech
Resolves the tests/test_litellm/test_main.py collision, where both sides appended a
new test at the end of the file, by keeping both.

Also carries the one-line fix from #39502: staging arrived with a duplicate
embedding_executor kwarg in the Bedrock KB fake handler, which ruff rejects as a
syntax error, so every commit here would otherwise fail lint. The change is byte
identical to #39502, so that PR merges cleanly once it lands.
2026-09-03 00:07:31 -07:00
mateo-berri
5120590890 docs(litellm-rust): fix the gateway run commands and point ADDING_A_PROVIDER at the one checks runbook
Both `cargo run` invocations in the ai-gateway README fail with "requires the
features: `server`", the same root cause as the missing CI coverage.
2026-09-03 00:07:19 -07:00
mateo-berri
0135e1d634 fix(oci): pin one response id per streamed completion, skip the [DONE] sentinel
OCIStreamWrapper.chunk_creator built every chunk straight from the apiFormat
handlers, so it never reached model_response_creator and OCI streams came back
with a fresh chatcmpl id, a drifting created value and no model on every chunk.
Both exits now go through the shared creator.

The GENERIC apiFormat also closes its stream with a literal `data: [DONE]` line,
which chunk_creator json-parsed and turned into a 500 on every OCI streaming
completion. It is skipped now.
2026-09-03 00:01:44 -07:00
yucheng-berri
ecabfbd5af
fix(guardrail): hide-secrets playground redaction and guardrail telemetry (#39398)
* Fix hide-secrets guardrail: playground redaction, UI dropdown entry, spend-log telemetry

The hide-secrets guardrail never implemented apply_guardrail, so the UI test
playground echoed secrets verbatim; it was missing from the Add Guardrail
dropdown; and it recorded no guardrail_information, so Spend Logs could not
distinguish a redacted request from a clean one.

- implement apply_guardrail (unified interface) with use_native_lifecycle_hooks
  so proxied traffic stays on async_pre_call_hook (per-key opt-out and
  data["prompt"] handling live only there)
- record standard_logging_guardrail_information (allow/mask + masked_entity_count)
  via _process_response/_process_error; opted-out keys and legacy nameless
  callback instances record nothing
- advertise hide-secrets in /guardrails/ui/add_guardrail_settings (pre_call only)
  and /guardrails/ui/provider_specific_params with a config model

Resolves LIT-3548

* Fix hide-secrets passthrough telemetry and JSON config input

* fix(guardrails): validate hide-secrets object config before submit

- apply_guardrail treats empty-string-only texts as no input, so no
  false allow is recorded
- the UI object field keeps raw text while editing and blocks submission
  until it parses to a JSON object, instead of posting a string to an
  object-only API
- supported_modes_by_provider keeps its dict[str, list[str]] value type

* fix(guardrails): record no hide-secrets telemetry when nothing was inspected

walk_user_text and the prompt redaction now report how many non-empty
strings they visited; when neither inspected anything (image-only
content, empty strings), the run records no guardrail entry instead of
an 'allow' row that counts a check which never saw any text.
2026-09-03 00:01:03 -07:00
mateo-berri
7b8cc0319e fix(proxy-extras): only spend a migrate-deploy attempt when a pass made no progress
The v2 migration resolver gave `prisma migrate deploy` four attempts, and
every recovery path ended in a bare `continue`, so each one burned an attempt.
A database first brought up with `--use_prisma_db_push` has a full schema and
no migrations ledger, so the baseline spent attempt one and the first three
migrations whose objects already existed spent the rest. The proxy then exited
before binding its port, and that database could never be moved onto the
resolver.

The retry budget now counts only attempts that got nowhere. Creating the
baseline, and each migration newly marked applied, leaves the budget alone, so
a push-created database works through its pre-existing objects one pass at a
time. Timeouts, deadlock rollbacks, advisory-lock waits, and a repeat of a
recovery that already ran still spend an attempt, so a run that stops making
progress gives up exactly as before.
2026-09-02 23:56:09 -07:00
mateo-berri
9c795e52f3 chore(lint): drop budget limits back to the values on the merged base
The staging merge resolved three budget conflicts by keeping this branch's
older, higher numbers, which turned budget-ratchet red. Nothing on the branch
adds violations for those rules, so the base's limits hold.
2026-09-02 23:55:32 -07:00
mateo-berri
e5c1133a79 chore(lint): re-ratchet lint budgets after merging staging 2026-09-03 06:53:19 +00:00