Commit graph

46318 commits

Author SHA1 Message Date
Yuneng Jiang
cfe247ebfe
fix(vector_stores): translate the two MongoDB driver errors that still reached callers as 500s
A connection string whose password holds an unescaped '/' makes pymongo's URI
parser raise a plain ValueError, not a PyMongoError, and a URI with no
credentials at all makes Atlas close the connection, which surfaces as
AutoReconnect. Neither was handled, so both fell through to litellm's generic
wrapper and were served as 500s with a traceback for what are routine typos.
Both now return a 400 naming the cause. The ConnectionFailure branch sits after
the ServerSelectionTimeoutError and NetworkTimeout branches, which subclass it,
and two ordering tests pin that.
2026-09-02 14:55:13 -07:00
Yuneng Jiang
52de1bb1d3
fix(vector_stores): reject MongoDB search params the provider cannot honour
filters was already refused, but ranking_options and rewrite_query were
accepted and then dropped. A caller asking for score_threshold 0.9 got results
scoring 0.5 with a 200 and no indication the threshold never ran, which is the
silent-wrong-answer case the filters check exists to prevent. Both now raise
the same 400 naming the parameter and what to do instead.
2026-09-02 14:19:28 -07:00
Yuneng Jiang
1b47486724
style(vector_stores): cut the explanatory comments down to one line each
The repo's rule allows a comment only where the logic stays confusing after the
code has been made as clear as it can be, and then only one concise line about
why. Three multi-line blocks did not meet that: the reason "connection" joins
the sensitive patterns belongs in the commit that added it, and the weakref and
Atlas error-code notes each say what they need to in a single line.
2026-09-02 13:50:45 -07:00
Yuneng Jiang
d4b0266192
fix(vector_stores): release MongoDB clients built on closed event loops
The async client cache is keyed per event loop, and pymongo's AsyncMongoClient
holds a reference to the loop it was built on, so an entry for a closed loop
kept that client and its sockets alive for the life of the process. A script
that calls asyncio.run once per search fills the cache to its cap this way and
then stops caching entirely. Measured live against Atlas over 40 loops: 32
pinned clients and 212 open descriptors before, 1 cached client and no
monotonic descriptor growth after.
2026-09-02 13:39:14 -07:00
Yuneng Jiang
ed8203757a
fix(vector_stores): refuse MongoDB vector store create with a 400, not a 500
litellm.exception_type passes only litellm's own exception types through
untouched, so the NotImplementedError the search-only refusal raised reached
the caller as APIConnectionError. The proxy served that as a 500 with a
traceback in the body for what is a plain client mistake. Raising
BadRequestError gives the caller a 400 and the message on its own.
2026-09-02 13:21:55 -07:00
Yuneng Jiang
32b501bf74
docs(vector_stores): register mongodb in the provider endpoint support matrix 2026-09-02 13:14:56 -07:00
Yuneng Jiang
3d0223b661 ci: install the mongodb extra for the unit test shard that runs the provider
tests/test_litellm/llms/mongodb imports pymongo's exception classes to check the
error translation against the real hierarchy, and the shard that runs it
(tests/test_litellm/llms, per test-unit.yml) synced --extra google, proxy,
semantic-router and saml but not mongodb, so 24 of 109 tests would have errored
with ModuleNotFoundError on the first CI run. CircleCI hid this because it syncs
--all-groups --all-extras.

  uv export --frozen ... --extra saml                  -> no pymongo
  uv export --frozen ... --extra saml --extra mongodb  -> pymongo==4.17.0

Also close the two gaps a mutation run found in the suite: nothing asserted that
a short request timeout shortens server selection as well as connect, and the
existing code 13 case carried "not authorized", which the message markers match
too, so it could not tell whether the code was still being checked. 28 of 28
mutants now die.
2026-09-02 12:51:19 -07:00
Yuneng Jiang
517e508700 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/mongodb-vector-store-e4ff63
# Conflicts:
#	uv.lock
2026-09-02 12:10:20 -07:00
Yuneng Jiang
fdbee3af25 refactor(vector_stores): build the MongoDB pipeline immutably and inject the client class
The type-discipline and test-quality gates blamed the branch for 4 LIT001, 12
LIT002 and 5 TQ008 violations. Rather than suppress them:

- the $vectorSearch and $project stages are MappingProxyType and the query
  vector a tuple, verified against live Atlas to encode identically. The outer
  pipeline stays a list because pymongo's common.validate_list raises
  "pipeline must be a list, not <class 'tuple'>", which a unit test now pins.
- the client caches are Final[dict[...]] and _client_kwargs returns a
  MappingProxyType.
- _field_value recurses over the dotted path instead of rebinding a local.
- _client_key declared Final locals in one branch and reassigned them in the
  others, so it is split into an early-returning _timeout_ms.
- the injected callables carry explicit Final[Callable[...]] annotations, which
  stops pyright resolving self.embedding_fn against litellm.embedding's
  overloads.
- get_sync_client and get_async_client take an optional client_class, so the
  cache tests inject a recording double instead of patching the importer, and
  can assert the connection string and timeouts the client was built with.

SensitiveDataMasker is public SDK surface, so extra_sensitive_patterns moves to
the end of the signature: in slot two it silently reinterpreted an existing
caller's positional override set as extra sensitive patterns.
2026-09-02 12:09:51 -07:00
Ali Ahmed
a677242d6f
fix(headroom): stop re-compressing retrieved CCR content in client tool loops (#38591)
When the headroom_retrieve tool is exposed to a client that runs its own
tool-execution loop (the LiteLLM MCP gateway path), the client executes the
retrieve call and sends the recovered original content back as a tool result
on the next turn. The guardrail then compressed that row again, and because
CCR is content-addressed it collapsed back to the exact same hash it was just
retrieved from. The model never saw the expansion and the agent looped.

Hold tool-result rows that carry headroom_retrieve output back from the
compression service, the same way the live turn and trailing tool exchange are
already protected, so the expansion survives. Retrieve calls are matched by the
direct headroom_retrieve name and the mcp__<server>__headroom_retrieve gateway
name. Because a long gateway name is truncated past 64 chars in the
OpenAI-translated view the guardrail scans, the pairing also falls back to the
tool-call id read from the request's own untranslated messages, which is never
truncated.

Fixes #38558
2026-09-02 12:07:13 -07:00
Mateo Wang
2f0f0685c0
Merge pull request #39188 from BerriAI/litellm_bump_tornado_658
fix(deps): raise the tornado and pypdf floors for six new advisories
2026-09-02 12:00:06 -07:00
Mateo Wang
ba2e5d2d8e
Merge pull request #39341 from BerriAI/litellm_azure_deepseek_v4_flash_0731
fix(models): key Azure DeepSeek V4 Flash 0731 by its Foundry catalog id
2026-09-02 11:52:22 -07:00
mateo-berri
f38a1ec129 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_azure_deepseek_v4_flash_0731
# Conflicts:
#	litellm/model_prices_and_context_window_backup.json
#	model_prices_and_context_window.json
2026-09-02 11:43:22 -07:00
Mateo Wang
b600f02fc2
Merge pull request #34788 from BerriAI/litellm_fix_s3_vectors_search
fix(vector_stores): s3 vectors search router bypass + rag query config drop + ui error swallow
2026-09-02 11:35:56 -07:00
Yuneng Jiang
5d7bf187a4
refactor(ui): pick the vector store id placeholder from a map
The chain had grown to four nested ternaries with a fifth level inside the
Vertex Search branch, which no-nested-ternary had two suppressions for. A lookup
keyed by provider drops both suppressions and leaves one condition, the Vertex
Search case that depends on whether an engine id has been entered.

Also hoists the MongoDB form fixtures in the tests, which the inline-object
budget counts.
2026-09-02 11:30:36 -07:00
Yuneng Jiang
9434e563f3
fix(vector_stores): translate MongoDB client construction failures too
Building the client parses the URI and, for mongodb+srv://, performs a DNS SRV
lookup, so it fails on exactly the inputs a user is most likely to get wrong. It
sat outside the try that translates driver errors, so a malformed URI or an
unresolvable cluster escaped as a raw pymongo exception and reached the caller as
a 500 with a traceback in the body.

The three DNS-shaped failures are also told apart now: a lookup that ran out of
time is a Timeout, a cluster name that is not in DNS says so and points at the
URI Atlas shows under Connect Drivers, and anything else keeps the generic
"not a usable MongoDB connection string".

Verified live: a tampered scheme, a nonexistent cluster and a 1ms timeout each
come back as their own message instead of a traceback.
2026-09-02 11:28:33 -07:00
Mateo Wang
a43228ef72
Merge pull request #39249 from BerriAI/litellm_router_settings_reject_unknown_keys
fix: apply optional_pre_call_checks and reject unsupported router settings on /config/update
2026-09-02 11:23:22 -07:00
mateo-berri
f7b37e4e6e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bump_tornado_658 2026-09-02 11:19:59 -07:00
Yuneng Jiang
1fed1029e0
fix(vector_stores): report a MongoDB text field that no matched document has
Atlas matches on the vector alone, so a mistyped mongodb_text_field still returns
confidently scored results whose content is empty, and the model is handed an
empty context with nothing to explain it. When every matched document lacks the
field the search now says which setting to fix; a sparse document among others
that do have it, and a document whose text is genuinely the empty string, both
still come back normally.

Unrecognised mongodb_* parameters are named too. The params model has to ignore
unrelated keys because litellm_params carries plenty of them, which turned a
mistyped mongodb_collection into "mongodb_collection is required" pointing the
reader at a key they can see they have set.
2026-09-02 11:12:51 -07:00
Mateo Wang
4049a075bd
Merge pull request #39170 from BerriAI/litellm_registry_audit_2026_09_01
fix(models): registry audit 2026-09-01: openai realtime and long-context tiers, mistral aliases, voyage, xai, fireworks, together, scaleway, azure ai, govcloud, azure gov, cloudflare whisper, deprecation dates
2026-09-02 11:02:43 -07:00
mateo-berri
e0be9a35e6 fix(deps): raise the pypdf floor to 6.16.1 for three new advisories
GHSA-jp53-mhqp-8xcg (fixed in 6.16.0), GHSA-23w6-3w8w-8484 and
GHSA-763m-79hh-57f2 (fixed in 6.16.1) flag pypdf 6.15.0 in uv.lock and
keep osv-scan red alongside the tornado advisories. The proxy-runtime
extra now requires pypdf>=6.16.1 and the lock resolves 6.16.2.
2026-09-02 11:02:43 -07:00
mateo-berri
a9d3a0746c fix(models): price Azure DeepSeek V4 Flash 0731 from its own meters under the catalog id 2026-09-02 10:59:11 -07:00
devin-ai-integration[bot]
ffc0a8e428
fix: run access group key sync UPDATEs on the writer, not the read replica (#39128)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 10:51:17 -07:00
mateo-berri
47804d0f1a Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_bump_tornado_658 2026-09-02 10:51:08 -07:00
Yuneng Jiang
a472484291
fix(vector_stores): stop MongoDB handing a new event loop a closed loop's client
The async client cache was keyed on id(loop). CPython recycles those ids so
aggressively that a fresh event loop nearly always lands on the id of one already
collected, measured at 37 of 40 rounds, so the cache handed the new loop an
AsyncMongoClient bound to a closed loop and every operation on it raised
"Event loop is closed".

The entry now carries a weak reference to the loop it was built on and a hit only
counts when that reference still points at the running loop, so a recycled id
misses and builds a fresh client. A stale entry can also be replaced once the
cache is full, which the old size check prevented.

pymongo's own client keeps its loop alive, which is why the sync proxy path never
saw this; a script calling asyncio.run() per search, or a test suite with a loop
per test, does.
2026-09-02 10:36:36 -07:00
mateo-berri
7a35c34303 fix(models): add the us-gov. geo inference profile keys for Claude Sonnet 5 and Opus 4.8 2026-09-02 10:30:14 -07:00
Mateo Wang
2ffe6a1dc8
Merge pull request #39160 from BerriAI/litellm_gemini_thinking_content
fix(gemini): return enabled thinking content by default
2026-09-02 10:22:28 -07:00
Yuneng Jiang
1f8cbee8aa
feat(ui): add MongoDB Atlas to the vector store provider dropdown
The create form now offers MongoDB Atlas with its connection string, database,
collection, embedding model, vector field, text field and candidate count. The
connection string renders as a password input because it carries the database
user's password, and the embedding model is picked from the proxy's own models,
matching how Milvus and Valkey do it.

The vector store id doubles as the Atlas Vector Search index name, so the
placeholder says so.
2026-09-02 10:22:01 -07:00
Yassin Kortam
de80e3afe4
fix(helm): scale the classic chart's HPA out at the documented 60 percent CPU (#35975)
* fix(helm): scale the classic chart's HPA out at the documented 60 percent CPU

The litellm-helm chart shipped targetCPUUtilizationPercentage: 80, which is
unexamined helm create scaffold rather than a chosen number. It arrived packaged
with the stock minReplicas: 1, maxReplicas: 100, a commented-out
targetMemoryUtilizationPercentage: 80, and the boilerplate "such as Minikube"
comment, the same provenance as the 128Mi resource example this file just
corrected.

60 is the documented recommendation. The mechanism behind it is scale-up lag:
the chart's own startupProbe is failureThreshold: 30 times periodSeconds: 10, so
a replica can take up to 300 seconds to become ready, and a pod added at 80
percent utilization arrives minutes after saturation.

The memory target stays commented out on purpose. The prisma query engine's
resident memory is a high-water mark that ratchets to the pod's worst-ever write
and is never returned, so a memory-target HPA reads the largest write a pod ever
did rather than what it is doing now, and replicas ratchet up without scaling
back in.

hpa_tests.yaml carried its second suite after a YAML document separator, and
helm-unittest loads only the first document per file, so that suite never ran;
an assertion planted in it still passed. Fold it into the one live suite and add
coverage pinning the rendered CPU target, the absence of a memory metric by
default, and that overrides still take effect.

Bump the chart to 1.1.2, since rendered output changes for anyone running with
autoscaling enabled.

* fix(helm): bump litellm-helm to 1.1.3 after rebase onto 1.1.2

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 10:19:03 -07:00
Mateo Wang
6ed2cdd428
Merge pull request #39238 from BerriAI/litellm_anthropic_models_display_name
feat(proxy): configurable display_name for the Anthropic-shaped /v1/models listing
2026-09-02 10:18:43 -07:00
Mateo Wang
921c1d1248
Merge pull request #39246 from BerriAI/litellm_fix_e2e_junit_properties_types
test(e2e): read JUnit properties off the real collected pytest Item
2026-09-02 10:18:18 -07:00
Mateo Wang
1710d977bf
Merge pull request #39176 from BerriAI/litellm_rerank_provider_error_body
fix(rerank): map provider errors with the resolved provider on sync and async paths
2026-09-02 10:18:11 -07:00
Mateo Wang
25c5b6ce43
Merge pull request #39069 from BerriAI/litellm_stream_usage_cost_default
feat(streaming): carry final response cost on streamed usage by default
2026-09-02 10:17:51 -07:00
Yuneng Jiang
85431297b9
fix(vector_stores): name the connection string when Atlas rejects MongoDB credentials
Atlas answers a wrong password with code 8000 "AtlasError" rather than the 18 a
self-hosted deployment returns, so the code-only check never fired and a bad
password came back as a generic "MongoDB rejected the vector search", pointing
the reader at the index instead of at their credentials. Verified live against
Atlas with a tampered password.
2026-09-02 10:03:23 -07:00
Yuneng Jiang
8374b34b81
fix(vector_stores): return 400 for MongoDB misconfiguration instead of 500
litellm.exception_type passes a litellm exception through untouched and wraps
anything else into APIConnectionError, so every bare ValueError this provider
raised reached the caller as HTTP 500 with a Python traceback in the response
body. "max_num_results must be between 1 and 50" is the caller's to fix, not a
connection failure.

Configuration and validation failures now raise BadRequestError (400) and the
two timeout cases raise Timeout (408). ExecutionTimeout subclasses
OperationFailure, so it is matched before it; previously an Atlas query that ran
out of time was reported as "MongoDB rejected the vector search".
2026-09-02 10:01:31 -07:00
Mateo Wang
e4b903d2cd
Merge pull request #39194 from BerriAI/litellm_fix_deepseek_ocr_model_prefix
fix(vertex): avoid duplicate DeepSeek OCR model namespace
2026-09-02 09:57:50 -07:00
Yuneng Jiang
22d34960e5
fix(vector_stores): redact wire-protocol connection strings in management responses
A MongoDB vector store's whole credential is its connection string, and
mongodb+srv://<user>:<password>@<cluster> embeds the database password. None of
the masker's default patterns (api_key, secret, token, credential) match a key
named mongodb_connection_string, so /vector_store/list and /vector_store/info
returned it verbatim to every caller that can read a vector store.

SensitiveDataMasker gains extra_sensitive_patterns, which unions onto the
defaults instead of replacing them, and the vector-store redactor adds
"connection" so the URI is masked while mongodb_database, mongodb_collection and
the field names stay readable.
2026-09-02 09:56:17 -07:00
mateo-berri
0c7fe53c27 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_registry_audit_2026_09_01 2026-09-02 09:25:36 -07:00
Mateo Wang
55e9e4ce2f
Merge pull request #39340 from BerriAI/litellm_gemini_3_8_flash
feat(gemini): day-0 pricing for gemini-3.8-flash
2026-09-02 09:24:24 -07:00
Yuneng Jiang
85bda43d63
fix(vector_stores): turn MongoDB's silent misconfiguration failures into errors
Driving the sad path against a live Atlas cluster showed four cases returning
an empty result set instead of failing: a missing index, a missing database, a
missing collection, and the async path for all three. $vectorSearch reports
none of these as errors, so a misconfigured store looked exactly like a query
that matched nothing, which is the worst shape for this to fail in.

An empty result set is now checked against the index catalogue, which does
report all three correctly, and a store that cannot work says so. The check
costs one extra round trip and only on the empty path, so a search that
returned hits is unaffected.

Atlas also reports a wrong vector path and a dimension mismatch under the same
error code. Both previously surfaced as "index not found", which sent the
reader looking in the wrong place; they are now told apart and each names the
setting that is actually wrong.
2026-09-02 09:10:51 -07:00
Yuneng Jiang
800cd17d17
test(vector_stores): cover the MongoDB Atlas vector store config
65 cases across pipeline construction, response mapping, parameter validation,
client caching, and driver-error translation. The sad-path cases assert on the
message the caller actually sees, since a vector search that fails quietly
returns an empty result set rather than an error.
2026-09-02 09:04:54 -07:00
Yuneng Jiang
a616b8aaed
feat(vector_stores): add MongoDB Atlas vector store provider
Atlas Vector Search has no HTTP query API, since the Data API and HTTPS
Endpoints are end-of-life, so this provider extends BaseDirectVectorStoreConfig
and runs the $vectorSearch aggregation through pymongo rather than shaping an
httpx request. That is the same seam Valkey uses for RESP.

vector_store_id names the Atlas Search index, matching Valkey, with the
database and collection supplied through litellm_params.

pymongo lives in a new optional `mongodb` extra and is imported lazily, so the
base install still pulls no MongoDB driver. The floor is 4.17 because that is
where dnspython became a core dependency instead of the `srv` extra, and Atlas
issues mongodb+srv:// URIs that will not resolve without it.

Clients are cached per connection rather than opened per search. Measured
against Atlas, a fresh client costs ~890ms versus ~80ms warm, so copying the
Valkey open-and-close-per-call pattern would have added ~810ms to every query.
2026-09-02 09:02:36 -07:00
mateo-berri
69cd1bada6 test(gemini): compare gemini-3.8-flash to 3.7 flash field by field 2026-09-02 08:46:28 -07:00
devin-ai-integration[bot]
2ce4e3f8a9
fix(guardrails): run apply_guardrail-only providers in logging_only mode (#39297)
* fix(guardrails): run apply_guardrail-only providers in logging_only mode

A CustomGuardrail that implements only apply_guardrail inherited the CustomLogger
no-op async_logging_hook, so mode: logging_only never scanned anything and never
recorded guardrail_information. CustomGuardrail.async_logging_hook now routes the
logged request and response through the call type's guardrail translation on
copies and appends the verdict to standard_logging_object.guardrail_information.

Resolves LIT-4876

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep logging_only scan copies inside the error boundary and return a fresh logging payload

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): cover embedding scan, native-hook bypass, and unmapped call type in logging_only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 08:32:49 -07:00
mateo
da23e0241d fix(models): add cloudflare whisper transcription pricing and pin govcloud pricing tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 15:28:53 +00:00
Yujong Lee
07cf9dc46f feat(models): add Azure DeepSeek V4 Flash 0731 2026-09-02 08:10:52 -07:00
mateo
b761277740 fix(models): drop inherited retirement dates from azure/us-gov entries pending a Government schedule source
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 15:09:49 +00:00
mateo-berri
6b83b16559 feat(gemini): day-0 pricing for gemini-3.8-flash
Gemini 3.8 Flash launches today with the same promotional pricing, limits,
and thinking settings as Gemini 3.7 Flash, so the gemini/, vertex_ai/, and
bare cost map entries mirror the 3.7 Flash ones. Regression tests lock the
launch prices, the 4096-token cache minimum, and the gemini-3 thought
signature gate in for the new model.
2026-09-02 08:04:14 -07:00
Oliver Jensen
8588a2ea42
fix(docker): install saml extra in litellm-backend image (#39291)
The monolithic images install the saml extra but the split backend image
did not, so /sso/saml/* returned 501 on Helm split-image deployments.
The gateway image is unchanged since /sso/ routes are backend-only.
2026-09-02 08:01:54 -07:00
mateo
a7836ede15 fix(models): absorb open registry PRs: govcloud bedrock and mantle, azure gov, openai tiered long-context, scaleway, together qwen3.8, azure ai cache and kimi k2.7 code, azure mai deprecations
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 14:48:51 +00:00