mongod serves $vectorSearch identically whether mongot runs under Atlas or beside
a self-managed deployment, so the provider already worked against on-prem. The
guidance did not: a refused connection told the operator to check their project's
IP access list and whether the cluster was paused, neither of which exists outside
Atlas, and the index errors claimed an "Atlas Vector Search index" they do not have.
Every message now names a remedy for both, keeping the Atlas-specific hint labelled
as such.
Also diagnoses unescaped credentials, which self-managed deployments hit more often
because the password is usually generated. pymongo reports those three different
ways and none of them mentions the password: '@', ':' and '%' raise an RFC 3986
complaint, '/' is read as the database separator and surfaces as Bad database name,
and an unescaped ':' looks like a bad port and comes back as a plain ValueError.
All three now point at the credentials. The ValueError branch's comment claimed it
fired on an unescaped '/', which pymongo actually reports as InvalidURI; corrected
to the port parse it really catches.
Verified against a self-managed mongod 8.0 with mongot, reached over plain
mongodb:// with no SRV and no TLS: 13 cases with live OpenAI embeddings, and 4
credential cases against an auth-enabled instance whose password holds % @ / and :.
list_search_indexes returns the same queryable and status fields there as on Atlas,
so the index-readiness check needed no change.
4.17 was picked on the belief that dnspython only became a core pymongo
dependency there, which is wrong: pymongo has declared dnspython>=1.16.0,<3.0.0
as a core requirement since well before that, so mongodb+srv:// URIs resolve at
4.9 too. The real floor is 4.9, the release AsyncMongoClient landed in, and 4.8
has no AsyncMongoClient at all.
Verified against live Atlas on 4.9: sync and async search, list_search_indexes,
same top hit and score as 4.17. Resolution is unchanged, pymongo 4.17.0 either
way, so this only widens what an existing environment is allowed to bring.
A connection string whose password holds an unescaped '/' makes pymongo's URI
parser raise a plain ValueError, not a PyMongoError, and a URI with no
credentials at all makes Atlas close the connection, which surfaces as
AutoReconnect. Neither was handled, so both fell through to litellm's generic
wrapper and were served as 500s with a traceback for what are routine typos.
Both now return a 400 naming the cause. The ConnectionFailure branch sits after
the ServerSelectionTimeoutError and NetworkTimeout branches, which subclass it,
and two ordering tests pin that.
filters was already refused, but ranking_options and rewrite_query were
accepted and then dropped. A caller asking for score_threshold 0.9 got results
scoring 0.5 with a 200 and no indication the threshold never ran, which is the
silent-wrong-answer case the filters check exists to prevent. Both now raise
the same 400 naming the parameter and what to do instead.
The repo's rule allows a comment only where the logic stays confusing after the
code has been made as clear as it can be, and then only one concise line about
why. Three multi-line blocks did not meet that: the reason "connection" joins
the sensitive patterns belongs in the commit that added it, and the weakref and
Atlas error-code notes each say what they need to in a single line.
The async client cache is keyed per event loop, and pymongo's AsyncMongoClient
holds a reference to the loop it was built on, so an entry for a closed loop
kept that client and its sockets alive for the life of the process. A script
that calls asyncio.run once per search fills the cache to its cap this way and
then stops caching entirely. Measured live against Atlas over 40 loops: 32
pinned clients and 212 open descriptors before, 1 cached client and no
monotonic descriptor growth after.
litellm.exception_type passes only litellm's own exception types through
untouched, so the NotImplementedError the search-only refusal raised reached
the caller as APIConnectionError. The proxy served that as a 500 with a
traceback in the body for what is a plain client mistake. Raising
BadRequestError gives the caller a 400 and the message on its own.
tests/test_litellm/llms/mongodb imports pymongo's exception classes to check the
error translation against the real hierarchy, and the shard that runs it
(tests/test_litellm/llms, per test-unit.yml) synced --extra google, proxy,
semantic-router and saml but not mongodb, so 24 of 109 tests would have errored
with ModuleNotFoundError on the first CI run. CircleCI hid this because it syncs
--all-groups --all-extras.
uv export --frozen ... --extra saml -> no pymongo
uv export --frozen ... --extra saml --extra mongodb -> pymongo==4.17.0
Also close the two gaps a mutation run found in the suite: nothing asserted that
a short request timeout shortens server selection as well as connect, and the
existing code 13 case carried "not authorized", which the message markers match
too, so it could not tell whether the code was still being checked. 28 of 28
mutants now die.
The type-discipline and test-quality gates blamed the branch for 4 LIT001, 12
LIT002 and 5 TQ008 violations. Rather than suppress them:
- the $vectorSearch and $project stages are MappingProxyType and the query
vector a tuple, verified against live Atlas to encode identically. The outer
pipeline stays a list because pymongo's common.validate_list raises
"pipeline must be a list, not <class 'tuple'>", which a unit test now pins.
- the client caches are Final[dict[...]] and _client_kwargs returns a
MappingProxyType.
- _field_value recurses over the dotted path instead of rebinding a local.
- _client_key declared Final locals in one branch and reassigned them in the
others, so it is split into an early-returning _timeout_ms.
- the injected callables carry explicit Final[Callable[...]] annotations, which
stops pyright resolving self.embedding_fn against litellm.embedding's
overloads.
- get_sync_client and get_async_client take an optional client_class, so the
cache tests inject a recording double instead of patching the importer, and
can assert the connection string and timeouts the client was built with.
SensitiveDataMasker is public SDK surface, so extra_sensitive_patterns moves to
the end of the signature: in slot two it silently reinterpreted an existing
caller's positional override set as extra sensitive patterns.
When the headroom_retrieve tool is exposed to a client that runs its own
tool-execution loop (the LiteLLM MCP gateway path), the client executes the
retrieve call and sends the recovered original content back as a tool result
on the next turn. The guardrail then compressed that row again, and because
CCR is content-addressed it collapsed back to the exact same hash it was just
retrieved from. The model never saw the expansion and the agent looped.
Hold tool-result rows that carry headroom_retrieve output back from the
compression service, the same way the live turn and trailing tool exchange are
already protected, so the expansion survives. Retrieve calls are matched by the
direct headroom_retrieve name and the mcp__<server>__headroom_retrieve gateway
name. Because a long gateway name is truncated past 64 chars in the
OpenAI-translated view the guardrail scans, the pairing also falls back to the
tool-call id read from the request's own untranslated messages, which is never
truncated.
Fixes#38558
The chain had grown to four nested ternaries with a fifth level inside the
Vertex Search branch, which no-nested-ternary had two suppressions for. A lookup
keyed by provider drops both suppressions and leaves one condition, the Vertex
Search case that depends on whether an engine id has been entered.
Also hoists the MongoDB form fixtures in the tests, which the inline-object
budget counts.
Building the client parses the URI and, for mongodb+srv://, performs a DNS SRV
lookup, so it fails on exactly the inputs a user is most likely to get wrong. It
sat outside the try that translates driver errors, so a malformed URI or an
unresolvable cluster escaped as a raw pymongo exception and reached the caller as
a 500 with a traceback in the body.
The three DNS-shaped failures are also told apart now: a lookup that ran out of
time is a Timeout, a cluster name that is not in DNS says so and points at the
URI Atlas shows under Connect Drivers, and anything else keeps the generic
"not a usable MongoDB connection string".
Verified live: a tampered scheme, a nonexistent cluster and a 1ms timeout each
come back as their own message instead of a traceback.
Atlas matches on the vector alone, so a mistyped mongodb_text_field still returns
confidently scored results whose content is empty, and the model is handed an
empty context with nothing to explain it. When every matched document lacks the
field the search now says which setting to fix; a sparse document among others
that do have it, and a document whose text is genuinely the empty string, both
still come back normally.
Unrecognised mongodb_* parameters are named too. The params model has to ignore
unrelated keys because litellm_params carries plenty of them, which turned a
mistyped mongodb_collection into "mongodb_collection is required" pointing the
reader at a key they can see they have set.
GHSA-jp53-mhqp-8xcg (fixed in 6.16.0), GHSA-23w6-3w8w-8484 and
GHSA-763m-79hh-57f2 (fixed in 6.16.1) flag pypdf 6.15.0 in uv.lock and
keep osv-scan red alongside the tornado advisories. The proxy-runtime
extra now requires pypdf>=6.16.1 and the lock resolves 6.16.2.
The async client cache was keyed on id(loop). CPython recycles those ids so
aggressively that a fresh event loop nearly always lands on the id of one already
collected, measured at 37 of 40 rounds, so the cache handed the new loop an
AsyncMongoClient bound to a closed loop and every operation on it raised
"Event loop is closed".
The entry now carries a weak reference to the loop it was built on and a hit only
counts when that reference still points at the running loop, so a recycled id
misses and builds a fresh client. A stale entry can also be replaced once the
cache is full, which the old size check prevented.
pymongo's own client keeps its loop alive, which is why the sync proxy path never
saw this; a script calling asyncio.run() per search, or a test suite with a loop
per test, does.
The create form now offers MongoDB Atlas with its connection string, database,
collection, embedding model, vector field, text field and candidate count. The
connection string renders as a password input because it carries the database
user's password, and the embedding model is picked from the proxy's own models,
matching how Milvus and Valkey do it.
The vector store id doubles as the Atlas Vector Search index name, so the
placeholder says so.
* fix(helm): scale the classic chart's HPA out at the documented 60 percent CPU
The litellm-helm chart shipped targetCPUUtilizationPercentage: 80, which is
unexamined helm create scaffold rather than a chosen number. It arrived packaged
with the stock minReplicas: 1, maxReplicas: 100, a commented-out
targetMemoryUtilizationPercentage: 80, and the boilerplate "such as Minikube"
comment, the same provenance as the 128Mi resource example this file just
corrected.
60 is the documented recommendation. The mechanism behind it is scale-up lag:
the chart's own startupProbe is failureThreshold: 30 times periodSeconds: 10, so
a replica can take up to 300 seconds to become ready, and a pod added at 80
percent utilization arrives minutes after saturation.
The memory target stays commented out on purpose. The prisma query engine's
resident memory is a high-water mark that ratchets to the pod's worst-ever write
and is never returned, so a memory-target HPA reads the largest write a pod ever
did rather than what it is doing now, and replicas ratchet up without scaling
back in.
hpa_tests.yaml carried its second suite after a YAML document separator, and
helm-unittest loads only the first document per file, so that suite never ran;
an assertion planted in it still passed. Fold it into the one live suite and add
coverage pinning the rendered CPU target, the absence of a memory metric by
default, and that overrides still take effect.
Bump the chart to 1.1.2, since rendered output changes for anyone running with
autoscaling enabled.
* fix(helm): bump litellm-helm to 1.1.3 after rebase onto 1.1.2
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Atlas answers a wrong password with code 8000 "AtlasError" rather than the 18 a
self-hosted deployment returns, so the code-only check never fired and a bad
password came back as a generic "MongoDB rejected the vector search", pointing
the reader at the index instead of at their credentials. Verified live against
Atlas with a tampered password.
litellm.exception_type passes a litellm exception through untouched and wraps
anything else into APIConnectionError, so every bare ValueError this provider
raised reached the caller as HTTP 500 with a Python traceback in the response
body. "max_num_results must be between 1 and 50" is the caller's to fix, not a
connection failure.
Configuration and validation failures now raise BadRequestError (400) and the
two timeout cases raise Timeout (408). ExecutionTimeout subclasses
OperationFailure, so it is matched before it; previously an Atlas query that ran
out of time was reported as "MongoDB rejected the vector search".
A MongoDB vector store's whole credential is its connection string, and
mongodb+srv://<user>:<password>@<cluster> embeds the database password. None of
the masker's default patterns (api_key, secret, token, credential) match a key
named mongodb_connection_string, so /vector_store/list and /vector_store/info
returned it verbatim to every caller that can read a vector store.
SensitiveDataMasker gains extra_sensitive_patterns, which unions onto the
defaults instead of replacing them, and the vector-store redactor adds
"connection" so the URI is masked while mongodb_database, mongodb_collection and
the field names stay readable.
Driving the sad path against a live Atlas cluster showed four cases returning
an empty result set instead of failing: a missing index, a missing database, a
missing collection, and the async path for all three. $vectorSearch reports
none of these as errors, so a misconfigured store looked exactly like a query
that matched nothing, which is the worst shape for this to fail in.
An empty result set is now checked against the index catalogue, which does
report all three correctly, and a store that cannot work says so. The check
costs one extra round trip and only on the empty path, so a search that
returned hits is unaffected.
Atlas also reports a wrong vector path and a dimension mismatch under the same
error code. Both previously surfaced as "index not found", which sent the
reader looking in the wrong place; they are now told apart and each names the
setting that is actually wrong.
65 cases across pipeline construction, response mapping, parameter validation,
client caching, and driver-error translation. The sad-path cases assert on the
message the caller actually sees, since a vector search that fails quietly
returns an empty result set rather than an error.
Atlas Vector Search has no HTTP query API, since the Data API and HTTPS
Endpoints are end-of-life, so this provider extends BaseDirectVectorStoreConfig
and runs the $vectorSearch aggregation through pymongo rather than shaping an
httpx request. That is the same seam Valkey uses for RESP.
vector_store_id names the Atlas Search index, matching Valkey, with the
database and collection supplied through litellm_params.
pymongo lives in a new optional `mongodb` extra and is imported lazily, so the
base install still pulls no MongoDB driver. The floor is 4.17 because that is
where dnspython became a core dependency instead of the `srv` extra, and Atlas
issues mongodb+srv:// URIs that will not resolve without it.
Clients are cached per connection rather than opened per search. Measured
against Atlas, a fresh client costs ~890ms versus ~80ms warm, so copying the
Valkey open-and-close-per-call pattern would have added ~810ms to every query.
* fix(guardrails): run apply_guardrail-only providers in logging_only mode
A CustomGuardrail that implements only apply_guardrail inherited the CustomLogger
no-op async_logging_hook, so mode: logging_only never scanned anything and never
recorded guardrail_information. CustomGuardrail.async_logging_hook now routes the
logged request and response through the call type's guardrail translation on
copies and appends the verdict to standard_logging_object.guardrail_information.
Resolves LIT-4876
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): keep logging_only scan copies inside the error boundary and return a fresh logging payload
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): cover embedding scan, native-hook bypass, and unmapped call type in logging_only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Gemini 3.8 Flash launches today with the same promotional pricing, limits,
and thinking settings as Gemini 3.7 Flash, so the gemini/, vertex_ai/, and
bare cost map entries mirror the 3.7 Flash ones. Regression tests lock the
launch prices, the 4096-token cache minimum, and the gemini-3 thought
signature gate in for the new model.