open-webui/backend/open_webui/retrieval/vector/dbs
Classic298 fe56ab24f3
fix: page Chroma get() so hybrid search works on collections over 32k chunks (#30368)
With Chroma as the vector DB, hybrid search on a knowledge base with more than 32766 chunks fails with HTTP 400 "Error querying knowledge base". The legacy hybrid path fetches the whole collection to build the BM25 index, and Chroma's unbounded collection.get() binds one SQLite variable per row, so any collection above SQLite's 32766 variable limit raises "too many SQL variables" (reproduced on chromadb 1.5.9 with both PersistentClient and HttpClient). Vector-only search on the same collection works, which makes it look like a hybrid-search bug.

The Chroma adapter now reads the collection in pages of 10000 rows via limit/offset and concatenates them into the same GetResult shape as before.

Verified on a 90000-row collection: every row returned exactly once with documents and metadata aligned to ids, page order stable across page sizes, empty and exactly-one-page collections unchanged, and query_doc_with_hybrid_search returns results where it previously raised. The tests repo unit suite is identical before and after.

Fixes #30351
2026-09-22 14:48:45 -04:00
..
chroma.py fix: page Chroma get() so hybrid search works on collections over 32k chunks (#30368) 2026-09-22 14:48:45 -04:00
elasticsearch.py refac 2026-08-25 16:27:17 -04:00
mariadb_vector.py refac 2026-07-31 17:41:14 -04:00
milvus.py refac 2026-08-25 16:27:17 -04:00
milvus_multitenancy.py fix: apply the shared metadata size cap to the last four vector DB backends (#29502) 2026-09-04 11:42:20 -04:00
opengauss.py fix: release the openGauss connection when a read returns no rows (#30144) 2026-09-18 19:31:01 -04:00
opensearch.py refac 2026-08-25 16:27:17 -04:00
oracle23ai.py fix: apply the shared metadata size cap to the last four vector DB backends (#29502) 2026-09-04 11:42:20 -04:00
pgvector.py fix: pgvector reads leak their connection and lose most of their neighbours (#30142) 2026-09-18 19:31:49 -04:00
pinecone.py refac 2026-08-25 16:27:17 -04:00
qdrant.py fix: apply the shared metadata size cap to the last four vector DB backends (#29502) 2026-09-04 11:42:20 -04:00
qdrant_multitenancy.py fix: apply the shared metadata size cap to the last four vector DB backends (#29502) 2026-09-04 11:42:20 -04:00
s3vector.py refac 2026-08-25 16:27:17 -04:00
valkey.py perf: build info log messages lazily so raising the log level actually saves work (#27837) 2026-08-02 15:39:10 -05:00
weaviate.py refac 2026-08-25 16:27:17 -04:00