open-webui/backend/open_webui/retrieval/vector
Classic298 fe56ab24f3
fix: page Chroma get() so hybrid search works on collections over 32k chunks (#30368)
With Chroma as the vector DB, hybrid search on a knowledge base with more than 32766 chunks fails with HTTP 400 "Error querying knowledge base". The legacy hybrid path fetches the whole collection to build the BM25 index, and Chroma's unbounded collection.get() binds one SQLite variable per row, so any collection above SQLite's 32766 variable limit raises "too many SQL variables" (reproduced on chromadb 1.5.9 with both PersistentClient and HttpClient). Vector-only search on the same collection works, which makes it look like a hybrid-search bug.

The Chroma adapter now reads the collection in pages of 10000 rows via limit/offset and concatenates them into the same GetResult shape as before.

Verified on a 90000-row collection: every row returned exactly once with documents and metadata aligned to ids, page order stable across page sizes, empty and exactly-one-page collections unchanged, and query_doc_with_hybrid_search returns results where it previously raised. The tests repo unit suite is identical before and after.

Fixes #30351
2026-09-22 14:48:45 -04:00
..
dbs fix: page Chroma get() so hybrid search works on collections over 32k chunks (#30368) 2026-09-22 14:48:45 -04:00
async_client.py refac 2026-09-06 16:48:30 -04:00
factory.py refac 2026-09-06 17:13:32 -04:00
main.py refac 2026-06-22 16:10:19 +02:00
type.py feat: add support for Valkey vector database (#24769) 2026-06-01 12:20:01 -07:00
utils.py refac 2026-08-25 16:27:17 -04:00