open-webui/backend/open_webui/retrieval
Classic298 fe56ab24f3
fix: page Chroma get() so hybrid search works on collections over 32k chunks (#30368)
With Chroma as the vector DB, hybrid search on a knowledge base with more than 32766 chunks fails with HTTP 400 "Error querying knowledge base". The legacy hybrid path fetches the whole collection to build the BM25 index, and Chroma's unbounded collection.get() binds one SQLite variable per row, so any collection above SQLite's 32766 variable limit raises "too many SQL variables" (reproduced on chromadb 1.5.9 with both PersistentClient and HttpClient). Vector-only search on the same collection works, which makes it look like a hybrid-search bug.

The Chroma adapter now reads the collection in pages of 10000 rows via limit/offset and concatenates them into the same GetResult shape as before.

Verified on a 90000-row collection: every row returned exactly once with documents and metadata aligned to ids, page order stable across page sizes, empty and exactly-one-page collections unchanged, and query_doc_with_hybrid_search returns results where it previously raised. The tests repo unit suite is identical before and after.

Fixes #30351
2026-09-22 14:48:45 -04:00
..
loaders fix: send only the file name to Docling instead of the full storage path (#30357) 2026-09-22 11:35:05 -04:00
models fix: repair two broken logging calls, one of which makes VECTOR_DB=opengauss unusable (#27838) 2026-08-10 22:19:41 -06:00
vector fix: page Chroma get() so hybrid search works on collections over 32k chunks (#30368) 2026-09-22 14:48:45 -04:00
web fix: surface searchapi errors, news results and redirect links (#30308) 2026-09-21 10:44:06 -04:00
external.py refac 2026-07-27 01:59:17 -04:00
utils.py feat: let operators expose chosen file metadata to the model in retrieved sources (#29696) 2026-09-19 17:02:59 -05:00