fix: pass web search results to model when embedding & retrieval enabled

Web search results stored as a vector collection were silently dropped
before retrieval when BYPASS_RETRIEVAL_ACCESS_CONTROL is False (the
default). The server-generated web_search file item carries a
'collection_name' but its 'web_search' type is not matched by any
explicit dispatch branch in get_sources_from_items, so it fell through
to the untrusted client-supplied collection_name branch and was ignored.

Add an explicit branch for type == 'web_search' items so the collection
is queried again. Access control is preserved: the collection still
passes through filter_accessible_collections, which already allowlists
web-search-* and bypasses only for admins.

Regression introduced when the retrieval access-control hardening gated
the bare collection_name fallback behind BYPASS_RETRIEVAL_ACCESS_CONTROL.
This commit is contained in:
Classic298 2026-06-02 11:26:52 +02:00 • committed by GitHub
parent 1a97751e37
commit 46d75bf040
No known key found for this signature in database
GPG key ID: B5690EEEBB952194

View file

@ -1374,6 +1374,10 @@ async def get_sources_from_items(
'documents': [[doc.get('content') for doc in item.get('docs')]],
'metadatas': [[doc.get('metadata') for doc in item.get('docs')]],
}
elif item.get('type') == 'web_search' and item.get('collection_name'):
# Trusted server-generated collection; authorized by
# filter_accessible_collections below (allowlists web-search-*).
collection_names.append(item['collection_name'])
elif item.get('collection_name'):
if BYPASS_RETRIEVAL_ACCESS_CONTROL:
collection_names.append(item['collection_name'])