Commit graph

724 commits

Author SHA1 Message Date
Timothy Jaeryang Baek
09dcb60887 refac 2026-10-09 21:42:26 +04:00
Classic298
c04c7ddae0
fix: reranker gets no user info headers in chat and knowledge search (#32061)
With ENABLE_FORWARD_USER_INFO_HEADERS on and hybrid search enabled, rerank requests went out without the user headers when a chat searched attached knowledge, when a model used the built-in knowledge search tool, and when the collection query API was called with hybrid turned off for that one request. The embedding requests of the same search did carry them, so external rerankers that use these headers for per-user auth, rate limits or auditing saw anonymous calls. Rerank requests now carry the signed-in user, the same as embedding requests.

Fixes #32060
2026-10-08 16:55:45 +04:00
Classic298
f6cbeb1a1c
fix: files added to a knowledge base on Qdrant can end up with only part of their content (#31961)
On Qdrant, uploading a file into a knowledge base or editing a file's content could leave the knowledge base with only part of the file, often exactly 64 chunks, while the file showed as completed and nothing was logged. The file's chunks were saved without waiting for Qdrant to make them searchable, and the knowledge base copied them straight after, so it only got the ones Qdrant had finished storing. Saving now waits until Qdrant has finished storing the chunks, so the knowledge base always gets the whole file. On a busy Qdrant this makes file processing somewhat slower, since each save now waits for the server.

Fixes #31959
2026-10-07 16:29:13 +04:00
Classic298
77febc65d6
fix: external knowledge citations show the domain instead of the page title (#31982)
The inline citation markers in answers from external knowledge bases (Qdrant, Milvus, pgvector) showed the site's domain, for example docs.example.test, instead of the page title saved with each document. They now show that title, the same way a regular knowledge base shows the file name, and still fall back to the domain when a document has no title. Two different pages can share a title, so each marker now looks up its title by the page it points to; before, a repeated title made every later marker show the next page's title and the last one show nothing.

Fixes https://github.com/open-webui/open-webui/issues/31929
2026-10-07 15:05:34 +04:00
Classic298
30dbf3256a
fix: large files fail to upload when Milvus is the vector database (#31990)
With Milvus as the vector database, all chunks of a file were sent to Milvus in one request. Large files exceeded Milvus's default 64 MB request size limit, so processing ran through the whole embedding step and then failed with the error RESOURCE_EXHAUSTED. With a 4096-dimension embedding model this already happened at about 4,000 chunks, which is a few MB of text. Chunks are now sent in batches of 128, with and without ENABLE_MILVUS_MULTITENANCY_MODE.

Fixes #31989
2026-10-07 15:05:25 +04:00
Classic298
d8177c4177
fix: send the query embedding as a vector in external pgvector retrieval (#31112)
External knowledge bases on the pgvector provider failed on every search with "operator does not exist: vector <=> double precision[]", so they looked empty to users. This happened regardless of the VECTOR_DB setting.

The query embedding was bound as a plain Python list. register_vector only adapts pgvector's own Vector type and numpy arrays, so psycopg sent the list as a float array, which the <=> operator does not accept. Wrapping the embedding in pgvector.Vector sends it as a real vector.

Vector is imported from the package root, which works on the pinned pgvector 0.4.2 and on 0.5.x, where the pgvector.psycopg re-export no longer exists.

Verified against a pgvector Postgres: before the fix the reported error reproduces; after it, results come back ranked by cosine distance and filtered to the collection, including schema-qualified tables, halfvec columns and 1536-dimension embeddings.

Fixes #26663
2026-10-01 07:51:32 +04:00
Classic298
7cf35bedc0
fix: editing a knowledge file on Elasticsearch keeps its old text searchable (#31530)
With VECTOR_DB=elasticsearch, changing a file's content in a knowledge base added the new text but never removed the old chunks, so searches and chats kept returning the old text next to the new. The old chunks are now found and removed after the edit, the same as on the default store.

Fixes #31523
2026-10-01 07:36:44 +04:00
Classic298
c3fbf36384
fix: uploaded files lose Chinese punctuation and quotes (#31655)
Text extracted from any uploaded file had full-width punctuation like :(),!? turned into ASCII :(),!? and curly quotes like “ ” turned into straight quotes, so both the file preview and the model saw altered text. Extracted text now keeps these characters as written, while garbled text from wrong encodings (like café becoming café) is still repaired. Files uploaded before this change keep the altered text until they are uploaded again.

Fixes #17087
2026-10-01 07:22:08 +04:00
Classic298
aee7c47277
feat: Milvus hybrid search without a second text-only copy of every collection (#31660)
Native hybrid search on Milvus kept a second, text-only collection beside every Milvus collection and searched both. Each Milvus collection now holds its vectors, its text and its BM25 keyword index together, and results rank exactly as before, with the BM25 weight setting working the same way it does on pgvector. Existing data moves on the first start with ENABLE_DB_MIGRATIONS on: startup waits while every collection is copied once (vectors included, nothing is re-embedded) and the originals are only dropped after every copy succeeded, so a failed run changes nothing and is retried on the next start. The copy needs free disk space for a second copy of the data until it finishes, and on one 16-thread machine with Milvus's official docker compose setup it ran at 11 to 22 MB/s, so 200 GB takes about 2.5 to 5 hours depending on chunk size. Milvus servers older than 2.5 are detected and keep the existing hybrid search.
2026-10-01 03:07:50 +04:00
Classic298
4ef7e35b88
fix: Playwright web loader returns only the site menu for pages with more than one <main> element (#31644)
Some pages, like the Ubiquiti tech specs pages, have more than one <main> element. The Playwright web loader only read the first one, so these pages came back as just their site menu and the actual content was lost. When a page has more than one, the loader now ignores those tags and reads the whole page. Pages with a single <main> load the same as before.

Fixes #28643
2026-09-30 21:06:01 +04:00
Classic298
75bff4bcd9
feat: native hybrid search for Milvus and Milvus multitenancy (#31645)
With hybrid search on, Milvus installs fetch every chunk of a collection and score BM25 in Python for each search. In both Milvus modes, one collection per knowledge base and multitenancy, Milvus now runs the keyword half itself with its built-in BM25 full-text search and merges it with the vector results, as pgvector already does. Milvus cannot add a BM25 index to an existing collection, so each collection gets a second, text-only collection next to it; existing installs build these once during startup when ENABLE_DB_MIGRATIONS is on, which copies the chunk text (extra storage roughly the size of that text) and leaves the original vectors and indexes untouched. On one standalone Milvus server the copy ran at about 11,000 chunks per second, around 40 minutes for 200 GB with 1536-dimension embeddings. Collections that cannot be copied, and Milvus servers older than 2.5, which have no BM25, keep using the existing hybrid search.

Fixes #26243
2026-09-30 21:03:47 +04:00
Classic298
bee06b08ba
fix: jina-colbert-v2 reranker fails to load and turns hybrid search off (#31532)
Choosing jinaai/jina-colbert-v2 as the reranking model failed on the current transformers release with "'HF_ColBERT' object has no attribute 'all_tied_weights_keys'", and saving the Documents settings quietly switched hybrid search back off. The ColBERT reranker now finishes loading, reranks search results and hybrid search stays on after saving.

Fixes #31522
2026-09-30 19:34:20 +04:00
Classic298
710b9f1e2c
fix: exact matches score as the worst result on Weaviate (#31531)
With Weaviate as the vector database, a chunk identical to the query (distance 0) was treated as having no distance and got a relevance score of 0. Perfect matches could land at the bottom of the results or fall below the relevance threshold. They now score 1 as expected.

Fixes #31527
2026-09-30 19:29:59 +04:00
Classic298
fc9ad75164
fix: admins can still read and change other users' chats with ENABLE_ADMIN_CHAT_ACCESS off (#31416)
With ENABLE_ADMIN_CHAT_ACCESS turned off, opening another user's chat was refused, but through direct API requests an admin could still get the whole chat back in the reply to editing or deleting one of its messages, grant themselves read access in the chat's share settings, clone a chat someone shared privately with another user, or delete the chat. They could also send messages into it, attach it as context to their own chat, approve its tool calls, and list or stop its running replies. All of these are now refused for an admin on another user's chat, the same as opening it. With the setting on, admins keep full access as before.

Fixes #31413
2026-09-28 01:51:59 +04:00
Classic298
6e5e5fe6e9
feat: add a Tavily search depth setting (#31308)
Tavily web search always ran at Tavily's default depth (basic), because the search request never sent `search_depth`. The only Tavily depth control in Admin > Settings > Web Search, "Tavily Extract Depth", applies to the Extract API used by the web loader, never to search.

This adds `TAVILY_SEARCH_DEPTH` (env var and persisted setting, default `basic`) and a "Tavily Search Depth" select (ultra-fast, fast, basic, advanced) under the Tavily search engine settings. The value is sent as `search_depth` on every Tavily search request, so admins can set search and extract depth independently, for example fast search with advanced extraction.

The default matches Tavily's own default, so existing setups keep the same behaviour until the setting is changed.

Fixes #29891
2026-09-27 23:15:09 +04:00
Classic298
59ea3b7c2c
fix: backslashes in uploaded HTML files turn into line breaks or break the upload (#31450)
With the default content extraction engine, backslashes in an uploaded .html or .htm file were read as escape sequences. A path like C:\new\table was saved with a line break and a tab in it, and a page containing C:\Users failed to upload with a 'unicodeescape' codec error. HTML files are now read the same way as .txt and .md uploads, so the saved text matches the page.

Fixes #31440
2026-09-27 23:08:14 +04:00
Classic298
eda8d85361
fix: hybrid search finds nothing when a collection cannot be read (#31460)
With hybrid search on and a vector database without built-in hybrid search, a collection that failed to load (for example Qdrant strict mode rejecting the request) was skipped quietly, so retrieval returned no documents and never fell back to normal vector search. A failed load now counts as a failed collection, so when every collection fails retrieval falls back to vector search, the same way it already does when the search itself fails. The retrieval API returns its usual error in that case.

Part of #31459
2026-09-27 22:59:02 +04:00
Classic298
8a90f0fc93
fix: file uploads and hybrid search fail on Qdrant with strict mode enabled (#31461)
With Qdrant strict mode on and a max_query_limit below 999999999, Qdrant rejects Open WebUI's reads with "Limit exceeded", so every file upload after the first fails in the default multitenancy mode, and hybrid search finds nothing. Reads now go in pages of 1000 points, so any strict-mode limit of 1000 or more works. Without strict mode the results are the same as before.

Tested against Qdrant 1.19.1 with max_query_limit 1000, for both multitenancy on and off: collections of up to 2500 points come back complete, limits are respected, and tenants stay separated.

Fixes #31459
2026-09-27 22:58:42 +04:00
Classic298
3c47f0d7e7
fix: stop decoding Korean and Japanese text uploads as Chinese (#31356)
Korean EUC-KR and Japanese Shift-JIS text files were stored as garbled Chinese characters, so retrieval, knowledge bases and the model context all worked on text that is not in the file.

Encoding detection puts chardet's guess in front of a fixed GB18030, Big5, EUC-KR, EUC-JP try order, and GB18030 decodes almost any double-byte text without an error. The guess map was written for chardet 5. Since the bump to chardet 7 in v0.10.0, Korean text is reported as CP949 and Japanese text as cp932 or SHIFT_JIS, which the map either did not know or dropped because the codec was not in the try order, so these files fell through to GB18030.

The map now covers CP949 and cp932, and a mapped guess is always tried first. SHIFT_JIS maps to cp932, the Windows superset, because chardet also reports SHIFT_JIS for ordinary Japanese files containing characters such as ① or ㈱ that plain Shift-JIS cannot decode; this is the same subset-to-superset rule the map already applies to GB2312.

Korean and Japanese files now decode correctly, and Chinese, EUC-JP, UTF-8 and Western files decode as before. The one trade-off of trusting the guess: chardet 7 labels some files holding only a few Chinese characters (a short label or a one-line comment) as CP949, and those now read as Korean. No regressions were found in files with more Chinese text than that.

Fixes #31352
2026-09-25 00:01:38 -04:00
Classic298
fe56ab24f3
fix: page Chroma get() so hybrid search works on collections over 32k chunks (#30368)
With Chroma as the vector DB, hybrid search on a knowledge base with more than 32766 chunks fails with HTTP 400 "Error querying knowledge base". The legacy hybrid path fetches the whole collection to build the BM25 index, and Chroma's unbounded collection.get() binds one SQLite variable per row, so any collection above SQLite's 32766 variable limit raises "too many SQL variables" (reproduced on chromadb 1.5.9 with both PersistentClient and HttpClient). Vector-only search on the same collection works, which makes it look like a hybrid-search bug.

The Chroma adapter now reads the collection in pages of 10000 rows via limit/offset and concatenates them into the same GetResult shape as before.

Verified on a 90000-row collection: every row returned exactly once with documents and metadata aligned to ids, page order stable across page sizes, empty and exactly-one-page collections unchanged, and query_doc_with_hybrid_search returns results where it previously raised. The tests repo unit suite is identical before and after.

Fixes #30351
2026-09-22 14:48:45 -04:00
Classic298
f956f7adc0
fix: send only the file name to Docling instead of the full storage path (#30357)
With the Docling content extraction engine, every conversion request
carried the server's full internal upload path (for example
/app/backend/data/uploads/<id>_report.pdf) as the multipart file name.
Docling only needs a bare file name, and a hosted Docling instance has
no business learning where Open WebUI keeps its files on disk.

The loader now sends the base name of the stored file, which is what
the MinerU, Datalab and Mistral loaders already do. Nothing else in
the request or the parsed result changes, verified against a capturing
mock server before and after.

Fixes #30352
2026-09-22 11:35:05 -04:00
Classic298
94eea41a75
fix: surface searchapi errors, news results and redirect links (#30308)
Web search via searchapi.io could come back empty or near-empty with no
hint of why: an invalid or expired API key turned into an empty result
set instead of an error, the google_news engine splits its results
between organic_results and top_stories and only the first block was
read, and google links came back as google.com/goto redirects the web
loader cannot fetch, so citations pointed at a redirect blob.

The search now reads both result blocks, asks google engines for
resolved destination links, raises on HTTP errors, carries a 30s request
timeout, skips result rows without a link, and logs the response body at
debug instead of dumping every search at info.

Fixes #30305
2026-09-21 10:44:06 -04:00
Classic298
9012bd153d
fix: stop sending count to the Staan search API (#30303)
Staan rejects any count other than its fixed page size of 10, so every search 400ed with the default result count of 3. The API always returns 10 results; the local results[:count] slice already applies the configured count, so the request parameter is simply dropped.
2026-09-21 08:04:26 -04:00
Classic298
7fa705f3b8
feat: let operators expose chosen file metadata to the model in retrieved sources (#29696)
Custom metadata attached to a file upload now reaches the vector DB, but the model still never sees it. Both prompt-assembly paths build their output from a fixed field set: the classic RAG <source> tag carries only id, name and resource type, and the retrieval tools return only content, source and file id per chunk. A scraper that records where each document came from therefore cannot get that origin in front of the model, so answers cannot state it.

RAG_SOURCE_METADATA_KEYS names the chunk metadata keys allowed through to the model. Configured keys are emitted as extra attributes on the <source> tag and as extra fields on tool result chunks, covering both retrieval paths. It is empty by default, so nothing changes for existing deployments.

An allowlist instead of passing everything through, because chunk metadata also carries file hashes, collection names, embedding config and relevance scores, which would then be added to every retrieved chunk of every request. Values are attacker-controllable through an uploaded file, so they are escaped before they go into the tag, and a configured key can never displace a field the tag or the chunk already defines.

Reported in open-webui/open-webui#29486.
2026-09-19 17:02:59 -05:00
Classic298
3d29548716
fix: pgvector reads leak their connection and lose most of their neighbours (#30142)
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
Two defects on the pgvector read path. A search, query or get that finds nothing returns before the rollback that ends its read-only transaction, so the session keeps the connection it checked out; retrieval fans out over worker threads and the session is thread-local, so every thread that runs an empty read holds a connection for the lifetime of that thread. Users see vector search die after a while with "QueuePool limit of size 5 overflow 10 reached" and no way back other than a restart. Rolling back before the three early returns puts the connection back: measured on PostgreSQL 16, twelve empty reads on a default-shaped pool left all 12 connections checked out before and 0 after.

All collections also share one table under one vector index, and the WHERE collection_name filter is applied after the index walk, so a knowledge base holding a small share of the rows keeps only a small share of its neighbours, silently. pgvector 0.8 added iterative scans for this: the scan keeps going until enough rows pass the filter. This sets it per search so it cannot leak into other sessions, and only when the installed extension supports it, since setting it on pgvector 0.7 would make every search raise. PGVECTOR_ITERATIVE_SCAN turns it off or picks strict_order; it defaults to relaxed_order because the shipped behaviour is silently wrong results, and the cost is about a millisecond per search.

Measured through PgvectorClient.search on PostgreSQL 17 and pgvector 0.8, 101500 rows of 384 dimensions with the knowledge base at 1.5% of the table, hnsw m=16, recall@10 against an exact scan over 40 queries: 0.070 at 4.9 ms before, 0.970 at 7.2 ms after.

Fixes #30133
Fixes #30135
2026-09-18 19:31:49 -04:00
Classic298
fffcb727d2
fix: release the openGauss connection when a read returns no rows (#30144)
An openGauss query or get that finds nothing returns before the rollback that ends its read-only transaction, so the session keeps the connection it checked out of the pool. Retrieval fans out over worker threads and the session is thread-local, so every thread that runs an empty read holds a connection for the lifetime of that thread, and vector search eventually fails with a pool checkout timeout that a restart is the only way out of. Empty reads are routine: an empty knowledge base, a file whose chunks were deleted, or a metadata filter that matches nothing all produce one.

Rolling back before the two early returns puts the connection back. Measured by driving OpenGaussClient itself with a pool of 5: three empty query reads left 3 connections checked out and 0 free before the change, and 0 checked out with 3 free after, same for get. search is already correct, it has no early return.

The pgvector client had the same defect, fixed separately in #30142.
2026-09-18 19:31:01 -04:00
Classic298
9ee810ad09
feat: add Staan as a native web search provider (#30138)
European deployments that need EU data residency have no hosted web search option out of the box: every plug-and-play provider shipped today is US-based, and the only sovereign alternative is self-hosting SearXNG, which means running and maintaining that infrastructure yourself. Staan (staan.ai) is a European search API with EU data residency, so adding it gives those deployments a drop-in choice.

It is configured like any other provider, through the admin UI or STAAN_API_KEY, STAAN_MARKET and STAAN_MAX_SNIPPETS. Market sets the region and language of the results and defaults to en-us. Max snippets asks Staan to fetch each result page and return semantically scored chunks of it, which get merged into that result's snippet so retrieval has more to work with; leaving it at 0 uses the plain search endpoint.

Wired the same way as Tavily and Exa, with domain filtering going through the shared get_filtered_results.

Requested in #26006.
2026-09-18 10:38:36 -05:00
Classic298
95406fd28d
fix: build the pgvector ivfflat index once there are rows to cluster on (#30143)
ivfflat places its centroids by clustering the rows it can see when the index is built, so an index built on an empty table gets centroids that mean nothing, and everything inserted afterwards is filed against them. On a fresh install the index is created immediately after the table, before a single chunk exists, and it is never rebuilt, so that install keeps a permanently untrained index and quietly retrieves the wrong chunks. An install that upgraded into the version introducing the index is unaffected, its table already had rows.

The index is now created once the table holds 50 rows per list, the sample size ivfflat itself aims for. Below that, and until the next start, searches fall back to an exact scan, which is correct and costs about a millisecond at that size. hnsw is untouched, it builds its graph as rows are inserted and has nothing to train on. One trade-off: the build moves from the first start to that later one, so an instance that has grown large in between pays a one-time index build during startup.

Measured on PostgreSQL 17 and pgvector 0.8 with the default lists=100 and probes=1, 384 dimensions, recall@10 against an exact scan, varying only how many rows existed when the index was built:

| rows at build | 200 | 1000 | 2500 | 5000 | 20000 |
|---|---|---|---|---|---|
| recall@10 | 0.180 | 0.563 | 0.967 | 1.000 | 1.000 |

Through PgvectorClient.search on a 20000-row table, an index built as it is today scores 0.480 against 1.000 built after the rows arrive, at the same 5 ms. An existing install can repair its index with REINDEX INDEX idx_document_chunk_vector, measured to take it from 0.480 back to 1.000.

Fixes #30134
2026-09-18 10:37:25 -05:00
Classic298
9923c53c10
fix: treat SVG uploads as documents instead of vision images (#30102)
Uploading an .svg to a chat attached it as a vision image input, so the model received a data URI it could not decode. PIL-backed servers answered "cannot identify image file" and OpenAI answered "The image data you provided does not represent a valid image". No setting made it work.

SVG now takes the ordinary file upload path, so its XML source is extracted and indexed and the model can answer questions about it. Rasterizing was the alternative and it would have discarded the part of an SVG a model reads best, the source itself. Raster formats are untouched and still go up as image inputs.

A shared helper replaces the ad hoc image/ prefix checks at the points that decide image input versus document, on both ends. It normalises the content type first, because a stored "image/SVG+xml" or a trailing charset parameter slipped past a plain comparison.

One behaviour change worth knowing: an SVG now needs the model to have the file upload capability, where before it rode in as an image.

Fixes #30100
2026-09-17 17:40:07 -04:00
Classic298
0837f310be
fix: reject Docling conversions that failed inside an HTTP 200 response (#30107)
Uploading a file that Docling declines or fails to convert either dies with `TypeError: argument of type 'NoneType' is not iterable`, or silently succeeds and stores the literal string `<No text content found>` as the document's text, which then gets indexed and handed to the model as if it were the file. Docling returns the conversion outcome inside the HTTP 200 body, so checking only the HTTP status made a refused conversion look identical to a successful one, and the `errors` array that says why in plain words was never read.

Failed and skipped conversions are now rejected with the messages Docling returned, so an unsupported format surfaces as "File format not allowed: example.dxf" and the traceback is gone. The markdown field is also read as nullable, because Docling returns JSON `null` for every content format it was not asked to produce, which any Docling Parameters setting `to_formats` without `md` will hit, and that null was what raised the TypeError.

Successful conversions with empty markdown keep the existing `<No text content found>` placeholder, matching what TikaLoader and the Mistral loader already do in the same package.

Fixes #29808
2026-09-17 17:39:35 -04:00
Classic298
ff7f35a30d
fix: stop logging a traceback twice per failed vector search (#29981)
A vector DB outage renders a full traceback twice for every collection and query pair. query_doc and query_doc_with_hybrid_search each log and then re-raise, and the fan-out handler that catches them logs the same exception again, so twenty knowledge bases and five expanded queries turn one outage into hundreds of identical stack traces per message on every replica.

Both helpers lose the try/except that only logged before re-raising. Each fan-out keeps reporting failures through its return value and now carries the collection name back, so one record after the gather names every collection that failed and attaches a single traceback. The hybrid path logs before its existing raise, so a total failure is still reported before the caller falls back to vector search.

Measured with every collection failing, the vector fan-out drops from 1000 traceback records for fifty collections across ten queries to one, and the hybrid fan-out from 100 for ten collections across five queries to one. Results are byte-identical across eighteen and twenty-six scenarios covering healthy, failing, partial and degenerate inputs.

Three limits worth stating. Nine of the fifteen vector clients swallow the error inside search() and return None, chroma and pgvector among them, so those backends see neither the old tracebacks nor the new record. get_sources_from_items queries one attached item at a time, so a chat with twenty knowledge bases still logs twenty records. And with hybrid search on a total outage logs once from each fan-out as the caller falls back, down from two per pair plus one.

The old "All collection queries failed" warning goes with this. The replacement cannot express that case, because a collection can fail one query and succeed another, so naming every failed collection no longer implies an empty payload.
2026-09-16 15:59:05 -04:00
Classic298
06d9d2e7c7
fix: sync Playwright loader spins a CPU core when a page opens a WebSocket (#30050)
Fetching a single URL through the Playwright loader (fetch_url tool, a URL attached to a
chat, the process/web endpoint) never returned when the page opened a WebSocket. The worker
thread stayed at 100% CPU for the life of the process, and every further hit cost another
core, so the whole instance got slow. Web search was unaffected, it uses the async loader.

The sync loader's websocket route handler called the synchronous close(). Playwright runs
websocket route handlers directly on its dispatcher fiber, so that call waited on the very
loop it was blocking and busy-spun forever. The handler is now a no-op: a routed socket only
reaches the network when the handler asks for it, so the page still cannot dial out, and
nothing in the handler waits on the dispatcher any more.

Aborting the upgrade request from the HTTP route handler instead does not work, page.route
never sees WebSocket handshakes and the connection goes through.

Fixes #30024
2026-09-15 22:51:37 -04:00
Timothy Jaeryang Baek
05484aa055 refac 2026-09-12 20:09:21 -04:00
G30
6f55514c5c
fix: let validate_url accept dotless hosts when local web fetch is enabled (#29945) 2026-09-12 17:56:16 -04:00
Timothy Jaeryang Baek
12b14124b9 refac 2026-09-08 23:18:48 -04:00
Timothy Jaeryang Baek
ba34bee2d1 refac 2026-09-08 12:45:11 -04:00
G30
2f2f4c872f
fix(retrieval): drain playwright route handlers before the page closes (#29325) 2026-09-07 13:20:59 -04:00
Classic298
ae01ef9c95
fix: skip image, media and font requests in the Playwright web loader (#29742)
Fetching a page with many media files through the Playwright web loader was extremely slow or timed out, while raw Playwright loaded the same page in a couple of seconds. The loader routes every request the page makes through the backend HTTP client and downloads the full body before the browser sees any of it, and the browser cancelling a media request once it has enough never reaches that download. On a page with a few dozen audio players every file was pulled in full for a text extraction that never reads it.

Image, media and font requests are now aborted in the interceptor before any fetch is made. None of them feed the text extraction. Measured on the page from the report with the default timeout on the same connection:

| | requests fetched | bytes downloaded | elapsed |
|---|---|---|---|
| before | 151 | 55.0 MB | 10.1 s |
| after | 45 | 3.9 MB | 2.6 s |

Fixes #29741
2026-09-06 17:01:00 -05:00
Timothy Jaeryang Baek
0fa4dea5ff refac 2026-09-06 17:21:57 -04:00
Timothy Jaeryang Baek
d27aa72ab4 refac 2026-09-06 17:13:32 -04:00
Timothy Jaeryang Baek
cb942bb94c refac 2026-09-06 16:48:30 -04:00
Classic298
68a74da70b
Fix: mps inference evaluations (#29735)
* fix(retrieval): serialize local embedding and reranking on MPS

On Apple Silicon the server process is killed outright (SIGSEGV or SIGTRAP, no traceback) partway through answering any question that retrieves from a knowledge base with hybrid search and a local reranking model. The client sees a dropped connection and the answer is lost.

Hybrid search fans its queries out concurrently and every task calls the same shared local model on a worker thread. Torch's Metal shader cache is a process-wide singleton whose lookup tables have no lock, so two of those threads racing inside it corrupt the cache and take the process down with it.

Guard the local SentenceTransformer and CrossEncoder calls with a shared lock that is only a real lock when the selected device is MPS. CPU and CUDA installs keep the concurrency they have today, and external reranking endpoints are untouched. Reranking several queries on a Mac now runs one at a time, which is the cost of the process staying alive.

Verified by driving the real hybrid-search fan-out with 16 concurrent queries: peak simultaneous entries into the local model drops from 16 to 1 on MPS, stays at 16 on CPU, and the returned documents, scores and ordering are byte-identical in every case.

Fixes #29722

* fix(evaluations): serialize the leaderboard embedder against retrieval on MPS

The leaderboard's tag-similarity search builds its own SentenceTransformer, and on Apple Silicon sentence-transformers places it on the MPS device. It runs on a worker thread, so an admin running a leaderboard search while anyone queries a knowledge base puts two threads into torch's Metal backend at the same time, which kills the server process outright with no traceback.

Move the lock added for the retrieval path into env.py, beside the device selection that decides whether MPS is used at all, and take it around the leaderboard's embedding calls as well. Sharing one lock between the two modules is the whole point, since two separate locks would still let a leaderboard search collide with a retrieval query.

Only inference is guarded, matching the retrieval path. Model construction stays as it is here and in the retrieval routers.

Verified by driving the leaderboard similarity path and retrieval reranking from six threads against one instrumented model: peak simultaneous entries drops from six to one on MPS, and the similarity scores are unchanged.

Related to #29722.
2026-09-06 15:20:41 -04:00
Classic298
1bfa59acbd
fix: keep HTML entities in extracted document text (#29736)
Uploading a file whose text contains literal HTML entities stored a rewritten copy of it: `&nbsp;` became a non-breaking space, `&gt;` became `>`, and `&amp;nbsp;` was decoded twice down to a bare non-breaking space. That stored text is what gets indexed and what the model reads, so notes, specs and source files reached the model differing from the file that was uploaded.

Every loaded document goes through `ftfy.fix_text`, which is there to repair mojibake left by the encoding-detection fallback. Its default configuration also decodes HTML entities, per line and sticky forward: entities are decoded on every line up to the first line holding a literal `<`, then left alone for the rest of the document. The same escape therefore survives or vanishes depending on where it sits in the file. This disables that one behaviour and leaves every other ftfy repair in place.

Text from a third-party extraction engine that returns escaped output now keeps those escapes. Guessing whether an escape is markup or content is the bug being fixed.

Fixes #29732
2026-09-06 15:13:45 -04:00
Classic298
acd03b147c
fix: treat .ino files as source text so knowledge uploads stop failing (#29673)
Uploading an Arduino sketch (`.ino`) to a knowledge base failed with `Expecting value: line 1 column 1 (char 0)` whenever the content extraction engine was Tika or Docling. Browsers send `.ino` as `application/octet-stream`, and the extension was missing from the known source extension list, so the file was handed to the extraction server instead of being read as plain text. The server answered with a non-JSON body and the loader crashed while decoding it. `.cpp` and `.h` sketches in the same folder uploaded fine, because those extensions are already on the list.

Adding `ino` to that list routes it to the plain text loader, the same way the yaml/toml gap was closed in 710320601a. A sketch is plain C++ text, so there is nothing for a document extraction server to do with it.

Verified by dispatch matrix over 35 extensions, 5 content types and all 8 engines against a stub server that reproduces the non-JSON response: the only rows that change are `.ino` under Tika and Docling, which now resolve to the text loader and extract the sketch verbatim. Every other row is unchanged.

Fixes #29670
2026-09-04 19:13:04 -04:00
Classic298
ea1b58eb6a
fix: apply the shared metadata size cap to the last four vector DB backends (#29502)
Chunk metadata inserted into the vector DB carries arbitrary client-
supplied fields (e.g. custom metadata from file uploads), so it needs
the same size cap, type coercion and null-byte sanitization every
other backend already applies through process_metadata(). Four
backends never called it: Qdrant, Qdrant multitenancy, Milvus
multitenancy and Oracle23ai, so an oversized or malformed metadata
blob went into those unfiltered.

Wire process_metadata() in at each backend's single insert/upsert
chokepoint, matching the pattern already used by the other eleven
backends.
2026-09-04 11:42:20 -04:00
Timothy Jaeryang Baek
1caf22b5a8 refac 2026-08-31 00:40:15 -04:00
Timothy Jaeryang Baek
873fb741c2 refac 2026-08-31 00:39:16 -04:00
Timothy Jaeryang Baek
5c62cc0517 chore: format 2026-08-25 16:53:53 -04:00
Timothy Jaeryang Baek
1d6d4e6e66 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-25 16:27:17 -04:00
Timothy Jaeryang Baek
6dcc2d5269 refac 2026-08-25 15:48:30 -04:00