Commit graph

5232 commits

Author SHA1 Message Date
Timothy Jaeryang Baek
3fc1146c13 refac 2026-09-21 10:44:37 -04:00
Classic298
94eea41a75
fix: surface searchapi errors, news results and redirect links (#30308)
Web search via searchapi.io could come back empty or near-empty with no
hint of why: an invalid or expired API key turned into an empty result
set instead of an error, the google_news engine splits its results
between organic_results and top_stories and only the first block was
read, and google links came back as google.com/goto redirects the web
loader cannot fetch, so citations pointed at a redirect blob.

The search now reads both result blocks, asks google engines for
resolved destination links, raises on HTTP errors, carries a 30s request
timeout, skips result rows without a link, and logs the response body at
debug instead of dumping every search at info.

Fixes #30305
2026-09-21 10:44:06 -04:00
Timothy Jaeryang Baek
e8bd0661d3 refac 2026-09-21 10:30:53 -04:00
Timothy Jaeryang Baek
754c4b5762 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-09-21 10:25:20 -04:00
G30
c864e3ae7c
fix: stop asking for chat variables a model's system prompt no longer declares (#30173) 2026-09-21 10:23:21 -04:00
Timothy Jaeryang Baek
a9541c18ca refac 2026-09-21 10:19:59 -04:00
Classic298
aceda892bd
perf: serve the shared model pool from a per-worker cache (#28176)
With the Redis websocket manager the shared model pool is a `RedisDict`, so resolving a model id fetches from Redis, and more than one of those fetches pulls the whole pool. Every chat request pays that latency and that traffic, and the cost grows with the number of models configured: at 120 models a request moves several hundred kilobytes to ask about models that have not changed since the last request.

The pool now keeps a per-worker cache of the hash and refetches it only when the signature key that `set()` already maintains changes. A cache may only trust a signature that describes the bytes it fetched, so every write invalidates the signature and no signature value is ever issued twice. Without a fresh token each time, a pool that flaps back to an earlier state returns to an earlier digest, so a reader that races the write keeps serving the old pool. `delete_many()` was a second hole in that: it updated a dead attribute and never touched the signature, so readers kept serving models it had already removed. A TTL cache would have been simpler, but it serves a pool it already knows may be stale and costs the same single round trip this check costs.

Replaying the real access pattern, a request that finds the pool unchanged moves about 1 KB no matter how many models are configured, over the same number of round trips as before, except after a write that leaves no signature, which holds reads at two commands until the next non-empty `set()` writes one. A request that first sees a changed pool costs one extra command, and so does rewriting the pool. Values stay cached serialized, so `items()` and `values()` still decode on every call, and each worker holds the whole pool in memory.
2026-09-21 10:15:49 -04:00
Classic298
49b25506dd
perf: stop round-tripping the whole chat to read or write one message (#28184)
Long conversations get progressively more expensive to stream into. Every message-level write reloads the entire chat JSON, walks every string in it for a null-byte sanitisation pass and rewrites the whole column, and reading a single message loads and validates the whole chat too. The websocket event emitter reads and then writes, so one message event round-trips the full conversation twice to change a few hundred bytes.

Reading one message now selects only history.messages out of the JSON column and indexes it in Python, so the rest of the chat is neither loaded nor model-validated. Writing one message now sanitises only what is actually entering the chat, plus the title column that is mirrored off the blob, instead of re-walking a conversation that was already sanitised when it was written. Legacy rows are still healed on read by get_chat_by_id, which is unchanged.

One behaviour change: chats written before the null-byte sanitisation existed can still hold null bytes in the stored JSON. Those rows used to be rewritten clean as a side effect of any message write, and are now cleaned when the chat is read instead. The visible consequence is that the message write and delete endpoints echo back the updated chat, and for such a legacy row that echo now carries the raw null bytes rather than stripped ones, until the next read of that chat heals it. A GET of the chat is unaffected, and title, the one text column that PostgreSQL cannot store a null byte in, is still sanitised on every write.

Measured on SQLite with a 10.8 MB chat (3000 messages): a single-message upsert is 227 ms before the branch and 173 ms after, and the legacy single-message read is 135 ms before and 79 ms after. The whole-blob scan the delete path used to run costs 5.40 ms on a 1.8 MB chat and 38.79 ms on the 10.8 MB one, and is gone.

Ref https://github.com/open-webui/open-webui/issues/28169

### Contributor License Agreement

<!--
🚨 DO NOT DELETE THE TEXT BELOW 🚨
Keep the "Contributor License Agreement" confirmation text intact.
Deleting it will trigger the CLA-Bot to INVALIDATE your PR.

Your PR will NOT be reviewed or merged until you check the box below confirming that you have read and agree to the terms of the CLA.
-->

- [x] By submitting this pull request, I confirm that I have read and fully agree to the [Contributor License Agreement (CLA)](https://github.com/open-webui/open-webui/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT), and I am providing my contributions under its terms.

> [!NOTE]
> Deleting the CLA section will lead to immediate closure of your PR and it will not be merged in.
2026-09-21 10:11:38 -04:00
Timothy Jaeryang Baek
5fb869db22 refac 2026-09-21 08:59:39 -04:00
Classic298
488b3c571c
perf: stop note co-editing echoing every remote update back to the server (#28185)
Every client in a note re-broadcasts each update it receives, with a full content snapshot attached, and the server appends each echo to the document log and writes the note again. Traffic and note writes therefore scale with the number of people who have the note open: every extra participant adds one more full echo of every keystroke.

Remote updates are now applied with the `'server'` origin the state path already uses, which the local listener ignores, so the echo stops. The echo did carry one thing worth keeping: the receiving client holds the merged document, which the sender had not seen yet, so each receiver now sends a content-only message once the edits settle, debounced 500ms, and flushes a pending one when the editor is torn down. The backend accepts an update message with no `update` field for that case, where it previously raised and dropped the save.

Replaying keystrokes at 120ms with a real Yjs document, socket messages fall 48.8% with two clients, 65.6% with three and 79.2% with five, with byte counts tracking the same on notes up to 50KB. Server-side update appends drop by a factor of the client count. The cost is one snapshot upload per receiving client per typing pause, which the server's own debounce then collapses into a single extra note write however many people are watching. A snapshot has to come from a client because the server cannot rebuild the markdown, HTML and JSON shape the note record stores. It also moves the merge 500ms later than the echo delivered it, so if two edits cross on the wire and every editor then loses its connection and closes inside that window, the note keeps what it was last sent and one of the two edits is lost. A client still connected at teardown flushes its pending snapshot, and any other editor left in the note closes the gap.
2026-09-21 08:53:27 -04:00
Classic298
6fb68e43cb
refac(images): request an unencoded body when fetching remote chat images (#29623)
The remote chat-image fetch now asks for an identity-encoded response and skips one that comes back content-encoded.

An image whose host stores and echoes a `Content-Encoding` regardless of what the client asks for (an S3 or MinIO object uploaded with that metadata) is no longer inlined; the message is forwarded with the original URL instead, the same way an unreachable image already behaves.
2026-09-21 08:53:06 -04:00
Classic298
438d9db8db
fix: correct recurrence rule parsing for schedules and calendar events (#29262)
* fix: correct recurrence rule parsing for schedules and calendar events

An automation set to repeat a limited number of times, say ten or a hundred, was treated as a one-shot and reported no repeat interval, because any count whose digits began with a one matched a text check for the one-shot case. The scheduler already answers that question correctly by asking the rule for its next two occurrences, so the text check is gone and the count is read as the number it is.

A recurrence rule that carries its start date on the same line as the repeat text kept that date when the automation was parsed, so the schedule ran from whatever date the rule happened to carry and ignored the start the user picked. The filter that drops the start date now splits the rule on any whitespace, the same way the rule parser itself does, so both agree on where one part of the rule ends and the next begins.

The same mismatch on the calendar path anchored a recurring event to the date inside its rule, so occurrences showed up before the event had begun and at the wrong time of day. That filter splits the rule the same way now, and the series starts at the event's own start.

All three come from one place, recurrence rules being matched and cut as text. Rules written across several lines, which is what the schedule and calendar editors produce, behave exactly as before.

* refac: name rrule token vars parts to match calendar.py
2026-09-21 08:44:13 -04:00
Timothy Jaeryang Baek
0180efecf3 refac 2026-09-21 08:43:45 -04:00
Classic298
7fa8673296
fix: resolve this instance's own file URLs in edit_image regardless of the URL host (#29691)
The native edit_image tool fails with "400: [ERROR: Error loading image]" whenever the model hands it an absolute URL for an image Open WebUI already stores. Such a URL is treated as local only when its host string matches the incoming request's host exactly, so a default-port form, a container name or any host the model composed itself falls through to an outbound HTTP fetch instead. That fetch asks /api/v1/files/{id}/content without a session, gets a 401, and the user sees the generic 400.

Match the file URL on its path and let the existing local branch resolve it. Fetching that endpoint over the network can never succeed for a local or a remote instance, because it requires an authenticated user, so the host comparison only decided which way the request failed. Access control is unchanged: the local branch still goes through get_file_content_by_id, which enforces owner, admin or shared access.

Fixes #29220
2026-09-21 08:32:38 -04:00
G30
f40318c1ff
fix: keep the live prompt unchanged when a version is saved without Set as Production (#30231) 2026-09-21 08:32:02 -04:00
Timothy Jaeryang Baek
7f703dcc98 refac 2026-09-21 08:31:27 -04:00
Timothy Jaeryang Baek
dc4f5da5a0 refac 2026-09-21 08:28:36 -04:00
Classic298
9012bd153d
fix: stop sending count to the Staan search API (#30303)
Staan rejects any count other than its fixed page size of 10, so every search 400ed with the default result count of 3. The API always returns 10 results; the local results[:count] slice already applies the configured count, so the request parameter is simply dropped.
2026-09-21 08:04:26 -04:00
G30
702da1e471
fix: stop injecting memories and memory tools when memory is switched off (#30228) 2026-09-21 01:00:08 -04:00
Timothy Jaeryang Baek
97e013a661 refac 2026-09-21 00:51:28 -04:00
G30
402a187e82
fix: return no chats when a folder search term matches no folder (#30273) 2026-09-20 23:39:54 -05:00
Classic298
8e4cc946ce
fix: emit the resolved file path in terminal file events (#30282)
When a model calls display_file, write_file or replace_file_content with a relative path, Open Terminal resolves it against the session working directory and returns the absolute path, but the event sent to the browser carried the raw argument instead. The file panel matches that string against the file browser root, a relative path never matches, so the preview never opens, the panel jumps to the root and the session working directory is rewritten to the root. With the root turned off (OPEN_TERMINAL_FILE_BROWSER_ROOT=filesystem) there is nothing to clamp to and the relative string is sent to the terminal as the new working directory, moving it silently. The tool call itself succeeds either way, so the failure only shows up as a panel that will not open the file the model just wrote.

Both events now carry the path from the tool result and fall back to the argument when the result cannot be read, which is what build_terminal_file_tool_result already does for the chat file attachment. The same one-line rule is applied to the direct tool server path in the frontend, where the browser runs the tool itself and the write_file branch beside it was already correct.

Checked against a live Open Terminal: relative arguments now emit the absolute path, absolute ones are unchanged, non-existent files and inline displays still emit nothing, unreadable or error results still fall back to the argument, and run_command is untouched.

Related to #30051
2026-09-21 00:31:14 -04:00
Classic298
7fa705f3b8
feat: let operators expose chosen file metadata to the model in retrieved sources (#29696)
Custom metadata attached to a file upload now reaches the vector DB, but the model still never sees it. Both prompt-assembly paths build their output from a fixed field set: the classic RAG <source> tag carries only id, name and resource type, and the retrieval tools return only content, source and file id per chunk. A scraper that records where each document came from therefore cannot get that origin in front of the model, so answers cannot state it.

RAG_SOURCE_METADATA_KEYS names the chunk metadata keys allowed through to the model. Configured keys are emitted as extra attributes on the <source> tag and as extra fields on tool result chunks, covering both retrieval paths. It is empty by default, so nothing changes for existing deployments.

An allowlist instead of passing everything through, because chunk metadata also carries file hashes, collection names, embedding config and relevance scores, which would then be added to every retrieved chunk of every request. Values are attacker-controllable through an uploaded file, so they are escaped before they go into the tag, and a configured key can never displace a field the tag or the chunk already defines.

Reported in open-webui/open-webui#29486.
2026-09-19 17:02:59 -05:00
Timothy Jaeryang Baek
3a6d0fd203 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-09-19 17:54:09 -04:00
G30
65a3a115eb
fix: keep the speech-to-text extension allowlist when the Audio settings are saved (#30208) 2026-09-19 17:36:38 -04:00
Timothy Jaeryang Baek
f6922a4c42 refac 2026-09-19 17:33:41 -04:00
Timothy Jaeryang Baek
10d1cfe637 refac 2026-09-19 17:26:18 -04:00
G30
b64cb06043
fix: let a user without the chat delete permission delete a folder while keeping its chats (#30163) 2026-09-19 16:07:34 -05:00
Timothy Jaeryang Baek
946be43237 refac 2026-09-19 17:05:57 -04:00
G30
57063aaad8
fix: dropping a folder onto the folder it is already in no longer fails with Folder already exists (#30169) 2026-09-19 15:41:10 -05:00
G30
2690d04cac
fix: clean up orphan tags for the chat's owner, not the admin, when an admin deletes another user's chat (#30171) 2026-09-19 15:40:58 -05:00
G30
9e293a58ea
fix: label the search modal's archive action Unarchive for archived chats and report what happened (#30177) 2026-09-19 10:02:11 -05:00
G30
dbe538c033
fix: keep a note's sharing when a collaborator saves it (#30175) 2026-09-19 10:00:39 -05:00
Classic298
1f8f1f61bb
fix: tell ask_user models that the first option is shown as Recommended (#30196)
The ask_user card badges the first option of every question as "Recommended", but nothing ever told the model that. The model picks whatever order it likes, so the badge really means "listed first" and users act on a recommendation the model never made.

The ask_user tool description now states that the first option is labelled Recommended and that the model should list the option it recommends first, so the badge reflects an actual choice.

Kept the badge and instructed the model instead of adding a per-option "recommended" flag: the flag would need a schema change, validation in the request normalizer and a frontend change, for the same result in the common case. Dropping the badge was the other option, but it removes a useful affordance rather than fixing it.

Verified that the added line reaches the model by running the docstring through the tool-spec builder and checking the generated OpenAI function schema.

Fixes #30195
2026-09-19 09:59:18 -05:00
Classic298
3d29548716
fix: pgvector reads leak their connection and lose most of their neighbours (#30142)
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
Two defects on the pgvector read path. A search, query or get that finds nothing returns before the rollback that ends its read-only transaction, so the session keeps the connection it checked out; retrieval fans out over worker threads and the session is thread-local, so every thread that runs an empty read holds a connection for the lifetime of that thread. Users see vector search die after a while with "QueuePool limit of size 5 overflow 10 reached" and no way back other than a restart. Rolling back before the three early returns puts the connection back: measured on PostgreSQL 16, twelve empty reads on a default-shaped pool left all 12 connections checked out before and 0 after.

All collections also share one table under one vector index, and the WHERE collection_name filter is applied after the index walk, so a knowledge base holding a small share of the rows keeps only a small share of its neighbours, silently. pgvector 0.8 added iterative scans for this: the scan keeps going until enough rows pass the filter. This sets it per search so it cannot leak into other sessions, and only when the installed extension supports it, since setting it on pgvector 0.7 would make every search raise. PGVECTOR_ITERATIVE_SCAN turns it off or picks strict_order; it defaults to relaxed_order because the shipped behaviour is silently wrong results, and the cost is about a millisecond per search.

Measured through PgvectorClient.search on PostgreSQL 17 and pgvector 0.8, 101500 rows of 384 dimensions with the knowledge base at 1.5% of the table, hnsw m=16, recall@10 against an exact scan over 40 queries: 0.070 at 4.9 ms before, 0.970 at 7.2 ms after.

Fixes #30133
Fixes #30135
2026-09-18 19:31:49 -04:00
Classic298
52cd298411
fix: allow forking a chat that holds a stale unfinished message (#30131)
Forking returned 409 "Wait for the current response to finish before forking." forever once any assistant message anywhere in the chat was left at done = false, with nothing generating and even when that message sat on a branch that was not being forked. Interrupted turns leave the flag behind and nothing clears it, so an affected chat could never be forked again.

The endpoint now refuses only while a task is actually running, which is the check /compact has always relied on by itself. A message still unfinished on the forked branch is marked done in the copy so the fork does not open showing a spinner, while a turn paused waiting on tool approval keeps its unfinished state and its pending call so the fork still shows the prompt. The source chat is left untouched either way.

The client-side check had the same shape of bug: it read history.currentId instead of the message actually being forked, so a paused last turn also blocked forking earlier, finished messages. It now reads the message it is about to fork.

Fixes #30128
2026-09-18 19:31:20 -04:00
Classic298
fffcb727d2
fix: release the openGauss connection when a read returns no rows (#30144)
An openGauss query or get that finds nothing returns before the rollback that ends its read-only transaction, so the session keeps the connection it checked out of the pool. Retrieval fans out over worker threads and the session is thread-local, so every thread that runs an empty read holds a connection for the lifetime of that thread, and vector search eventually fails with a pool checkout timeout that a restart is the only way out of. Empty reads are routine: an empty knowledge base, a file whose chunks were deleted, or a metadata filter that matches nothing all produce one.

Rolling back before the two early returns puts the connection back. Measured by driving OpenGaussClient itself with a pool of 5: three empty query reads left 3 connections checked out and 0 free before the change, and 0 checked out with 3 free after, same for get. search is already correct, it has no early return.

The pgvector client had the same defect, fixed separately in #30142.
2026-09-18 19:31:01 -04:00
G30
4d01f1ebed
fix: keep the archived state and chat variables when importing a chat export (#30155) 2026-09-18 19:29:30 -04:00
Timothy Jaeryang Baek
64bbdf7a73 refac 2026-09-18 19:28:40 -04:00
Classic298
d5cacb3c0a
fix: convert forced tool_choice and non-streaming tool calls for Responses API connections (#30095)
On a connection with API type "responses", a forced tool choice sent in
the Chat Completions shape was forwarded to the provider unchanged, so
providers that validate the Responses API schema rejected the request.
A non-streaming reply that carried a function call was also flattened
to empty text with finish_reason "stop", so API clients never saw the
tool call.

The request converter now flattens a forced function choice to the
Responses API shape, and the result converter turns every function_call
output item into a Chat Completions tool_calls entry, mirroring the
existing Responses output mapping in utils/misc.py.

Fixes #30085
2026-09-18 19:18:26 -04:00
Classic298
9ee810ad09
feat: add Staan as a native web search provider (#30138)
European deployments that need EU data residency have no hosted web search option out of the box: every plug-and-play provider shipped today is US-based, and the only sovereign alternative is self-hosting SearXNG, which means running and maintaining that infrastructure yourself. Staan (staan.ai) is a European search API with EU data residency, so adding it gives those deployments a drop-in choice.

It is configured like any other provider, through the admin UI or STAAN_API_KEY, STAAN_MARKET and STAAN_MAX_SNIPPETS. Market sets the region and language of the results and defaults to en-us. Max snippets asks Staan to fetch each result page and return semantically scored chunks of it, which get merged into that result's snippet so retrieval has more to work with; leaving it at 0 uses the plain search endpoint.

Wired the same way as Tavily and Exa, with domain filtering going through the shared get_filtered_results.

Requested in #26006.
2026-09-18 10:38:36 -05:00
Classic298
95406fd28d
fix: build the pgvector ivfflat index once there are rows to cluster on (#30143)
ivfflat places its centroids by clustering the rows it can see when the index is built, so an index built on an empty table gets centroids that mean nothing, and everything inserted afterwards is filed against them. On a fresh install the index is created immediately after the table, before a single chunk exists, and it is never rebuilt, so that install keeps a permanently untrained index and quietly retrieves the wrong chunks. An install that upgraded into the version introducing the index is unaffected, its table already had rows.

The index is now created once the table holds 50 rows per list, the sample size ivfflat itself aims for. Below that, and until the next start, searches fall back to an exact scan, which is correct and costs about a millisecond at that size. hnsw is untouched, it builds its graph as rows are inserted and has nothing to train on. One trade-off: the build moves from the first start to that later one, so an instance that has grown large in between pays a one-time index build during startup.

Measured on PostgreSQL 17 and pgvector 0.8 with the default lists=100 and probes=1, 384 dimensions, recall@10 against an exact scan, varying only how many rows existed when the index was built:

| rows at build | 200 | 1000 | 2500 | 5000 | 20000 |
|---|---|---|---|---|---|
| recall@10 | 0.180 | 0.563 | 0.967 | 1.000 | 1.000 |

Through PgvectorClient.search on a 20000-row table, an index built as it is today scores 0.480 against 1.000 built after the rows arrive, at the same 5 ms. An existing install can repair its index with REINDEX INDEX idx_document_chunk_vector, measured to take it from 0.480 back to 1.000.

Fixes #30134
2026-09-18 10:37:25 -05:00
Timothy Jaeryang Baek
1ddba7e2c6 refac 2026-09-18 11:31:59 -04:00
Timothy Jaeryang Baek
ca1eefe293 refac 2026-09-18 11:30:15 -04:00
Timothy Jaeryang Baek
ad9da98168 refac 2026-09-17 20:11:05 -04:00
Timothy Jaeryang Baek
dbb17a5725 refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
2026-09-17 19:49:35 -04:00
Classic298
9923c53c10
fix: treat SVG uploads as documents instead of vision images (#30102)
Uploading an .svg to a chat attached it as a vision image input, so the model received a data URI it could not decode. PIL-backed servers answered "cannot identify image file" and OpenAI answered "The image data you provided does not represent a valid image". No setting made it work.

SVG now takes the ordinary file upload path, so its XML source is extracted and indexed and the model can answer questions about it. Rasterizing was the alternative and it would have discarded the part of an SVG a model reads best, the source itself. Raster formats are untouched and still go up as image inputs.

A shared helper replaces the ad hoc image/ prefix checks at the points that decide image input versus document, on both ends. It normalises the content type first, because a stored "image/SVG+xml" or a trailing charset parameter slipped past a plain comparison.

One behaviour change worth knowing: an SVG now needs the model to have the file upload capability, where before it rode in as an image.

Fixes #30100
2026-09-17 17:40:07 -04:00
Classic298
460f2e7634
fix: surface initial chat title generation failures in the log (#30106)
When title generation for a brand new chat fails, the chat silently keeps the provisional title (the full first user message) and nothing is written to the log at the default level, so there is no way to tell that the feature is broken rather than disabled. The failure was caught by a broad `except Exception` and reported with `log.debug`, which is invisible unless GLOBAL_LOG_LEVEL is set to DEBUG. Issue #29533 describes an outage of this path that survived four releases for exactly that reason.

This logs it with `log.exception` instead, matching how the rest of main.py reports background task failures, so an admin sees one ERROR line plus the traceback naming the real cause.

Nothing else changes: the success path is untouched, the exception is still swallowed so the detached task cannot take anything down with it, and `background_tasks_handler` does not log this exception itself, so there is no duplicate traceback.

Fixes #29533
2026-09-17 17:39:43 -04:00
Classic298
0837f310be
fix: reject Docling conversions that failed inside an HTTP 200 response (#30107)
Uploading a file that Docling declines or fails to convert either dies with `TypeError: argument of type 'NoneType' is not iterable`, or silently succeeds and stores the literal string `<No text content found>` as the document's text, which then gets indexed and handed to the model as if it were the file. Docling returns the conversion outcome inside the HTTP 200 body, so checking only the HTTP status made a refused conversion look identical to a successful one, and the `errors` array that says why in plain words was never read.

Failed and skipped conversions are now rejected with the messages Docling returned, so an unsupported format surfaces as "File format not allowed: example.dxf" and the traceback is gone. The markdown field is also read as nullable, because Docling returns JSON `null` for every content format it was not asked to produce, which any Docling Parameters setting `to_formats` without `md` will hit, and that null was what raised the TypeError.

Successful conversions with empty markdown keep the existing `<No text content found>` placeholder, matching what TikaLoader and the Mistral loader already do in the same package.

Fixes #29808
2026-09-17 17:39:35 -04:00
Timothy Jaeryang Baek
58b36765a7 refac 2026-09-16 22:47:37 -04:00