After Redis restarted or closed Open WebUI's connections, the first few requests that needed Redis failed. Chat requests answered with a 500 error, and a browser tab connecting at that moment got no live updates, such as streamed replies, until it was reloaded. Open WebUI now reconnects and retries the failed Redis call once when Redis has closed the connection, so requests go through as soon as Redis is back. While Redis is down, requests still fail as before.
With Stream Chat Response turned off, a reply where the model called a tool was never saved or marked as finished, so the chat kept showing it as still generating until the page was reloaded. The reply is now saved and finished like any other non-streamed reply. Tools themselves still only run with streaming on, as the docs describe.
Outlet filters see the finished reply with its content and token usage, but had no way to tell whether the model ended on its own, ran into the token limit or stopped to call a tool, short of reading every chunk in a stream filter. For OpenAI-compatible and Ollama models, the assistant message handed to outlet now carries finish_reason with the value the provider reported on the last model call of the turn, for streaming and non-streaming replies. Ollama replies cut off by the token limit now report length as well, where they always said stop before.
With ENABLE_FORWARD_USER_INFO_HEADERS on and hybrid search enabled, rerank requests went out without the user headers when a chat searched attached knowledge, when a model used the built-in knowledge search tool, and when the collection query API was called with hybrid turned off for that one request. The embedding requests of the same search did carry them, so external rerankers that use these headers for per-user auth, rate limits or auditing saw anonymous calls. Rerank requests now carry the signed-in user, the same as embedding requests.
Fixes#32060
When a model streams its images in separate chunks, each new chunk saved all earlier images again, so the copies doubled with every image: four images were stored as fifteen and the chat showed duplicates. New replies now keep each image once. Chats that already hold duplicates stay as they are.
Fixes#32053
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
* fix: Direct Connection replies are saved scrambled or empty
With a Direct Connection the browser tab forwards the model's reply to the server piece by piece. Since the server started checking the tab's session again for every one of those pieces, the pieces could overtake each other while the checks ran, so the saved reply came out in the wrong order, or empty when the end of the stream arrived first. Pieces from one tab are now handled one after another in the order they arrived, and each one is still checked.
Fixes#31953
* fix: check Direct Connection reply pieces side by side while keeping them in order
Handling one tab's reply pieces strictly one after another also made each piece wait for the previous piece's session check, so a fast 2000-piece reply took 12 to 22% longer to save than without the ordering. Each piece's check now starts as soon as it arrives and only the hand-over waits its turn, so replies save as fast as before and still in order. Once a check fails, for example after a sign-out, no later piece from that tab gets through either.
On Qdrant, uploading a file into a knowledge base or editing a file's content could leave the knowledge base with only part of the file, often exactly 64 chunks, while the file showed as completed and nothing was logged. The file's chunks were saved without waiting for Qdrant to make them searchable, and the knowledge base copied them straight after, so it only got the ones Qdrant had finished storing. Saving now waits until Qdrant has finished storing the chunks, so the knowledge base always gets the whole file. On a busy Qdrant this makes file processing somewhat slower, since each save now waits for the server.
Fixes#31959
When a model file was downloaded by URL in Manage Ollama, the end of the file could still be on its way to disk when Open WebUI sent it to Ollama, so Ollama got a cut-off copy. The download is now fully on disk before it is sent, so Ollama gets the whole file.
Fixes#31956
When a model returned an image, only PNG was saved as a file. JPEG and WebP images stayed as raw base64 data inside the chat in the database, so a 1.5 MB JPEG made the stored chat 1.5 MB larger. These images are now saved as files and the chat keeps only a link to them, the same way PNG already worked.
Fixes#31916
The inline citation markers in answers from external knowledge bases (Qdrant, Milvus, pgvector) showed the site's domain, for example docs.example.test, instead of the page title saved with each document. They now show that title, the same way a regular knowledge base shows the file name, and still fall back to the domain when a document has no title. Two different pages can share a title, so each marker now looks up its title by the page it points to; before, a repeated title made every later marker show the next page's title and the last one show nothing.
Fixes https://github.com/open-webui/open-webui/issues/31929
With Milvus as the vector database, all chunks of a file were sent to Milvus in one request. Large files exceeded Milvus's default 64 MB request size limit, so processing ran through the whole embedding step and then failed with the error RESOURCE_EXHAUSTED. With a 4096-dimension embedding model this already happened at about 4,000 chunks, which is a few MB of text. Chunks are now sent in batches of 128, with and without ENABLE_MILVUS_MULTITENANCY_MODE.
Fixes#31989
When a reply was paused and then picked up again, either by answering a question from the built-in Ask User tool or by pressing Continue, every tool round from the second one on sent everything from before the pause to the model twice. Providers that reject repeated tool calls, like DeepSeek, then fail with "Duplicate 'call_id'" and the chat stops, while others quietly see the earlier tool calls and text twice. Everything from before the pause is now sent exactly once. Tested against a mock provider that records every request: three tool rounds after an answered question and after Continue send everything once, and replies that were never paused send exactly what they sent before.
Fixes#31991
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Some providers, like Kimi K3 on OpenRouter, number their tool calls from zero again each time the model calls tools within the same reply. A later batch of calls then overwrote the earlier calls with the same id, so the saved chat showed the earlier calls with the later arguments, and the model was sent its earlier results next to the wrong arguments. A new tool call whose id is already used in the same reply now gets a fresh id, so every call keeps its own arguments and result, and providers that send unique ids are untouched. Tested before and after against a mock provider that reuses ids on every round.
Fixes#28305