With Kagi as the web search engine, any entry in the Domain Filter List made every search fail with "'SearchResult' object has no attribute 'get'", so the model got no search results. The Domain Filter List now works with Kagi the same way it does with the other search engines, and the remaining results reach the model.
Fixes#32004
With Perplexity Search as the web search engine, the Domain Filter List had no effect: blocked domains still showed up in the results the model got, and an allowlist let every result through. The filter now applies to Perplexity Search results as it does for the other search engines.
Fixes#32005
When a skill was saved several times within one second, its version history in the skill editor listed those versions in a shuffled order. This happens when a model edits a skill several times in one reply, or when an import replaces a skill right after it was created. Each new version now gets a save time at least one second later than the one before it, so the history keeps the order the versions were saved in. During such a burst, the shown save time runs ahead by about one second per extra save.
Fixes#32134
Searching Workspace > Skills only matched a skill's original name, description and id, so typing the German name given to a skill in its editor found nothing, even with the interface in German. The search now also matches the names and descriptions translated in the skill editor, in any language, so both the translated and the original name find the skill, as the Translations page in the docs says.
Fixes#32018
When a chat is named after its first message (title generation off, or the automatic title is blank), a message that starts with a skill picked by typing $ in the message box gave the chat a title like `<$tides-1234|Tides> when is high tide`. The title now shows each picked skill by its name, so the sidebar reads "Tides when is high tide".
Fixes#32019
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
After Redis restarted or closed Open WebUI's connections, the first few requests that needed Redis failed. Chat requests answered with a 500 error, and a browser tab connecting at that moment got no live updates, such as streamed replies, until it was reloaded. Open WebUI now reconnects and retries the failed Redis call once when Redis has closed the connection, so requests go through as soon as Redis is back. While Redis is down, requests still fail as before.
With Stream Chat Response turned off, a reply where the model called a tool was never saved or marked as finished, so the chat kept showing it as still generating until the page was reloaded. The reply is now saved and finished like any other non-streamed reply. Tools themselves still only run with streaming on, as the docs describe.
Outlet filters see the finished reply with its content and token usage, but had no way to tell whether the model ended on its own, ran into the token limit or stopped to call a tool, short of reading every chunk in a stream filter. For OpenAI-compatible and Ollama models, the assistant message handed to outlet now carries finish_reason with the value the provider reported on the last model call of the turn, for streaming and non-streaming replies. Ollama replies cut off by the token limit now report length as well, where they always said stop before.
With ENABLE_FORWARD_USER_INFO_HEADERS on and hybrid search enabled, rerank requests went out without the user headers when a chat searched attached knowledge, when a model used the built-in knowledge search tool, and when the collection query API was called with hybrid turned off for that one request. The embedding requests of the same search did carry them, so external rerankers that use these headers for per-user auth, rate limits or auditing saw anonymous calls. Rerank requests now carry the signed-in user, the same as embedding requests.
Fixes#32060
When a model streams its images in separate chunks, each new chunk saved all earlier images again, so the copies doubled with every image: four images were stored as fifteen and the chat showed duplicates. New replies now keep each image once. Chats that already hold duplicates stay as they are.
Fixes#32053
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
* fix: Direct Connection replies are saved scrambled or empty
With a Direct Connection the browser tab forwards the model's reply to the server piece by piece. Since the server started checking the tab's session again for every one of those pieces, the pieces could overtake each other while the checks ran, so the saved reply came out in the wrong order, or empty when the end of the stream arrived first. Pieces from one tab are now handled one after another in the order they arrived, and each one is still checked.
Fixes#31953
* fix: check Direct Connection reply pieces side by side while keeping them in order
Handling one tab's reply pieces strictly one after another also made each piece wait for the previous piece's session check, so a fast 2000-piece reply took 12 to 22% longer to save than without the ordering. Each piece's check now starts as soon as it arrives and only the hand-over waits its turn, so replies save as fast as before and still in order. Once a check fails, for example after a sign-out, no later piece from that tab gets through either.
On Qdrant, uploading a file into a knowledge base or editing a file's content could leave the knowledge base with only part of the file, often exactly 64 chunks, while the file showed as completed and nothing was logged. The file's chunks were saved without waiting for Qdrant to make them searchable, and the knowledge base copied them straight after, so it only got the ones Qdrant had finished storing. Saving now waits until Qdrant has finished storing the chunks, so the knowledge base always gets the whole file. On a busy Qdrant this makes file processing somewhat slower, since each save now waits for the server.
Fixes#31959
When a model file was downloaded by URL in Manage Ollama, the end of the file could still be on its way to disk when Open WebUI sent it to Ollama, so Ollama got a cut-off copy. The download is now fully on disk before it is sent, so Ollama gets the whole file.
Fixes#31956
When a model returned an image, only PNG was saved as a file. JPEG and WebP images stayed as raw base64 data inside the chat in the database, so a 1.5 MB JPEG made the stored chat 1.5 MB larger. These images are now saved as files and the chat keeps only a link to them, the same way PNG already worked.
Fixes#31916