When a chat is named after its first message (title generation off, or the automatic title is blank), a message that starts with a skill picked by typing $ in the message box gave the chat a title like `<$tides-1234|Tides> when is high tide`. The title now shows each picked skill by its name, so the sidebar reads "Tides when is high tide".
Fixes#32019
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
After Redis restarted or closed Open WebUI's connections, the first few requests that needed Redis failed. Chat requests answered with a 500 error, and a browser tab connecting at that moment got no live updates, such as streamed replies, until it was reloaded. Open WebUI now reconnects and retries the failed Redis call once when Redis has closed the connection, so requests go through as soon as Redis is back. While Redis is down, requests still fail as before.
With Stream Chat Response turned off, a reply where the model called a tool was never saved or marked as finished, so the chat kept showing it as still generating until the page was reloaded. The reply is now saved and finished like any other non-streamed reply. Tools themselves still only run with streaming on, as the docs describe.
Outlet filters see the finished reply with its content and token usage, but had no way to tell whether the model ended on its own, ran into the token limit or stopped to call a tool, short of reading every chunk in a stream filter. For OpenAI-compatible and Ollama models, the assistant message handed to outlet now carries finish_reason with the value the provider reported on the last model call of the turn, for streaming and non-streaming replies. Ollama replies cut off by the token limit now report length as well, where they always said stop before.
When a model streams its images in separate chunks, each new chunk saved all earlier images again, so the copies doubled with every image: four images were stored as fifteen and the chat showed duplicates. New replies now keep each image once. Chats that already hold duplicates stay as they are.
Fixes#32053
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
When a model returned an image, only PNG was saved as a file. JPEG and WebP images stayed as raw base64 data inside the chat in the database, so a 1.5 MB JPEG made the stored chat 1.5 MB larger. These images are now saved as files and the chat keeps only a link to them, the same way PNG already worked.
Fixes#31916
When a reply was paused and then picked up again, either by answering a question from the built-in Ask User tool or by pressing Continue, every tool round from the second one on sent everything from before the pause to the model twice. Providers that reject repeated tool calls, like DeepSeek, then fail with "Duplicate 'call_id'" and the chat stops, while others quietly see the earlier tool calls and text twice. Everything from before the pause is now sent exactly once. Tested against a mock provider that records every request: three tool rounds after an answered question and after Continue send everything once, and replies that were never paused send exactly what they sent before.
Fixes#31991
Some providers, like Kimi K3 on OpenRouter, number their tool calls from zero again each time the model calls tools within the same reply. A later batch of calls then overwrote the earlier calls with the same id, so the saved chat showed the earlier calls with the later arguments, and the model was sent its earlier results next to the wrong arguments. A new tool call whose id is already used in the same reply now gets a fresh id, so every call keeps its own arguments and result, and providers that send unique ids are untouched. Tested before and after against a mock provider that reuses ids on every round.
Fixes#28305
Images stored inside the chat record itself (chats from older versions, images that could not be saved as files, temporary chats) had their encoded data counted as text when estimating how full the context is, so a single 3 MB image counted as about a million tokens. That pushed the chat over the compaction threshold, summarizing older messages while the chat was well under the limit, and made the context usage indicator jump. The encoded image data is now left out of the token estimate, both for compaction and for the indicator.
Fixes#31913
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
When a reply was stopped or failed before the model sent any text, the chat kept an empty reply, and every later message in that chat sent it back to the model as part of the conversation. Providers that reject empty assistant messages then refused every new request, so the chat stayed unusable unless the empty reply was deleted from the database by hand. Empty replies are now skipped when the conversation is sent to the model, as replies that ended in an error already were. Chats already stuck like this work again on their next message, since nothing stored has to change.
Fixes#25083
Connecting an MCP tool server over OAuth 2.1 failed after signing in at the provider with an "invalid or expired state" error whenever the server advertises a long list of scopes, such as the Google Workspace MCP server with its 42 Google scopes. Open WebUI saved the full authorization link it sends to the provider in the session cookie, which pushed the cookie past the 4096-byte browser limit, so the browser dropped it and Open WebUI could not recognise the user when the provider sent them back. The link is no longer saved there, so the cookie stays at a few hundred bytes even with long scope lists.
Fixes#26382
Open WebUI puts the MCP server ID in front of every MCP tool name it offers to the model. When the two together are longer than 64 characters, providers that cap tool names at 64 (the OpenAI API, AWS Bedrock) reject the whole chat request, even if the model never picks that tool. Names over the limit are now cut to 64 characters ending in a short hash, so the same tool always gets the same name and the MCP server is still called with the original tool name. Names that already fit are unchanged.
Fixes#31821
Every chat that got a reply left a small lock behind in the server process, so a long-running server kept one for every chat it had ever answered and only a restart gave the memory back. The lock is now dropped as soon as nothing is using it, so a finished reply leaves nothing behind per chat, while replies, sub-agent results and timers for the same chat still wait for each other as before.
Fixes#31521
When a model called one tool, got its result and then called a second tool with no text in between, the next request merged both calls into an earlier assistant message that the provider had already seen. That changed the conversation's beginning, so the provider's prompt cache stopped matching from there for the rest of the chat. Each tool call and its result are now sent as their own messages, so the start of the conversation stays identical from one request to the next.
Fixes#31588
In a chat inside a folder with knowledge attached, the chat only searches the folder's knowledge, but a sub-agent it started could list and search every knowledge base the user can read. Sub-agents now get the folder's knowledge like the chat that started them. They already receive the folder's system prompt as part of the chat's instructions, so it is not added a second time.
Fixes#31569
* fix: audit log shows passwords that contain a double quote
With AUDIT_LOG_LEVEL set to REQUEST or REQUEST_RESPONSE, a password containing a double quote was only masked up to that quote, so a new password like Q"secret was logged as "********"secret. A request with whitespace before the colon, such as "new_password" : "secret", was not masked at all. Any field whose name ends in "password" is now masked through its closing quote in both cases.
* fix: audit log records passwords sent back in responses
With AUDIT_LOG_LEVEL set to REQUEST_RESPONSE, fields whose name ends in "password" were masked in request bodies, but response bodies were logged unmasked. Saving or opening the LDAP server settings therefore wrote the Application DN Password to the audit log in plain text, because the settings come back in the response, and the Jupyter passwords in the code execution settings leaked the same way. Responses now get the same masking as requests.
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Since 0.11.1, when a filter function wrote part of the reply before the model answered (for example some text and an HTML block), the artifact preview opened and then closed straight away. The page showed the filter text but was never sent it as part of the reply, so the model's first output overwrote it, the page briefly saw no HTML block and closed the preview. The filter text is now sent to the page before the model's output, so the preview stays open.
Fixes#31643
MCP tool servers that sign in through Google, such as Google's hosted Gmail, Drive and Calendar servers, never got a refresh token, because Google only issues one when the sign-in explicitly asks for offline access. When the one-hour access token ran out the refresh failed and the connection was removed, so every user had to sign in again every hour. When the server's sign-in page is Google's, the sign-in now asks for offline access and a fresh consent, so the token renews on its own. Other providers get the same sign-in request as before, and existing Google connections pick this up the next time the user signs in.
Fixes#28319
An automation whose schedule ends after a fixed number of runs (COUNT in its RRULE) showed extra future runs in the Scheduled Tasks calendar after each run. With COUNT=3, the calendar kept showing three upcoming runs after the first and second run, although only the remaining ones execute. The calendar now counts runs from the schedule's own start date, the same way automations are actually run, so it shows exactly the runs that are still going to happen. The 5000-entry display limit now only counts entries inside the visible range, so very frequent schedules that started shortly before it no longer show up short or empty.
Fixes#31600