Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
When a model returned an image, only PNG was saved as a file. JPEG and WebP images stayed as raw base64 data inside the chat in the database, so a 1.5 MB JPEG made the stored chat 1.5 MB larger. These images are now saved as files and the chat keeps only a link to them, the same way PNG already worked.
Fixes#31916
When a reply was paused and then picked up again, either by answering a question from the built-in Ask User tool or by pressing Continue, every tool round from the second one on sent everything from before the pause to the model twice. Providers that reject repeated tool calls, like DeepSeek, then fail with "Duplicate 'call_id'" and the chat stops, while others quietly see the earlier tool calls and text twice. Everything from before the pause is now sent exactly once. Tested against a mock provider that records every request: three tool rounds after an answered question and after Continue send everything once, and replies that were never paused send exactly what they sent before.
Fixes#31991
Some providers, like Kimi K3 on OpenRouter, number their tool calls from zero again each time the model calls tools within the same reply. A later batch of calls then overwrote the earlier calls with the same id, so the saved chat showed the earlier calls with the later arguments, and the model was sent its earlier results next to the wrong arguments. A new tool call whose id is already used in the same reply now gets a fresh id, so every call keeps its own arguments and result, and providers that send unique ids are untouched. Tested before and after against a mock provider that reuses ids on every round.
Fixes#28305
Images stored inside the chat record itself (chats from older versions, images that could not be saved as files, temporary chats) had their encoded data counted as text when estimating how full the context is, so a single 3 MB image counted as about a million tokens. That pushed the chat over the compaction threshold, summarizing older messages while the chat was well under the limit, and made the context usage indicator jump. The encoded image data is now left out of the token estimate, both for compaction and for the indicator.
Fixes#31913
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
When a reply was stopped or failed before the model sent any text, the chat kept an empty reply, and every later message in that chat sent it back to the model as part of the conversation. Providers that reject empty assistant messages then refused every new request, so the chat stayed unusable unless the empty reply was deleted from the database by hand. Empty replies are now skipped when the conversation is sent to the model, as replies that ended in an error already were. Chats already stuck like this work again on their next message, since nothing stored has to change.
Fixes#25083
Connecting an MCP tool server over OAuth 2.1 failed after signing in at the provider with an "invalid or expired state" error whenever the server advertises a long list of scopes, such as the Google Workspace MCP server with its 42 Google scopes. Open WebUI saved the full authorization link it sends to the provider in the session cookie, which pushed the cookie past the 4096-byte browser limit, so the browser dropped it and Open WebUI could not recognise the user when the provider sent them back. The link is no longer saved there, so the cookie stays at a few hundred bytes even with long scope lists.
Fixes#26382
Open WebUI puts the MCP server ID in front of every MCP tool name it offers to the model. When the two together are longer than 64 characters, providers that cap tool names at 64 (the OpenAI API, AWS Bedrock) reject the whole chat request, even if the model never picks that tool. Names over the limit are now cut to 64 characters ending in a short hash, so the same tool always gets the same name and the MCP server is still called with the original tool name. Names that already fit are unchanged.
Fixes#31821
Every chat that got a reply left a small lock behind in the server process, so a long-running server kept one for every chat it had ever answered and only a restart gave the memory back. The lock is now dropped as soon as nothing is using it, so a finished reply leaves nothing behind per chat, while replies, sub-agent results and timers for the same chat still wait for each other as before.
Fixes#31521
When a model called one tool, got its result and then called a second tool with no text in between, the next request merged both calls into an earlier assistant message that the provider had already seen. That changed the conversation's beginning, so the provider's prompt cache stopped matching from there for the rest of the chat. Each tool call and its result are now sent as their own messages, so the start of the conversation stays identical from one request to the next.
Fixes#31588
In a chat inside a folder with knowledge attached, the chat only searches the folder's knowledge, but a sub-agent it started could list and search every knowledge base the user can read. Sub-agents now get the folder's knowledge like the chat that started them. They already receive the folder's system prompt as part of the chat's instructions, so it is not added a second time.
Fixes#31569
* fix: audit log shows passwords that contain a double quote
With AUDIT_LOG_LEVEL set to REQUEST or REQUEST_RESPONSE, a password containing a double quote was only masked up to that quote, so a new password like Q"secret was logged as "********"secret. A request with whitespace before the colon, such as "new_password" : "secret", was not masked at all. Any field whose name ends in "password" is now masked through its closing quote in both cases.
* fix: audit log records passwords sent back in responses
With AUDIT_LOG_LEVEL set to REQUEST_RESPONSE, fields whose name ends in "password" were masked in request bodies, but response bodies were logged unmasked. Saving or opening the LDAP server settings therefore wrote the Application DN Password to the audit log in plain text, because the settings come back in the response, and the Jupyter passwords in the code execution settings leaked the same way. Responses now get the same masking as requests.
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Since 0.11.1, when a filter function wrote part of the reply before the model answered (for example some text and an HTML block), the artifact preview opened and then closed straight away. The page showed the filter text but was never sent it as part of the reply, so the model's first output overwrote it, the page briefly saw no HTML block and closed the preview. The filter text is now sent to the page before the model's output, so the preview stays open.
Fixes#31643
MCP tool servers that sign in through Google, such as Google's hosted Gmail, Drive and Calendar servers, never got a refresh token, because Google only issues one when the sign-in explicitly asks for offline access. When the one-hour access token ran out the refresh failed and the connection was removed, so every user had to sign in again every hour. When the server's sign-in page is Google's, the sign-in now asks for offline access and a fresh consent, so the token renews on its own. Other providers get the same sign-in request as before, and existing Google connections pick this up the next time the user signs in.
Fixes#28319
An automation whose schedule ends after a fixed number of runs (COUNT in its RRULE) showed extra future runs in the Scheduled Tasks calendar after each run. With COUNT=3, the calendar kept showing three upcoming runs after the first and second run, although only the remaining ones execute. The calendar now counts runs from the schedule's own start date, the same way automations are actually run, so it shows exactly the runs that are still going to happen. The 5000-entry display limit now only counts entries inside the visible range, so very frequent schedules that started shortly before it no longer show up short or empty.
Fixes#31600
When a background sub-agent finished while the model was still writing the answer that started it, the sub-agent's report was placed as a reply to the previous answer instead. Once the answer ended, the chat switched to that other branch of the conversation, which hid the latest question and answer, and the model replied to the report without seeing that question and answer. The report now follows the answer that started the sub-agent (or the newest completed reply after it), so the conversation stays on one branch.
Fixes#31507
With ENABLE_OTEL, ENABLE_OTEL_TRACES and ENABLE_OTEL_LOGS all on, the collector received each log line twice, once with code location attributes and once without, because the logging instrumentation for traces now attaches its own log exporter next to Open WebUI's. It now keeps trace context on log lines without adding that second exporter, so each log line reaches the collector once, and OTEL_PYTHON_LOG_AUTO_INSTRUMENTATION=false is no longer needed as a workaround.
Fixes#31524
With ENABLE_ORJSON off (the default), non-English letters were saved to the database as escape codes, so "Ü" was stored as \u00dc. Searches that ignore upper and lower case compare against that saved text, so they missed any match that differs only in the case of a non-English letter: filtering models by the tag "Überblick" found nothing on Postgres, and searching automations for "отчёт" missed a prompt containing "Отчёт" on SQLite and Postgres. The 100,000 character size limit for user and chat variables counted the escape codes too, so Cyrillic or Chinese variables were refused as too large (or chat variables silently came out empty in the system prompt) at about a sixth of that size. Non-English text is now saved as written, which is how it is already saved with ENABLE_ORJSON on, so nothing changes for those instances, and the limit counts real characters. Anything saved before this keeps the escape codes until it is next edited.
With request auditing turned on, the audit log only masked fields named exactly "password". The new password from a password change, and passwords entered in admin settings such as YaCy or Jupyter, were written to the log as-is. Any field whose name ends in "password", in any letter case, is now replaced with asterisks.
When signing in through an OAuth/OIDC provider failed, for example because the provider denied access, the account had no email or its email domain was not allowed, the login page told the user their email or password was wrong, even though they never typed one. Every such failure now shows "Sign-in with your identity provider failed. Please contact your administrator for assistance." The text is the same for every cause so it does not reveal which check failed, and the exact reason is still written to the server log as a warning.
Fixes#31627
The link in "chat finished" and "chat failed" webhook notifications was missing the `/c/` part of the chat address, so clicking it opened a 404 page. The link is sent as `/c/<chat id>` again and opens the chat.
Fixes#31565
In a temporary chat, when the model handed a task to a sub-agent, the sub-agent's conversation with the task and its answer was saved on the server, although a temporary chat should leave nothing behind. The model is no longer offered sub-agents in temporary chats.
Fixes#31567
When a sub-agent running in the background finished, or a timer the model set went off, the result and the model's follow-up reply were saved but did not show in the chat the user had open. They only appeared after a manual page reload. They now show in the open chat as soon as they arrive.
Fixes#31566
A tool call to an MCP server had no time limit, so a server that hung kept the chat waiting until a reverse proxy or the server itself closed the connection, even with AIOHTTP_CLIENT_TIMEOUT_TOOL_SERVER set. The call now stops after the configured number of seconds and the model sees the timeout as a tool error, which is how OpenAPI tool servers already behave. When the variable is unset it falls back to AIOHTTP_CLIENT_TIMEOUT, and with neither set (or a value of 0 or below) MCP tool calls still have no time limit.
Fixes#31640
A repeating event that was still running when the visible dates began was left out of the calendar. A weekly event from 23:00 to 01:00, for example, did not show on the following day, while the same event without a repeat did. Repeating events are now shown whenever any part of them falls within the visible dates.
Fixes#31605
With tool approval set to ask, the request sent to the model after approving a tool call left out the system prompt. In a compacted chat it also left out the conversation summary and sent the whole history again. The request after approval now has the same system prompt and compacted context as the one before it, plus the tool call and its result. The system prompt is picked the same way as for any other message: the chat Controls prompt, else your personal Settings prompt, else the admin default. A system prompt sent only in an API request is not kept by the server, so it is still missing after approval.
Fixes#31499
With ENABLE_ADMIN_CHAT_ACCESS turned off, opening another user's chat was refused, but through direct API requests an admin could still get the whole chat back in the reply to editing or deleting one of its messages, grant themselves read access in the chat's share settings, clone a chat someone shared privately with another user, or delete the chat. They could also send messages into it, attach it as context to their own chat, approve its tool calls, and list or stop its running replies. All of these are now refused for an admin on another user's chat, the same as opening it. With the setting on, admins keep full access as before.
Fixes#31413
With native function calling, citations produced by query_knowledge_files and query_chat_files never showed the relevance percentage badge, while the same knowledge base queried through classic RAG did.
The tools already return a distance per chunk, but the step that groups tool results into citation sources dropped it. Each grouped source now carries a distances list aligned with its documents, the same shape the classic RAG path emits, so the existing citation UI shows the badge without frontend changes. Chunks without a score (notes) leave the list empty, which the UI already treats as no score.
Fixes#29776
When the provider failed partway through a streaming request to the Anthropic Messages endpoint, the stream still ended like a normally finished answer, so Claude Code and the Anthropic SDKs took the cut-off text as complete. The stream now ends with an Anthropic error event, with the provider's error message if it sent one, so clients raise an error. Successful streams are unchanged.
Fixes#31403
Saving an arena model in Admin Settings > Models (for example to set default tools or capabilities) creates a model entry with the arena id. That entry replaced the arena model's metadata wholesale, dropping the access grants, model_ids and filter_mode configured in Admin Settings > Evaluations. From then on every non-admin user lost the arena model, even when it was public, and chats through it ignored the configured model pool.
The override now keeps those three keys from the evaluation config, which is where arena access and the model pool are managed. Everything else set in Settings > Models (tools, capabilities, description, profile image) still applies.
Verified end to end on base and patched: after the override a user sees and can chat with a public arena model (base: hidden, 400), private arena models stay hidden, and 20 admin chats all route to the configured pool (base: spread across all models).
Fixes#29564