When a model wraps its answer in <|begin_of_solution|> and <|end_of_solution|>, only the opening marker was removed. The closing marker stayed visible in the reply and was saved with the message, and anything the model wrote after it was glued onto the answer. Now both markers are removed and text after the answer shows up as a normal part of the reply.
Fixes#31434
Sometimes a reply where the model used tools gets saved with a tool result that no call in that reply asked for, or with a tool call that never got its result. The chat history was then sent to the provider unchanged, Anthropic and Bedrock rejected it, and every following message in that chat failed until the user deleted the broken reply. Now each tool call is only kept together with its own result from the same reply, and the unmatched calls and results are left out of what gets sent to the model. The chat itself is not changed, and correctly saved chats are sent exactly as before.
Fixes#28937
When an API request to an Ollama model set max_tokens, Open WebUI passed it on in a place Ollama does not read, so Ollama ignored it and replies ran to full length. The limit now reaches Ollama as its own output length setting, so replies stop at the requested length. It also wins over a max_tokens value saved in the model's advanced parameters, as the API docs describe. Chats in the web UI were not affected, since their limit already reached Ollama correctly.
Fixes#31432
With reasoning models, when the first part of the answer arrived together with the end of the thinking block and ended with a space, that space went missing, so "The answer is 4." was shown and saved as "The answeris 4.". That space is now kept.
Fixes#31435
With a connection set to the Responses API, when the provider reported a reply as failed, the error showed while streaming but was gone after a reload, leaving an empty reply. Some other provider errors never showed up at all, not even while streaming. Both kinds of error now show up and are still there after a reload, the same as on Chat Completions connections.
Fixes#31433
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
On the default SQLite setup (session sharing off), a POST to /api/v1/models/sync containing any model that already exists answered 200 with an empty list and stored nothing. After a 5 second stall, the only trace was "database is locked" in the server log.
The sync wrote each model's access grants while the model update was still uncommitted. Without session sharing the grant writes run on a second database session, and SQLite allows one writer at a time, so that write waited on the same request's uncommitted update until the busy timeout expired and the whole sync was dropped.
Grants are now written after the model changes are committed, the same order model create and update already use. A failed model commit now also leaves every grant untouched. PostgreSQL and setups with session sharing on behave as before.
Fixes#31346
With ENABLE_OAUTH_GROUP_CREATION on, every group in a user's OAuth claim was created on login, including groups matching OAUTH_BLOCKED_GROUPS. Membership sync already ignored those groups, so the result was empty groups nobody could join. With IdPs that send a user's full directory membership (Keycloak backed by LDAP/AD), one login could fill the group table with thousands of them.
Group creation now applies the same blocklist check as the membership add and remove steps, so a blocked group is never created, joined or left through OAuth. Groups that are not blocked are created as before.
Fixes#29558
Korean EUC-KR and Japanese Shift-JIS text files were stored as garbled Chinese characters, so retrieval, knowledge bases and the model context all worked on text that is not in the file.
Encoding detection puts chardet's guess in front of a fixed GB18030, Big5, EUC-KR, EUC-JP try order, and GB18030 decodes almost any double-byte text without an error. The guess map was written for chardet 5. Since the bump to chardet 7 in v0.10.0, Korean text is reported as CP949 and Japanese text as cp932 or SHIFT_JIS, which the map either did not know or dropped because the codec was not in the try order, so these files fell through to GB18030.
The map now covers CP949 and cp932, and a mapped guess is always tried first. SHIFT_JIS maps to cp932, the Windows superset, because chardet also reports SHIFT_JIS for ordinary Japanese files containing characters such as ① or ㈱ that plain Shift-JIS cannot decode; this is the same subset-to-superset rule the map already applies to GB2312.
Korean and Japanese files now decode correctly, and Chinese, EUC-JP, UTF-8 and Western files decode as before. The one trade-off of trusting the guess: chardet 7 labels some files holding only a few Chinese characters (a short label or a one-line comment) as CP949, and those now read as Korean. No regressions were found in files with more Chinese text than that.
Fixes#31352
Fixes#31340
When the attached Open Terminal runs on Windows, the AGENTS.md in its home directory was never handed to the model. The home check only accepted POSIX absolute paths, so a drive-letter or UNC home such as C:\ProgramData\OpenTerminal\inst was treated as invalid and the file was skipped without any log line.
The check now also accepts Windows absolute paths. Relative and drive-relative homes are still skipped, and POSIX homes send byte-identical requests.
The file path keeps its forward-slash join. Open Terminal normalises the path on the host, so C:\Users\bob/AGENTS.md opens C:\Users\bob\AGENTS.md. Picking ntpath.join for Windows homes would give native separators on the wire but adds a second branch for no change in which file gets read.
Verified against Open Terminal's own path resolution with Windows semantics for drive-letter, forward-slash, drive-root, trailing-backslash and UNC homes, over both the backend request and the browser direct-connection path.
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Fixes#31348
With title generation turned off (per user or by the admin), the chat is saved with the first user message as its title, but the live title event sent the assistant message instead. A new chat therefore kept showing "New Chat" in the header and tab until a reload, because the assistant message is still empty at that point. In a note's chat panel the whole model reply showed up as the title.
The fallback branch now sends the title it just saved, the same way the generated-title branch above it already does. The sidebar was never affected because it reloads titles from the database.
On a phone with UI Scale at 1.2x or higher, the model selector in the chat composer covers the Integrations button, so it cannot be seen or tapped. From 1.3x it covers the + button too, which leaves no way to attach files or open integrations.
The model selector's width cap is rem based, so it grows with UI Scale, and the left button group could shrink to nothing, so the selector took the space first. The + and Integrations buttons now sit in a group that keeps its width, and only the toggled chips stay in the scrolling strip. When space runs out the model name truncates, and chips scroll as before.
Measured in the real app at a 360px mobile viewport: before, Integrations is covered at 1.2x and both buttons at 1.3x and 1.5x; after, both are tappable at every scale, with the model name truncated (114px at 1.3x). Button positions at 1x and on desktop are unchanged to the pixel, with and without chips.
Fixes#29989
In voice mode, pressing M after a turn typed an "m" into the chat box instead of muting. After each voice message the chat input pulled keyboard focus back to itself, and the call overlay ignores M while focus is in a text field so typing still works.
The chat input now leaves focus alone while a call is open: clearing it after a send, the remount on the first message of a new chat and the editor's autofocus all skip focusing during a call. Clicking into the box and typing during a call still works, and outside a call the input focuses exactly as before.
Checked in a browser with a fake microphone: M mutes after the first, second and third turn, in new and existing chats; "m" typed into a focused input during a call still types.
Fixes#30406
When a ComfyUI workflow ends in the core "Save Image (Advanced)" node, ComfyUI finishes the job and saves the image, but Open WebUI returns an empty result, so the chat shows nothing. Image editing workflows such as the Qwen Image Edit template use this node by default.
Open WebUI only collects images from output nodes of type SaveImage and PreviewImage. This adds SaveImageAdvanced to that list. The node reports its files in the same format as SaveImage, so the rest of the download and storage path works unchanged. Generation and editing share this code, so both are fixed.
Fixes#30404
When an MCP tool returns an image or audio item, the file is uploaded to storage, but the whole MCP item, including its base64 payload, was also passed as upload metadata. That metadata is persisted in the file table's meta column, so every such result was stored twice: once in storage and once as base64 in the database, growing the DB and every file query that loads meta.
The MCP path now passes only chat_id, message_id and session_id, the same metadata the non-MCP tool image path already stores. Nothing reads the removed key.
Fixes#30411
MCP tool servers using OAuth 2.1 with dynamic client registration authorized without any scope when the authorization server left scope out of its registration response, which RFC 7591 allows (Atlassian and Notion do). Consent completed and the tool showed as connected, but the issued token lacked the scopes the resource requires, so every tool call was refused. Discovered scopes and the custom OAuth Scopes field were both affected.
The stored client now falls back to the scope sent in the registration request when the response has none. A scope the server does return is kept as is.
Connections registered before this fix already have a null scope stored. The protected resource metadata recovery that static-credential clients already use now also runs for dynamically registered clients, so those connections pick up the discovered scopes on the next load without registering again.
Fixes#29967
On iOS and iPadOS, auto-playback of a finished reply is never heard and the speaker button stays stuck in "speaking", so the first tap only stops a playback that never started. WebKit rejects play() started from a network event, and the audio queue ignored that rejection, so its state never reset. Call mode on the same devices was silent too.
A rejected play() now resets the queue, returns the button to idle and shows a toast (an aborted play from stop or a message switch is ignored). The first user tap or keypress plays a 10 ms silent clip on the shared audio element, which WebKit then allows to play later without a gesture; if that attempt fails it retries on the next gesture. Call mode plays unmuted: WebKit pauses an element that is unmuted after play() outside a gesture, even once unlocked.
The unlock and call mode change follow the reporter's on-device tests (iPhone iOS 27, iPad iPadOS 26.6.2). Verified in Chromium with autoplay restricted: the base queue wedges and drops later chunks; with the fix it reports the error, recovers, the unlock plays once and never interrupts audio already playing, and chunks queued during the unlock still play.
Fixes#30262
Adding a second model to a chat turned off Web Search, Image Generation and Code Interpreter even when every selected model has them as Default Features, so the comparison ran without them. Changing the selection clears the feature toggles and then reapplies the model defaults, but defaults were only reapplied for a single model.
In compare mode a default feature is now turned on when every selected model supports it and has it as a default, since the toggles are shared by all models in the comparison. It only ever turns features on, so flags passed in the URL (?web-search=true) are not overwritten. Single-model defaults are unchanged.
The input reset now waits for the model-selection update to finish before applying defaults. Before, the Web Search value actually sent could disagree with the toggle on screen: a single model with Web Search on by default showed the toggle on but sent it off, and switching to a model without it showed it off but still sent it on (see #29326). The toggle and the request now match.
Fixes#30310
Any Python code that only mentioned "matplotlib" (a comment, a string, or an importlib.util.find_spec("matplotlib") check) failed with ModuleNotFoundError before a single line of it ran. The plt.show() patch was applied on a plain substring match, but matplotlib is only installed when the code actually imports it, so the patch's own import crashed the run. The user's try/except could not catch it, because their code never started. The same thing broke the code editor's Python formatter on any code that imports matplotlib, since that run only installs black.
The patch now also requires matplotlib to be present in the runtime's loaded packages, in both the worker and the sandboxed iframe host. Real matplotlib code still gets inline PNG output from plt.show(), and code that merely mentions matplotlib runs unchanged.
Gating on the callers' import regex instead was also tested and still fails the formatter case, because there the code containing the import sits inside a string while only black is installed.
Fixes#29894
Automations lost their model's tool bindings (including MCP servers), default features such as web search, default filters and terminal on the first run after a restart. The model then answered that it had no tools. Later runs and a manual Regenerate worked. Automations on the base models cache were not affected.
The run read the model from app.state.MODELS before anything had loaded it. After a restart or a connection settings save, that cache stays empty until a browser loads the model list or a chat completion runs. The completion runs only after the automation has already built its request.
execute_automation now loads the models when the cache is empty, using the same guard chat_completion uses, before either the chat or the channel target reads it. This also fixes channel automations showing the raw model ID in place of the model name on a cold cache.
Verified end to end on a restarted instance with a mock upstream: before, both chat and channel runs reached the pipeline without tool_ids or features. After, both carry the model's tools and web search, and the upstream receives the tool.
Fixes#27694
Asked for a PowerPoint or Word file, the model had no library for it, so it followed the prompt's "use an alternative approach" and wrote the OOXML zip by hand. The sandbox reported success and Office refused to open the result.
python-pptx and python-docx are now vendored the same way as openpyxl: their wheels (plus xlsxwriter) go through the PyPI wheel path, lxml joins the Pyodide distribution list so its wasm wheel is cached, and importing pptx or docx installs them from the bundled wheels. No prompt change is needed, since the app installs on import.
static/pyodide grows by about 3 MB (58 to 61 MB) and verifyBundledWheels() passes.
Verified in headless Chromium with pypi.org, files.pythonhosted.org and the jsDelivr CDN blocked: both packages install from the local wheels only, and a deck and a document built in the sandbox reopen with the native libraries. openpyxl, seaborn, black, pandas, matplotlib and requests still install offline. Without the change both installs fail offline.
Fixes#30361
Attach Webpage fails with 403 Forbidden on sites that reject the bare aiohttp user agent, Wikipedia among them, even when USER_AGENT is set. The web loader sends USER_AGENT, but the request that runs first to decide whether the URL is a page or a file does not, so the attachment fails before the loader is ever reached.
The pre-check now sends USER_AGENT as the request User-Agent when it is set. With it unset the request is unchanged and keeps the aiohttp default.
Verified against the real _fetch_url with https://en.wikipedia.org/wiki/OpenAI: 403 before, page detected after; USER_AGENT unset still returns the same 403 as before, and a direct PDF URL is still detected as a file.
Fixes#29617
On Admin Settings > Models the drag handle was disabled as soon as a search, view or tag filter was active, so on a long list the only way to move a model was to clear everything and hunt for it by eye.
Dragging now works in any filtered list. The move is applied to the full order: the dragged model is placed directly after the visible model it was dropped below (or directly before the one it was dropped above), and every model hidden by the filter keeps its place. That anchoring is what makes reordering a subset safe, which is why the filters previously blocked it.
The tag filter now filters the loaded list client-side like search and view already do. Before, it reloaded the page data and rebuilt the order from only the tagged models, so saving under a tag would have dropped every other model from the order, and switching tags discarded unsaved moves. Export keeps its existing tag-filtered behaviour.
Verified in the browser with search, view (enabled/disabled) and tag filters, dragging up and down, multiple moves before one save and switching tags with unsaved moves; the saved order always contains every model.
Closes#29634
#30426 stopped concurrent requests from refreshing the same OAuth session twice, but its lock only lives inside one process. With several uvicorn workers or replicas, two requests on different workers still send the same refresh token, a rotating provider rejects the second with invalid_grant, and the session gets deleted, so the user's OAuth session is logged out again.
When Redis is configured, which multi-worker and multi-replica deployments require, the refresh now takes a Redis lock per session instead of the in-process one. Single-process deployments without Redis keep the in-process lock. The waiter re-reads the session inside the lock as before and uses the token that was just stored.
It uses redis-py's own async lock because the existing RedisLock is synchronous and never waits. The Sentinel proxy now passes `lock` through unwrapped like `pipeline` and `pubsub`; otherwise it returned a coroutine and every refresh behind Sentinel would fail.
Tested with separate OS processes on one sqlite DB, a real Redis and a rotating mock provider: 2 and 5 processes (and 5 processes x 3 requests) now cause 1 refresh, every caller gets the new token and the session is kept (before: one refresh per process, session deleted every run). Single refresh, failed refresh, valid token and the single-process path without Redis are unchanged.
Follow-up to #30426, refs #30416