Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Native hybrid search on Milvus kept a second, text-only collection beside every Milvus collection and searched both. Each Milvus collection now holds its vectors, its text and its BM25 keyword index together, and results rank exactly as before, with the BM25 weight setting working the same way it does on pgvector. Existing data moves on the first start with ENABLE_DB_MIGRATIONS on: startup waits while every collection is copied once (vectors included, nothing is re-embedded) and the originals are only dropped after every copy succeeded, so a failed run changes nothing and is retried on the next start. The copy needs free disk space for a second copy of the data until it finishes, and on one 16-thread machine with Milvus's official docker compose setup it ran at 11 to 22 MB/s, so 200 GB takes about 2.5 to 5 hours depending on chunk size. Milvus servers older than 2.5 are detected and keep the existing hybrid search.
Some pages, like the Ubiquiti tech specs pages, have more than one <main> element. The Playwright web loader only read the first one, so these pages came back as just their site menu and the actual content was lost. When a page has more than one, the loader now ignores those tags and reads the whole page. Pages with a single <main> load the same as before.
Fixes#28643
With hybrid search on, Milvus installs fetch every chunk of a collection and score BM25 in Python for each search. In both Milvus modes, one collection per knowledge base and multitenancy, Milvus now runs the keyword half itself with its built-in BM25 full-text search and merges it with the vector results, as pgvector already does. Milvus cannot add a BM25 index to an existing collection, so each collection gets a second, text-only collection next to it; existing installs build these once during startup when ENABLE_DB_MIGRATIONS is on, which copies the chunk text (extra storage roughly the size of that text) and leaves the original vectors and indexes untouched. On one standalone Milvus server the copy ran at about 11,000 chunks per second, around 40 minutes for 200 GB with 1536-dimension embeddings. Collections that cannot be copied, and Milvus servers older than 2.5, which have no BM25, keep using the existing hybrid search.
Fixes#26243
When files were attached to a chat message and one of them was already attached to that chat (on the same message or another one), none of the new files were recorded as part of the chat, and nothing showed up in the logs. The sender still saw every file in their own chat, but anyone opening a shared copy of the chat could not open the new ones. Files already attached to the chat are now skipped and the new ones are recorded normally.
Fixes#31648
iPhone photos were only converted to JPEG when the browser reported their type as exactly image/heic. Firefox reports image/heif and some browsers report no type at all, so those photos were uploaded as HEIC and vision models failed with an error. When conversion did run, the JPEG was still uploaded with the HEIC file name and type, so the model was told it was HEIC anyway. HEIC and HEIF photos are now converted whatever type the browser reports and are uploaded as a .jpg with the JPEG type, in chats, channels and notes.
Fixes#28411
Since 0.11.1, when a filter function wrote part of the reply before the model answered (for example some text and an HTML block), the artifact preview opened and then closed straight away. The page showed the filter text but was never sent it as part of the reply, so the model's first output overwrote it, the page briefly saw no HTML block and closed the preview. The filter text is now sent to the page before the model's output, so the preview stays open.
Fixes#31643
When a reply contained an artifact and the preview opened automatically, it could check for artifacts before they had been picked up from that reply, find none and close again. This happens occasionally in normal chats, with or without filters, and is older than 0.11.1. The preview now looks for artifacts again at the moment it opens, so it stays open. It is a separate cause from #31643, found while looking into that issue.
When someone mentioned a user or channel in a channel message, the notification toast and the browser notification showed the raw mention markup, including the internal id. They now show the mention the same way the message does in the channel, for example "@Alex".
Fixes#31586
On an Amazon Bedrock OpenAI-compatible connection, GPT-5.6 and GPT-6 models have ids like `us.openai.gpt-6-sol` or `openai.gpt-6-luna`. These were not recognised as new OpenAI models, so `max_tokens` went upstream unchanged and Bedrock rejected it with a 400. Title and emoji generation failed on every chat, and any request with a token limit failed too. Setting `max_completion_tokens` by hand did not help, because non-OpenAI URLs convert it back to `max_tokens`.
`is_openai_new_model()` now drops a leading `openai.` or `<region>.openai.` (`us.`, `eu.`, `global.`, `us-gov.`) before matching, so these ids get the same handling as bare `gpt-5` ids. Ids that are not new models, such as `openai.gpt-oss-120b-1:0` and `gpt-4o`, are unchanged, and so is the LiteLLM `openai/` prefix.
Fixes#30510
Tools that return inline HTML showed no embed at all when the HTML contained an entity such as `"` or `"`, and HTML containing `&` or `<` showed up with those turned into real characters. The chat view unescaped HTML entities in a tool's embeds, arguments and file links, which is only right for chats saved by older versions, so on current chats it changed the tool's own HTML before displaying it. Embeds, tool arguments and file links now show exactly what the tool returned, and older chats render as before.
Fixes#28085
Ejecting a loaded model from the model selector failed with a certificate error on llama.cpp and Ollama connections served over HTTPS with a self-signed or internal CA, even though chatting with the same connection worked. The unload request skipped the AIOHTTP_CLIENT_SESSION_SSL setting and always verified against the default system certificates, so setting it to a CA bundle or to false had no effect there. It now uses that setting, the same way chat requests to the connection already do.
Fixes#31371
After a reply, the first suggested follow-up question sits as grey ghost text in the empty chat input. Typing into it with an input method (Chinese, Japanese or Korean) broke the word being composed: on iOS the input lost focus after the first character, and in Chromium the first letter was left behind, so typing "ni" and picking "你" gave "n你". Clearing the ghost text rebuilt the input's line of text while the keyboard was still composing the word. The ghost text is now drawn over the empty input without being written into it, so input methods work again, and Tab to accept it and the follow-up buttons under the reply behave as before.
Fixes#31372
MCP tool servers that sign in through Google, such as Google's hosted Gmail, Drive and Calendar servers, never got a refresh token, because Google only issues one when the sign-in explicitly asks for offline access. When the one-hour access token ran out the refresh failed and the connection was removed, so every user had to sign in again every hour. When the server's sign-in page is Google's, the sign-in now asks for offline access and a fresh consent, so the token renews on its own. Other providers get the same sign-in request as before, and existing Google connections pick this up the next time the user signs in.
Fixes#28319
Every member listed in a SCIM group response came back with "$ref": null, so identity providers had no link from a group member to that user's SCIM resource. Members now carry the URL of the user they point to, on both the single group and the group list responses.
Fixes#31525
An automation whose schedule ends after a fixed number of runs (COUNT in its RRULE) showed extra future runs in the Scheduled Tasks calendar after each run. With COUNT=3, the calendar kept showing three upcoming runs after the first and second run, although only the remaining ones execute. The calendar now counts runs from the schedule's own start date, the same way automations are actually run, so it shows exactly the runs that are still going to happen. The 5000-entry display limit now only counts entries inside the visible range, so very frequent schedules that started shortly before it no longer show up short or empty.
Fixes#31600
With the Artifacts pane open, the Copy Last Code Block shortcut put the pane's full HTML document on the clipboard instead of the last code block in the chat. The pane opens on its own whenever a reply contains an HTML block, so the shortcut was wrong in every such chat.
The shortcut clicks the last Copy button of chat code blocks on the page, and the Artifacts Copy button carried the same marker class. The pane renders after the messages, so it always won. The marker class is now removed from the Artifacts button. It had no styling attached, so the button looks and works as before.
Fixes#31476
When a background sub-agent finished while the model was still writing the answer that started it, the sub-agent's report was placed as a reply to the previous answer instead. Once the answer ended, the chat switched to that other branch of the conversation, which hid the latest question and answer, and the model replied to the report without seeing that question and answer. The report now follows the answer that started the sub-agent (or the newest completed reply after it), so the conversation stays on one branch.
Fixes#31507
ENABLE_ORJSON has shipped as an option since v0.11.0 (2026-07-27), five releases and two months ago, and orjson is already installed with every instance. The only two problems ever found with it (rare line break characters splitting a stream, and extra encoding options being ignored) were fixed in v0.11.1 and nothing has come up since. The regression suite at https://github.com/open-webui/tests now runs 222 tests with the option on and off side by side, on SQLite, Postgres, several workers sharing one Redis with some on and some off, and in the browser: chats, completions for every provider format, tool calls, citations, all workspace and admin data, exports and imports, notes and live socket updates behave the same, and every API response is byte for byte identical. The only differences were in how non-English text gets saved to the database, where the standard encoder is the one with bugs (missed searches and too small size limits). Turning it on by default gives every instance the speedup measured in #27583 (live socket updates encode 17x and decode 3x faster), and ENABLE_ORJSON=false keeps the old encoder.
With ENABLE_OTEL, ENABLE_OTEL_TRACES and ENABLE_OTEL_LOGS all on, the collector received each log line twice, once with code location attributes and once without, because the logging instrumentation for traces now attaches its own log exporter next to Open WebUI's. It now keeps trace context on log lines without adding that second exporter, so each log line reaches the collector once, and OTEL_PYTHON_LOG_AUTO_INSTRUMENTATION=false is no longer needed as a workaround.
Fixes#31524
Starting a new chat from the search dialog with "Start a new conversation" sent a different message than the one typed: everything from a # or & onwards was dropped and + turned into a space. The typed text now reaches the new chat unchanged.
Fixes#31469
Choosing jinaai/jina-colbert-v2 as the reranking model failed on the current transformers release with "'HF_ColBERT' object has no attribute 'all_tied_weights_keys'", and saving the Documents settings quietly switched hybrid search back off. The ColBERT reranker now finishes loading, reranks search results and hybrid search stays on after saving.
Fixes#31522
With Weaviate as the vector database, a chunk identical to the query (distance 0) was treated as having no distance and got a relevance score of 0. Perfect matches could land at the bottom of the results or fall below the relevance threshold. They now score 1 as expected.
Fixes#31527
Since v0.11.4 a new visitor always got their browser's language even when an admin set DEFAULT_LOCALE, because the browser language was remembered as if the user had picked it, so the configured default never applied. The configured default now applies on the first visit again, as it did up to v0.11.3. A language the user picks in Settings and a ?lang= link still take precedence, and instances without DEFAULT_LOCALE keep following the browser language.
Fixes#31548