With VECTOR_DB=elasticsearch, changing a file's content in a knowledge base added the new text but never removed the old chunks, so searches and chats kept returning the old text next to the new. The old chunks are now found and removed after the edit, the same as on the default store.
Fixes#31523
When a model called one tool, got its result and then called a second tool with no text in between, the next request merged both calls into an earlier assistant message that the provider had already seen. That changed the conversation's beginning, so the provider's prompt cache stopped matching from there for the rest of the chat. Each tool call and its result are now sent as their own messages, so the start of the conversation stays identical from one request to the next.
Fixes#31588
In a chat inside a folder with knowledge attached, the chat only searches the folder's knowledge, but a sub-agent it started could list and search every knowledge base the user can read. Sub-agents now get the folder's knowledge like the chat that started them. They already receive the folder's system prompt as part of the chat's instructions, so it is not added a second time.
Fixes#31569
* fix: audit log shows passwords that contain a double quote
With AUDIT_LOG_LEVEL set to REQUEST or REQUEST_RESPONSE, a password containing a double quote was only masked up to that quote, so a new password like Q"secret was logged as "********"secret. A request with whitespace before the colon, such as "new_password" : "secret", was not masked at all. Any field whose name ends in "password" is now masked through its closing quote in both cases.
* fix: audit log records passwords sent back in responses
With AUDIT_LOG_LEVEL set to REQUEST_RESPONSE, fields whose name ends in "password" were masked in request bodies, but response bodies were logged unmasked. Saving or opening the LDAP server settings therefore wrote the Application DN Password to the audit log in plain text, because the settings come back in the response, and the Jupyter passwords in the code execution settings leaked the same way. Responses now get the same masking as requests.
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Native hybrid search on Milvus kept a second, text-only collection beside every Milvus collection and searched both. Each Milvus collection now holds its vectors, its text and its BM25 keyword index together, and results rank exactly as before, with the BM25 weight setting working the same way it does on pgvector. Existing data moves on the first start with ENABLE_DB_MIGRATIONS on: startup waits while every collection is copied once (vectors included, nothing is re-embedded) and the originals are only dropped after every copy succeeded, so a failed run changes nothing and is retried on the next start. The copy needs free disk space for a second copy of the data until it finishes, and on one 16-thread machine with Milvus's official docker compose setup it ran at 11 to 22 MB/s, so 200 GB takes about 2.5 to 5 hours depending on chunk size. Milvus servers older than 2.5 are detected and keep the existing hybrid search.
Some pages, like the Ubiquiti tech specs pages, have more than one <main> element. The Playwright web loader only read the first one, so these pages came back as just their site menu and the actual content was lost. When a page has more than one, the loader now ignores those tags and reads the whole page. Pages with a single <main> load the same as before.
Fixes#28643
With hybrid search on, Milvus installs fetch every chunk of a collection and score BM25 in Python for each search. In both Milvus modes, one collection per knowledge base and multitenancy, Milvus now runs the keyword half itself with its built-in BM25 full-text search and merges it with the vector results, as pgvector already does. Milvus cannot add a BM25 index to an existing collection, so each collection gets a second, text-only collection next to it; existing installs build these once during startup when ENABLE_DB_MIGRATIONS is on, which copies the chunk text (extra storage roughly the size of that text) and leaves the original vectors and indexes untouched. On one standalone Milvus server the copy ran at about 11,000 chunks per second, around 40 minutes for 200 GB with 1536-dimension embeddings. Collections that cannot be copied, and Milvus servers older than 2.5, which have no BM25, keep using the existing hybrid search.
Fixes#26243
When files were attached to a chat message and one of them was already attached to that chat (on the same message or another one), none of the new files were recorded as part of the chat, and nothing showed up in the logs. The sender still saw every file in their own chat, but anyone opening a shared copy of the chat could not open the new ones. Files already attached to the chat are now skipped and the new ones are recorded normally.
Fixes#31648
Since 0.11.1, when a filter function wrote part of the reply before the model answered (for example some text and an HTML block), the artifact preview opened and then closed straight away. The page showed the filter text but was never sent it as part of the reply, so the model's first output overwrote it, the page briefly saw no HTML block and closed the preview. The filter text is now sent to the page before the model's output, so the preview stays open.
Fixes#31643
On an Amazon Bedrock OpenAI-compatible connection, GPT-5.6 and GPT-6 models have ids like `us.openai.gpt-6-sol` or `openai.gpt-6-luna`. These were not recognised as new OpenAI models, so `max_tokens` went upstream unchanged and Bedrock rejected it with a 400. Title and emoji generation failed on every chat, and any request with a token limit failed too. Setting `max_completion_tokens` by hand did not help, because non-OpenAI URLs convert it back to `max_tokens`.
`is_openai_new_model()` now drops a leading `openai.` or `<region>.openai.` (`us.`, `eu.`, `global.`, `us-gov.`) before matching, so these ids get the same handling as bare `gpt-5` ids. Ids that are not new models, such as `openai.gpt-oss-120b-1:0` and `gpt-4o`, are unchanged, and so is the LiteLLM `openai/` prefix.
Fixes#30510
Ejecting a loaded model from the model selector failed with a certificate error on llama.cpp and Ollama connections served over HTTPS with a self-signed or internal CA, even though chatting with the same connection worked. The unload request skipped the AIOHTTP_CLIENT_SESSION_SSL setting and always verified against the default system certificates, so setting it to a CA bundle or to false had no effect there. It now uses that setting, the same way chat requests to the connection already do.
Fixes#31371
MCP tool servers that sign in through Google, such as Google's hosted Gmail, Drive and Calendar servers, never got a refresh token, because Google only issues one when the sign-in explicitly asks for offline access. When the one-hour access token ran out the refresh failed and the connection was removed, so every user had to sign in again every hour. When the server's sign-in page is Google's, the sign-in now asks for offline access and a fresh consent, so the token renews on its own. Other providers get the same sign-in request as before, and existing Google connections pick this up the next time the user signs in.
Fixes#28319
Every member listed in a SCIM group response came back with "$ref": null, so identity providers had no link from a group member to that user's SCIM resource. Members now carry the URL of the user they point to, on both the single group and the group list responses.
Fixes#31525
An automation whose schedule ends after a fixed number of runs (COUNT in its RRULE) showed extra future runs in the Scheduled Tasks calendar after each run. With COUNT=3, the calendar kept showing three upcoming runs after the first and second run, although only the remaining ones execute. The calendar now counts runs from the schedule's own start date, the same way automations are actually run, so it shows exactly the runs that are still going to happen. The 5000-entry display limit now only counts entries inside the visible range, so very frequent schedules that started shortly before it no longer show up short or empty.
Fixes#31600
When a background sub-agent finished while the model was still writing the answer that started it, the sub-agent's report was placed as a reply to the previous answer instead. Once the answer ended, the chat switched to that other branch of the conversation, which hid the latest question and answer, and the model replied to the report without seeing that question and answer. The report now follows the answer that started the sub-agent (or the newest completed reply after it), so the conversation stays on one branch.
Fixes#31507
ENABLE_ORJSON has shipped as an option since v0.11.0 (2026-07-27), five releases and two months ago, and orjson is already installed with every instance. The only two problems ever found with it (rare line break characters splitting a stream, and extra encoding options being ignored) were fixed in v0.11.1 and nothing has come up since. The regression suite at https://github.com/open-webui/tests now runs 222 tests with the option on and off side by side, on SQLite, Postgres, several workers sharing one Redis with some on and some off, and in the browser: chats, completions for every provider format, tool calls, citations, all workspace and admin data, exports and imports, notes and live socket updates behave the same, and every API response is byte for byte identical. The only differences were in how non-English text gets saved to the database, where the standard encoder is the one with bugs (missed searches and too small size limits). Turning it on by default gives every instance the speedup measured in #27583 (live socket updates encode 17x and decode 3x faster), and ENABLE_ORJSON=false keeps the old encoder.
With ENABLE_OTEL, ENABLE_OTEL_TRACES and ENABLE_OTEL_LOGS all on, the collector received each log line twice, once with code location attributes and once without, because the logging instrumentation for traces now attaches its own log exporter next to Open WebUI's. It now keeps trace context on log lines without adding that second exporter, so each log line reaches the collector once, and OTEL_PYTHON_LOG_AUTO_INSTRUMENTATION=false is no longer needed as a workaround.
Fixes#31524
Choosing jinaai/jina-colbert-v2 as the reranking model failed on the current transformers release with "'HF_ColBERT' object has no attribute 'all_tied_weights_keys'", and saving the Documents settings quietly switched hybrid search back off. The ColBERT reranker now finishes loading, reranks search results and hybrid search stays on after saving.
Fixes#31522
With Weaviate as the vector database, a chunk identical to the query (distance 0) was treated as having no distance and got a relevance score of 0. Perfect matches could land at the bottom of the results or fall below the relevance threshold. They now score 1 as expected.
Fixes#31527
With WEBSOCKET_MANAGER=redis and several instances or workers, an instance only subscribed to Redis once a browser tab had connected to it. Until then, when a tool or Function asked the user something (a confirmation or an input dialog) and the user's tab was connected to another instance, the user's reply never reached the tool and it waited until it timed out. Every instance now subscribes at startup, so the reply arrives whichever instance the tab is on.
With ENABLE_ORJSON off (the default), non-English letters were saved to the database as escape codes, so "Ü" was stored as \u00dc. Searches that ignore upper and lower case compare against that saved text, so they missed any match that differs only in the case of a non-English letter: filtering models by the tag "Überblick" found nothing on Postgres, and searching automations for "отчёт" missed a prompt containing "Отчёт" on SQLite and Postgres. The 100,000 character size limit for user and chat variables counted the escape codes too, so Cyrillic or Chinese variables were refused as too large (or chat variables silently came out empty in the system prompt) at about a sixth of that size. Non-English text is now saved as written, which is how it is already saved with ENABLE_ORJSON on, so nothing changes for those instances, and the limit counts real characters. Anything saved before this keeps the escape codes until it is next edited.
With request auditing turned on, the audit log only masked fields named exactly "password". The new password from a password change, and passwords entered in admin settings such as YaCy or Jupyter, were written to the log as-is. Any field whose name ends in "password", in any letter case, is now replaced with asterisks.
When an image generation backend returns a link instead of the image itself, the download now goes through the same safety checks used for other external image downloads. Links on the configured ComfyUI address are still trusted as before, so a ComfyUI server on a local network keeps working.
When signing in through an OAuth/OIDC provider failed, for example because the provider denied access, the account had no email or its email domain was not allowed, the login page told the user their email or password was wrong, even though they never typed one. Every such failure now shows "Sign-in with your identity provider failed. Please contact your administrator for assistance." The text is the same for every cause so it does not reveal which check failed, and the exact reason is still written to the server log as a warning.
Fixes#31627
The link in "chat finished" and "chat failed" webhook notifications was missing the `/c/` part of the chat address, so clicking it opened a 404 page. The link is sent as `/c/<chat id>` again and opens the chat.
Fixes#31565
In a temporary chat, when the model handed a task to a sub-agent, the sub-agent's conversation with the task and its answer was saved on the server, although a temporary chat should leave nothing behind. The model is no longer offered sub-agents in temporary chats.
Fixes#31567
The person who started a direct message could add or remove people through the API, although the app only offers this in group channels. Someone added this way could read the whole earlier conversation, and because the original pair no longer matched the conversation, their next message opened a second, empty direct message. Changing the members of a direct message now answers with a 403.
Fixes#31570
When a sub-agent running in the background finished, or a timer the model set went off, the result and the model's follow-up reply were saved but did not show in the chat the user had open. They only appeared after a manual page reload. They now show in the open chat as soon as they arrive.
Fixes#31566
* fix: automation still shows "Last run Never" after "Run now"
Running an automation with "Run now" added the run to its history, but the automation page and the Automations list kept showing "Last run Never" (or the time of the last scheduled run). Only scheduled runs recorded a last run time. A manual run now records it too, so the automation page and the Automations list show the time of the run you just started.
Fixes#31580
* fix: show the new last run time right after Run now
The automation page now takes the automation the server returns after Run now, so the last run time updates on the spot and no reload is needed.
A tool call to an MCP server had no time limit, so a server that hung kept the chat waiting until a reverse proxy or the server itself closed the connection, even with AIOHTTP_CLIENT_TIMEOUT_TOOL_SERVER set. The call now stops after the configured number of seconds and the model sees the timeout as a tool error, which is how OpenAPI tool servers already behave. When the variable is unset it falls back to AIOHTTP_CLIENT_TIMEOUT, and with neither set (or a value of 0 or below) MCP tool calls still have no time limit.
Fixes#31640
A repeating event that was still running when the visible dates began was left out of the calendar. A weekly event from 23:00 to 01:00, for example, did not show on the following day, while the same event without a repeat did. Repeating events are now shown whenever any part of them falls within the visible dates.
Fixes#31605
With tool approval set to ask, the request sent to the model after approving a tool call left out the system prompt. In a compacted chat it also left out the conversation summary and sent the whole history again. The request after approval now has the same system prompt and compacted context as the one before it, plus the tool call and its result. The system prompt is picked the same way as for any other message: the chat Controls prompt, else your personal Settings prompt, else the admin default. A system prompt sent only in an API request is not kept by the server, so it is still missing after approval.
Fixes#31499
With tool approval set to ask, the result of an approved tool call was dropped from the chat once the reply finished, so the model no longer saw it in later turns. When the model then asked for a second tool, the first call went back to waiting for approval, and approving it again ran the tool a second time. Approved results now stay in the chat and each tool runs once.
Related to #31499
With ENABLE_ADMIN_CHAT_ACCESS turned off, opening another user's chat was refused, but through direct API requests an admin could still get the whole chat back in the reply to editing or deleting one of its messages, grant themselves read access in the chat's share settings, clone a chat someone shared privately with another user, or delete the chat. They could also send messages into it, attach it as context to their own chat, approve its tool calls, and list or stop its running replies. All of these are now refused for an admin on another user's chat, the same as opening it. With the setting on, admins keep full access as before.
Fixes#31413
Uploading or downloading a GGUF model to an Ollama connection sent neither the key nor the connection's custom headers, so it failed behind gateways such as Cloudflare Access and on servers that need a key. Unloading a model dropped the custom headers and sent the key as a Bearer token even with the authentication type set to None, for Ollama and llama.cpp connections alike. These requests now use the connection's headers and authentication type the same as chatting and the Manage Ollama dialog already do. Follow-up to #31489.
Listing, pulling, creating, copying and deleting models from the Manage Ollama dialog ignored the connection's custom headers and authentication type, so the dialog failed behind gateways such as Cloudflare Access and sent the key as a Bearer token even with the authentication type set to None. Checking a single connection's version sent no key at all. All of these, and the other requests to an Ollama connection such as text generation and embeddings, now use the connection's headers and authentication type, matching what verifying the connection and chatting already do.
Fixes#31487
With native function calling, citations produced by query_knowledge_files and query_chat_files never showed the relevance percentage badge, while the same knowledge base queried through classic RAG did.
The tools already return a distance per chunk, but the step that groups tool results into citation sources dropped it. Each grouped source now carries a distances list aligned with its documents, the same shape the classic RAG path emits, so the existing citation UI shows the badge without frontend changes. Chunks without a score (notes) leave the list empty, which the UI already treats as no score.
Fixes#29776
When the provider failed partway through a streaming request to the Anthropic Messages endpoint, the stream still ended like a normally finished answer, so Claude Code and the Anthropic SDKs took the cut-off text as complete. The stream now ends with an Anthropic error event, with the provider's error message if it sent one, so clients raise an error. Successful streams are unchanged.
Fixes#31403
Tavily web search always ran at Tavily's default depth (basic), because the search request never sent `search_depth`. The only Tavily depth control in Admin > Settings > Web Search, "Tavily Extract Depth", applies to the Extract API used by the web loader, never to search.
This adds `TAVILY_SEARCH_DEPTH` (env var and persisted setting, default `basic`) and a "Tavily Search Depth" select (ultra-fast, fast, basic, advanced) under the Tavily search engine settings. The value is sent as `search_depth` on every Tavily search request, so admins can set search and extract depth independently, for example fast search with advanced extraction.
The default matches Tavily's own default, so existing setups keep the same behaviour until the setting is changed.
Fixes#29891
Saving an arena model in Admin Settings > Models (for example to set default tools or capabilities) creates a model entry with the arena id. That entry replaced the arena model's metadata wholesale, dropping the access grants, model_ids and filter_mode configured in Admin Settings > Evaluations. From then on every non-admin user lost the arena model, even when it was public, and chats through it ignored the configured model pool.
The override now keeps those three keys from the evaluation config, which is where arena access and the model pool are managed. Everything else set in Settings > Models (tools, capabilities, description, profile image) still applies.
Verified end to end on base and patched: after the override a user sees and can chat with a public arena model (base: hidden, 400), private arena models stay hidden, and 20 admin chats all route to the configured pool (base: spread across all models).
Fixes#29564