Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
With ENABLE_CUSTOM_MODEL_FALLBACK on, a workspace model whose base model is gone should be answered by the first default model. That worked for plain API calls, but every chat sent from the web UI failed with "Model not found" for users and "Model '' was not found" for admins, and no model was called.
Web UI chats carry a chat id and a socket session, so the request is split into one task per selected model. Each task was rebuilt with the originally requested model id, which dropped the fallback chosen earlier in the handler. The task for the requested model now keeps the fallback model when one was chosen. The chat still records the workspace model the user picked.
Tested end to end against a mock upstream, as user and admin, in new and existing chats: before, every web UI send with such a model errored; after, the default model answers. Healthy models, workspace models with a valid base and multi-model sends behave as before, and with the fallback disabled the chat still fails with "Model not found".
Fixes#31345
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Ejecting a loaded model from the model selector failed with a certificate error on llama.cpp and Ollama connections served over HTTPS with a self-signed or internal CA, even though chatting with the same connection worked. The unload request skipped the AIOHTTP_CLIENT_SESSION_SSL setting and always verified against the default system certificates, so setting it to a CA bundle or to false had no effect there. It now uses that setting, the same way chat requests to the connection already do.
Fixes#31371
ENABLE_ORJSON has shipped as an option since v0.11.0 (2026-07-27), five releases and two months ago, and orjson is already installed with every instance. The only two problems ever found with it (rare line break characters splitting a stream, and extra encoding options being ignored) were fixed in v0.11.1 and nothing has come up since. The regression suite at https://github.com/open-webui/tests now runs 222 tests with the option on and off side by side, on SQLite, Postgres, several workers sharing one Redis with some on and some off, and in the browser: chats, completions for every provider format, tool calls, citations, all workspace and admin data, exports and imports, notes and live socket updates behave the same, and every API response is byte for byte identical. The only differences were in how non-English text gets saved to the database, where the standard encoder is the one with bugs (missed searches and too small size limits). Turning it on by default gives every instance the speedup measured in #27583 (live socket updates encode 17x and decode 3x faster), and ENABLE_ORJSON=false keeps the old encoder.
With WEBSOCKET_MANAGER=redis and several instances or workers, an instance only subscribed to Redis once a browser tab had connected to it. Until then, when a tool or Function asked the user something (a confirmation or an input dialog) and the user's tab was connected to another instance, the user's reply never reached the tool and it waited until it timed out. Every instance now subscribes at startup, so the reply arrives whichever instance the tab is on.
With tool approval set to ask, the result of an approved tool call was dropped from the chat once the reply finished, so the model no longer saw it in later turns. When the model then asked for a second tool, the first call went back to waiting for approval, and approving it again ran the tool a second time. Approved results now stay in the chat and each tool runs once.
Related to #31499
With ENABLE_ADMIN_CHAT_ACCESS turned off, opening another user's chat was refused, but through direct API requests an admin could still get the whole chat back in the reply to editing or deleting one of its messages, grant themselves read access in the chat's share settings, clone a chat someone shared privately with another user, or delete the chat. They could also send messages into it, attach it as context to their own chat, approve its tool calls, and list or stop its running replies. All of these are now refused for an admin on another user's chat, the same as opening it. With the setting on, admins keep full access as before.
Fixes#31413
Uploading or downloading a GGUF model to an Ollama connection sent neither the key nor the connection's custom headers, so it failed behind gateways such as Cloudflare Access and on servers that need a key. Unloading a model dropped the custom headers and sent the key as a Bearer token even with the authentication type set to None, for Ollama and llama.cpp connections alike. These requests now use the connection's headers and authentication type the same as chatting and the Manage Ollama dialog already do. Follow-up to #31489.
MCP tool servers using OAuth 2.1 with dynamic client registration authorized without any scope when the authorization server left scope out of its registration response, which RFC 7591 allows (Atlassian and Notion do). Consent completed and the tool showed as connected, but the issued token lacked the scopes the resource requires, so every tool call was refused. Discovered scopes and the custom OAuth Scopes field were both affected.
The stored client now falls back to the scope sent in the registration request when the response has none. A scope the server does return is kept as is.
Connections registered before this fix already have a null scope stored. The protected resource metadata recovery that static-credential clients already use now also runs for dynamically registered clients, so those connections pick up the discovered scopes on the next load without registering again.
Fixes#29967
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
When title generation for a brand new chat fails, the chat silently keeps the provisional title (the full first user message) and nothing is written to the log at the default level, so there is no way to tell that the feature is broken rather than disabled. The failure was caught by a broad `except Exception` and reported with `log.debug`, which is invisible unless GLOBAL_LOG_LEVEL is set to DEBUG. Issue #29533 describes an outage of this path that survived four releases for exactly that reason.
This logs it with `log.exception` instead, matching how the rest of main.py reports background task failures, so an admin sees one ERROR line plus the traceback naming the real cause.
Nothing else changes: the success path is untouched, the exception is still swallowed so the detached task cannot take anything down with it, and `background_tasks_handler` does not log this exception itself, so there is no duplicate traceback.
Fixes#29533
GET /api/version/updates returned the running version as latest whenever
the GitHub request failed, so an instance that cannot reach GitHub
reported itself up to date however far behind it was. The exception was
logged at debug, below the default level, so nothing recorded that the
check never happened.
The failure path now returns latest: None and logs at warning.
A null latest cannot be passed to compareVersion as it stood.
current.localeCompare(null) coerces to the string "null", and "0.10.2"
sorts before it, so the function returned true. The backend change alone
would have turned a false (latest) into a false update-available plus a
toast, so the guard is part of the fix.
The three callers stop substituting the running version in their catch,
and the two badge surfaces gain a third state. When latest is unknown
the badge is plain text, since there is no release to link to.
Admin Settings > General was wrong in a worse way than reported: it
initialised updateAvailable false with latest set to the running
version, and never checked on mount, so it claimed (latest) having made
no request at all. It now matches About.svelte, which starts unknown and
checks on mount.
Closes#29580
On Windows hosts the built-in code interpreter fails immediately with "Failed to fetch dynamically imported module: .../pyodide/pyodide.asm.mjs", and the browser console shows the server answered with a MIME type of "text/plain". Code execution is unusable for those users.
Python's mimetypes module reads the Windows registry after loading its own table, so a stray registry entry silently replaces the correct type for an extension and Starlette then labels the file with it. Browsers enforce strict MIME checking for module scripts and streaming WASM compilation, so the pyodide loader gets refused. The same workaround already existed for .js; this extends it to the two other extensions pyodide ships, and moves it out of the frontend-build branch so the unconditionally mounted /static assets are covered as well.
Fixes#29133
* fix: use the pooled client timeout for the Anthropic Messages passthrough
The native `/api/v1/messages` passthrough still referenced `openai.AIOHTTP_CLIENT_TIMEOUT`, which stopped existing when `routers/openai.py` moved onto `session_pool.get_client_timeout()`. Every passthrough request therefore raised `AttributeError: module 'open_webui.routers.openai' has no attribute 'AIOHTTP_CLIENT_TIMEOUT'` before it was sent, and the surrounding handler turned that into a 502 "Open WebUI: Server Connection Error", so Anthropic-format clients such as Cline could not reach any model at all.
Use `get_client_timeout(stream=...)` like the OpenAI and Ollama proxies do, so the configured `AIOHTTP_CLIENT_TIMEOUT` applies and streaming requests additionally get the idle-read timeout.
Fixes#27595
* fix: authenticate native Anthropic requests with x-api-key
The Anthropic Messages passthrough and the token-count forwarding both build their upstream request through `get_anthropic_request_target`, which sends the connection key as `Authorization: Bearer <key>`. Anthropic's OpenAI-compatible `/chat/completions` endpoint accepts that, which is why the model works in the chat UI, but the native `/v1/messages` and `/v1/messages/count_tokens` endpoints do not: they require the key in `x-api-key` and reject a bearer token with 401 `Invalid bearer token` (and `jwt auth is not yet supported on count_tokens`). They also require an `anthropic-version` header, which was never sent.
For `api.anthropic.com` connections, send `anthropic-version` and move the key into `x-api-key`, dropping the bearer header. Connections using session, OAuth or Entra ID auth keep their token untouched, LiteLLM passthrough connections are unaffected, and admin-configured custom headers still win over both defaults.
Fixes#27695
periodic_usage_pool_cleanup, periodic_session_pool_cleanup and
scheduler_worker_loop were started with asyncio.create_task and their
handles discarded. The event loop keeps only a weak reference to a task,
so a task with no other referent can be garbage collected while it is
suspended at an await. All three are while True loops meant to run for
the process lifetime, and if one is collected the failure is silent:
pool entries stop being cleaned up, or automations and calendar alerts
stop firing, with nothing logged.
Six lines above, redis_task_command_listener is already stored on
app.state and cancelled on shutdown. This applies the same treatment to
the other three.
Closes#28052
The gate now applies the same authorship condition the channel message update route already uses, so both paths agree on which messages a caller may modify.