Commit graph

7172 commits

Author SHA1 Message Date
Classic298
583f5a66d2
fix: max_tokens sent through the API is ignored for Ollama models (#31437)
When an API request to an Ollama model set max_tokens, Open WebUI passed it on in a place Ollama does not read, so Ollama ignored it and replies ran to full length. The limit now reaches Ollama as its own output length setting, so replies stop at the requested length. It also wins over a max_tokens value saved in the model's advanced parameters, as the API docs describe. Chats in the web UI were not affected, since their limit already reached Ollama correctly.

Fixes #31432
2026-09-26 07:18:43 +04:00
Classic298
8081ac299f
fix: missing space in reasoning model answers right after the thinking block (#31438)
With reasoning models, when the first part of the answer arrived together with the end of the thinking block and ended with a space, that space went missing, so "The answer is 4." was shown and saved as "The answeris 4.". That space is now kept.

Fixes #31435
2026-09-26 07:18:35 +04:00
Classic298
9d6b17ffbc
fix: errors on Responses API connections are not shown or not kept after a reload (#31439)
With a connection set to the Responses API, when the provider reported a reply as failed, the error showed while streaming but was gone after a reload, leaving an empty reply. Some other provider errors never showed up at all, not even while streaming. Both kinds of error now show up and are still there after a reload, the same as on Chat Completions connections.

Fixes #31433
2026-09-26 07:18:23 +04:00
Classic298
ac00d40e36
fix: model sync no longer fails with "database is locked" on SQLite (#31349)
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
On the default SQLite setup (session sharing off), a POST to /api/v1/models/sync containing any model that already exists answered 200 with an empty list and stored nothing. After a 5 second stall, the only trace was "database is locked" in the server log.

The sync wrote each model's access grants while the model update was still uncommitted. Without session sharing the grant writes run on a second database session, and SQLite allows one writer at a time, so that write waited on the same request's uncommitted update until the busy timeout expired and the whole sync was dropped.

Grants are now written after the model changes are committed, the same order model create and update already use. A failed model commit now also leaves every grant untouched. PostgreSQL and setups with session sharing on behave as before.

Fixes #31346
2026-09-25 01:06:56 -04:00
Classic298
fcb0af3fd4
fix: skip blocked OAuth groups when auto-creating groups (#31316)
With ENABLE_OAUTH_GROUP_CREATION on, every group in a user's OAuth claim was created on login, including groups matching OAUTH_BLOCKED_GROUPS. Membership sync already ignored those groups, so the result was empty groups nobody could join. With IdPs that send a user's full directory membership (Keycloak backed by LDAP/AD), one login could fill the group table with thousands of them.

Group creation now applies the same blocklist check as the membership add and remove steps, so a blocked group is never created, joined or left through OAuth. Groups that are not blocked are created as before.

Fixes #29558
2026-09-25 00:08:01 -04:00
Classic298
3c47f0d7e7
fix: stop decoding Korean and Japanese text uploads as Chinese (#31356)
Korean EUC-KR and Japanese Shift-JIS text files were stored as garbled Chinese characters, so retrieval, knowledge bases and the model context all worked on text that is not in the file.

Encoding detection puts chardet's guess in front of a fixed GB18030, Big5, EUC-KR, EUC-JP try order, and GB18030 decodes almost any double-byte text without an error. The guess map was written for chardet 5. Since the bump to chardet 7 in v0.10.0, Korean text is reported as CP949 and Japanese text as cp932 or SHIFT_JIS, which the map either did not know or dropped because the codec was not in the try order, so these files fell through to GB18030.

The map now covers CP949 and cp932, and a mapped guess is always tried first. SHIFT_JIS maps to cp932, the Windows superset, because chardet also reports SHIFT_JIS for ordinary Japanese files containing characters such as ① or ㈱ that plain Shift-JIS cannot decode; this is the same subset-to-superset rule the map already applies to GB2312.

Korean and Japanese files now decode correctly, and Chinese, EUC-JP, UTF-8 and Western files decode as before. The one trade-off of trusting the guess: chardet 7 labels some files holding only a few Chinese characters (a short label or a one-line comment) as CP949, and those now read as Korean. No regressions were found in files with more Chinese text than that.

Fixes #31352
2026-09-25 00:01:38 -04:00
Classic298
3f5881c520
fix: load terminal AGENTS.md on Windows Open Terminal hosts (#31342)
Fixes #31340

When the attached Open Terminal runs on Windows, the AGENTS.md in its home directory was never handed to the model. The home check only accepted POSIX absolute paths, so a drive-letter or UNC home such as C:\ProgramData\OpenTerminal\inst was treated as invalid and the file was skipped without any log line.

The check now also accepts Windows absolute paths. Relative and drive-relative homes are still skipped, and POSIX homes send byte-identical requests.

The file path keeps its forward-slash join. Open Terminal normalises the path on the host, so C:\Users\bob/AGENTS.md opens C:\Users\bob\AGENTS.md. Picking ntpath.join for Windows homes would give native separators on the wire but adds a second branch for no change in which file gets read.

Verified against Open Terminal's own path resolution with Windows semantics for drive-letter, forward-slash, drive-root, trailing-backslash and UNC homes, over both the backend request and the browser direct-connection path.
2026-09-25 00:00:59 -04:00
Classic298
fccd755684
fix: send the saved title in the chat:title event when title generation is off (#31355)
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
Fixes #31348

With title generation turned off (per user or by the admin), the chat is saved with the first user message as its title, but the live title event sent the assistant message instead. A new chat therefore kept showing "New Chat" in the header and tab until a reload, because the assistant message is still empty at that point. In a note's chat panel the whole model reply showed up as the title.

The fallback branch now sends the title it just saved, the same way the generated-title branch above it already does. The sidebar was never affected because it reloads titles from the database.
2026-09-24 15:58:50 -04:00
Timothy Jaeryang Baek
7ad0ae4687 refac 2026-09-24 12:34:04 -04:00
G30
9836367ab2
fix: apply the group filter to the Analytics hourly chart (#30978) 2026-09-24 12:11:13 -04:00
Classic298
6c6d35c72e
chore: bump hiredis (#31328)
* Update pyproject.toml

* Update requirements-slim.txt

* Update requirements.txt
2026-09-24 12:10:46 -04:00
Timothy Jaeryang Baek
f412538756 refac 2026-09-24 12:10:30 -04:00
G30
a046e07044
fix: save a file edit made through write access to a shared knowledge base (#30448) 2026-09-23 23:05:37 -05:00
Classic298
d9dff09c30
refac: build the SQLite LIKE regex from wildcard-split segments (#30393)
Each '%'-separated segment of the LIKE pattern is now matched with an atomic group.
2026-09-23 23:55:57 -04:00
Classic298
14c5b65704
refac: tighten details and image patterns in background task message cleanup (#30394)
The patterns that remove details blocks and inline images from task messages now stop at the next details tag, bracket or parenthesis.
2026-09-23 23:55:07 -04:00
Classic298
f7294161fa
fix: show ComfyUI images saved by the Save Image (Advanced) node (#30420)
When a ComfyUI workflow ends in the core "Save Image (Advanced)" node, ComfyUI finishes the job and saves the image, but Open WebUI returns an empty result, so the chat shows nothing. Image editing workflows such as the Qwen Image Edit template use this node by default.

Open WebUI only collects images from output nodes of type SaveImage and PreviewImage. This adds SaveImageAdvanced to that list. The node reports its files in the same format as SaveImage, so the rest of the download and storage path works unchanged. Generation and editing share this code, so both are fixed.

Fixes #30404
2026-09-23 23:54:58 -04:00
Classic298
b6191a0510
fix: stop storing MCP image/audio base64 in file metadata (#30419)
When an MCP tool returns an image or audio item, the file is uploaded to storage, but the whole MCP item, including its base64 payload, was also passed as upload metadata. That metadata is persisted in the file table's meta column, so every such result was stored twice: once in storage and once as base64 in the database, growing the DB and every file query that loads meta.

The MCP path now passes only chat_id, message_id and session_id, the same metadata the non-MCP tool image path already stores. Nothing reads the removed key.

Fixes #30411
2026-09-23 23:54:23 -04:00
Classic298
f7ce4024d6
fix: keep requested MCP OAuth scope when DCR response omits it (#30384)
MCP tool servers using OAuth 2.1 with dynamic client registration authorized without any scope when the authorization server left scope out of its registration response, which RFC 7591 allows (Atlassian and Notion do). Consent completed and the tool showed as connected, but the issued token lacked the scopes the resource requires, so every tool call was refused. Discovered scopes and the custom OAuth Scopes field were both affected.

The stored client now falls back to the scope sent in the registration request when the response has none. A scope the server does return is kept as is.

Connections registered before this fix already have a null scope stored. The protected resource metadata recovery that static-credential clients already use now also runs for dynamically registered clients, so those connections pick up the discovered scopes on the next load without registering again.

Fixes #29967
2026-09-23 23:52:29 -04:00
Classic298
60ded561a9
fix: load models before resolving automation model defaults (#30379)
Automations lost their model's tool bindings (including MCP servers), default features such as web search, default filters and terminal on the first run after a restart. The model then answered that it had no tools. Later runs and a manual Regenerate worked. Automations on the base models cache were not affected.

The run read the model from app.state.MODELS before anything had loaded it. After a restart or a connection settings save, that cache stays empty until a browser loads the model list or a chat completion runs. The completion runs only after the automation has already built its request.

execute_automation now loads the models when the cache is empty, using the same guard chat_completion uses, before either the chat or the channel target reads it. This also fixes channel automations showing the raw model ID in place of the model name on a cold cache.

Verified end to end on a restarted instance with a mock upstream: before, both chat and channel runs reached the pipeline without tool_ids or features. After, both carry the model's tools and web search, and the upstream receives the tool.

Fixes #27694
2026-09-23 23:51:29 -04:00
Classic298
744ce6cbfe
fix: send the configured USER_AGENT on the Attach Webpage pre-check (#30385)
Attach Webpage fails with 403 Forbidden on sites that reject the bare aiohttp user agent, Wikipedia among them, even when USER_AGENT is set. The web loader sends USER_AGENT, but the request that runs first to decide whether the URL is a page or a file does not, so the attachment fails before the loader is ever reached.

The pre-check now sends USER_AGENT as the request User-Agent when it is set. With it unset the request is unchanged and keeps the aiohttp default.

Verified against the real _fetch_url with https://en.wikipedia.org/wiki/OpenAI: 403 before, page detected after; USER_AGENT unset still returns the same 403 as before, and a direct PDF URL is still detected as a file.

Fixes #29617
2026-09-23 23:50:08 -04:00
Classic298
2867f7225b
refac: sync channel room on member removal (#30446)
Removing members from a group or DM channel now also removes their sessions from the channel room, matching how access grant changes are handled.
2026-09-23 23:43:20 -04:00
Classic298
443736e1fe
refactor: use secrets module for generated secret key (#30441)
Generates the default WEBUI_SECRET_KEY file with secrets.token_bytes, matching how the start scripts read from the OS random source.
2026-09-23 23:36:26 -04:00
Classic298
a973e77cfa
refac: folder file checks (#30442)
Folder file entries are also checked against the user making the change.
2026-09-23 23:36:15 -04:00
Classic298
4caf255389
fix: refresh an expiring OAuth token once across workers and replicas (#30450)
#30426 stopped concurrent requests from refreshing the same OAuth session twice, but its lock only lives inside one process. With several uvicorn workers or replicas, two requests on different workers still send the same refresh token, a rotating provider rejects the second with invalid_grant, and the session gets deleted, so the user's OAuth session is logged out again.

When Redis is configured, which multi-worker and multi-replica deployments require, the refresh now takes a Redis lock per session instead of the in-process one. Single-process deployments without Redis keep the in-process lock. The waiter re-reads the session inside the lock as before and uses the token that was just stored.

It uses redis-py's own async lock because the existing RedisLock is synchronous and never waits. The Sentinel proxy now passes `lock` through unwrapped like `pipeline` and `pubsub`; otherwise it returned a coroutine and every refresh behind Sentinel would fail.

Tested with separate OS processes on one sqlite DB, a real Redis and a rotating mock provider: 2 and 5 processes (and 5 processes x 3 requests) now cause 1 refresh, every caller gets the new token and the session is kept (before: one refresh per process, session deleted every run). Single refresh, failed refresh, valid token and the single-process path without Redis are unchanged.

Follow-up to #30426, refs #30416
2026-09-23 23:33:09 -04:00
G30
0ab3a2b335
fix: restore tag rows after unarchiving all chats and remove unused ones after deleting all chats (#30453) 2026-09-23 23:32:54 -04:00
G30
7e2bea8253
fix: load the leaderboard activity chart for model ids that contain a slash (#30456) 2026-09-23 23:32:42 -04:00
G30
95d6af44d7
fix: mark imported and cloned chats as read (#30754) 2026-09-23 23:32:14 -04:00
Classic298
9e51fe421c
fix: read S3 files with long non-ASCII names (#30418)
With S3 storage, a file whose stored name is close to the 255-byte filename limit and contains non-ASCII characters (for example a Cyrillic name of about 210-218 bytes) uploads fine, but every later read fails with "File name too long". Processing never gets the content, so the file shows as attached while the model receives no text.

The read path used boto3's download_file, which first writes to a temporary name with 9 extra characters. boto3 caps that temporary name by characters, not bytes, so multibyte names end up over the limit even though the final name fits.

The download now streams straight into the local path with download_fileobj, the same way the Azure provider already writes its local copy. That path is the one the upload just wrote successfully, so it always fits. ASCII names, key prefixes and multipart downloads behave as before.

Fixes #30409
2026-09-23 23:30:55 -04:00
Classic298
b3bb82e776
refac: scope ydoc update save scheduling to note documents (#30395)
The ydoc document update handler now schedules the debounced save only for note documents, since notes are the only ydoc documents with a save handler.
2026-09-23 23:30:21 -04:00
Classic298
abd60d33fe
fix: evaluate automation schedules on Windows with PostgreSQL (#30424)
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Since 0.11.4, creating or editing an automation on Windows with PostgreSQL fails with a 400, and the scheduler logs NotImplementedError on every tick, so automations do not work at all on that setup.

Schedules are now evaluated in a worker subprocess so a pathological rule can be killed after the 2s budget. On Windows with PostgreSQL, Open WebUI switches to the selector event loop that psycopg needs, and that loop cannot spawn subprocesses.

When spawning fails there, the evaluation now reruns on a Proactor event loop in a worker thread. The subprocess, the 2s budget and the kill on timeout all stay the same, and the global loop policy psycopg depends on is untouched. Falling back to a plain thread was considered and rejected: a thread cannot be stopped, so a costly rule would keep burning CPU after the timeout.

Verified with a loop that refuses subprocesses: base raises NotImplementedError, the fix returns the same results as base, still times out a pathological rule at 2s with the worker killed, and leaves no processes or loops behind under repeated and concurrent calls. Other platforms take the unchanged path.

Fixes #30400
2026-09-23 08:47:04 -05:00
Classic298
9db1a518d3
fix: refresh an expiring OAuth token once when requests race (#30426)
With an OIDC provider that rotates refresh tokens, sending a chat to a system_oauth connection often logged the user's OAuth session out. Two requests reached the refresh at the same time and both sent the same refresh token. The provider rejected the second one with invalid_grant, and Open WebUI deleted the session, so every following request lost its token until the user logged in again.

Refreshes now take a per-session lock. A request that waited for another one re-reads the session and uses the token that was just stored, so the provider sees one refresh per rotation.

Tested with real sqlite sessions and a rotating mock provider: 2 and 5 concurrent callers now cause 1 refresh, all callers get the new token and the session is kept (before: one refresh per caller, all callers got nothing, session deleted). Single refresh, failed refresh and valid-token paths are unchanged.

The lock is per process, so deployments with several workers or replicas can still race across processes.

Fixes #30416
2026-09-23 08:46:31 -05:00
Classic298
2f92635409
fix: keep the model system prompt on Ollama tool-call follow-ups (#30375)
With native function calling on an Ollama model, every request sent after a tool result was missing the model's system prompt, so the final answer ignored the model's instructions. Only other system content, such as the attached knowledge tag, was left. OpenAI connections were not affected.

Tool-call follow-ups are rebuilt from the chat's message list and skip the router's system prompt step, because the first request is expected to have already added it to that list. The OpenAI path does add it there, but the Ollama path converts the messages into a copy first and adds the prompt only to the copy, so the follow-ups never see it.

The model system prompt is now applied to the messages before the Ollama conversion, and the Ollama router is told to skip it for that request so it is not added twice. Ollama now behaves the same as the OpenAI path. Direct calls to /ollama/api/chat still get the prompt from the router as before.

Checked baseline against patched: first request and follow-up for plain, custom and arena Ollama models, with and without a chat system prompt, with template variables and on the OpenAI path. The prompt is now present exactly once on every Ollama follow-up, and nothing else changed.

Fixes #30161
2026-09-22 16:47:20 -04:00
Classic298
1a74a9f46c
perf: per-room channel delivery for the socket.io Redis manager (#28818)
Reworked from the ground up after the feedback that the registry implementation did not land. The Redis room registry and its whole recovery protocol (heartbeats, liveness keys, pruning, distrust windows, cache invalidation) are gone; the change is now ~105 lines with no state kept outside the process.

With WEBSOCKET_MANAGER=redis every emit is published on one shared channel and every instance JSON-decodes every message: a 16 instance fleet decodes each streamed token delta 16 times and 15 discard it. py-spy across a loaded fleet (16 instances, ~4000 users) puts ~31% of all active CPU samples in the pubsub listener parse chain, the largest bucket.

Room-targeted emits are now published on a per-room channel instead; every instance keeps one static pattern subscription covering all room channels and drops messages for rooms without local members by channel name, paying a set lookup instead of a JSON parse. No state leaves the process, so recovery paths and loss windows are identical to the stock manager; acks and control messages stay on the shared channel and sio.call works across instances unchanged. This is the delivery scheme the official socket.io Redis adapter for Node.js ships by default.

Enabled by default; WEBSOCKET_REDIS_ROOM_CHANNELS=false restores shared-channel-only delivery. All instances must run the same mode, so the switch rides the full-stop upgrade this release already requires for its migration; in a mixed fleet, room emits from updated instances would not reach not-yet-updated ones. Verified end to end with two instances on a real Redis: cross-instance token streams delivered with the shared channel completely silent. Ref #28173.

<!--
🚨 DO NOT DELETE THE TEXT BELOW 🚨
Keep the "Contributor License Agreement" confirmation text intact.
Deleting it will trigger the CLA-Bot to INVALIDATE your PR.

Your PR will NOT be reviewed or merged until you check the box below confirming that you have read and agree to the terms of the CLA.
-->

- [x] By submitting this pull request, I confirm that I have read and fully agree to the [Contributor License Agreement (CLA)](https://github.com/open-webui/open-webui/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT), and I am providing my contributions under its terms.

> [!NOTE]
> Deleting the CLA section will lead to immediate closure of your PR and it will not be merged in.
2026-09-22 14:48:56 -04:00
Classic298
fe56ab24f3
fix: page Chroma get() so hybrid search works on collections over 32k chunks (#30368)
With Chroma as the vector DB, hybrid search on a knowledge base with more than 32766 chunks fails with HTTP 400 "Error querying knowledge base". The legacy hybrid path fetches the whole collection to build the BM25 index, and Chroma's unbounded collection.get() binds one SQLite variable per row, so any collection above SQLite's 32766 variable limit raises "too many SQL variables" (reproduced on chromadb 1.5.9 with both PersistentClient and HttpClient). Vector-only search on the same collection works, which makes it look like a hybrid-search bug.

The Chroma adapter now reads the collection in pages of 10000 rows via limit/offset and concatenates them into the same GetResult shape as before.

Verified on a 90000-row collection: every row returned exactly once with documents and metadata aligned to ids, page order stable across page sizes, empty and exactly-one-page collections unchanged, and query_doc_with_hybrid_search returns results where it previously raised. The tests repo unit suite is identical before and after.

Fixes #30351
2026-09-22 14:48:45 -04:00
Timothy Jaeryang Baek
e93a59f4dd refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-09-22 13:13:22 -04:00
Classic298
0a2e9a42e7
fix: detect the real image type of bare base64 generated images (#30359)
Generated images that come back as bare base64 (OpenAI b64_json, Gemini
bytesBase64Encoded and inlineData, Automatic1111) were always stored as
generated-image.png with content type image/png, even when the provider
returned JPEG or WebP, for example with {"output_format": "jpeg"} in the
OpenAI extra params. The image still rendered because browsers read the
bytes, but the download name, the served Content-Type and the type sent
along on a later image edit were wrong.

Bare base64 carries no format, so the type is now read from the bytes with
Pillow, the same way the file already inspects images elsewhere. A response
that is not an image at all now fails the generation instead of storing a
broken png. The file extension comes from the module's own extension map
first, because the Python 3.11 Docker image has no mime database entry for
WebP and would otherwise name the file generated-imageNone.

Fixes #29948
2026-09-22 11:35:59 -04:00
Classic298
6b36f620cf
fix: stop KeyError 'model' traceback on every new chat from initial title generation (#30356)
Every new saved chat logged "Error generating initial chat title" with a
KeyError: 'model' traceback. The title itself was already generated and
saved by then, so the log was misleading, and the memory settings played no
part in it. The error was silently logged at debug level since v0.10.0 and
became visible in v0.11.4 when the title path switched to log.exception.

The initial title task runs the shared background handler with a context
that has no resolved model, which the memory review step read
unconditionally. It now reads it with a default of None, which the memory
review already accepts. The title path never carries an assistant message,
so the review stops before doing any work there, and the main completion
path keeps reviewing memory exactly once per turn with the real model.

Fixes #30339
2026-09-22 11:35:14 -04:00
Classic298
f956f7adc0
fix: send only the file name to Docling instead of the full storage path (#30357)
With the Docling content extraction engine, every conversion request
carried the server's full internal upload path (for example
/app/backend/data/uploads/<id>_report.pdf) as the multipart file name.
Docling only needs a bare file name, and a hosted Docling instance has
no business learning where Open WebUI keeps its files on disk.

The loader now sends the base name of the stored file, which is what
the MinerU, Datalab and Mistral loaders already do. Nothing else in
the request or the parsed result changes, verified against a capturing
mock server before and after.

Fixes #30352
2026-09-22 11:35:05 -04:00
Classic298
6f6792d484
fix: pass MCP tool images to the model, not only to the UI (#30358)
When an MCP tool returns an image (e.g. a Home Assistant camera snapshot), the
snapshot shows up in the tool call section but the model never sees it: it
answers that there is no image. Only images arriving as inline data URIs were
attached to the model request; MCP images are uploaded to Files first and their
file URL was treated as display-only.

Now an image file item with a file URL is attached to the model request as an
input_image part in addition to staying in the tool call's displayed files. The
existing URL-to-base64 step already resolves file URLs, so the model receives
the image bytes; verified on a running instance with a mock MCP server and a
mock upstream (the second upstream request carries the byte-identical JPEG).
Inline data-URI images keep their existing model-only handling.

Fixes #30327
2026-09-22 11:34:57 -04:00
Timothy Jaeryang Baek
ee3ece1e2b refac 2026-09-21 14:43:04 -04:00
Timothy Jaeryang Baek
bf476452a8 chore: format 2026-09-21 11:35:30 -04:00
Classic298
e8c26f8394
fix: stop the background memory review when memory is switched off (#30309)
With memories disabled instance-wide, or for a user barred from the feature,
the background review still ran every interval turn: it spent a task-model
call drafting memory operations and only then failed at the write, because
the router's permission check rejected it. The review now checks the
'memories.enable' switch and re-checks the 'features.memories' permission
the same way the context-injection path already does, so a model whose memory
capability is on no longer triggers memory work that can never land.

The permission lookup costs a groups query, so it runs last, after the free
config and interval gates; those stay on every turn's hot path.
2026-09-21 11:22:21 -04:00
Timothy Jaeryang Baek
478d1785fd refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-09-21 11:21:42 -04:00
Timothy Jaeryang Baek
30881dbcc9 refac 2026-09-21 11:09:09 -04:00
Classic298
419d248093
refac: check the timer owner's role before running a due timer (#30220)
The due-timer executor rehydrates the owner from the database and now
verifies that the owner is still a user or an admin before entering the
chat completion pipeline, mirroring the check the scheduled-automation
executor already performs. A timer whose owner no longer qualifies is
recorded as an error instead of being run.
2026-09-21 10:46:48 -04:00
Timothy Jaeryang Baek
3fc1146c13 refac 2026-09-21 10:44:37 -04:00
Classic298
94eea41a75
fix: surface searchapi errors, news results and redirect links (#30308)
Web search via searchapi.io could come back empty or near-empty with no
hint of why: an invalid or expired API key turned into an empty result
set instead of an error, the google_news engine splits its results
between organic_results and top_stories and only the first block was
read, and google links came back as google.com/goto redirects the web
loader cannot fetch, so citations pointed at a redirect blob.

The search now reads both result blocks, asks google engines for
resolved destination links, raises on HTTP errors, carries a 30s request
timeout, skips result rows without a link, and logs the response body at
debug instead of dumping every search at info.

Fixes #30305
2026-09-21 10:44:06 -04:00
Timothy Jaeryang Baek
e8bd0661d3 refac 2026-09-21 10:30:53 -04:00
Timothy Jaeryang Baek
754c4b5762 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-09-21 10:25:20 -04:00
G30
c864e3ae7c
fix: stop asking for chat variables a model's system prompt no longer declares (#30173) 2026-09-21 10:23:21 -04:00