Commit graph

7274 commits

Author SHA1 Message Date
Timothy Jaeryang Baek
0f5a58f5fb refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
2026-10-07 01:02:33 +04:00
Timothy Jaeryang Baek
fa4c7fe5e8 refac 2026-10-07 00:57:06 +04:00
Timothy Jaeryang Baek
efd94fde63 refac 2026-10-07 00:54:53 +04:00
Timothy Jaeryang Baek
093bfce2b6 refac 2026-10-07 00:39:16 +04:00
Timothy Jaeryang Baek
e9cca320b4 refac 2026-10-06 22:02:23 +04:00
Timothy Jaeryang Baek
106aae70e9 refac 2026-10-06 21:45:10 +04:00
Timothy Jaeryang Baek
b0bcd94519 refac 2026-10-06 16:41:33 +04:00
Timothy Jaeryang Baek
b8738494cf refac 2026-10-06 16:35:49 +04:00
Timothy Jaeryang Baek
398c37c73c refac 2026-10-05 14:35:15 +04:00
Timothy Jaeryang Baek
d4c561d9f2 refac 2026-10-05 14:01:02 +04:00
Timothy Jaeryang Baek
24e30d1cbd refac 2026-10-05 12:17:25 +04:00
Timothy Jaeryang Baek
250f63175e refac 2026-10-05 12:03:25 +04:00
Timothy Jaeryang Baek
425da8b6cc refac 2026-10-05 12:03:19 +04:00
Classic298
88720c6928
fix: tool calls get mixed up between rounds when the model reuses tool call ids (#31887)
Some providers, like Kimi K3 on OpenRouter, number their tool calls from zero again each time the model calls tools within the same reply. A later batch of calls then overwrote the earlier calls with the same id, so the saved chat showed the earlier calls with the later arguments, and the model was sent its earlier results next to the wrong arguments. A new tool call whose id is already used in the same reply now gets a fresh id, so every call keeps its own arguments and result, and providers that send unique ids are untouched. Tested before and after against a mock provider that reuses ids on every round.

Fixes #28305
2026-10-05 11:19:45 +04:00
Timothy Jaeryang Baek
4e1c52696f refac 2026-10-05 11:11:42 +04:00
Timothy Jaeryang Baek
fb741ebcd2 refac 2026-10-05 10:56:06 +04:00
Timothy Jaeryang Baek
cf5755f949 refac 2026-10-05 10:48:41 +04:00
Timothy Jaeryang Baek
d8659c237c refac 2026-10-05 10:46:08 +04:00
Timothy Jaeryang Baek
dd576bade3 refac 2026-10-05 10:40:09 +04:00
Timothy Jaeryang Baek
77e6bc2893 refac 2026-10-05 09:56:10 +04:00
Classic298
f25708484b
fix: Daily Messages chart in Analytics fails to load on All time when a message has no timestamp (#31888)
With All time selected, the Daily Messages chart in Admin Panel > Analytics failed to load (the daily analytics request returned a 500) as soon as one message had been saved without a timestamp, for example a chat created through the API with a null timestamp. A date range still worked because those messages were left out of it. Such messages now get the time they were saved, and ones already stored without a timestamp are counted on today's date instead of breaking the chart. Messages that came after the missing timestamp in the same chat were also left out of analytics before and are now counted.

Fixes #27316
2026-10-05 06:46:49 +04:00
G30
1b2ceedd61
fix: create the model when a GGUF is uploaded to Ollama, in File Mode and URL Mode (#31862) 2026-10-05 06:43:06 +04:00
G30
8c34af3033
fix: keep private arena models with no access grants visible to admins without the admin bypass (#31858) 2026-10-05 06:42:55 +04:00
G30
2e43d66985
fix: report a web search where no page loaded instead of blaming the embedding settings (#31859) 2026-10-05 06:42:48 +04:00
Classic298
95dd3321af
fix: images stored inside a chat trigger needless context compaction (#31915)
Images stored inside the chat record itself (chats from older versions, images that could not be saved as files, temporary chats) had their encoded data counted as text when estimating how full the context is, so a single 3 MB image counted as about a million tokens. That pushed the chat over the compaction threshold, summarizing older messages while the chat was well under the limit, and made the context usage indicator jump. The encoded image data is now left out of the token estimate, both for compaction and for the indicator.

Fixes #31913
2026-10-05 06:41:50 +04:00
Classic298
743a46bdcd
fix: chat stays broken after a reply is stopped or fails before any text (#31892)
Some checks failed
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
When a reply was stopped or failed before the model sent any text, the chat kept an empty reply, and every later message in that chat sent it back to the model as part of the conversation. Providers that reject empty assistant messages then refused every new request, so the chat stayed unusable unless the empty reply was deleted from the database by hand. Empty replies are now skipped when the conversation is sent to the model, as replies that ended in an error already were. Chats already stuck like this work again on their next message, since nothing stored has to change.

Fixes #25083
2026-10-03 17:06:03 +04:00
Classic298
914ff1f116
fix: large pastes into a note are lost after a reload (#31893)
Pasting a large block of text (around 150 KB or more) into a note made the note's live connection drop and reconnect, and the text was never saved, so a reload showed the note without it. Each edit message carries the whole note in several formats, which pushed it past the server's 1 MB cap on a single message. That cap is now 16 MiB, the same size the server already accepts for a websocket message, so notes with up to about 2 MB of text save again.

Fixes #26140
2026-10-03 17:05:00 +04:00
Classic298
cd64930c05
fix: MCP OAuth sign-in fails with a state error when the server asks for many scopes (#31894)
Connecting an MCP tool server over OAuth 2.1 failed after signing in at the provider with an "invalid or expired state" error whenever the server advertises a long list of scopes, such as the Google Workspace MCP server with its 42 Google scopes. Open WebUI saved the full authorization link it sends to the provider in the session cookie, which pushed the cookie past the 4096-byte browser limit, so the browser dropped it and Open WebUI could not recognise the user when the provider sent them back. The link is no longer saved there, so the cookie stays at a few hundred bytes even with long scope lists.

Fixes #26382
2026-10-03 17:04:50 +04:00
Classic298
e1248e5cfd
fix: chat fails when an MCP server ID plus tool name is longer than 64 characters (#31822)
Open WebUI puts the MCP server ID in front of every MCP tool name it offers to the model. When the two together are longer than 64 characters, providers that cap tool names at 64 (the OpenAI API, AWS Bedrock) reject the whole chat request, even if the model never picks that tool. Names over the limit are now cut to 64 characters ending in a short hash, so the same tool always gets the same name and the MCP server is still called with the original tool name. Names that already fit are unchanged.

Fixes #31821
2026-10-01 20:27:58 +04:00
Timothy Jaeryang Baek
015dbc8619 refac 2026-10-01 08:38:21 +04:00
Classic298
0f46d6096c
fix: server memory grows with every chat until a restart (#31534)
Every chat that got a reply left a small lock behind in the server process, so a long-running server kept one for every chat it had ever answered and only a restart gave the memory back. The lock is now dropped as soon as nothing is using it, so a finished reply leaves nothing behind per chat, while replies, sub-agent results and timers for the same chat still wait for each other as before.

Fixes #31521
2026-10-01 08:14:33 +04:00
Classic298
ecd0ff67e1
fix: apply the duplicate-content check to knowledge batch add (#31336)
Adding files through `POST /api/v1/knowledge/{id}/files/batch/add` accepted a file whose extracted text was already in the knowledge base under another file, and linked both, while the single-file add rejects the same file with "Duplicate content detected". The same text was then embedded twice and retrieval returned the same passages twice.

The batch path now runs each file's content hash through the same check the single-file path uses, now shared by both, and also against the earlier files of the same batch. A duplicate is reported as a failed file in the batch result and is not linked, while the other files of the batch still go through. Batch-added chunks now carry the content hash in their metadata, so later adds through either endpoint detect them.

Chunks written by batch add before this change have no hash, so content added that way earlier is still not detected as a duplicate.

Verified on a running instance: two files with identical text now end up with exactly one linked in every order and combination (one batch call, separate batch calls, batch mixed with single add), and re-adding the same file is still accepted.

Fixes #31333
2026-10-01 07:53:06 +04:00
Classic298
5b0a889a74
fix: apply the custom model fallback to chats sent from the web UI (#31353)
With ENABLE_CUSTOM_MODEL_FALLBACK on, a workspace model whose base model is gone should be answered by the first default model. That worked for plain API calls, but every chat sent from the web UI failed with "Model not found" for users and "Model '' was not found" for admins, and no model was called.

Web UI chats carry a chat id and a socket session, so the request is split into one task per selected model. Each task was rebuilt with the originally requested model id, which dropped the fallback chosen earlier in the handler. The task for the requested model now keeps the fallback model when one was chosen. The chat still records the workspace model the user picked.

Tested end to end against a mock upstream, as user and admin, in new and existing chats: before, every web UI send with such a model errored; after, the default model answers. Healthy models, workspace models with a valid base and multi-model sends behave as before, and with the fallback disabled the chat still fails with "Model not found".

Fixes #31345
2026-10-01 07:52:25 +04:00
Classic298
5e1f4d8c33
fix: subscribe members added to a channel to its live feed (#31113)
A user added to an existing group or DM channel saw nothing from it until they reloaded the page: the channel did not appear in their sidebar, and opening it by URL showed the history but no new messages, edits, pins or reactions.

Adding members now does what channel creation already does for its participants: the newly inserted members get a `channel:created` event so their sidebar refreshes, and their open sessions join the channel room so live updates reach them. Standard channels are skipped because their access comes from access grants, matching the membership-removal path.

Verified against a running server with two live Socket.IO clients: the added user's open session now gets the sidebar refresh and the next message immediately; re-adding an existing member, removal and standard channels behave as before.

Fixes #30432
2026-10-01 07:52:07 +04:00
Classic298
d8177c4177
fix: send the query embedding as a vector in external pgvector retrieval (#31112)
External knowledge bases on the pgvector provider failed on every search with "operator does not exist: vector <=> double precision[]", so they looked empty to users. This happened regardless of the VECTOR_DB setting.

The query embedding was bound as a plain Python list. register_vector only adapts pgvector's own Vector type and numpy arrays, so psycopg sent the list as a float array, which the <=> operator does not accept. Wrapping the embedding in pgvector.Vector sends it as a real vector.

Vector is imported from the package root, which works on the pinned pgvector 0.4.2 and on 0.5.x, where the pgvector.psycopg re-export no longer exists.

Verified against a pgvector Postgres: before the fix the reported error reproduces; after it, results come back ranked by cosine distance and filtered to the collection, including schema-qualified tables, halfvec columns and 1536-dimension embeddings.

Fixes #26663
2026-10-01 07:51:32 +04:00
Classic298
7cf35bedc0
fix: editing a knowledge file on Elasticsearch keeps its old text searchable (#31530)
With VECTOR_DB=elasticsearch, changing a file's content in a knowledge base added the new text but never removed the old chunks, so searches and chats kept returning the old text next to the new. The old chunks are now found and removed after the edit, the same as on the default store.

Fixes #31523
2026-10-01 07:36:44 +04:00
Classic298
e88e1f8119
fix: prompt cache misses after a model calls tools one after another (#31593)
When a model called one tool, got its result and then called a second tool with no text in between, the next request merged both calls into an earlier assistant message that the provider had already seen. That changed the conversation's beginning, so the provider's prompt cache stopped matching from there for the rest of the chat. Each tool call and its result are now sent as their own messages, so the start of the conversation stays identical from one request to the next.

Fixes #31588
2026-10-01 07:34:59 +04:00
Classic298
1c233f1be6
fix: sub-agents in folder chats are not limited to the folder's knowledge (#31574)
In a chat inside a folder with knowledge attached, the chat only searches the folder's knowledge, but a sub-agent it started could list and search every knowledge base the user can read. Sub-agents now get the folder's knowledge like the chat that started them. They already receive the folder's system prompt as part of the chat's instructions, so it is not added a second time.

Fixes #31569
2026-10-01 07:34:37 +04:00
Classic298
c3fbf36384
fix: uploaded files lose Chinese punctuation and quotes (#31655)
Text extracted from any uploaded file had full-width punctuation like :(),!? turned into ASCII :(),!? and curly quotes like “ ” turned into straight quotes, so both the file preview and the model saw altered text. Extracted text now keeps these characters as written, while garbled text from wrong encodings (like café becoming café) is still repaired. Files uploaded before this change keep the altered text until they are uploaded again.

Fixes #17087
2026-10-01 07:22:08 +04:00
Classic298
3ef0d15433
fix: audit log shows passwords that contain a double quote (#31659)
* fix: audit log shows passwords that contain a double quote

With AUDIT_LOG_LEVEL set to REQUEST or REQUEST_RESPONSE, a password containing a double quote was only masked up to that quote, so a new password like Q"secret was logged as "********"secret. A request with whitespace before the colon, such as "new_password" : "secret", was not masked at all. Any field whose name ends in "password" is now masked through its closing quote in both cases.

* fix: audit log records passwords sent back in responses

With AUDIT_LOG_LEVEL set to REQUEST_RESPONSE, fields whose name ends in "password" were masked in request bodies, but response bodies were logged unmasked. Saving or opening the LDAP server settings therefore wrote the Application DN Password to the audit log in plain text, because the settings come back in the response, and the Jupyter passwords in the code execution settings leaked the same way. Responses now get the same masking as requests.
2026-10-01 07:21:58 +04:00
Timothy Jaeryang Baek
f50f9e6252 refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
2026-10-01 03:31:20 +04:00
Classic298
aee7c47277
feat: Milvus hybrid search without a second text-only copy of every collection (#31660)
Native hybrid search on Milvus kept a second, text-only collection beside every Milvus collection and searched both. Each Milvus collection now holds its vectors, its text and its BM25 keyword index together, and results rank exactly as before, with the BM25 weight setting working the same way it does on pgvector. Existing data moves on the first start with ENABLE_DB_MIGRATIONS on: startup waits while every collection is copied once (vectors included, nothing is re-embedded) and the originals are only dropped after every copy succeeded, so a failed run changes nothing and is retried on the next start. The copy needs free disk space for a second copy of the data until it finishes, and on one 16-thread machine with Milvus's official docker compose setup it ran at 11 to 22 MB/s, so 200 GB takes about 2.5 to 5 hours depending on chunk size. Milvus servers older than 2.5 are detected and keep the existing hybrid search.
2026-10-01 03:07:50 +04:00
Classic298
4ef7e35b88
fix: Playwright web loader returns only the site menu for pages with more than one <main> element (#31644)
Some pages, like the Ubiquiti tech specs pages, have more than one <main> element. The Playwright web loader only read the first one, so these pages came back as just their site menu and the actual content was lost. When a page has more than one, the loader now ignores those tags and reads the whole page. Pages with a single <main> load the same as before.

Fixes #28643
2026-09-30 21:06:01 +04:00
Classic298
75bff4bcd9
feat: native hybrid search for Milvus and Milvus multitenancy (#31645)
With hybrid search on, Milvus installs fetch every chunk of a collection and score BM25 in Python for each search. In both Milvus modes, one collection per knowledge base and multitenancy, Milvus now runs the keyword half itself with its built-in BM25 full-text search and merges it with the vector results, as pgvector already does. Milvus cannot add a BM25 index to an existing collection, so each collection gets a second, text-only collection next to it; existing installs build these once during startup when ENABLE_DB_MIGRATIONS is on, which copies the chunk text (extra storage roughly the size of that text) and leaves the original vectors and indexes untouched. On one standalone Milvus server the copy ran at about 11,000 chunks per second, around 40 minutes for 200 GB with 1536-dimension embeddings. Collections that cannot be copied, and Milvus servers older than 2.5, which have no BM25, keep using the existing hybrid search.

Fixes #26243
2026-09-30 21:03:47 +04:00
Classic298
3d43a497b5
fix: shared chats can't open new files attached together with a file already in the chat (#31650)
When files were attached to a chat message and one of them was already attached to that chat (on the same message or another one), none of the new files were recorded as part of the chat, and nothing showed up in the logs. The sender still saw every file in their own chat, but anyone opening a shared copy of the chat could not open the new ones. Files already attached to the chat are now skipped and the new ones are recorded normally.

Fixes #31648
2026-09-30 21:03:32 +04:00
Classic298
1058444d74
fix: artifact preview closes right after opening when a filter writes the reply (#31652)
Since 0.11.1, when a filter function wrote part of the reply before the model answered (for example some text and an HTML block), the artifact preview opened and then closed straight away. The page showed the filter text but was never sent it as part of the reply, so the model's first output overwrote it, the page briefly saw no HTML block and closed the preview. The filter text is now sent to the page before the model's output, so the preview stays open.

Fixes #31643
2026-09-30 21:02:27 +04:00
Classic298
9d2c3965ff
fix: send max_completion_tokens for Bedrock-prefixed OpenAI models (#30976)
On an Amazon Bedrock OpenAI-compatible connection, GPT-5.6 and GPT-6 models have ids like `us.openai.gpt-6-sol` or `openai.gpt-6-luna`. These were not recognised as new OpenAI models, so `max_tokens` went upstream unchanged and Bedrock rejected it with a 400. Title and emoji generation failed on every chat, and any request with a token limit failed too. Setting `max_completion_tokens` by hand did not help, because non-OpenAI URLs convert it back to `max_tokens`.

`is_openai_new_model()` now drops a leading `openai.` or `<region>.openai.` (`us.`, `eu.`, `global.`, `us-gov.`) before matching, so these ids get the same handling as bare `gpt-5` ids. Ids that are not new models, such as `openai.gpt-oss-120b-1:0` and `gpt-4o`, are unchanged, and so is the LiteLLM `openai/` prefix.

Fixes #30510
2026-09-30 20:09:06 +04:00
G30
d4d04dca0d
fix: keep a cloned chat in a shared folder the user can write to (#31370) 2026-09-30 19:59:11 +04:00
Classic298
fab58bd35f
fix: ejecting a model ignores AIOHTTP_CLIENT_SESSION_SSL (#31391)
Ejecting a loaded model from the model selector failed with a certificate error on llama.cpp and Ollama connections served over HTTPS with a self-signed or internal CA, even though chatting with the same connection worked. The unload request skipped the AIOHTTP_CLIENT_SESSION_SSL setting and always verified against the default system certificates, so setting it to a CA bundle or to false had no effect there. It now uses that setting, the same way chat requests to the connection already do.

Fixes #31371
2026-09-30 19:49:59 +04:00
Classic298
a5176f4cda
fix: Google MCP connections drop about an hour after signing in (#31395)
MCP tool servers that sign in through Google, such as Google's hosted Gmail, Drive and Calendar servers, never got a refresh token, because Google only issues one when the sign-in explicitly asks for offline access. When the one-hour access token ran out the refresh failed and the connection was removed, so every user had to sign in again every hour. When the server's sign-in page is Google's, the sign-in now asks for offline access and a fresh consent, so the token renews on its own. Other providers get the same sign-in request as before, and existing Google connections pick this up the next time the user signs in.

Fixes #28319
2026-09-30 19:49:25 +04:00