Compare commits

...

900 commits

Author SHA1 Message Date
Tim Baek
8bd8b4fac5
Merge pull request #29960 from open-webui/dev
Some checks failed
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Python CI / Ruff Format (3.11) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Release to PyPI / release (push) Has been cancelled
Release / publish (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
0.11.4
2026-09-21 15:25:06 -04:00
Timothy Jaeryang Baek
344ea53068 doc: changelog
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
2026-09-21 15:21:52 -04:00
Timothy Jaeryang Baek
ca75b4de23 refac 2026-09-21 15:20:43 -04:00
Classic298
33df4a75bd
chore: changelog (#29874)
* chore: open the 0.11.4 changelog

Starts the section for the commits landed on dev after the 0.11.3 entry was
merged. Fixed records the direct connection listener left behind on every
request that ended badly, the metadata attached to an uploaded file that never
reached the retrieved sources, the slash command menu showing white text on
white in the light theme, and image editing refusing an image the instance
already held. Added carries the general improvements placeholder and the
Traditional Chinese, Korean and Finnish catalogues.

The tailwind reflow and the raw string escapes are omitted, the first being
whitespace and the second changing nothing that is visible today, though it
keeps the knowledge filesystem tools importable on a later Python.

* chore: extend the metadata and translation entries in 0.11.4

The metadata entry gains the follow-up that trims what travels with each piece
of a file, keeping the oversized and internal fields out of every chunk and
picking up metadata nested a level down, and is recorded there rather than as a
second entry because it continues the same story. The translation entry gains
the Russian and Ukrainian catalogues.

The settings group headings and the keys added for them are omitted: the new
keys carry no translation in any catalogue yet, so every heading still reads as
it did, and the change only makes them translatable later.

* chore: add the security advisory notice to 0.11.4

Adds the standard notice as the first item in the Fixed section. The release
touches the guard that keeps image editing from reaching private addresses, and
carries a listener leak a caller could drive without limit, so it meets the
trigger the format sets for the notice. The notice is the fixed wording and
takes no reference links of its own; the individual entries keep theirs.

* chore: record the file metadata entry as an addition in 0.11.4

Metadata attached to an upload reaching the retrieved sources is a capability
that was not there before rather than a correction, so the entry moves to Added
and leads that section, ahead of the general improvements placeholder and the
translation entry in their reserved places. The wording drops the framing that
described it as something no longer going missing; its pull request, issue and
follow-up commit travel with it unchanged.

* chore: give the file metadata entry a noun phrase label

The label read as an outcome, which suits a correction rather than an addition
and left the entry sounding like something that had stopped going wrong. It is
now the plain noun phrase the format asks of a label, with the sentence beneath
carrying what the metadata now does.

* chore: add the settings group heading entry to 0.11.4

The groundwork that made these headings translatable was held back last time,
the keys having landed empty in every catalogue so nothing read differently.
Russian and Ukrainian now carry the wording, so the change shows and the entry
joins Fixed at the foot of the section, below the corrections that reach
further. The two commits that made the headings translatable are referenced;
the catalogues that filled them sit in the translation entry, which names both
languages and carries no links.

* chore: add the code formatter entry and widen the translation coverage entry in 0.11.4

Fixed leads with the built-in code formatter, which pulled in the formatter but
none of the packages it depends on, so saving any tool or function ended in a
missing module error however plain the code was. It sits directly below the
advisory, being the correction that stops the most people getting their work
saved.

The settings heading entry widens into one covering every string that was fixed
in English, the headings among them, now that labels and messages across the
admin pages, the workspace and notifications have been opened up the same way
and Spanish, Russian and Ukrainian have filled them in. Its five commits travel
together. German and Spanish join the translation entry.

* chore: add the notes shortcut and web citation entries to 0.11.4

Added records the shortcut that turns a note's menu button into a delete button
while Shift is held, ahead of the placeholder and translation entries in their
reserved places.

Fixed records the pages a web search only listed no longer being offered as
things to cite, which had produced citations pointing at the wrong result under
a title that read as correct. It sits directly below the code formatter, a
citation that is wrong being worse to a reader than one that is missing.

* chore: add the version check, ligature and search entries to 0.11.4

Fixed gains the version check that reported whatever it was running as the
newest release whenever it could not reach the listing, and logged nothing
about it, which is placed high because an administrator reads it as an answer
rather than a failure. It also gains the arrow sequences drawn as glyphs, which
had text look altered although every character was stored and sent as typed;
that one sits with the other reading corrections.

Added gains the text search moved off the path the server answers on, so a long
search no longer keeps other requests waiting.

The metadata entry gains the four vector stores that never applied the shared
size cap, coercion and sanitising, bringing them in line with the other eleven
rather than standing as an entry of its own.

* chore: extend the quick delete entry to automations in 0.11.4

The shortcut that turns a note's trailing controls into a delete button now does
the same for a row on the automations page, so the entry names both places and
carries both pull requests with the issues they close, rather than repeating
itself for the second surface.

The pass blanking redundant English values in the en-US catalogue is omitted:
the key stands in for the value there, so nothing reads differently.

* chore: add the note from search entry to 0.11.4

Fixed records the three ways starting a note from the search box went wrong, as
one entry because they are one action: the browser back button making a further
note on every press, the action doing nothing at all when the notes page was
already open, and the typed text losing everything from the first ampersand or
hash onwards. It sits below the listener leak, unwanted notes being a nuisance
rather than something that grows on the server.

* chore: add the arena, relevance and dialog entries to 0.11.4

Fixed gains the arena model that failed on something unrelated whenever its
provider answered with an error, so the reason never reached the chat and
titles and tags broke alongside it; it is placed high, a message that fails
outright costing more than one that reads oddly. It also gains the relevance
figure that a reply drawing on a single source never showed, appearing only
once a second source joined it, and the grey band the attach webpage dialog
carried between its address box and its button.

* chore: add the per-language model wording entries to 0.11.4

Added leads with a model carrying its name, description and starter prompts in
each language, written in the workspace editor or the admin defaults and chosen
by the language the interface is set to, falling back to the plain wording where
that language has none. Every part of it arrived together, so it is one entry.

Changed opens the section for the saving call that now answers with the
suggestions and their per-language wording together where it returned the bare
list, which anything reading that response directly has to account for. No
schema change travels with it, so the section carries no migration warning.

* chore: widen the per-language wording entry in 0.11.4

The mechanism that gave a model its name, description and starter prompts in
each language now covers tools, skills, functions, banners and arena entries
too, each written in the editor that owns it, so the entry names them rather
than standing as a second one beside it. Both commits travel with it.

The response shape entry is unchanged: it is the saving call for the admin
defaults, which this second commit does not touch.

* chore: carry the valve labels and translation table into the wording entry

The entry named only what a resource is called and says about itself, leaving
out the labels on its valves, which are translatable down to the options in a
dropdown, and the table each editor now offers in place of the code when a
language is picked, with the JSON file it reads and writes. An import that
drops a placeholder the original fills in is refused, which is worth the words
it takes to say.

It stays one entry: every part of this arrived together and none of it existed
in another form before.

* chore: add the interface wording and sketch upload entries to 0.11.4

Added records an administrator being able to replace the interface's own
wording language by language, which is a separate thing from a model or a tool
carrying wording of its own and so stands beside that entry rather than joining
it. A replacement that drops a placeholder the original fills in is refused, as
it is for the resources.

Fixed records the Arduino sketch that failed to upload to a knowledge base
because its extension was missing from the list of files read as plain text, so
it went to a document extraction server that answered with something the loader
could not read.

* chore: carry the wording editor follow-ups into the 0.11.4 entry

Two follow-ups tidy the editor the per-language wording is written in, dropping
the line that said whether a field was translated or using the default and
holding back the controls beside a skill or tool name. Neither changes what the
entry describes, so they join its links rather than standing on their own.

The terminal work is left out: the dock it adds is not rendered anywhere yet,
and the rest of that commit rewrites the existing terminal in place and sends
two identifying headers upstream, none of which shows to anyone using it.

* chore: add the document escape sequence entry to 0.11.4

Fixed records the escape sequences in an uploaded file being rewritten before
the text was stored, so what was indexed and what the model read differed from
the file as written, and inconsistently at that, since the rewriting stopped at
the first line holding an angle bracket. It sits beside the sketch upload, both
being corrections to what a file turns into on its way into knowledge.

* chore: add the Apple Silicon and WebKit entries to 0.11.4

Fixed gains the question against a knowledge base that killed the server
outright on a Mac when a reranking model ran locally, placed high, an answer
lost with the connection costing more than one that reads oddly.

It also gains the two separate things the WebKit pull request corrects, kept
apart because a reader meets them as different problems: an assistant reply
coming up blank on Apple devices whose browser never says Safari, and the
sidebar hover preview laid out at the width of a full conversation there since
0.11.0.

* chore: add the terminal dock entry to 0.11.4

Added leads with the terminal pane carrying a tab for every command a model is
running beside the shell, which is the most visible thing in the release for
anyone who watches an agent work.

It was held back earlier on the reading that the dock was built but never
rendered. That reading came from searching a checkout that predated the commit;
the commit swaps the single terminal for the dock in the files pane, so it has
been in front of people since it landed.

* chore: add the note date entry to 0.11.4

Fixed records the date beneath a note title, which was wrapped in a button left
over from an earlier styling pass, so it took a pointing hand and announced
itself as something to press while doing nothing at all. It sits at the foot of
the section beside the dialog banding, both being small corrections to how
something looks rather than what it does.

* chore: add the token logging, model picture and image size entries to 0.11.4

Fixed leads, below the advisory, with the credentials an identity provider
handed over being written into the application log when a sign-in failed part
way through, and with the model picture that any signed-in person could fetch
whatever their access, which also told them whether a model id existed at all.
Both sit above the rest of the section, a credential in a log file and a way to
enumerate what a server holds costing more than anything below them.

Changed records the consequence a viewer will notice: a model picture falling
back to the standard logo wherever no grant is held, which is a change to what
people see rather than a correction, so it is kept apart from the entry above.

Added records the container image losing the second Python it never used.

* chore: gather the image size work and add the docx preview entries to 0.11.4

The image size entry now carries the fonts nothing loaded and the packages
nothing imports alongside the second Python it already named, with all four
pull requests, and states the running total rather than the one measurement it
started from. They are one entry because a reader meets them as one thing: less
to pull.

Changed records what that costs anyone whose tool or function imported a
package it never declared, which worked only while the package happened to be
in the image.

Fixed records the Word document preview no longer rendering an HTML
sub-document embedded in the file, and refusing a link that points anywhere
but a web address, a mail address or a telephone number.

* chore: write the image size figure in digits

The file gives every measurement in digits and spells out only counts of
things, so the size the image loses is written the same way as the ones
recorded before it.

* chore: add the slim image entry to 0.11.4

Changed records the slim image giving up the local models and converters it
carried, so an instance built that way now needs a service configured for
embeddings, vector storage, document extraction past plain text, and speech,
and refuses to build alongside the graphics card or Ollama options. The admin
pages mark what is unavailable and the server answers with what to configure,
both part of the same change.

It is kept out of the image size entry above, which gathers what was carried
and never used; this one removes things that were used and asks for them from
somewhere else.

* chore: correct the slim entry and add the citation link entry to 0.11.4

The slim entry read as though a service had to be configured before the image
would run. It does not: the vector client is built on first use behind a lock,
the speech engine raises only when one is asked for, and extraction raises per
file, so an instance starts and holds a conversation with none of it in place
and each feature asks for what it needs when it is reached. The entry now says
that.

The image size entry gains the slim build leaving behind the tool that
installed its packages and the source maps built beside the interface.

Fixed gains the source attached to a reply opening only where it points at a
web address, which sits with the document preview correction, both being about
where a link in something the model handed back is allowed to go. Portuguese
(Brazil) joins the translation entry.

* chore: add the BuildKit entry and fold the install tool into the size entry

The install tool now leaves every image rather than only the slim one, since it
is mounted for the install step instead of being installed and removed, so the
size entry names it among what the image no longer carries and the slim clause
keeps only the source maps.

Changed records what building the image yourself now takes: the BuildKit
builder, which the syntax line alone did not require before, and reach to the
registry that tool is fetched from.

* chore: add the webhook picture and nameless tool call entries to 0.11.4

Fixed gains the channel webhook picture that anyone signed in could fetch, or
be sent onward to, without belonging to the channel, which sits directly below
the model picture entry, the two being the same gap on two routes.

It also gains the tool call arriving with no name at all, which was kept as it
came, written into the stored message and handed back on the next turn for the
endpoint to refuse, and now fails once where it starts.

The line uninstalling the install tool is not recorded: it went in the same
breath as the mount that made it unnecessary, and the size entry already says
the tool no longer ships.

* chore: add the tool server request entry to 0.11.4

Fixed records a tool call to an OpenAPI tool server repeating in the body the
arguments already placed in the address, which a server checking its input
strictly refused, so reading worked and writing through any such address did
not. It is placed high, a whole class of tool call failing outright costing
more than the corrections below it.

* chore: bring the 0.11.4 date up to the newest change it records

The date was set when the section was opened and left there while entries kept
arriving, so it sat three days behind the last commit it covers.

* chore: add the sign-in form setting entry to 0.11.4

Added records the sign-in form becoming something an administrator can switch
off from the authentication settings. The setting itself is not new, having
been available since before this release as a value read at startup; what is
new is reaching it without restarting the server, so it is recorded as an
addition rather than a change to how signing in works.

The date on the section already stands at the day of this change.

* chore: add the slim backend narrowing entry to 0.11.4

Changed gains the set of things slim will work with at all, which is a separate
matter from the entry above it: that one says the local models are gone and a
service is asked for when a feature reaches for one, this one says which
databases, vector stores and file stores remain, and that being pointed at
anything else stops the server before it starts rather than at first use.
The browser-driven page reader goes with them.

It sits directly below the entry it qualifies.

* chore: say which slim limits stop the server and which do not

The entry put refusing to start against all three limits. Only two behave that
way: the database check raises while the module is read, and the storage
provider is chosen by a call made as the module is read. The vector store is
built on first use, so pointing slim at anything but PostgreSQL for it fails
when something searches, not at startup. The sentence now separates them.

* chore: soften the slim backend entry and cover the browser speech runtime

* chore: lead the 0.11.4 sections with the slimming entries

* chore: cover the slim web search and code interpreter changes

* chore: split the slim limits into their own entries and cover the PDF removal

* chore: fold the dropped Gemini SDK into the tool imports entry

* chore: rewrite the changed entries and cover git leaving the slim image

* chore: drop the entry for the unused PDF endpoint

* chore: cover the tool export scoping and the terminal command echo

* chore: lead with the slim image download size

* chore: cover the model registry rewrite fix

* chore: cover the playwright media skip and lead the registry entry with the cost

* chore: lead the playwright entry with the speed gain

* chore: drop the build and picture entries and retitle the tool imports one

* chore: write the measured image sizes and retitle the page loader entry

* chore: state only how much smaller the images are

* chore: retitle the per-language wording entry

* chore: retitle the terminal tabs entry

* chore: fold the translation coverage entry into translation updates

* chore: date the 0.11.4 section 2026-09-07

* chore: cover continuing a reply in a temporary chat

* chore: write the connection listing role check as a security fix

* chore: keep the connection listing fix brief

* chore: fold the continue reply refinements into its entry

* chore: note the llama.cpp continuation flags

* chore: cover cookie forwarding and the note save on navigation

* chore: cover the model access denial logging

* chore: cover the tool step merge and the model list filters

* chore: cover three fixes from 29494, 29493 and 29325

* chore: note the terminal pane open and collapse behaviour

* chore: cover the settings save rework

* chore: write the settings clobbering fix and trim the api note

* chore: name what the interface permission was blocking

* chore: cover the account menu row highlight

* chore: cover diff blocks, note formatting, tika 4 and the oauth session expiry

* chore: cover folder upload and the file browser selection fixes

* chore: cover file comparison and the query link submit

* chore: cover the task list height cap

* chore: add the server side of the tool call match fix

* chore: cover the exa result length cap

* chore: cover the modal shortcut, build stamp and sticky code headings

* chore: cover knowledge listing order, terminal proxy headers and stale tool links

* chore: cover the forwarded auth type header and terminal attachment name clashes

* chore: cover the misleading model editor permission error

* chore: cover the thread notification and overlapping select chevrons

* chore: cover the OpenAI and Ollama proxy header stripping

* chore: cover stopping every task attached to a chat

* chore: cover code blocks in channel structured output replies

* chore: cover whitespace in model ids and knowledge folder deletion

* chore: cover automation schedule times and mention IDs

* chore: cover the latest translation fills

* chore: note the further pt-BR translation pass

* chore: cover the follow-up suggestion in the message box

* chore: cover notification target names, MCP tool paging and the stop follow-up

* chore: cover thinking blocks, channel terminals and custom header encoding

* chore: cover link schemes, spellcheck, dotless hosts and webhook pictures

* chore: cover tool images, stopped replies, sync errors and terminal proxy limits

* chore: cover malformed tool calls and the direct connection note

* chore: cover the langchain community removal and broken streams

* chore: cover the empty page upload message

* chore: cover embed rendering and two memory leaks

* chore: cover the sign-in limiter, socket cleanup and stored tool images

* docs: changelog entries for terminal skills

* docs: cover the skills:create command and the streaming append option

* docs: cover the model editor save button fix

* docs: fold the skill save location refinements into the skills:create entry

* docs: cover model backgrounds, note range edits and the playwright websocket fix

* docs: link the autoscroll follow-up commit

* docs: cover the responses tool strictness and attachment-only fixes

* docs: widen the non-blocking search entry to the knowledge grep tool

* docs: cover branch descent, checkbox defaults, fork folders, search logging and file access

* docs: cover the direct connection guidance note

* docs: cover tag text leaking into streamed replies

* docs: SVG attachments, Docling conversion failures, title generation logging

* docs: chat unblocked after a failed reply

* docs: directory provisioning patch handling

* docs: settings search matches individual settings

* docs: Staan search, Excel in the code interpreter, pgvector index training, terminal tool gating

* docs: Responses API tool calling

* docs: vector search recall and connections, forking, bundled packages, chat import, image connection checks

* docs: message pair shortcut no longer sends the input

* docs: shortcut recording, note sharing, unarchiving from search, skill saves, recommended badge

* docs: workspace exports cover every prompt and model

* docs: terminal instructions file, sign-in roles, folders, feedback, admin forms

* docs: session revocation, sidebar bulk actions, speech file types, source metadata keys

* docs: model fallback, shared note traffic, recurrence counts, RE2 search, Redis task expiry, direct connection events

Add the entries for the commits since the last pass: the chat model fallback
(#29757) and the shared-note co-editing traffic drop (#28185) in Added, with the
Staan count fix (#30303) folded into the Staan entry and Italian added to the
translation entry. Fixed covers the schedule recurrence and count parsing
(#29262), the RE2-backed file searches, the Redis task expiry and stalled
recurrence guards, the direct-connection events travelling between workers,
the IME guard while renaming a chat, the S3 content-encoding image handling
(#29623) and the sixteen interface and workspace fixes.

* docs: issue references on the fix entries, general improvements without links

Append the issue or discussion each fix PR carries in its body to its
changelog entry, and drop every link from the general improvements entry, which
never carries any.

* docs: model pool cache, message round trips, folder backgrounds, delete shortcut, chat variables

Cover the eight newest dev commits: the per-worker model pool cache (#28176)
and the single-message chat read and write (#28184) in Added, and in Fixed
the knowledge directory scoping, model sharing surviving a save, the folder
background on creation (#30218), the Delete Chat shortcut working wherever
the chat was opened from (#30165) and chat variables a prompt no longer
declares (#30173).

* docs: searchapi results and errors, blocked groups saving, timer owner checks, terminal access rechecks

Cover the four commits still missing from the last pass: the searchapi.io fix
(#30308, with its issue), the OAuth blocked groups save round trip (3fc1146c1,
typed comma lists never took effect), the due-timer owner role check (#30220)
and the terminal session access recheck, which folds its in-cycle introduction
(a1189a2d7) and the cross-worker rework (e8bd0661d) into one entry.

* docs: issue and pull references on the knowledge directory and model sharing entries

The knowledge directory scoping is the landed form of #29887 and the model
sharing fix traces to #30093, which closes #30087; both entries carried only
their commit links.

* docs: memory review gating, channel message quotes, knowledge filter, tool result flushing, localized suggestions

Cover the last dev batch: the background memory review now stops when memory
is switched off or the account is barred (#30309), a deleted channel message
clears the quotes pointing at it (#30314, #30313), the knowledge File content
filter applies on the first click (#30211, #30210), tool results are published
when the round finishes instead of waiting on the stream throttle, with channel
and continuation replies no longer carrying picture data a tool handed the
model, and a fresh install now suggests starter prompts in the interface
language, with a way back to the defaults from the model defaults panel.

* docs: playground image edit request shape

The Images playground edit call wrapped its fields under a heading the
endpoint never read, so every edit came back rejected; the request now carries
them flat (c07fa08b9).

* docs: same-origin diagrams and SVG, read-only folders on the empty-chat page, placeholder translations

Cover the latest dev batch: diagrams and SVG previews only follow same-origin
and data references and a diagram reaching outside is refused (#30271), the
empty-chat page no longer shows folder edit controls for a folder shared
read-only (ee3ece1e2), and translated strings that had their placeholder
names localized or their braces lost are restored across the thirteen locales
(#30326).
2026-09-21 15:19:47 -04:00
Nas
9530cc15c1
i18n: restore placeholder names that were translated or lost their braces (#30326)
{{ models }} -> {{ modelli }} etc. in 10 locales, {{name}} -> {{nombre}} in
gl-ES, {{file}} -> {{arxiu}} in ca-ES, and single/unbalanced braces in kab-DZ,
nb-NO and ca-ES. 22 strings in 13 locales, tokens only.
2026-09-21 15:16:04 -04:00
Classic298
46ea826e01
refac: keep rendered diagrams and SVG on same-origin resources (#30271)
Mermaid diagrams and the shared SVG sanitizer now accept only same-origin and data: references. Image URLs, class styles and directive config are checked on the parsed diagram before it renders, and the sanitizer drops attribute values and stylesheet rules that point at another origin. use elements keep local #id references only.

SVGPanZoom and the SVG file preview call the shared sanitizer instead of keeping their own configs, so SVG artifacts and uploaded SVG files follow the same rule. Uploaded SVGs that reference external sprites now render those parts blank, which is the point of the change.

A diagram that references an external resource now reports an error instead of rendering, and an SVG artifact keeps everything except the rules that reference one.
2026-09-21 15:15:28 -04:00
Timothy Jaeryang Baek
ee3ece1e2b refac 2026-09-21 14:43:04 -04:00
Timothy Jaeryang Baek
c07fa08b99 refac 2026-09-21 11:45:49 -04:00
Timothy Jaeryang Baek
bf476452a8 chore: format 2026-09-21 11:35:30 -04:00
G30
b988f06ce0
fix: drop every local reference to a channel message that was deleted (#30314) 2026-09-21 11:32:14 -04:00
G30
4b61b86a83
fix: apply the knowledge File content filter on the first click (#30211) 2026-09-21 11:30:42 -04:00
Timothy Jaeryang Baek
9fba2843b1 refac 2026-09-21 11:24:40 -04:00
Classic298
e8c26f8394
fix: stop the background memory review when memory is switched off (#30309)
With memories disabled instance-wide, or for a user barred from the feature,
the background review still ran every interval turn: it spent a task-model
call drafting memory operations and only then failed at the write, because
the router's permission check rejected it. The review now checks the
'memories.enable' switch and re-checks the 'features.memories' permission
the same way the context-injection path already does, so a model whose memory
capability is on no longer triggers memory work that can never land.

The permission lookup costs a groups query, so it runs last, after the free
config and interval gates; those stay on every turn's hot path.
2026-09-21 11:22:21 -04:00
Timothy Jaeryang Baek
478d1785fd refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-09-21 11:21:42 -04:00
Timothy Jaeryang Baek
30881dbcc9 refac 2026-09-21 11:09:09 -04:00
Classic298
419d248093
refac: check the timer owner's role before running a due timer (#30220)
The due-timer executor rehydrates the owner from the database and now
verifies that the owner is still a user or an admin before entering the
chat completion pipeline, mirroring the check the scheduled-automation
executor already performs. A timer whose owner no longer qualifies is
recorded as an error instead of being run.
2026-09-21 10:46:48 -04:00
Timothy Jaeryang Baek
3fc1146c13 refac 2026-09-21 10:44:37 -04:00
Classic298
94eea41a75
fix: surface searchapi errors, news results and redirect links (#30308)
Web search via searchapi.io could come back empty or near-empty with no
hint of why: an invalid or expired API key turned into an empty result
set instead of an error, the google_news engine splits its results
between organic_results and top_stories and only the first block was
read, and google links came back as google.com/goto redirects the web
loader cannot fetch, so citations pointed at a redirect blob.

The search now reads both result blocks, asks google engines for
resolved destination links, raises on HTTP errors, carries a 30s request
timeout, skips result rows without a link, and logs the response body at
debug instead of dumping every search at info.

Fixes #30305
2026-09-21 10:44:06 -04:00
G30
9688d32764
fix: keep the background image chosen while creating a folder from the sidebar (#30218) 2026-09-21 10:36:49 -04:00
Timothy Jaeryang Baek
e8bd0661d3 refac 2026-09-21 10:30:53 -04:00
Timothy Jaeryang Baek
754c4b5762 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-09-21 10:25:20 -04:00
G30
c864e3ae7c
fix: stop asking for chat variables a model's system prompt no longer declares (#30173) 2026-09-21 10:23:21 -04:00
G30
bc50026f7b
fix: make the Delete Chat shortcut work whenever a chat is open, not only while its sidebar row is rendered (#30165) 2026-09-21 10:21:19 -04:00
Timothy Jaeryang Baek
a9541c18ca refac 2026-09-21 10:19:59 -04:00
Classic298
aceda892bd
perf: serve the shared model pool from a per-worker cache (#28176)
With the Redis websocket manager the shared model pool is a `RedisDict`, so resolving a model id fetches from Redis, and more than one of those fetches pulls the whole pool. Every chat request pays that latency and that traffic, and the cost grows with the number of models configured: at 120 models a request moves several hundred kilobytes to ask about models that have not changed since the last request.

The pool now keeps a per-worker cache of the hash and refetches it only when the signature key that `set()` already maintains changes. A cache may only trust a signature that describes the bytes it fetched, so every write invalidates the signature and no signature value is ever issued twice. Without a fresh token each time, a pool that flaps back to an earlier state returns to an earlier digest, so a reader that races the write keeps serving the old pool. `delete_many()` was a second hole in that: it updated a dead attribute and never touched the signature, so readers kept serving models it had already removed. A TTL cache would have been simpler, but it serves a pool it already knows may be stale and costs the same single round trip this check costs.

Replaying the real access pattern, a request that finds the pool unchanged moves about 1 KB no matter how many models are configured, over the same number of round trips as before, except after a write that leaves no signature, which holds reads at two commands until the next non-empty `set()` writes one. A request that first sees a changed pool costs one extra command, and so does rewriting the pool. Values stay cached serialized, so `items()` and `values()` still decode on every call, and each worker holds the whole pool in memory.
2026-09-21 10:15:49 -04:00
Classic298
49b25506dd
perf: stop round-tripping the whole chat to read or write one message (#28184)
Long conversations get progressively more expensive to stream into. Every message-level write reloads the entire chat JSON, walks every string in it for a null-byte sanitisation pass and rewrites the whole column, and reading a single message loads and validates the whole chat too. The websocket event emitter reads and then writes, so one message event round-trips the full conversation twice to change a few hundred bytes.

Reading one message now selects only history.messages out of the JSON column and indexes it in Python, so the rest of the chat is neither loaded nor model-validated. Writing one message now sanitises only what is actually entering the chat, plus the title column that is mirrored off the blob, instead of re-walking a conversation that was already sanitised when it was written. Legacy rows are still healed on read by get_chat_by_id, which is unchanged.

One behaviour change: chats written before the null-byte sanitisation existed can still hold null bytes in the stored JSON. Those rows used to be rewritten clean as a side effect of any message write, and are now cleaned when the chat is read instead. The visible consequence is that the message write and delete endpoints echo back the updated chat, and for such a legacy row that echo now carries the raw null bytes rather than stripped ones, until the next read of that chat heals it. A GET of the chat is unaffected, and title, the one text column that PostgreSQL cannot store a null byte in, is still sanitised on every write.

Measured on SQLite with a 10.8 MB chat (3000 messages): a single-message upsert is 227 ms before the branch and 173 ms after, and the legacy single-message read is 135 ms before and 79 ms after. The whole-blob scan the delete path used to run costs 5.40 ms on a 1.8 MB chat and 38.79 ms on the 10.8 MB one, and is gone.

Ref https://github.com/open-webui/open-webui/issues/28169

### Contributor License Agreement

<!--
🚨 DO NOT DELETE THE TEXT BELOW 🚨
Keep the "Contributor License Agreement" confirmation text intact.
Deleting it will trigger the CLA-Bot to INVALIDATE your PR.

Your PR will NOT be reviewed or merged until you check the box below confirming that you have read and agree to the terms of the CLA.
-->

- [x] By submitting this pull request, I confirm that I have read and fully agree to the [Contributor License Agreement (CLA)](https://github.com/open-webui/open-webui/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT), and I am providing my contributions under its terms.

> [!NOTE]
> Deleting the CLA section will lead to immediate closure of your PR and it will not be merged in.
2026-09-21 10:11:38 -04:00
Timothy Jaeryang Baek
5fb869db22 refac 2026-09-21 08:59:39 -04:00
Classic298
488b3c571c
perf: stop note co-editing echoing every remote update back to the server (#28185)
Every client in a note re-broadcasts each update it receives, with a full content snapshot attached, and the server appends each echo to the document log and writes the note again. Traffic and note writes therefore scale with the number of people who have the note open: every extra participant adds one more full echo of every keystroke.

Remote updates are now applied with the `'server'` origin the state path already uses, which the local listener ignores, so the echo stops. The echo did carry one thing worth keeping: the receiving client holds the merged document, which the sender had not seen yet, so each receiver now sends a content-only message once the edits settle, debounced 500ms, and flushes a pending one when the editor is torn down. The backend accepts an update message with no `update` field for that case, where it previously raised and dropped the save.

Replaying keystrokes at 120ms with a real Yjs document, socket messages fall 48.8% with two clients, 65.6% with three and 79.2% with five, with byte counts tracking the same on notes up to 50KB. Server-side update appends drop by a factor of the client count. The cost is one snapshot upload per receiving client per typing pause, which the server's own debounce then collapses into a single extra note write however many people are watching. A snapshot has to come from a client because the server cannot rebuild the markdown, HTML and JSON shape the note record stores. It also moves the merge 500ms later than the echo delivered it, so if two edits cross on the wire and every editor then loses its connection and closes inside that window, the note keeps what it was last sent and one of the two edits is lost. A client still connected at teardown flushes its pending snapshot, and any other editor left in the note closes the gap.
2026-09-21 08:53:27 -04:00
Classic298
6fb68e43cb
refac(images): request an unencoded body when fetching remote chat images (#29623)
The remote chat-image fetch now asks for an identity-encoded response and skips one that comes back content-encoded.

An image whose host stores and echoes a `Content-Encoding` regardless of what the client asks for (an S3 or MinIO object uploaded with that metadata) is no longer inlined; the message is forwarded with the original URL instead, the same way an unreachable image already behaves.
2026-09-21 08:53:06 -04:00
Classic298
438d9db8db
fix: correct recurrence rule parsing for schedules and calendar events (#29262)
* fix: correct recurrence rule parsing for schedules and calendar events

An automation set to repeat a limited number of times, say ten or a hundred, was treated as a one-shot and reported no repeat interval, because any count whose digits began with a one matched a text check for the one-shot case. The scheduler already answers that question correctly by asking the rule for its next two occurrences, so the text check is gone and the count is read as the number it is.

A recurrence rule that carries its start date on the same line as the repeat text kept that date when the automation was parsed, so the schedule ran from whatever date the rule happened to carry and ignored the start the user picked. The filter that drops the start date now splits the rule on any whitespace, the same way the rule parser itself does, so both agree on where one part of the rule ends and the next begins.

The same mismatch on the calendar path anchored a recurring event to the date inside its rule, so occurrences showed up before the event had begun and at the wrong time of day. That filter splits the rule the same way now, and the series starts at the event's own start.

All three come from one place, recurrence rules being matched and cut as text. Rules written across several lines, which is what the schedule and calendar editors produce, behave exactly as before.

* refac: name rrule token vars parts to match calendar.py
2026-09-21 08:44:13 -04:00
Timothy Jaeryang Baek
0180efecf3 refac 2026-09-21 08:43:45 -04:00
Classic298
7fa8673296
fix: resolve this instance's own file URLs in edit_image regardless of the URL host (#29691)
The native edit_image tool fails with "400: [ERROR: Error loading image]" whenever the model hands it an absolute URL for an image Open WebUI already stores. Such a URL is treated as local only when its host string matches the incoming request's host exactly, so a default-port form, a container name or any host the model composed itself falls through to an outbound HTTP fetch instead. That fetch asks /api/v1/files/{id}/content without a session, gets a 401, and the user sees the generic 400.

Match the file URL on its path and let the existing local branch resolve it. Fetching that endpoint over the network can never succeed for a local or a remote instance, because it requires an authenticated user, so the host comparison only decided which way the request failed. Access control is unchanged: the local branch still goes through get_file_content_by_id, which enforces owner, admin or shared access.

Fixes #29220
2026-09-21 08:32:38 -04:00
G30
d9ea46f92d
fix: hide Insert into note on a note the viewer may only read (#30223) 2026-09-21 08:32:17 -04:00
G30
f40318c1ff
fix: keep the live prompt unchanged when a version is saved without Set as Production (#30231) 2026-09-21 08:32:02 -04:00
G30
440d13e198
fix: refresh the prompt editor after a version is set as production (#30233) 2026-09-21 08:31:51 -04:00
Timothy Jaeryang Baek
7f703dcc98 refac 2026-09-21 08:31:27 -04:00
Timothy Jaeryang Baek
dc4f5da5a0 refac 2026-09-21 08:28:36 -04:00
G30
c2f11edb2f
fix: keep a workspace model's base model when cloning it from the admin Models settings and hide Clone for arena models (#30080) 2026-09-21 08:15:12 -04:00
Classic298
f5967aebd7
fix: restore chat input draft after reload in chats started from the home page (#29762)
Typing a message, uploading a file or toggling web search in a chat started from the home page was lost on reload: the input came back empty and the uploaded files were gone for good.

Drafts are kept per chat in sessionStorage, but the key was built from the route prop, which stays empty for these chats because creating the chat only swaps the URL with history.replaceState and never re-runs the route load. Drafts were written under the new-chat key while a reload of /c/<id> read the per-chat key. Falling back to the active chat id lines the two up, the same fallback loadChat already uses for this reason.

The Placeholder submit handler also cleared the bare key rather than the current one, which left a stale draft behind once a chat's messages had all been deleted.

Verified against the base commit on a local instance: on base the draft lands under chat-input and is lost on reload, with the change it lands under chat-input-<id> and both the text and the uploaded file come back. Image attachments stay excluded from drafts by design.

The model selection half of #29760 has a different cause and is not fixed here.

Refs #29760
2026-09-21 08:08:45 -04:00
Classic298
f16e9eabe3
feat: fall back to an available model when a chat's models are gone (#29757)
Retiring a model leaves every chat that used it stuck. The chat still stores the old model id, so it loads with nothing usable selected and refuses to send until the user picks a replacement by hand, in every old conversation. There is no admin-side way to move those chats across.

Loading a chat now drops model ids that no longer exist, and an empty selection falls into the fallback this function already applies to new chats: the user's default model, then the admin-configured default, then the first available model. A chat that still has one live model keeps it. The filter runs before the single-model permission clamp so a restricted user keeps a live model the chat already has instead of losing it to the clamp.

The model list cannot distinguish a retired model from one whose connection is momentarily unreachable, so during such an outage a chat pinned to that connection will open on the fallback model and record it when saved. Models defined in the workspace are unaffected, since they stay in the list while their connection is down.
2026-09-21 08:08:27 -04:00
G30
5399e3977d
fix: enforce the account form's required fields the way the admin user dialog does (#30296) 2026-09-21 08:08:02 -04:00
G30
aa65f55692
fix: insert the whole emoji when its shortcode is a codepoint sequence (#30213) 2026-09-21 08:07:45 -04:00
Classic298
9012bd153d
fix: stop sending count to the Staan search API (#30303)
Staan rejects any count other than its fixed page size of 10, so every search 400ed with the default result count of 3. The API always returns 10 results; the local results[:count] slice already applies the configured count, so the request parameter is simply dropped.
2026-09-21 08:04:26 -04:00
G30
7eebbbb891
fix: stop the MinerU API key blocking a save in local mode (#30299) 2026-09-21 08:04:19 -04:00
andreadegiovine
e9017c7da4
i18n: update and improve it-IT (#30297)
* i18n: update and improve it-IT

* i18n: update and improve it-IT

---------

Co-authored-by: andreadegiovine <andreadegiovine@gmail.com>
2026-09-21 08:04:08 -04:00
G30
702da1e471
fix: stop injecting memories and memory tools when memory is switched off (#30228) 2026-09-21 01:00:08 -04:00
Timothy Jaeryang Baek
85146206f6 refac 2026-09-21 00:59:46 -04:00
G30
14be67348a
fix: open the new group dialog empty after a group is created (#30280) 2026-09-21 00:56:53 -04:00
Timothy Jaeryang Baek
97e013a661 refac 2026-09-21 00:51:28 -04:00
G30
394fcc7e24
fix: keep an unsaved code block edit visible after it is collapsed (#30284) 2026-09-20 23:44:16 -05:00
G30
9758f33b6c
fix: save a model whose description is switched from default back to custom (#30249) 2026-09-20 23:42:44 -05:00
G30
669d37bea1
fix: stop offering reactions on a channel the viewer may only read (#30241) 2026-09-20 23:42:32 -05:00
G30
b256115fa2
fix: save the context compaction threshold from general settings (#30270) 2026-09-20 23:42:22 -05:00
G30
402a187e82
fix: return no chats when a folder search term matches no folder (#30273) 2026-09-20 23:39:54 -05:00
G30
f2220d2e40
fix: keep citation chips inside bold, italic and linked text (#30278) 2026-09-21 00:33:47 -04:00
G30
4c00e08aa1
fix: keep a message's attachments when its edit is cancelled (#30281) 2026-09-21 00:31:26 -04:00
Classic298
8e4cc946ce
fix: emit the resolved file path in terminal file events (#30282)
When a model calls display_file, write_file or replace_file_content with a relative path, Open Terminal resolves it against the session working directory and returns the absolute path, but the event sent to the browser carried the raw argument instead. The file panel matches that string against the file browser root, a relative path never matches, so the preview never opens, the panel jumps to the root and the session working directory is rewritten to the root. With the root turned off (OPEN_TERMINAL_FILE_BROWSER_ROOT=filesystem) there is nothing to clamp to and the relative string is sent to the terminal as the new working directory, moving it silently. The tool call itself succeeds either way, so the failure only shows up as a panel that will not open the file the model just wrote.

Both events now carry the path from the tool result and fall back to the argument when the result cannot be read, which is what build_terminal_file_tool_result already does for the chat file attachment. The same one-line rule is applied to the direct tool server path in the frontend, where the browser runs the tool itself and the write_file branch beside it was already correct.

Checked against a live Open Terminal: relative arguments now emit the absolute path, absolute ones are unchanged, non-existent files and inline displays still emit nothing, unreadable or error results still fall back to the argument, and run_command is untouched.

Related to #30051
2026-09-21 00:31:14 -04:00
G30
17e7e8f5e9
fix: show a code run's error output alongside what it printed (#30286) 2026-09-21 00:30:55 -04:00
G30
f847c6a588
fix: apply the new text to speech engine's default voice and model when the engine changes (#30289) 2026-09-21 00:30:26 -04:00
Classic298
043784c2d0
fix: route external webhook avatar URLs through the profile image endpoint (#29892)
Some checks failed
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Python CI / Ruff Format (3.11) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
The channel webhooks modal rendered a webhook's stored profile_image_url straight into an img tag, so opening it sent every viewer's browser to whatever external host that URL named, leaking client IP, User-Agent and Referer no matter how the server was configured.

External URLs now render through the same webhook profile image endpoint the message list already uses, leaving it to the server to decide whether the browser is sent to that host. Locally picked images and the paths Open WebUI assigns itself still render inline, so previewing an upload before saving is unchanged, and the value written back on save is untouched.

This only takes effect together with the companion backend change that gates the endpoint on ENABLE_PROFILE_IMAGE_URL_FORWARDING. Until that lands the endpoint still redirects and the browser still reaches the external host. Also worth knowing: the endpoint matches the scheme case-sensitively, so a URL stored as HTTPS:// falls back to the default image here, which is already what the message list shows for it.

Verified in a browser against a local origin standing in for the external host: with forwarding on the avatar still renders through the endpoint in both places it appears, with forwarding off the browser makes no request to that origin, and picking a new file still previews immediately before saving.
2026-09-19 17:03:15 -05:00
Classic298
7fa705f3b8
feat: let operators expose chosen file metadata to the model in retrieved sources (#29696)
Custom metadata attached to a file upload now reaches the vector DB, but the model still never sees it. Both prompt-assembly paths build their output from a fixed field set: the classic RAG <source> tag carries only id, name and resource type, and the retrieval tools return only content, source and file id per chunk. A scraper that records where each document came from therefore cannot get that origin in front of the model, so answers cannot state it.

RAG_SOURCE_METADATA_KEYS names the chunk metadata keys allowed through to the model. Configured keys are emitted as extra attributes on the <source> tag and as extra fields on tool result chunks, covering both retrieval paths. It is empty by default, so nothing changes for existing deployments.

An allowlist instead of passing everything through, because chunk metadata also carries file hashes, collection names, embedding config and relevance scores, which would then be added to every retrieved chunk of every request. Values are attacker-controllable through an uploaded file, so they are escaped before they go into the tag, and a configured key can never displace a field the tag or the chunk already defines.

Reported in open-webui/open-webui#29486.
2026-09-19 17:02:59 -05:00
Timothy Jaeryang Baek
3a6d0fd203 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-09-19 17:54:09 -04:00
Timothy Jaeryang Baek
8b3ee28272 refac 2026-09-19 17:45:05 -04:00
G30
65a3a115eb
fix: keep the speech-to-text extension allowlist when the Audio settings are saved (#30208) 2026-09-19 17:36:38 -04:00
G30
881f7c9400
fix: save only the changed key when pinning, reordering or picking a default outside the Settings modal (#30183) 2026-09-19 17:34:02 -04:00
Timothy Jaeryang Baek
f6922a4c42 refac 2026-09-19 17:33:41 -04:00
Timothy Jaeryang Baek
67adde3193 refac
Co-Authored-By: G30 <50341825+silentoplayz@users.noreply.github.com>
2026-09-19 17:26:40 -04:00
Timothy Jaeryang Baek
10d1cfe637 refac 2026-09-19 17:26:18 -04:00
G30
c1a35b4691
fix: sort the feedback history by user when the User column is clicked (#30191) 2026-09-19 17:21:51 -04:00
G30
76bd90c832
fix: discard unsaved edits when the admin Edit User dialog is closed (#30193) 2026-09-19 17:21:44 -04:00
G30
4922b91c07
fix: stop the Model Defaults save from reverting the selected, pinned and ordered models (#30206) 2026-09-19 17:17:57 -04:00
G30
27dd191332
fix: let a calendar event's repeat, description and location be cleared again (#30204) 2026-09-19 17:17:43 -04:00
G30
757ee8132e
fix: stop collapsing the task list from sending the message being typed (#30202) 2026-09-19 17:17:30 -04:00
G30
5e1f995f01
fix: show a custom gender back in the account form instead of an empty dropdown (#30215) 2026-09-19 17:16:56 -04:00
G30
b64cb06043
fix: let a user without the chat delete permission delete a folder while keeping its chats (#30163) 2026-09-19 16:07:34 -05:00
Timothy Jaeryang Baek
946be43237 refac 2026-09-19 17:05:57 -04:00
G30
25948ea213
fix: block saving a prompt's input form when a required dropdown has no option chosen (#30078) 2026-09-19 15:42:39 -05:00
G30
3fbc39a9dc
fix: save a new feedback entry when rating a reply whose entry was deleted (#30077) 2026-09-19 15:42:25 -05:00
G30
a63ddb36d2
fix: clear a reply's score and reason when its rating switches thumbs (#30075) 2026-09-19 15:42:15 -05:00
G30
57063aaad8
fix: dropping a folder onto the folder it is already in no longer fails with Folder already exists (#30169) 2026-09-19 15:41:10 -05:00
G30
2690d04cac
fix: clean up orphan tags for the chat's owner, not the admin, when an admin deletes another user's chat (#30171) 2026-09-19 15:40:58 -05:00
G30
af9ef1dcbc
fix: hide the group picker from users who may not share with groups (#30189) 2026-09-19 15:34:33 -05:00
G30
621ed6de7f
fix: export every prompt and model from the workspace, not only the page on screen (#30187) 2026-09-19 10:03:17 -05:00
G30
9e293a58ea
fix: label the search modal's archive action Unarchive for archived chats and report what happened (#30177) 2026-09-19 10:02:11 -05:00
G30
dbe538c033
fix: keep a note's sharing when a collaborator saves it (#30175) 2026-09-19 10:00:39 -05:00
G30
d16fad585d
fix: keep a shortcut recording from firing the action already bound to the chord (#30160) 2026-09-19 10:00:18 -05:00
G30
2af297bfcd
fix: keep the pinned flag when importing a chat export (#30151) 2026-09-19 09:59:55 -05:00
G30
36276c9ead
fix: keep a skill's translations and disabled state when saving it from the editor (#30185) 2026-09-19 09:59:36 -05:00
Classic298
1f8f1f61bb
fix: tell ask_user models that the first option is shown as Recommended (#30196)
The ask_user card badges the first option of every question as "Recommended", but nothing ever told the model that. The model picks whatever order it likes, so the badge really means "listed first" and users act on a recommendation the model never made.

The ask_user tool description now states that the first option is labelled Recommended and that the model should list the option it recommends first, so the badge reflects an actual choice.

Kept the badge and instructed the model instead of adding a per-option "recommended" flag: the flag would need a schema change, validation in the request normalizer and a frontend change, for the same result in the common case. Dropping the badge was the other option, but it removes a useful affordance rather than fixing it.

Verified that the added line reaches the model by running the docstring through the tool-spec builder and checking the generated OpenAI function schema.

Fixes #30195
2026-09-19 09:59:18 -05:00
G30
5d7c17f41a
fix: stop the Generate Message Pair shortcut from also submitting the message input form (#30167) 2026-09-18 23:54:10 -04:00
Classic298
3d29548716
fix: pgvector reads leak their connection and lose most of their neighbours (#30142)
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
Two defects on the pgvector read path. A search, query or get that finds nothing returns before the rollback that ends its read-only transaction, so the session keeps the connection it checked out; retrieval fans out over worker threads and the session is thread-local, so every thread that runs an empty read holds a connection for the lifetime of that thread. Users see vector search die after a while with "QueuePool limit of size 5 overflow 10 reached" and no way back other than a restart. Rolling back before the three early returns puts the connection back: measured on PostgreSQL 16, twelve empty reads on a default-shaped pool left all 12 connections checked out before and 0 after.

All collections also share one table under one vector index, and the WHERE collection_name filter is applied after the index walk, so a knowledge base holding a small share of the rows keeps only a small share of its neighbours, silently. pgvector 0.8 added iterative scans for this: the scan keeps going until enough rows pass the filter. This sets it per search so it cannot leak into other sessions, and only when the installed extension supports it, since setting it on pgvector 0.7 would make every search raise. PGVECTOR_ITERATIVE_SCAN turns it off or picks strict_order; it defaults to relaxed_order because the shipped behaviour is silently wrong results, and the cost is about a millisecond per search.

Measured through PgvectorClient.search on PostgreSQL 17 and pgvector 0.8, 101500 rows of 384 dimensions with the knowledge base at 1.5% of the table, hnsw m=16, recall@10 against an exact scan over 40 queries: 0.070 at 4.9 ms before, 0.970 at 7.2 ms after.

Fixes #30133
Fixes #30135
2026-09-18 19:31:49 -04:00
Classic298
52cd298411
fix: allow forking a chat that holds a stale unfinished message (#30131)
Forking returned 409 "Wait for the current response to finish before forking." forever once any assistant message anywhere in the chat was left at done = false, with nothing generating and even when that message sat on a branch that was not being forked. Interrupted turns leave the flag behind and nothing clears it, so an affected chat could never be forked again.

The endpoint now refuses only while a task is actually running, which is the check /compact has always relied on by itself. A message still unfinished on the forked branch is marked done in the copy so the fork does not open showing a spinner, while a turn paused waiting on tool approval keeps its unfinished state and its pending call so the fork still shows the prompt. The source chat is left untouched either way.

The client-side check had the same shape of bug: it read history.currentId instead of the message actually being forked, so a paused last turn also blocked forking earlier, finished messages. It now reads the message it is about to fork.

Fixes #30128
2026-09-18 19:31:20 -04:00
Classic298
fffcb727d2
fix: release the openGauss connection when a read returns no rows (#30144)
An openGauss query or get that finds nothing returns before the rollback that ends its read-only transaction, so the session keeps the connection it checked out of the pool. Retrieval fans out over worker threads and the session is thread-local, so every thread that runs an empty read holds a connection for the lifetime of that thread, and vector search eventually fails with a pool checkout timeout that a restart is the only way out of. Empty reads are routine: an empty knowledge base, a file whose chunks were deleted, or a metadata filter that matches nothing all produce one.

Rolling back before the two early returns puts the connection back. Measured by driving OpenGaussClient itself with a pool of 5: three empty query reads left 3 connections checked out and 0 free before the change, and 0 checked out with 3 free after, same for get. search is already correct, it has no early return.

The pgvector client had the same defect, fixed separately in #30142.
2026-09-18 19:31:01 -04:00
Classic298
6d73780fd1
fix: bundle seaborn and give black its dependencies in the Pyodide lock (#30148)
seaborn was listed among the Pyodide distribution packages, but it is not part of that distribution, so the build wrote no wheel for it. Since seaborn is on the code interpreter's package list, it was quietly downloaded from pypi.org in the user's browser instead, which is what bundling the wheels is meant to avoid. It now takes the PyPI wheel path, like openpyxl.

The code editor's Format button was broken for non-admin users, with or without internet, because black's bundled lock entry declared no dependencies and the hand-written list next to the caller was missing packaging. black's dependencies are now declared in the lock, so that list goes back to black alone and every future caller gets them too.

Both of these shipped unnoticed because nothing checked that a listed package produced a wheel. The build now fails when one did not, naming it.

Verified offline with the network cut at undici's dispatcher: seaborn, black, the Excel roundtrip and the other listed packages all install from the bundled wheels, and a deleted wheel, a missing lock entry or a non-canonical key each fail the build.

Fixes #30145
2026-09-18 19:30:54 -04:00
G30
677341578d
fix: stop the model editor's Builtin Tools section from reading its labels before they exist (#30153) 2026-09-18 19:29:41 -04:00
G30
4d01f1ebed
fix: keep the archived state and chat variables when importing a chat export (#30155) 2026-09-18 19:29:30 -04:00
Timothy Jaeryang Baek
64bbdf7a73 refac 2026-09-18 19:28:40 -04:00
joaoback
f3cf833e54
i18n: add pt-BR translations for newly added UI items and consistency pass (#30158)
New **pt-BR** translations for items introduced in the latest releases, plus a consistency/quality pass across existing strings (grammar, tone, capitalization, pluralization). Placeholders and hotkeys preserved. No logic changes.
2026-09-18 19:18:45 -04:00
Classic298
d5cacb3c0a
fix: convert forced tool_choice and non-streaming tool calls for Responses API connections (#30095)
On a connection with API type "responses", a forced tool choice sent in
the Chat Completions shape was forwarded to the provider unchanged, so
providers that validate the Responses API schema rejected the request.
A non-streaming reply that carried a function call was also flattened
to empty text with finish_reason "stop", so API clients never saw the
tool call.

The request converter now flattens a forced function choice to the
Responses API shape, and the result converter turns every function_call
output item into a Chat Completions tool_calls entry, mirroring the
existing Responses output mapping in utils/misc.py.

Fixes #30085
2026-09-18 19:18:26 -04:00
Timothy Jaeryang Baek
c82634b9d0 refac 2026-09-18 11:49:56 -04:00
Timothy Jaeryang Baek
b513a6ad3b refac 2026-09-18 11:40:44 -04:00
Classic298
9ee810ad09
feat: add Staan as a native web search provider (#30138)
European deployments that need EU data residency have no hosted web search option out of the box: every plug-and-play provider shipped today is US-based, and the only sovereign alternative is self-hosting SearXNG, which means running and maintaining that infrastructure yourself. Staan (staan.ai) is a European search API with EU data residency, so adding it gives those deployments a drop-in choice.

It is configured like any other provider, through the admin UI or STAAN_API_KEY, STAAN_MARKET and STAAN_MAX_SNIPPETS. Market sets the region and language of the results and defaults to en-us. Max snippets asks Staan to fetch each result page and return semantically scored chunks of it, which get merged into that result's snippet so retrieval has more to work with; leaving it at 0 uses the plain search endpoint.

Wired the same way as Tavily and Exa, with domain filtering going through the shared get_filtered_results.

Requested in #26006.
2026-09-18 10:38:36 -05:00
Classic298
655337f7bf
fix: vendor openpyxl so Excel I/O works in the Pyodide code interpreter (#30140)
openpyxl has been listed as a Pyodide package since March, but it is not part of the Pyodide distribution, so the build never wrote a wheel for it and never warned about it. In the browser code interpreter every Excel operation failed with ModuleNotFoundError: import openpyxl, pd.read_excel() and DataFrame.to_excel() alike.

It now takes the PyPI wheel path together with its dependency et-xmlfile, and the code interpreter installs it when a snippet imports openpyxl or calls pandas' Excel helpers, which need it without importing it by name.

Injected lock entries are now keyed by the canonical dashed package name, the only spelling pyodide resolves them by. As a side effect black's mypy-extensions now resolves from the bundled wheel too, so formatting Python in the editor no longer reaches out to PyPI at runtime.

Verified offline against a built static/pyodide with every network call blocked: a to_excel then read_excel roundtrip loads openpyxl and et-xmlfile from the local directory and returns the frame.

Fixes #30130
2026-09-18 10:38:02 -05:00
Classic298
95406fd28d
fix: build the pgvector ivfflat index once there are rows to cluster on (#30143)
ivfflat places its centroids by clustering the rows it can see when the index is built, so an index built on an empty table gets centroids that mean nothing, and everything inserted afterwards is filed against them. On a fresh install the index is created immediately after the table, before a single chunk exists, and it is never rebuilt, so that install keeps a permanently untrained index and quietly retrieves the wrong chunks. An install that upgraded into the version introducing the index is unaffected, its table already had rows.

The index is now created once the table holds 50 rows per list, the sample size ivfflat itself aims for. Below that, and until the next start, searches fall back to an exact scan, which is correct and costs about a millisecond at that size. hnsw is untouched, it builds its graph as rows are inserted and has nothing to train on. One trade-off: the build moves from the first start to that later one, so an instance that has grown large in between pays a one-time index build during startup.

Measured on PostgreSQL 17 and pgvector 0.8 with the default lists=100 and probes=1, 384 dimensions, recall@10 against an exact scan, varying only how many rows existed when the index was built:

| rows at build | 200 | 1000 | 2500 | 5000 | 20000 |
|---|---|---|---|---|---|
| recall@10 | 0.180 | 0.563 | 0.967 | 1.000 | 1.000 |

Through PgvectorClient.search on a 20000-row table, an index built as it is today scores 0.480 against 1.000 built after the rows arrive, at the same 5 ms. An existing install can repair its index with REINDEX INDEX idx_document_chunk_vector, measured to take it from 0.480 back to 1.000.

Fixes #30134
2026-09-18 10:37:25 -05:00
Timothy Jaeryang Baek
1ddba7e2c6 refac 2026-09-18 11:31:59 -04:00
Timothy Jaeryang Baek
ca1eefe293 refac 2026-09-18 11:30:15 -04:00
Timothy Jaeryang Baek
9d98ffcddf refac 2026-09-18 10:55:55 -04:00
Timothy Jaeryang Baek
7cbaabe02f refac 2026-09-18 10:51:00 -04:00
Timothy Jaeryang Baek
ad9da98168 refac 2026-09-17 20:11:05 -04:00
Timothy Jaeryang Baek
dbb17a5725 refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
2026-09-17 19:49:35 -04:00
Classic298
9923c53c10
fix: treat SVG uploads as documents instead of vision images (#30102)
Uploading an .svg to a chat attached it as a vision image input, so the model received a data URI it could not decode. PIL-backed servers answered "cannot identify image file" and OpenAI answered "The image data you provided does not represent a valid image". No setting made it work.

SVG now takes the ordinary file upload path, so its XML source is extracted and indexed and the model can answer questions about it. Rasterizing was the alternative and it would have discarded the part of an SVG a model reads best, the source itself. Raster formats are untouched and still go up as image inputs.

A shared helper replaces the ad hoc image/ prefix checks at the points that decide image input versus document, on both ends. It normalises the content type first, because a stored "image/SVG+xml" or a trailing charset parameter slipped past a plain comparison.

One behaviour change worth knowing: an SVG now needs the model to have the file upload capability, where before it rode in as an image.

Fixes #30100
2026-09-17 17:40:07 -04:00
Classic298
d08b69025d
issue-template: require reproducing on latest AND dev right before submitting (#30104)
Reports for bugs already fixed on dev, sometimes weeks earlier, are common
because reporters check an old version once and never retest. Merges the
two soft checkboxes into one strict, time-bound requirement and warns in
the template body that unreproduced reports get closed without discussion.
2026-09-17 17:39:54 -04:00
Classic298
460f2e7634
fix: surface initial chat title generation failures in the log (#30106)
When title generation for a brand new chat fails, the chat silently keeps the provisional title (the full first user message) and nothing is written to the log at the default level, so there is no way to tell that the feature is broken rather than disabled. The failure was caught by a broad `except Exception` and reported with `log.debug`, which is invisible unless GLOBAL_LOG_LEVEL is set to DEBUG. Issue #29533 describes an outage of this path that survived four releases for exactly that reason.

This logs it with `log.exception` instead, matching how the rest of main.py reports background task failures, so an admin sees one ERROR line plus the traceback naming the real cause.

Nothing else changes: the success path is untouched, the exception is still swallowed so the detached task cannot take anything down with it, and `background_tasks_handler` does not log this exception itself, so there is no duplicate traceback.

Fixes #29533
2026-09-17 17:39:43 -04:00
Classic298
0837f310be
fix: reject Docling conversions that failed inside an HTTP 200 response (#30107)
Uploading a file that Docling declines or fails to convert either dies with `TypeError: argument of type 'NoneType' is not iterable`, or silently succeeds and stores the literal string `<No text content found>` as the document's text, which then gets indexed and handed to the model as if it were the file. Docling returns the conversion outcome inside the HTTP 200 body, so checking only the HTTP status made a refused conversion look identical to a successful one, and the `errors` array that says why in plain words was never read.

Failed and skipped conversions are now rejected with the messages Docling returned, so an unsupported format surfaces as "File format not allowed: example.dxf" and the traceback is gone. The markdown field is also read as nullable, because Docling returns JSON `null` for every content format it was not asked to produce, which any Docling Parameters setting `to_formats` without `md` will hit, and that null was what raised the TypeError.

Successful conversions with empty markdown keep the existing `<No text content found>` placeholder, matching what TikaLoader and the Mistral loader already do in the same package.

Fixes #29808
2026-09-17 17:39:35 -04:00
Timothy Jaeryang Baek
58b36765a7 refac 2026-09-16 22:47:37 -04:00
Classic298
fbc4897269
refac: resolve knowledge file access from the file association (#29937)
Some checks are pending
Python CI / Ruff Format (3.12) (push) Waiting to run
Python CI / Ruff Format (3.11) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
`has_access_to_file` now derives knowledge-base access purely from the knowledge-file association, instead of also consulting the `collection_name` value stored on the file record.
2026-09-16 15:59:25 -04:00
Classic298
ff7f35a30d
fix: stop logging a traceback twice per failed vector search (#29981)
A vector DB outage renders a full traceback twice for every collection and query pair. query_doc and query_doc_with_hybrid_search each log and then re-raise, and the fan-out handler that catches them logs the same exception again, so twenty knowledge bases and five expanded queries turn one outage into hundreds of identical stack traces per message on every replica.

Both helpers lose the try/except that only logged before re-raising. Each fan-out keeps reporting failures through its return value and now carries the collection name back, so one record after the gather names every collection that failed and attaches a single traceback. The hybrid path logs before its existing raise, so a total failure is still reported before the caller falls back to vector search.

Measured with every collection failing, the vector fan-out drops from 1000 traceback records for fifty collections across ten queries to one, and the hybrid fan-out from 100 for ten collections across five queries to one. Results are byte-identical across eighteen and twenty-six scenarios covering healthy, failing, partial and degenerate inputs.

Three limits worth stating. Nine of the fifteen vector clients swallow the error inside search() and return None, chroma and pgvector among them, so those backends see neither the old tracebacks nor the new record. get_sources_from_items queries one attached item at a time, so a chat with twenty knowledge bases still logs twenty records. And with hybrid search on a total outage logs once from each fan-out as the caller falls back, down from two per pair plus one.

The old "All collection queries failed" warning goes with this. The replacement cannot express that case, because a collection can fail one query and succeed another, so naming every failed collection no longer implies an empty payload.
2026-09-16 15:59:05 -04:00
Classic298
556e48e058
fix: honor false defaults on checkbox prompt variables (#30037)
A prompt template with `{{flag | checkbox:default=false}}` opened the input variables form with the checkbox already ticked, and the same happened for `False`, `"false"`, `0` and `"0"`. Template defaults arrive as strings and the checkbox was bound straight to that value, so any non-empty string counted as checked and there was no way to define a checkbox that starts unchecked once a default was given.

The checkbox now reads its ticked state from the value explicitly: it is checked only for `true` (boolean or string) and stays unchecked for anything else, while clicking it still writes a boolean like before. The seeded value itself is left alone, so what gets substituted into the prompt for an untouched form is unchanged, and typing `true` or `false` into the text field next to the checkbox now moves the tick accordingly.

Fixes #30036
2026-09-16 15:44:29 -04:00
Classic298
48fb2b84ba
refac: derive a forked chat's folder from the caller's write access (#30069)
A fork copied the source chat's folder id unchanged. It now keeps that folder only when the caller has write access to it, matching what chat creation and chat moves already do, and is created outside any folder otherwise.
2026-09-16 15:42:03 -04:00
Classic298
befdd86eb1
refac: share one guarded helper for chat branch descent (#30070)
The branch navigation handlers each carried their own copy of the loop that walks a message's childrenIds down to the deepest child, eleven copies in total across the single-response view, the multi-response view and Chat.svelte. They now call one getDeepestChildId helper in the frontend utils, which tracks the ids it has already visited and stops when one repeats, the same way the message list build already does.

The helper also absorbs the "start from the last root message" fallback that two of the call sites repeated inline, so every site is now a single call.
2026-09-16 15:41:56 -04:00
Timothy Jaeryang Baek
d9c8de9c39 refac 2026-09-16 10:38:46 -04:00
Timothy Jaeryang Baek
fd4fc80536 refac 2026-09-16 10:27:38 -04:00
qwist1233-cpu
a0c8593c92
i18n(tr-TR): translate 1056 missing Turkish strings (#30034)
Co-authored-by: qwist1233-cpu <267938209+qwist1233-cpu@users.noreply.github.com>
2026-09-16 10:26:16 -04:00
tsumon
e669f8aefe
i18n: improve Simplified Chinese translations (#30043) 2026-09-16 10:19:44 -04:00
Classic298
3348f68778
fix: keep attachment-only messages intact when skills are bound (#30045)
Sending an image or file with no typed text to a model with skills
bound replaced the (empty) message text with the list of skill names,
so the model answered with its skill catalog instead of handling the
attachment. The empty-message guard added for #24929 fired on any
empty last user text without checking for attachments.

The guard now skips the fallback when the current message carries
files, read from the user_message object the frontend sends with the
request. Attachment-only messages then go out exactly as they do on
models without skills. A bare skill selection with no attachment still
gets the fallback, so the provider 400 from #24929 stays fixed.

Request-level files were not usable as the signal: that list carries
the model's knowledge and folder files, so a bare skill selection on a
knowledge model would have gone out empty again.

Fixes #30040
2026-09-16 10:19:28 -04:00
Classic298
82da9093aa
fix: default converted Responses API function tools to strict=false (#30046)
Fixes #27750

When a model runs on the Responses API, every Chat Completions style tool (workspace tools, MCP and OpenAPI servers) is converted to the Responses tool shape. Only an explicit strict value was carried over, so tools without one were sent with strict omitted. Chat Completions defaults strict to false, the Responses API defaults it to true, so the conversion silently switched loose schemas into strict mode. The model then filled every optional property (empty strings, zeros, false, empty arrays) and combined mutually exclusive filters, and those arguments went to the tool unchanged. Search and filter tools received meaningful values where omission was intended and rejected the call or returned wrong results.

The converter now defaults strict to false when the source tool does not set it. Explicit true or false values are still preserved and already native Responses tools are passed through untouched, matching what the same tool definition gets on Chat Completions.
2026-09-16 10:19:08 -04:00
Classic298
4611394fa6
refac: correct the ENABLE_CHAT_RESPONSE_STREAM_INPLACE_APPEND comments (#30066)
The comments on the opt-in in-place append made it sound like the fast path
can lose streamed text under normal operation. The only way the field can be
emptied is an allocation failure, which means the host is already out of
memory, and the default path (which copies the whole accumulated string on
every chunk) raises in that situation as well. Reword both comments to state
that condition so the flag is not read as unsafe.
2026-09-16 10:08:46 -04:00
Timothy Jaeryang Baek
3394a10b76 refac 2026-09-16 08:51:02 -04:00
Timothy Jaeryang Baek
88e78b7819 refac 2026-09-16 08:35:01 -04:00
Timothy Jaeryang Baek
e9a0164690 refac 2026-09-16 00:34:24 -04:00
Timothy Jaeryang Baek
6d8e63e366 refac 2026-09-16 00:02:48 -04:00
Timothy Jaeryang Baek
dfde08aa73 refac 2026-09-15 23:53:41 -04:00
Classic298
3d6598fccb
fix: steer models to replace_range in the replace_note_content tool description (#30048)
Asked to add a section or change a few lines, models answer with a whole-note
replace_note_content call and the rest of the note is gone, with nothing to
undo because the editor only records versions on chat inserts. The tool
already supports replace_range operations with an expected guard, but its
description never said when to use them or how the offsets work, so models
defaulted to sending the whole note back.

The docstring, which is the description every model receives for this tool
in note chats and normal chats alike, now states the preference for range
edits and the offset, overlap and expected rules the handler enforces.
Verified the text lands in the generated tool spec unchanged and that range
edits, the expected mismatch rejection and whole-note replace behave as
described against a sqlite data dir.
2026-09-15 22:51:49 -04:00
Classic298
06d9d2e7c7
fix: sync Playwright loader spins a CPU core when a page opens a WebSocket (#30050)
Fetching a single URL through the Playwright loader (fetch_url tool, a URL attached to a
chat, the process/web endpoint) never returned when the page opened a WebSocket. The worker
thread stayed at 100% CPU for the life of the process, and every further hit cost another
core, so the whole instance got slow. Web search was unaffected, it uses the async loader.

The sync loader's websocket route handler called the synchronous close(). Playwright runs
websocket route handlers directly on its dispatcher fiber, so that call waited on the very
loop it was blocking and busy-spun forever. The handler is now a no-op: a routed socket only
reaches the network when the handler asks for it, so the page still cannot dial out, and
nothing in the handler waits on the dispatcher any more.

Aborting the upgrade request from the HTTP route handler instead does not work, page.route
never sees WebSocket handshakes and the connection goes through.

Fixes #30024
2026-09-15 22:51:37 -04:00
Timothy Jaeryang Baek
1cdd7aa459 refac 2026-09-15 22:51:15 -04:00
Timothy Jaeryang Baek
66addbd6b4 refac 2026-09-15 22:47:03 -04:00
Timothy Jaeryang Baek
a096961a31 refac
Some checks failed
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Python CI / Ruff Format (3.11) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
2026-09-14 17:57:09 -04:00
Timothy Jaeryang Baek
d25f6c7135 refac 2026-09-14 17:19:00 -04:00
Timothy Jaeryang Baek
9a93b44495 refac 2026-09-14 17:18:34 -04:00
Timothy Jaeryang Baek
cc5479d16d refac 2026-09-14 17:18:26 -04:00
Timothy Jaeryang Baek
dc98e3023f refac 2026-09-14 17:18:19 -04:00
Timothy Jaeryang Baek
58078ab304 refac 2026-09-14 15:47:50 -04:00
Timothy Jaeryang Baek
924a4a10fb refac 2026-09-14 15:46:31 -04:00
Timothy Jaeryang Baek
113c56fc8c refac 2026-09-14 15:46:23 -04:00
Timothy Jaeryang Baek
b6cf23f332 refac 2026-09-13 23:34:35 -04:00
Timothy Jaeryang Baek
e69236bccb refac 2026-09-13 23:33:22 -04:00
Classic298
55b7343be8
fix: keep a failed lock release from masking cancellation (#29979)
When Redis is unreachable at shutdown, the socket cleanup tasks do not stop. release_lock is a bare eval called from the finally of both periodic_session_pool_cleanup and periodic_usage_pool_cleanup, so it raises there and replaces the CancelledError already in flight. The usage task's except Exception then catches the Redis error and carries on reaping after shutdown cancelled it, and the session task ends with a ConnectionError in place of its cancellation.

release_lock now logs the failure and returns. aquire_lock sets the key with ex=self.timeout_secs and renew_lock re-expires it with the same value, so a release that never lands costs at most one lock timeout before another node can take over.

The except names both RedisClusterException and RedisError because the cluster-only types subclass Exception directly, and redis_cluster is a supported configuration, so RedisError alone would miss the outage on a cluster. Genuine bugs still propagate.

Verified against a real Redis: acquire, refusal while held, renew, compare-and-delete release and non-owner release are unchanged, a populated pool reaps identically with identical lock TTL lifecycles, and both tasks now cancel cleanly where before one kept running and the other died with the wrong exception.
2026-09-13 20:37:51 -05:00
Timothy Jaeryang Baek
d372bec704 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-09-13 21:37:12 -04:00
Classic298
601e0e4345
fix: stop tracking tasks under an empty item id (#29980)
An authenticated socket client can grow Open WebUI's memory without bound by sending ydoc updates for an empty document id. create_task files every task under item_tasks[id] whatever the id, while cleanup_task removes it only for a truthy id, so each update leaves a uuid behind for the process lifetime. normalize_document_id passes an empty id through, and such an id also skips the note access check.

create_task now files the task only when an id is provided, which is what its own comment already described and what the rest of the file does: redis_save_task and redis_cleanup_task both guard on a truthy item id, and stop_task normalizes a falsy one away. Nothing is filed, so nothing leaks, and cleanup_task's existing guard correctly no-ops.

This also settles a disagreement between the two backends. For an empty id, list_task_ids_by_item_id, has_active_tasks and stop_item_tasks answered one way with Redis and another without it; they now match. The visible consequence is that an instance without Redis no longer cancels a pending save for an empty document id, which is how Redis instances already behaved.

Measured over 300 calls with an empty id: 300 stale entries before, none after. Behaviour for a normal id is unchanged, including ordering, cancellation and key removal when the last task finishes.
2026-09-13 20:28:54 -05:00
Classic298
d2e62db69b
fix: stop the sign-in rate limiter blocking the loop and leaking memory (#29977)
A slow Redis freezes the whole worker during sign-in, not just the user signing in. RateLimiter held a synchronous redis-py client and signin called is_limited inline from a coroutine, so every attempt did blocking round trips on the event-loop thread, with REDIS_SOCKET_TIMEOUT defaulting to None so nothing bounded the wait. Its Redis methods are now async and take the handle as their first argument, and both handlers pass request.app.state.redis, the async client the lifespan already creates. Building one in the limiter instead would pin its pooled connection to the first event loop that used it.

Without Redis, which is the default single-instance setup, the fallback store leaked. It was keyed by the rate-limit key and pruned a key's expired buckets only when that same key was checked again, so a login email never seen again was never reclaimed, and that email comes straight from an unauthenticated request body. It is now keyed by bucket, so one prune drops every key an expired bucket held, and it lives on the instance: pruning uses the per-instance num_buckets, so a shared store would let a limiter with a short window delete buckets a longer-windowed one still needs.

With a Redis costing a second per call, the widest event-loop tick gap drops from 2.010s to 0.010s and a concurrent request is answered at 0.05s instead of 2.05s, at no cost to the caller's own latency. Across 20,000 distinct keys the store goes from 40,000 entries and 6.4 MB, growing linearly, to a flat 1,004 entries and 100 KB. Rate-limiting decisions are unchanged across 700,000 randomised calls over 14 window, bucket and limit combinations, against a real Redis and the in-memory fallback alike, and sign-in still returns its first 429 on attempt 16.

Two behaviour changes worth naming. Pruning is now global rather than per key, so a wall clock that jumps forward past a full window and back forgets a hit it previously kept. The two limiters also stop sharing a store, which previously let a sign-in attempt with an IP-shaped email touch the token-exchange limiter's counters.
2026-09-13 20:28:41 -05:00
Classic298
263e56e272
fix: keep the session pool reaper alive through a Redis error (#29976)
A single Redis blip permanently stops orphaned websocket sessions from being reaped. periodic_session_pool_cleanup acquires its lock outside the try, and that try has only a finally, so the first timeout or connection reset ends the coroutine for the life of the process. The session pool then only grows, and the sole trace is one "Task exception was never retrieved" at shutdown.

The loop body gets the same try/except Exception its sibling periodic_usage_pool_cleanup already has, which also brings the lock acquire inside the guarded region. The task now logs, releases the lock and retries after the existing delay, so another node can take the lock over meanwhile.

The diff reads long because the body is re-indented one level; nothing changes beyond indentation and the four added lines. Only Redis deployments are affected, since the lock functions are lambda: True otherwise.

Verified by injecting a ConnectionError at each of the four failure points (acquire, renew, the batch scan, the reaping delete), against a real Redis as well: the task survives all four and keeps retrying, where it previously died on the first. Reaping results, lock acquire and release counts, and cancellation at shutdown are unchanged.
2026-09-13 20:28:10 -05:00
Classic298
38a8dc9f32
fix: stop retaining every rejected profile image URL in memory (#29971)
Any signed-in user can grow Open WebUI's memory without bound by posting model entries with invalid profile image URLs. ModelMeta's validator keeps a set of every distinct rejected value so it warns about each one once, storing the full string with no size cap. FastAPI validates the request body before create_new_model reaches its workspace.models permission check, so the value is retained even when the caller is refused with a 401.

The set and the warning it served both go away; the validator clears the value exactly as before. The warning named no model and truncated the value at 80 characters, so it identified nothing. models/users.py swallows the identical ValueError and substitutes a fallback with no logging, so silence matches the neighbouring code.

Measured over the real create route with distinct 4KB invalid values: 8.9 MB retained at 2,000 values and 33.7 MB at 8,000 before, flat at 1.1 MB after.
2026-09-13 20:22:18 -04:00
Classic298
22d522c558
fix: drop a deleted plugin's source from the content cache (#29983)
Deleting a tool or a function leaves its full source text in memory for the life of the process. Each delete handler pops the module cache and leaves the matching content cache untouched, so the code of every plugin ever deleted stays resident.

Both handlers now pop the content cache next to the module cache. Measured over 200 deletes of an 8.8KB plugin: 1.78 MB of source retained per kind before, nothing after.

This is memory only, never a stale module. A cache hit needs the id present in both caches and delete already popped the module cache, so the leftover source could not have produced one.

The two routers are one change because it is the same pair of lines at the same point in sibling handlers, with no per-site reasoning and no reason to revert one without the other.
2026-09-13 20:21:40 -04:00
Classic298
5af01fe604
refac: make markdown embed rendering opt-in (#29985)
The embed flag defaults to off and is now forwarded through every markdown render path, including the standalone tool-call branch, colon fences and alerts, so embeds render only where a call site opts in. Assistant responses opt in for the viewer's own chat; channel messages pass it explicitly off.
2026-09-13 20:21:00 -04:00
Timothy Jaeryang Baek
6786ae1797 refac 2026-09-13 20:00:29 -04:00
Timothy Jaeryang Baek
c0fb36c9b8 refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
2026-09-12 20:23:50 -04:00
Timothy Jaeryang Baek
05484aa055 refac 2026-09-12 20:09:21 -04:00
Timothy Jaeryang Baek
fed94c9f5a refac 2026-09-12 19:46:26 -04:00
Timothy Jaeryang Baek
6a85abb3f5 refac 2026-09-12 19:45:42 -04:00
Timothy Jaeryang Baek
daf2295932 refac 2026-09-12 19:23:27 -04:00
Timothy Jaeryang Baek
15e2259a2a refac 2026-09-12 19:18:36 -04:00
Timothy Jaeryang Baek
5d66dd6ad2 refac 2026-09-12 19:17:11 -04:00
Classic298
f0a383af0f
refac: run the admin check for an explicit backend index inside the listing handlers (#29622)
The model listing routes accept the backend index as a path segment and as a query parameter on the index-less sibling route, so the check now runs in the handler and applies to both forms.
2026-09-12 18:10:07 -05:00
Classic298
ee25b7d42d
fix: cancelled responses no longer keep rendering as still running (#29495)
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
Frontend Build / Format & Build (push) Waiting to run
Pressing stop saved the assistant message as finished while its output items were still marked as running, so the affected block kept showing "Thinking...", "Analyzing..." or "Executing...". That state was written to the database, so it came back after every reload. A tool call cancelled mid-execution was affected as well: it is marked completed as soon as its arguments finish streaming, well before the tool returns, and the client reads that as still executing until the result item arrives.

Cancelling now closes those items before the message is saved, marking them incomplete, which is the status the client already treats as finished without flagging an error. Tool calls waiting for approval are left untouched so their approval prompt survives a reload. The corrected output is sent on the chat:tasks:cancel event that was already being emitted, so an open tab settles immediately instead of only after a reload.

Emitting a chat:completion event instead would have fired response auto-copy, text to speech playback and the chat finished listeners on a response the user had just cancelled.

Verified on a running instance against a mock upstream: cancelling mid-response leaves no item marked as running in the saved message, and the identical run on unpatched code leaves one.

Fixes #29281
2026-09-12 17:57:58 -05:00
Timothy Jaeryang Baek
c1615bec2f refac 2026-09-12 18:57:18 -04:00
Classic298
ca8f151a12
fix: show which file and why knowledge directory sync failed (#29507)
Syncing a knowledge base from a local directory only ever showed a
generic "Error accessing directory" toast when a per-file read
failed, and production builds strip console.error, so nothing else
recorded the cause either. On Windows this hides a real Chromium bug
where files with very long absolute paths throw a NotFoundError from
the File System Access API, leaving no way to tell which file failed
or why.

Both per-file read paths (the directory picker and drag-and-drop) now
wrap read failures with the file's relative path and the underlying
browser error before they reach the toast, so the message actually
names the file and the reason. Reuses the existing translated toast
text instead of adding a new i18n key, so the fix doesn't orphan
existing translations of that string.
2026-09-12 18:55:31 -04:00
Timothy Jaeryang Baek
51bb8cb142 refac 2026-09-12 18:53:22 -04:00
Timothy Jaeryang Baek
8de26e5c4c refac 2026-09-12 18:21:15 -04:00
Timothy Jaeryang Baek
e112a5252f refac 2026-09-12 18:16:37 -04:00
Timothy Jaeryang Baek
199490eadb refac 2026-09-12 18:16:19 -04:00
Classic298
afda094544
fix: stop base64 images in tool results from reaching the model as text (#29665)
When a tool returns a base64 image, Open WebUI only moves it out of the model's context if the entire result is that image, or if the tool is MCP. An image sitting inside a returned object or a list was serialised into the tool message instead, so a single screenshot could cost hundreds of thousands of tokens, push out the rest of the conversation and leave the model answering from garbage.

Any string in a tool result that is exactly one image data URI is now moved into the result's files wherever it sits in the structure, and replaced with a short marker. The model gets the image as an attachment rather than as text, and it renders for the user instead of being dropped.

Detection is deliberately limited to values that are entirely a data URI. Scanning inside longer strings was tried and abandoned: a payload wrapped over several lines, or one followed by text, gets cut short, and shipping a truncated image is worse than the bloat because the provider rejects the whole request.

This also fixes the OpenAPI branch removing entries from the list it was iterating over, which skipped every second data URI and left it in the model's context.

Fixes #29208
2026-09-12 17:15:05 -05:00
G30
6f55514c5c
fix: let validate_url accept dotless hosts when local web fetch is enabled (#29945) 2026-09-12 17:56:16 -04:00
Timothy Jaeryang Baek
5bb321de06 refac 2026-09-12 17:43:01 -04:00
Classic298
08578557de
refac: share one community origin allowlist across window message handlers (#29918)
The three community origins were repeated inline in five window message handlers and now come from a single COMMUNITY_ORIGINS constant in constants.ts.

The sync stats modal uses that same list in both directions: it reads messages only from a community origin, replies to the origin the message came from, and names the community origins as the targets of the messages it sends. Its chat id goes into the request as one encoded path segment.
2026-09-12 16:35:16 -05:00
G30
887372aca5
fix: apply the selection highlight only when the editor is unfocused and keep it out of the chat input (#29952) 2026-09-12 16:19:09 -05:00
Timothy Jaeryang Baek
0edd731c74 refac 2026-09-12 17:11:37 -04:00
Classic298
2a44c9d384
refac: validate the URL scheme before rendering a link href (#29890)
Link URLs rendered from markdown, citations and web search results now pass a scheme check before they reach an anchor. A new safeLinkUrl helper sits beside isValidHttpUrl and keeps http, https, mailto, tel and relative URLs, returning undefined for anything else so the label renders without a link. Markdown links with an unusual scheme (ftp, sms, and application deep links such as obsidian or vscode) render as plain text from now on.

The citation checks move off a substring test for "http" onto isValidHttpUrl, which is what Citations.svelte already uses for the same question. That also drops two long-standing quirks: an uppercase HTTP:// source never rendered as a link, and a filename merely containing "http" rendered as a dead external one.
2026-09-12 16:06:58 -05:00
Classic298
746caa7c78
fix: honor ENABLE_PROFILE_IMAGE_URL_FORWARDING for channel webhook profile images (#29889)
Setting ENABLE_PROFILE_IMAGE_URL_FORWARDING=false stops the user and model profile image endpoints from redirecting browsers to external avatar URLs, but channel webhook avatars kept redirecting regardless. An operator who turned the setting off precisely to stop clients leaking their IP, User-Agent and Referer to outside origins still leaked all three whenever anyone viewed a channel message posted by a webhook with an external profile image URL.

The webhook profile image endpoint now reads the same setting the user and model endpoints already read, and serves the bundled default image instead of the redirect when forwarding is off. Stored URLs are untouched, so turning the setting back on restores the previous behaviour.

Verified against the real handler with seeded webhook rows: with the setting unset or true the endpoint still returns the 302 with the original Location, with it false it returns the default favicon as image/png with no Location and no header carrying the external host, and the data URI, no image and unknown webhook responses are byte identical in both states.
2026-09-12 16:06:34 -05:00
G30
ee4834e299
fix: only log a reranking model change when the request carries one (#29922) 2026-09-12 16:04:14 -05:00
Timothy Jaeryang Baek
a910b0d8f6 refac 2026-09-12 17:02:44 -04:00
Classic298
75b1836322
fix: drop thinking blocks when converting Anthropic Messages requests to Chat Completions (#29849)
* fix: drop thinking blocks when converting Anthropic Messages requests to Chat Completions

Claude Code and other Anthropic SDK clients pointed at /api/v1/messages send the assistant's earlier thinking blocks back with every follow-up request. Since 0.11.0 those blocks were copied into the OpenAI assistant message as content parts of type thinking, a part type Chat Completions does not define. Strict OpenAI-compatible servers such as NVIDIA Dynamo reject the whole request with 400 "data did not match any variant of untagged enum ChatCompletionRequestAssistantMessageContent", so a conversation with a reasoning model died on its second turn. The error reports its position at the very end of the body, which made the request look cut off; it was complete.

Thinking and redacted_thinking blocks are now skipped in the conversion, which is what happened before 0.11.0. An assistant turn that held only thinking blocks is kept as an empty assistant message so the turn order survives. Native Anthropic and LiteLLM connections are unaffected because they receive the request untouched, and thinking blocks in responses are still produced.

Fixes #29799

* fix: keep signed thinking blocks when converting Anthropic Messages requests

Dropping every thinking block also removed the signed ones. Those are the blocks a gateway such as LiteLLM forwards to Anthropic, which needs the signed thinking block of the previous assistant turn when a tool-use turn continues with extended thinking. Only the unsigned blocks are the problem: Open WebUI creates them itself from reasoning_content, and strict Chat Completions backends reject them because thinking is not a content part type they know.

Unsigned thinking blocks are now dropped while signed thinking and redacted_thinking blocks are kept, the same rule LiteLLM applies before forwarding to Anthropic and the rule the native chat path already uses for Anthropic reasoning details. Backends that never emit a signature, such as NVIDIA Dynamo and vLLM, keep receiving requests without thinking parts, so the original failure stays fixed.
2026-09-12 15:56:59 -05:00
Timothy Jaeryang Baek
c78ad89934 refac 2026-09-12 16:56:26 -04:00
Timothy Jaeryang Baek
7a4a4b93dc refac 2026-09-12 15:41:57 -04:00
Timothy Jaeryang Baek
ffae4116a8 refac 2026-09-12 14:57:22 -04:00
Classic298
0313ea0238
fix: stop every task of a chat when stopping a response without Redis (#29844)
Pressing Stop on a chat that has more than one running task (a multi-model
response, or a follow-up sent from another tab or device while a response
is still streaming) reported success but only cancelled the first task.
The survivors kept streaming and kept executing tool calls until the
iteration limit, which is the runaway reported in the issue.

Without Redis the stop loop iterates the live task-id list of the chat.
Each awaited cancellation runs that task's cleanup, which removes its id
from the same list mid-iteration, so the loop runs out one element early
and the last task is never cancelled. Returning a snapshot of the list to
callers keeps the loop on the ids it started with. Redis deployments
already got a fresh list from the set and were not affected.

Verified against a mock upstream that calls a tool on every turn: with three
tasks on one chat, stop left one or two alive before the change and cancels
all of them after it.

Fixes #29816
2026-09-12 13:55:13 -05:00
Dolores
af54ea6283
i18n: update zh-TW (#29885) 2026-09-12 13:54:00 -05:00
G30
4ee53e051e
fix: sanitize user-typed notification target ids the same way generated ones are (#29947) 2026-09-12 13:50:19 -05:00
Timothy Jaeryang Baek
7aaa4a692e refac
Some checks are pending
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Co-Authored-By: G30 <50341825+silentoplayz@users.noreply.github.com>
2026-09-10 19:05:36 -04:00
joaoback
5e7e38c035
i18n: add pt-BR translations for newly added UI items and consistency pass (#29870)
Some checks are pending
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
New **pt-BR** translations for items introduced in the latest releases, plus a consistency/quality pass across existing strings (grammar, tone, capitalization, pluralization). Placeholders and hotkeys preserved. No logic changes.
2026-09-10 13:57:43 -04:00
Timothy Jaeryang Baek
4948842bea chore: i18n 2026-09-09 23:34:25 -04:00
G30
eb9b7b659f
fix: render mentions whose ID contains any non-whitespace character as chips (#29864)
* fix: render channel mentions whose ID contains spaces or parentheses, such as workspace model IDs

* fix: accept any non-whitespace mention ID, matching the model ID rule on dev
2026-09-09 19:27:35 -04:00
Timothy Jaeryang Baek
540467b90a refac 2026-09-09 19:16:16 -04:00
Timothy Jaeryang Baek
17dbc6f001 refac
Some checks failed
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
Frontend Build / Format & Build (push) Waiting to run
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
2026-09-09 17:09:53 -04:00
Timothy Jaeryang Baek
8a19e2f867 refac 2026-09-09 17:09:45 -04:00
G30
14e30933af
fix: render code blocks in channel model replies with the same editor as every other message (#29861) 2026-09-09 16:56:48 -04:00
Timothy Jaeryang Baek
e35b907f73 refac 2026-09-09 16:39:44 -04:00
Classic298
44f9a4f7f9
fix: stop forwarding upstream Server and Date headers from the OpenAI and Ollama proxies (#29843)
Streamed chat completions and the other proxied OpenAI and Ollama responses
went out with two Server and two Date headers: the upstream's copies, forwarded
verbatim, plus uvicorn's own. nginx in front of Open WebUI logs "upstream sent
duplicate header line" for both on every streamed request.

Server and Date belong to whoever terminates the connection, so both proxies now
drop the upstream's copies next to the encoding headers they already stripped.
The filter also compares header names case-insensitively. Before, it matched
title-case names only, so uvicorn-based upstreams such as vLLM and LiteLLM,
which send lowercase header names, had none of their headers stripped at all,
including the Content-Encoding entry the filter exists for.

Same fix as #29824 for the terminal proxy, applied to the other two proxy paths.
2026-09-09 16:35:48 -04:00
G30
b198c94efb
fix: open the thread when a notification toast for a thread reply is clicked (#29856) 2026-09-09 16:24:10 -04:00
G30
5b0f8a73ca
fix: keep the select chevron clear of the text in shrink-to-fit selects (#29866) 2026-09-09 16:22:00 -04:00
Classic298
8383bc708c
fix: show the right error when a model cannot be loaded for editing (#29694)
Opening the workspace model editor with an id that cannot be loaded redirected back to the model list and then fell through into the write-access check on the same empty model, so a plain not-found id produced "You do not have permission to edit this model". The message pointed people at their permissions for something that was never a permission problem.

The editor now stops after the redirect, so that message only appears when the model loaded and the user really cannot edit it.

Refs #29629
2026-09-09 12:55:38 -04:00
Timothy Jaeryang Baek
ee46e2664a refac 2026-09-09 12:51:59 -04:00
Timothy Jaeryang Baek
d70053e449 refac 2026-09-09 12:51:29 -04:00
Classic298
c955cbd2c5
fix: stop forwarding upstream Server and Date headers from the terminal proxy (#29841)
Proxied terminal responses (GET /api/v1/terminals/{id}/ports and every other
proxied route) went out with two Server and two Date headers: the terminal
server's copies, forwarded verbatim, plus uvicorn's own. nginx in front of Open
WebUI logs "upstream sent duplicate header line" for both on every request, and
the ports route is polled often enough to fill gigabytes of error log per day.

Server and Date belong to whoever terminates the connection, so the proxy now
drops the upstream's copies next to the framing headers it already stripped.
uvicorn's own values still go out, once. Everything else, including custom
upstream headers and TERMINAL_PROXY_HEADERS, passes through unchanged.

The OpenAI and Ollama proxies forward upstream Server and Date the same way and
are left for a separate change.

Fixes #29824
2026-09-09 12:31:02 -04:00
Classic298
1b67da7004
feat: sort flags for kb_exec file listings (#29840)
kb_exec silently ignored ls -t. It accepted the flag, dropped it and
returned the same undefined database order as a plain ls, so the model
believed it had a newest-first list when it did not. With a few thousand
files in a knowledge base there was no way to ask what changed recently
without reading the whole listing and comparing dates by eye.

ls, tree and find now sort their file lines by name by default, so the
same knowledge base always lists the same way. -t sorts newest first,
-S largest first and -r reverses, combinable like -at or -tr, matching
the flags the model already knows from a shell. Directories keep their
existing name order and stay grouped first. No query changes: the
timestamps and sizes were already loaded for the date and size columns.
2026-09-09 12:30:47 -04:00
G30
a253bf0c32
fix: drop tool IDs from the tools and tool-ids URL parameters that do not match a known tool (#29803) 2026-09-09 12:30:32 -04:00
G30
b38755b21c
fix: keep sticky code block headers at the top of channel scroll containers (#29836)
* fix: keep sticky code block headers at the top of channel scroll containers

* fix: clip the sticky code block header to the block's rounded corners

* fix: give channel scroll containers their own layer so firefox clips sticky code headers
2026-09-09 12:12:15 -04:00
G30
babf08e301
fix: stamp docker builds with the build hash so stale clients reload (#29832) 2026-09-09 12:11:52 -04:00
G30
e4d65d0351
fix: close modal shortcut closes only the top modal (#29830) 2026-09-09 12:11:38 -04:00
Timothy Jaeryang Baek
12b14124b9 refac 2026-09-08 23:18:48 -04:00
Timothy Jaeryang Baek
31b272d3c9 refac 2026-09-08 22:20:33 -04:00
Timothy Jaeryang Baek
307b9b9133 refac 2026-09-08 21:16:10 -04:00
Timothy Jaeryang Baek
a23b579233 refac 2026-09-08 21:01:26 -04:00
Timothy Jaeryang Baek
3808eace6c refac 2026-09-08 20:44:23 -04:00
Timothy Jaeryang Baek
8556033c6b refac 2026-09-08 20:34:42 -04:00
Timothy Jaeryang Baek
674760bfc1 refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Unit Tests (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
2026-09-08 16:54:15 -04:00
Timothy Jaeryang Baek
4cc0d48b4d refac 2026-09-08 16:49:54 -04:00
Timothy Jaeryang Baek
f80ef8bd00 refac 2026-09-08 16:47:01 -04:00
spoofy
94073b1f16
i18n: complete Japanese translations (ja-JP) (#29798)
* i18n: fill empty Japanese translations

* i18n: restore Japanese translation placeholders

* i18n: polish Japanese UI wording for clarity
2026-09-08 15:59:23 -04:00
Timothy Jaeryang Baek
d8f27e745b refac 2026-09-08 13:25:16 -04:00
spoofy
56de6dd792
i18n: es es new UI strings (#29795)
* i18n: translate new UI strings into Spanish

* i18n: use precise Spanish term for Shell
2026-09-08 13:23:51 -04:00
spoofy
02efa6ed10
i18n: ru uk new UI strings (#29794)
* i18n: translate new UI strings into Russian and Ukrainian

* i18n: preserve established Ukrainian Valves terminology

* i18n: shorten Russian Modified label

* i18n: shorten Ukrainian Modified label

* i18n: shorten Russian and Ukrainian default labels
2026-09-08 13:23:41 -04:00
Timothy Jaeryang Baek
254e29b9af refac 2026-09-08 13:19:14 -04:00
Timothy Jaeryang Baek
ba34bee2d1 refac 2026-09-08 12:45:11 -04:00
Timothy Jaeryang Baek
aaaf26fb8e refac 2026-09-08 12:38:07 -04:00
Timothy Jaeryang Baek
9a4b643130 chore: i18n 2026-09-08 11:39:57 -04:00
Timothy Jaeryang Baek
e723dcda58 refac 2026-09-08 01:45:35 -04:00
Timothy Jaeryang Baek
3572e01a71 refac 2026-09-08 01:04:12 -04:00
Timothy Jaeryang Baek
98a920168e refac 2026-09-07 14:46:52 -04:00
Timothy Jaeryang Baek
675b9839f1 refac 2026-09-07 13:26:30 -04:00
G30
d49e424874
fix: enable ask_user submit when the last question uses the other field (#29494) 2026-09-07 13:21:12 -04:00
G30
2f2f4c872f
fix(retrieval): drain playwright route handlers before the page closes (#29325) 2026-09-07 13:20:59 -04:00
G30
2823bf64e1
fix: scale dropdown menu item text with the UI scale setting (#29493) 2026-09-07 13:20:43 -04:00
Timothy Jaeryang Baek
9e634c0c56 refac 2026-09-07 13:19:16 -04:00
Timothy Jaeryang Baek
e1bfefdf9f refac 2026-09-07 13:13:25 -04:00
Timothy Jaeryang Baek
f5fcf4c89f refac 2026-09-07 12:40:44 -04:00
Timothy Jaeryang Baek
f9f815c862 refac 2026-09-07 12:21:23 -04:00
Timothy Jaeryang Baek
b141fdc49e refac 2026-09-07 12:18:23 -04:00
Timothy Jaeryang Baek
3cda47cdb4 refac 2026-09-07 12:16:09 -04:00
Timothy Jaeryang Baek
b71744b178 refac 2026-09-07 12:15:20 -04:00
G30
3aabbfebfe
fix: wait for a pending note save before navigating away (#29745) 2026-09-07 11:26:45 -04:00
Timothy Jaeryang Baek
57acc2b68f refac
Some checks failed
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
2026-09-06 23:23:59 -04:00
Timothy Jaeryang Baek
3795d5b292 refac 2026-09-06 19:37:59 -04:00
Timothy Jaeryang Baek
d418840aa9 refac 2026-09-06 19:36:17 -04:00
Timothy Jaeryang Baek
7eefeef4f1 refac 2026-09-06 19:32:31 -04:00
Timothy Jaeryang Baek
6c7aa3543d refac 2026-09-06 19:02:44 -04:00
Classic298
4b10190096
refac: move the connection index role check into the listing handlers (#29619)
The Ollama and OpenAI model listing handlers now check the caller's role themselves instead of declaring it as a route-level dependency.
2026-09-06 18:30:31 -04:00
Timothy Jaeryang Baek
57fc344873 refac 2026-09-06 18:30:00 -04:00
Timothy Jaeryang Baek
77d2000eb7 refac 2026-09-06 18:29:46 -04:00
Classic298
ae01ef9c95
fix: skip image, media and font requests in the Playwright web loader (#29742)
Fetching a page with many media files through the Playwright web loader was extremely slow or timed out, while raw Playwright loaded the same page in a couple of seconds. The loader routes every request the page makes through the backend HTTP client and downloads the full body before the browser sees any of it, and the browser cancelling a media request once it has enough never reaches that download. On a page with a few dozen audio players every file was pulled in full for a text extraction that never reads it.

Image, media and font requests are now aborted in the interceptor before any fetch is made. None of them feed the text extraction. Measured on the page from the report with the default timeout on the same connection:

| | requests fetched | bytes downloaded | elapsed |
|---|---|---|---|
| before | 151 | 55.0 MB | 10.1 s |
| after | 45 | 3.9 MB | 2.6 s |

Fixes #29741
2026-09-06 17:01:00 -05:00
Timothy Jaeryang Baek
649c012ecf refac 2026-09-06 17:53:02 -04:00
Timothy Jaeryang Baek
de1203f9b5 refac 2026-09-06 17:46:35 -04:00
Classic298
2e16a761ee
refactor: scope tool export to tools the caller can write (#29310)
The bulk tool export returned every tool the caller could read, while the per-tool export path returns only what the caller can write. This aligns the two, matching how model export already scopes its query.

Callers still export their own tools and any tool shared with them for writing; admins running with BYPASS_ADMIN_ACCESS_CONTROL are unaffected.
2026-09-06 17:44:45 -04:00
Timothy Jaeryang Baek
f3eade42ae refac 2026-09-06 17:43:17 -04:00
Timothy Jaeryang Baek
30eed12513 refac 2026-09-06 17:38:21 -04:00
Timothy Jaeryang Baek
a1c02098aa refac 2026-09-06 17:35:49 -04:00
Timothy Jaeryang Baek
508de20779 refac 2026-09-06 17:27:30 -04:00
Timothy Jaeryang Baek
98fcb844e1 refac 2026-09-06 17:24:23 -04:00
Timothy Jaeryang Baek
0fa4dea5ff refac 2026-09-06 17:21:57 -04:00
Timothy Jaeryang Baek
f1c803d36b refac 2026-09-06 17:18:16 -04:00
Timothy Jaeryang Baek
d27aa72ab4 refac 2026-09-06 17:13:32 -04:00
Timothy Jaeryang Baek
c4a349651e refac 2026-09-06 17:10:01 -04:00
Classic298
1932ca649e
fix: stop sending OpenAPI tool server path and query parameters in the request body (#29717)
Tool calls to an OpenAPI tool server put every argument the model returned into the JSON request body, including the parameters that were already substituted into the URL. Servers that validate their input strictly (additionalProperties: false) answered 422 "unexpected property", so reads worked and every write through an endpoint with a path or query parameter failed.

The body is now built from the model's arguments minus the operation's declared parameters, keeping any name the requestBody schema declares as a property of its own, so an endpoint that wants the resource id in the body as well as in the path still gets it.

The filter only runs when the resolved body schema lists its properties. A free-form, composed or non-JSON body offers nothing to check a name against, so those requests go out exactly as before.

src/lib/apis/index.ts carries the same request builder for direct tool server connections and had the same bug, so it gets the same fix.

Fixes #29716
2026-09-06 17:01:26 -04:00
Classic298
66e021a926
fix: normalize a tool call name sent as null (#29690)
Some endpoints stream a tool call whose function name is JSON null instead of a string. Nothing normalized it, so the null stayed on the tool call, was written into the stored message, and was sent back to the endpoint in the assistant message on the next turn, where a null is not a valid function name.

The delta accumulator now replaces a null name with an empty string, at the same point it already normalizes the arguments field. The call still fails as an unknown tool, which is the right outcome for a call that has no name, so the result is one failed tool call instead of a follow-up request the endpoint has to reject.

This is done where the delta enters the accumulator rather than at the consumers, because the name is emitted to the client and persisted while the response is still streaming, before anything downstream could clean it up.

Checked against 1261 streaming delta sequences: behaviour is unchanged except where the delta that creates the tool call carries a null name.

Seen in #29686 with a custom sglang build.
2026-09-06 17:00:03 -04:00
Timothy Jaeryang Baek
3361a972b3 refac 2026-09-06 16:59:40 -04:00
Classic298
cd68a66fba
Gate the channel webhook profile image endpoint on channel access (#29703)
Any authenticated user could fetch a channel webhook's avatar, or be redirected to its external profile image URL, without belonging to the channel or holding any read access to it. This was the only webhook route with neither a channel check nor the channels feature gate.

The route now applies the same read gate every other route in this router uses: active membership for group and direct message channels, admin or a channel read grant otherwise, answering with 403 on denial and 404 when the webhook's channel row no longer exists. It also runs the channels feature and permission gate, so with channels disabled, or the permission withdrawn from regular users, the endpoint now refuses where it previously served the image.

Avatars keep rendering for channel members, and a denied request shows the default logo rather than a broken image, because the avatar component already falls back on an image error.
2026-09-06 16:58:50 -04:00
Classic298
517617b601
chore: stop shipping uv in the Docker image (#29728)
The Docker image gets about 50 MB smaller on disk (about 20 MB off the pull). uv is only needed to run the requirements install, but it is pip-installed into the image and stays there. It is now bind-mounted from the uv image for that RUN only, pinned to 0.12.10, so the runtime image never contains it. This is the pattern uv's own Docker guide recommends for the case.

Uninstalling uv at the end of the same RUN would also keep it out of the layer; the mount was preferred because it fixes the uv version and the mounted layer is cached by the builder across rebuilds.

pip stays available in the container, so hand-installing optional packages the way requirements.txt describes keeps working; only running uv inside the container goes away. The build now requires BuildKit (the syntax directive alone was a comment to the classic builder, this line makes it mandatory), needs access to ghcr.io next to PyPI, and the uv version is a pin to bump by hand, like the base images.

Part of #29721.
2026-09-06 16:58:05 -04:00
Timothy Jaeryang Baek
91f8775b28 refac 2026-09-06 16:55:28 -04:00
joaoback
53cc969e64
i18n: add pt-BR translations for newly added UI items and consistency pass (#29740)
New **pt-BR** translations for items introduced in the latest releases, plus a consistency/quality pass across existing strings (grammar, tone, capitalization, pluralization). Placeholders and hotkeys preserved. No logic changes.
2026-09-06 16:55:06 -04:00
Classic298
831b3b0df2
refac: validate the citation embed URL before use (#29701)
A citation's embed_url now has to be an http or https URL, or protocol-relative, before it is opened or handed to the embed panel. Anything else falls back to the citation modal, which already renders the source.
2026-09-06 16:54:44 -04:00
Timothy Jaeryang Baek
cb942bb94c refac 2026-09-06 16:48:30 -04:00
Classic298
7fc979f95b
fix: skip embedded HTML parts and restrict link targets in docx preview (#29699)
* fix: skip embedded HTML parts in docx preview

The docx preview no longer renders altChunk parts, so an embedded HTML sub-document is omitted from the rendered output instead of being handed to the renderer.

* fix: restrict docx preview link targets to safe schemes

The docx preview kept whatever link target the document supplied, so a document could point a link at any scheme the browser understands.

After rendering, a link target is now kept only when it resolves to http, https, mailto or tel. Anything else has its target removed and the link renders as plain text. Targets resolve against the page URL, so internal bookmark links and relative targets are unaffected, while an empty target, which the renderer emits for a hyperlink with no external relationship, is dropped instead of reloading the app.

DOMPurify was not used because running it over the rendered document would strip the renderer's own markup and styling, so the check stays limited to link targets. Links using file: or Office application schemes no longer resolve.
2026-09-06 16:40:01 -04:00
Classic298
08ca3d859f
chore: drop the unused Noto Sans variable fonts (#29723)
The four Noto Sans variable fonts under open_webui/static/fonts have never been loaded. Removing them makes an installed package 40 MB smaller on disk, the wheel about 24 MB, and the Docker image about 80 MB, because the image currently stores the static directory twice.

The PDF generator registers only the static faces through add_font, and the stylesheet that names the variable fonts, pdf-style.css, is read into a variable that nothing ever uses, so no code path can reach them. The frontend never fetches these files either.

Only the four @font-face blocks that pointed at the deleted files are removed; the rest of the stylesheet, font stack included, is left exactly as it is.

The files were reachable under the public /static mount, so anything outside this repository that hotlinked them gets a 404 from now on.

Part of #29721.
2026-09-06 16:39:36 -04:00
Classic298
e07e8ed0d4
chore: drop the Google Drive client and other unimported dependency pins (#29726)
The Docker image and a fresh pip install get about 100 MB smaller uncompressed. The removed pins are ones nothing in the backend imports and that no installed package requires, with the one exception of async-timeout, which redis still requires below Python 3.11.3; the resolver installs it there as a transitive dependency, just no longer pinned to 5.0.1 for plain pip installs.

google-api-python-client, google-auth-httplib2 and google-auth-oauthlib were added for Google Drive in 2024, but the picker is frontend-only and loads gapi from apis.google.com; together they are about 95 MB on disk. pymongo, the langchain meta package, pymdown-extensions, pytube, APScheduler and RestrictedPython have no importer; the YouTube loader, the scheduler and the tool sandbox are all hand-written in this repository.

google-genai stays: nothing imports it either, but the tool-call path carries accommodations written for python-genai callers, so Gemini pipes are expected to find it in the shared environment.

uv.lock is regenerated, deletions only.

One user-visible consequence: a tool or function that imported one of the removed packages without declaring it in its frontmatter requirements has worked only because the package was preinstalled. Declaring it fixes that where frontmatter installs are enabled and the instance can reach PyPI; an offline instance needs the package installed into the image instead.

Part of #29721.
2026-09-06 16:39:24 -04:00
Classic298
ca9ec06c7e
chore: drop nltk, unused at the pinned versions (#29725)
The main and CUDA Docker images get about 21 MB smaller (the nltk package, the punkt_tab data and its zip); slim images, which never downloaded the data, about 6 MB.

nltk was in the image for unstructured, which used it to tokenize documents. The Dockerfile download was added for airgapped containers failing on the missing punkt_tab data (#21150; the same request in #16260), the same lookup failed on first use in other setups (#17594, #4642), and the download in start.sh and start_windows.bat came with the Playwright web loader mode and sits in that branch.

unstructured 0.22.31, the pinned version, has no nltk references at all and tokenizes with spaCy, nothing else installed requires nltk outside transformers' testing and dev extras, and nothing in the backend imports it, so the pin and both downloads go together.

One user-visible consequence: a tool or function that imports nltk inside the container stops working unless it declares nltk in its frontmatter requirements. On an offline instance the package, and any nltk data such as punkt_tab, have to be installed into the image instead.

Part of #29721.
2026-09-06 16:39:05 -04:00
Classic298
55a6198e26
Gate the model profile image endpoint on model read access (#29700)
Any authenticated user could fetch the profile image of a model they have no access to, and could tell an existing model id from an unknown one by whether the response carried the image or the default logo.

The endpoint now serves an image only to callers who can see the model itself: the owner, an admin under the admin bypass, or the holder of a read grant, with the same rule applied to arena models defined in config. Everyone else gets the default logo, byte for byte the response an unknown id already returned, so ids can no longer be probed. BYPASS_MODEL_ACCESS_CONTROL is honoured here because it is what decides which models reach a user's model list to begin with.

Avatars now fall back to the default logo wherever a viewer meets a model id without holding a grant on it: a model reply in a channel shown to the other members, and the admin analytics and evaluation pages when the admin bypass is switched off.
2026-09-06 16:20:13 -04:00
Classic298
39c1e86e95
chore: keep OAuth token payloads out of logs (#29709)
Two OAuth failure paths interpolated the raw token object into their log message. On the callback path that object is a live credential set, so a provider returning no user data wrote an access token, and usually a refresh token, straight into the application log.

Both messages now log without the payload. The token-exchange failure keeps its error level and its client_id binding and reports the provider's error description instead of the raw response body, which by that branch's own condition never contained an access token anyway. The callback failure keeps its warning level and identifies the provider, matching the other failure logs in that handler.
2026-09-06 16:19:23 -04:00
Classic298
039c4d665e
chore: drop python3-dev from the Docker image (#29731)
The apt layer of the Docker image shrinks by about 60 MB. python3-dev installs Debian's own interpreter with its headers, and nothing in the image uses it.

The image's Python is the /usr/local build from the base image, which ships its own headers, and it is also the interpreter that installs tool and function requirements at runtime, so Debian's headers were never on the include path of anything built in the container. The zlib headers python3-dev pulled in stay through libmariadb-dev, which also brings the OpenSSL headers, so the optional mariadb connector still builds; the only other header package that goes with it is libexpat1-dev, which nothing in the pinned tree builds against.

Part of #29721.
2026-09-06 16:19:10 -04:00
G30
ffe6ab9869
fix: stop the note editor date reading as a button (#29708) 2026-09-06 15:31:47 -04:00
Classic298
a3c90e5aac
fix: blank chat messages on iOS home screen apps and in-app browsers (#29734)
Chat messages are virtualized with content-visibility: auto, which WebKit paints incorrectly and can leave blank. #26805 skipped that on Safari by looking for the Safari token in the user agent, but iOS in-app browsers, home screen apps and iPadOS desktop-class standalone windows send no such token, so those users still get empty assistant responses.

Check the navigator vendor string as well, which every WebKit surface reports regardless of user agent. The user agent check stays, because non-Apple WebKit ports can compile a different vendor string while shipping the same paint bug.

That earlier fix also withheld the message-listitem class entirely, and the class doubles as the styling hook the sidebar hover preview reaches through, so hover previews have rendered at full chat spacing and width on Safari since v0.11.0. Only the content-visibility rule is gated now, on its own class, and the hook stays on every message. Safari hover previews become compact like every other engine.

Verified across fourteen engine cases: virtualization is off on every Apple WebKit surface, unchanged on Chromium, Firefox and Android, the hover preview overrides apply again on Safari, and the screenshot export still captures every message on both.

Refs #26712, #29688
2026-09-06 15:21:09 -04:00
Classic298
68a74da70b
Fix: mps inference evaluations (#29735)
* fix(retrieval): serialize local embedding and reranking on MPS

On Apple Silicon the server process is killed outright (SIGSEGV or SIGTRAP, no traceback) partway through answering any question that retrieves from a knowledge base with hybrid search and a local reranking model. The client sees a dropped connection and the answer is lost.

Hybrid search fans its queries out concurrently and every task calls the same shared local model on a worker thread. Torch's Metal shader cache is a process-wide singleton whose lookup tables have no lock, so two of those threads racing inside it corrupt the cache and take the process down with it.

Guard the local SentenceTransformer and CrossEncoder calls with a shared lock that is only a real lock when the selected device is MPS. CPU and CUDA installs keep the concurrency they have today, and external reranking endpoints are untouched. Reranking several queries on a Mac now runs one at a time, which is the cost of the process staying alive.

Verified by driving the real hybrid-search fan-out with 16 concurrent queries: peak simultaneous entries into the local model drops from 16 to 1 on MPS, stays at 16 on CPU, and the returned documents, scores and ordering are byte-identical in every case.

Fixes #29722

* fix(evaluations): serialize the leaderboard embedder against retrieval on MPS

The leaderboard's tag-similarity search builds its own SentenceTransformer, and on Apple Silicon sentence-transformers places it on the MPS device. It runs on a worker thread, so an admin running a leaderboard search while anyone queries a knowledge base puts two threads into torch's Metal backend at the same time, which kills the server process outright with no traceback.

Move the lock added for the retrieval path into env.py, beside the device selection that decides whether MPS is used at all, and take it around the leaderboard's embedding calls as well. Sharing one lock between the two modules is the whole point, since two separate locks would still let a leaderboard search collide with a retrieval query.

Only inference is guarded, matching the retrieval path. Model construction stays as it is here and in the retrieval routers.

Verified by driving the leaderboard similarity path and retrieval reranking from six threads against one instrumented model: peak simultaneous entries drops from six to one on MPS, and the similarity scores are unchanged.

Related to #29722.
2026-09-06 15:20:41 -04:00
Classic298
1bfa59acbd
fix: keep HTML entities in extracted document text (#29736)
Uploading a file whose text contains literal HTML entities stored a rewritten copy of it: `&nbsp;` became a non-breaking space, `&gt;` became `>`, and `&amp;nbsp;` was decoded twice down to a bare non-breaking space. That stored text is what gets indexed and what the model reads, so notes, specs and source files reached the model differing from the file that was uploaded.

Every loaded document goes through `ftfy.fix_text`, which is there to repair mojibake left by the encoding-detection fallback. Its default configuration also decodes HTML entities, per line and sticky forward: entities are decoded on every line up to the first line holding a literal `<`, then left alone for the rest of the document. The same escape therefore survives or vanishes depending on where it sits in the file. This disables that one behaviour and leaves every other ftfy repair in place.

Text from a third-party extraction engine that returns escaped output now keeps those escapes. Guessing whether an escape is markup or content is the bug being fixed.

Fixes #29732
2026-09-06 15:13:45 -04:00
Timothy Jaeryang Baek
85b11a4f35 refac
Some checks failed
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
2026-09-04 23:34:32 -04:00
Timothy Jaeryang Baek
54a7a7a7ce refac 2026-09-04 22:56:00 -04:00
Timothy Jaeryang Baek
3facfa61d4 refac 2026-09-04 22:53:37 -04:00
Tim Baek
8652227eab
ci: run the external regression suite on release pull requests (#29313) (#29678)
* ci: run the external regression suite on release pull requests

Adds a workflow that runs the open-webui/tests unit suite against release
candidates, so a release that reintroduces a fixed bug is caught before it is cut
rather than after users report it. The suite is roughly 4500 source-level tests
pinned to specific past issues and PRs, and takes about three minutes; the
dependency install dominates the run and is cached.

It runs only on pull requests into main whose title starts with a version, which
is how releases are titled here, or which touch package.json. Everything else
into main, and every pull request into dev, skips it and reports green.

Two settings are needed for this to block anything, both outside the diff:
require the Regression / Result check on main, and require branches to be up to
date before merging so the suite covers what actually lands.

The reusable workflow is referenced at @main so a release always runs the current
tests. Pinning it to a tag instead is a reasonable call to make here.

* ci: cancel superseded regression runs

A queued run on a release PR meant a stale commit's suite kept blocking
the required check after newer commits shipped, wasting a runner slot
and the author's time waiting on a result nobody needed. Cancel it
instead so the suite always runs against the latest push.

* ci: rename the Regression workflow to Tests

* Update regression.yaml

* ci: gate the test suite with a job condition instead of a gate job

Replaces the gate job with a condition on the suite job itself. The job existed
to look for a version title or a change to package.json, and the package.json
check is redundant: a release bumps the version in that file and carries it in
the title, so the title alone identifies one. That removes a runner, an API call
and the pull-requests read permission.

The suite now runs on version-titled pull requests from dev into main, and on
version-titled pull requests into dev so it can be exercised outside a release.
An edit only re-runs it when the title itself changed, and an edit no longer
cancels a suite that is already running, which would otherwise leave the check
green with nothing behind it.

* ci: match only the version prefixes releases actually use

Release pull requests are titled 0.11.3, not v0.11.3, so the leading v never
matched. The remaining digits are dropped with it and the dot is kept, so a
title that merely starts with a digit does not run the suite.

Signed-off-by: Adam Tao <tcx4c70@gmail.com>
Signed-off-by: Riley Des <riley.desserre@improving.com>
Signed-off-by: James Liounis <james.liounis@perplexity.ai>
Co-authored-by: joaoback <156559121+joaoback@users.noreply.github.com>
Co-authored-by: Algorithm5838 <108630393+Algorithm5838@users.noreply.github.com>
Co-authored-by: Kylapaallikko <Kylapaallikko@users.noreply.github.com>
Co-authored-by: Teay <pythontogoplease@gmail.com>
Co-authored-by: tcx4c70 <tcx4c70@gmail.com>
Co-authored-by: goodbey857 <76645482+goodbey857@users.noreply.github.com>
Co-authored-by: Jacob Leksan <63938553+jmleksan@users.noreply.github.com>
Co-authored-by: RomualdYT <romuald@gameurnews.fr>
Co-authored-by: Lucas <lucas@vanosenbruggen.com>
Co-authored-by: Classic298 <27028174+Classic298@users.noreply.github.com>
Co-authored-by: Constantine <Runixer@gmail.com>
Co-authored-by: Athanasios Oikonomou <athoik@gmail.com>
Co-authored-by: Shirasawa <764798966@qq.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Aleix Dorca <aleixdorca@mac.com>
Co-authored-by: Vincent Agra <agravj007@gmail.com>
Co-authored-by: Shamil <ashm.tech@proton.me>
Co-authored-by: Cyp <cypher9715@naver.com>
Co-authored-by: looselyhuman <fieldian@gmail.com>
Co-authored-by: Circe (Claude Code Sonnet 4.6) <circe@athena-council.org>
Co-authored-by: _00_ <131402327+rgaricano@users.noreply.github.com>
Co-authored-by: Daniel Nylander <po@danielnylander.se>
Co-authored-by: Daniel Nylander <daniel@danielnylander.se>
Co-authored-by: Asbjørn Dyhrberg Thegler <asbjoern@dyhrbergthegler.dk>
Co-authored-by: ShigekiTsuchiyama <ShigekiTsuchiyama@users.noreply.github.com>
Co-authored-by: berkant-koc <berkant-koc@users.noreply.github.com>
Co-authored-by: mayamsin <58122670+mayamsin@users.noreply.github.com>
Co-authored-by: G30 <50341825+silentoplayz@users.noreply.github.com>
Co-authored-by: Aindriú Mac Giolla Eoin <aindriu80@gmail.com>
Co-authored-by: POV9en <POV9en@users.noreply.github.com>
Co-authored-by: oxsignal <oxsignal@users.noreply.github.com>
Co-authored-by: Dara Adib <dara@quietapple.org>
Co-authored-by: Sergey Zinchenko <sergey.zinchenko.rnd@gmail.com>
Co-authored-by: maco <gosarmarcel7@gmail.com>
Co-authored-by: Sakıp Han Dursun <100518315+sakiphan@users.noreply.github.com>
Co-authored-by: anishgirianish <161533316+anishgirianish@users.noreply.github.com>
Co-authored-by: Mateusz Hajder <6783135+mhajder@users.noreply.github.com>
Co-authored-by: Amir Subhi <amirsubhi@hotmail.com>
Co-authored-by: sermikr0 <230672901+sermikr0@users.noreply.github.com>
Co-authored-by: bannert <58707896+bannert1337@users.noreply.github.com>
Co-authored-by: Boris Rybalkin <ribalkin@gmail.com>
Co-authored-by: HW <hwinkler@first-it-consulting.de>
Co-authored-by: sfwani <sfwani@users.noreply.github.com>
Co-authored-by: rileydes-improving <riley.desserre@improving.com>
Co-authored-by: Zixin Yu <183055163+ivvi0927@users.noreply.github.com>
Co-authored-by: Lukáš Kucharczyk <lukas@kucharczyk.xyz>
Co-authored-by: russelg <russelg@users.noreply.github.com>
Co-authored-by: Mr. Meowgi <ovehbe@gmail.com>
Co-authored-by: Zaid Marji <91486926+zaid-marji@users.noreply.github.com>
Co-authored-by: Craig <66838006+TipKnuckle@users.noreply.github.com>
Co-authored-by: James Liounis <james.liounis@perplexity.ai>
Co-authored-by: Chane Lu <2522992009@qq.com>
Co-authored-by: Justin Williams <justinjohnwilliams@gmail.com>
Co-authored-by: Syed Mustafa Quadri <175467872+code-quad3@users.noreply.github.com>
Co-authored-by: hungryBird <pantazis.marina@gmail.com>
Co-authored-by: Marina Pantazis <marina.pantazis@bit.admin.ch>
Co-authored-by: cwanglab <cwanglab@users.noreply.github.com>
Co-authored-by: Sicknine <156204309+SNaytiP@users.noreply.github.com>
Co-authored-by: JuanMa Diaz <torgus@gmail.com>
Co-authored-by: Syed Osama Ali Shah <86572800+osamaali313@users.noreply.github.com>
2026-09-04 19:44:21 -04:00
Classic298
c2e6766ed8
Move the vitest tests to the open-webui/tests suite (#29677)
The colocated vitest files (shortcuts, the colon fence marked extension) now
live in open-webui/tests under frontend/, next to the rest of the regression
suite, where they run against the source of any ref through the shared
regression workflow. This removes the copies here; the unit-tests workflow job
and the test:frontend script stay and pass with no tests.

src/lib/utils/_template_old.ts goes as well: it imports vitest but its name
never matched the test glob, so those tests have not run since they were added.
2026-09-04 19:43:48 -04:00
Classic298
0a7c15832f
ci: run the external regression suite on release pull requests (#29313)
Some checks failed
Release / publish (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Release to PyPI / release (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
* ci: run the external regression suite on release pull requests

Adds a workflow that runs the open-webui/tests unit suite against release
candidates, so a release that reintroduces a fixed bug is caught before it is cut
rather than after users report it. The suite is roughly 4500 source-level tests
pinned to specific past issues and PRs, and takes about three minutes; the
dependency install dominates the run and is cached.

It runs only on pull requests into main whose title starts with a version, which
is how releases are titled here, or which touch package.json. Everything else
into main, and every pull request into dev, skips it and reports green.

Two settings are needed for this to block anything, both outside the diff:
require the Regression / Result check on main, and require branches to be up to
date before merging so the suite covers what actually lands.

The reusable workflow is referenced at @main so a release always runs the current
tests. Pinning it to a tag instead is a reasonable call to make here.

* ci: cancel superseded regression runs

A queued run on a release PR meant a stale commit's suite kept blocking
the required check after newer commits shipped, wasting a runner slot
and the author's time waiting on a result nobody needed. Cancel it
instead so the suite always runs against the latest push.

* ci: rename the Regression workflow to Tests

* Update regression.yaml

* ci: gate the test suite with a job condition instead of a gate job

Replaces the gate job with a condition on the suite job itself. The job existed
to look for a version title or a change to package.json, and the package.json
check is redundant: a release bumps the version in that file and carries it in
the title, so the title alone identifies one. That removes a runner, an API call
and the pull-requests read permission.

The suite now runs on version-titled pull requests from dev into main, and on
version-titled pull requests into dev so it can be exercised outside a release.
An edit only re-runs it when the title itself changed, and an edit no longer
cancels a suite that is already running, which would otherwise leave the check
green with nothing behind it.

* ci: match only the version prefixes releases actually use

Release pull requests are titled 0.11.3, not v0.11.3, so the leading v never
matched. The remaining digits are dropped with it and the dot is kept, so a
title that merely starts with a digit does not run the suite.
2026-09-04 19:43:32 -04:00
Timothy Jaeryang Baek
7bf898f461 refac 2026-09-04 19:21:49 -04:00
Classic298
acd03b147c
fix: treat .ino files as source text so knowledge uploads stop failing (#29673)
Uploading an Arduino sketch (`.ino`) to a knowledge base failed with `Expecting value: line 1 column 1 (char 0)` whenever the content extraction engine was Tika or Docling. Browsers send `.ino` as `application/octet-stream`, and the extension was missing from the known source extension list, so the file was handed to the extraction server instead of being read as plain text. The server answered with a non-JSON body and the loader crashed while decoding it. `.cpp` and `.h` sketches in the same folder uploaded fine, because those extensions are already on the list.

Adding `ino` to that list routes it to the plain text loader, the same way the yaml/toml gap was closed in 710320601a. A sketch is plain C++ text, so there is nothing for a document extraction server to do with it.

Verified by dispatch matrix over 35 extensions, 5 content types and all 8 engines against a stub server that reproduces the non-JSON response: the only rows that change are `.ino` under Tika and Docling, which now resolve to the text loader and extract the sketch verbatim. Every other row is unchanged.

Fixes #29670
2026-09-04 19:13:04 -04:00
Timothy Jaeryang Baek
67ac1a4e93 refac 2026-09-04 19:12:41 -04:00
Timothy Jaeryang Baek
858ab727d2 refac 2026-09-04 18:31:59 -04:00
Timothy Jaeryang Baek
7b6562e335 refac 2026-09-04 17:39:34 -04:00
Classic298
323f8da202
fix: show relevance badge for single-source citations (#29647)
A response with exactly one citation never showed its relevance badge, even though the source carried a perfectly valid score. Adding a second citation made the badge appear for that same source, so the score looked like it came and went at random.

calculateShowRelevance hides relevance when a result set mixes distance metrics (cosine in -1..1 next to unbounded L2), since those numbers are not comparable. With a single distance that check degenerates: distances.length - 1 is 0, so the "all but one are out of range" clause matches whatever the value is, and the score is hidden. A lone distance cannot be a mix of two metrics, so it now returns before the outlier check runs.

Sets of two or more distances behave exactly as before, mixed-metric sets are still suppressed.

Fixes #29646
2026-09-04 17:17:14 -04:00
Classic298
d75a7aaf54
fix: surface upstream errors on arena models instead of crashing (#29662)
Sending a message to an arena model failed with "'JSONResponse' object has no attribute 'body_iterator'" whenever the backing provider answered with an HTTP error, so the real error never reached the user. Non-streaming requests on an arena model, such as title and tag generation, broke the same way with "'JSONResponse' object is not a mapping".

The arena wrapper assumed the sub-model call always returns a stream for a streaming request and a dict for everything else, but the OpenAI-compatible router returns a plain response object as soon as the provider answers 4xx or 5xx. Both arms now hand that response straight back, which is exactly what the non-arena path already does, so the existing error handling turns it into the usual error message in the chat.

Verified against a matrix of streaming and non-streaming requests with the sub-model returning a stream, a dict, a JSONResponse and a PlainTextResponse: both crashes are gone and the two success paths are unchanged, including the selected_model_id prelude on the stream.

Fixes #29658
2026-09-04 17:17:02 -04:00
G30
4f419627f8
fix: remove the stray background band from the Attach Webpage modal footer (#29664) 2026-09-04 17:16:47 -04:00
G30
b49db6f287
fix: stop the browser back button creating extra notes (#29645) 2026-09-04 15:54:10 -04:00
G30
bd0b3875ba
i18n: blank redundant en-US translation values (#29318) 2026-09-04 11:58:50 -04:00
G30
78c4f9f83d
feat: shift quick delete for automations (#29640) 2026-09-04 11:58:05 -04:00
spoofy
bfed5f8950
i18n: unify es-ES terminology on common loanwords (#29593)
Adopts the loanwords normally used in Spanish-language developer and AI
interfaces for three terms the locale currently calques, and applies each
one consistently across the file.

- prompt:   Indicador -> Prompt   (74 strings)
- pipeline: Tubería   -> Pipeline (14 strings)
- delete:   Borrar    -> Eliminar (62 strings), matching the rendering the
            locale already uses for "remove" and closing the Borrar /
            Eliminar split that applied both words to the same action
- Admin -> Administrador, Email -> Correo electrónico (3 strings)

"Indicador" means indicator or gauge and is not how Spanish-language AI
tools refer to a prompt. "Tubería" is a physical pipe; the file already
hedged once with "Tuberías (Pipelines)". Gender agreement is updated where
it changes, since Pipeline is masculine and Tubería feminine.

Before this change the locale rendered "delete" as Borrar in 34 strings and
Eliminar in 28, with no distinction in meaning between them.

This changes wording contributed by the Spanish community rather than
fixing defects, so it is a terminology decision rather than a correctness
fix, and is kept separate from the translation work for that reason.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 11:51:51 -04:00
Alexey Shabalin
ffd4c48b94
i18n: fix incorrect Russian translations (#29625)
Two entries lost their interpolation placeholder, so they render
literally instead of being substituted:

- "{{ models }}" had the placeholder name itself translated
- the 50-word summary prompt replaced {{topic}} with a bracketed note

The rest change the meaning or contradict neighbouring strings:

- "Enter Key Behavior" was read as the verb "enter", turning the label
  for how the Enter key behaves into "type in the behaviour key", while
  its own description already explains the setting correctly
- "Regenerate Menu" was read as "refresh the menu" rather than the menu
  of regenerate actions its description describes
- "Enter reasoning effort" was not meaningful Russian
- "Enter Bing Search V7 Subscription Key" mentioned an endpoint that the
  source string does not
- "Web Search in Chat" called the feature a search engine
- "Enter Application DN" said ND; "Enter Bocha Search API Key" said APi
- "Toast Notifications for New Updates" dropped "new" and used the wrong
  form of the preposition
- "Enter Top K Reranker" used a term its own label does not, and
  "Select a reranking model engine" differed from every other reranking
  string
- two of the four iframe sandbox switches did not follow the wording of
  the other two
- "Follow-Up Auto-Generation" capitalised a word mid-phrase

Only values changed; the keys and their order are untouched.
2026-09-04 11:49:24 -04:00
Classic298
ea1b58eb6a
fix: apply the shared metadata size cap to the last four vector DB backends (#29502)
Chunk metadata inserted into the vector DB carries arbitrary client-
supplied fields (e.g. custom metadata from file uploads), so it needs
the same size cap, type coercion and null-byte sanitization every
other backend already applies through process_metadata(). Four
backends never called it: Qdrant, Qdrant multitenancy, Milvus
multitenancy and Oracle23ai, so an oversized or malformed metadata
blob went into those unfiltered.

Wire process_metadata() in at each backend's single insert/upsert
chokepoint, matching the pattern already used by the other eleven
backends.
2026-09-04 11:42:20 -04:00
Classic298
82f11b14c9
refac: run the built-in file grep off the event loop (#29621)
The built-in chat and knowledge grep tools awaited their matching helper directly, so the search ran on the event loop and held it for as long as the match took. Both call sites now hand the helper to a worker thread.

Output and error handling are unchanged: 17 cases (literal, regex, case-insensitive, count-only, no-match, invalid pattern, rejected quantifiers, missing file data, result truncation) compare byte-for-byte against the previous behaviour. The matcher's time budget lives in a contextvar, which asyncio.to_thread copies into the worker, so budget scoping still behaves as before.
2026-09-04 11:41:06 -04:00
G30
cc4d84b93d
fix: show literal ascii sequences instead of arrow ligatures (#29595)
Inter ships arrow substitutions as OpenType contextual alternates, and
Open WebUI sets no font-variant-ligatures, so a sequence like <--- is
stored and sent exactly as typed while being drawn as an arrow. Reported
as text conversion in #29594, where the memory rule had in fact worked
and only the display was misleading.

Reading the DOM back with the bundled Inter file loaded in isolation
confirms every field holds U+3C U+2D U+2D U+2D. Only the glyphs differ.

no-contextual rather than none: none, no-contextual, and
font-feature-settings "calt" 0 all fix it, which places the arrow in
contextual alternates. no-contextual disables that one feature and
leaves standard ligatures intact, so ordinary prose is unaffected.

Set on html, next to the existing document-wide typography, so it
inherits to editable fields and rendered content without a new selector.
Scoping it to inputs would fix the field you type into and leave the
memory list and chat messages still drawing arrows.
2026-09-04 11:40:51 -04:00
G30
cc14f3dac0
fix: report version check failures instead of claiming latest (#29626)
GET /api/version/updates returned the running version as latest whenever
the GitHub request failed, so an instance that cannot reach GitHub
reported itself up to date however far behind it was. The exception was
logged at debug, below the default level, so nothing recorded that the
check never happened.

The failure path now returns latest: None and logs at warning.

A null latest cannot be passed to compareVersion as it stood.
current.localeCompare(null) coerces to the string "null", and "0.10.2"
sorts before it, so the function returned true. The backend change alone
would have turned a false (latest) into a false update-available plus a
toast, so the guard is part of the fix.

The three callers stop substituting the running version in their catch,
and the two badge surfaces gain a third state. When latest is unknown
the badge is plain text, since there is no release to link to.

Admin Settings > General was wrong in a worse way than reported: it
initialised updateAvailable false with latest set to the running
version, and never checked on mount, so it claimed (latest) having made
no request at all. It now matches About.svelte, which starts unknown and
checks on mount.

Closes #29580
2026-09-04 11:39:52 -04:00
Classic298
c7cce962d2
fix: stop citing web search results as sources (#29631)
A web search is not something that can be cited. The search engine
returns a title, a link and a one line snippet for each hit, and the
model never opens any of those pages. Emitting them as citation sources
produced one <source> tag per result, all named search_web with empty
bodies, and the citation template then instructed the model to cite them
by id. Models either hesitated visibly or attached an id to content from
a different result, which the citations panel then resolved to a title
that looked plausible, so the misattribution read as correct.

Web search results now stay in the tool output the model reads, and stop
being offered as things to cite. Where the model needs to cite a page it
calls fetch_url, whose citation names the URL and already works.

Web search results no longer appear in the citations panel. That is the
point of the change: the panel was offering pages that nothing had read.

Scoped to the native tool-calling path. The legacy handler cites every
tool result as one opaque source and does not single out web search, so
it is left alone rather than special-cased.
2026-09-04 11:28:18 -04:00
G30
0b804c5316
feat: shift quick delete for notes in list and grid views (#29635) 2026-09-04 11:27:57 -04:00
spoofy
18a48cffba
i18n: add Russian and Ukrainian translations for new UI strings (#29601) 2026-09-04 00:57:21 -04:00
spoofy
eb870cd9e1
i18n: translate newly added Spanish UI strings (#29599) 2026-09-04 00:57:09 -04:00
Classic298
59f9502240
fix: install black's transitive deps in the pyodide code formatter (#29503)
Saving any tool or function in the admin UI threw "ModuleNotFoundError:
No module named 'click'" in the browser, even on a fresh tool with no
custom code, because the click name never appears in user code at all.

The in-browser formatter runs black through a Pyodide/micropip worker.
black needs click, mypy_extensions, pathspec, platformdirs and pytokens
at import time, but the vendored pyodide-lock.json records none of
these as black's dependencies, and micropip only walks a locked
package's declared deps when install() is given explicit constraints,
which this worker never does. So only black itself ever got installed.

Requesting the five packages explicitly alongside black fixes it
without touching the worker or the lock file.
2026-09-04 00:45:58 -04:00
Timothy Jaeryang Baek
8aa25dd358 chore: i18n 2026-09-03 20:36:27 -04:00
Timothy Jaeryang Baek
237b11c6d9 refac 2026-09-03 20:26:40 -04:00
Classic298
7ab8b76449
i18n: complete the German (de-DE) translation (#29590)
Ten strings were still untranslated in de-DE, so German users saw raw English keys in the settings sidebar and in the interface settings. The settings group headings (Basics, Services, Preferences, Data, AI, Quality, Experience) all rendered in English above otherwise German tab names, plus the font family label and its description, and the multi-file upload failure toast in a knowledge base.

All ten now have German values, chosen to match the terminology the catalog already uses: AI as KI, Services as Dienste (matching Dienstkonto and the existing service endpoint strings), Interface as Benutzeroberfläche in the font description, and the third-person descriptive voice the neighbouring setting descriptions use. Experience is rendered as Benutzererlebnis rather than Darstellung because that group covers audio as well as interface and images. Both interpolation placeholders in the upload toast are preserved.

Only de-DE is touched. No key was added, removed or reordered, and no already-translated string was changed. The catalog now has no empty values left.
2026-09-03 20:23:38 -04:00
spoofy
63c61b7d39
i18n: complete es-ES translation and fix defective strings (#29592)
Fills the 917 empty strings in the es-ES locale and repairs 94 existing
entries that rendered incorrectly. No keys are added or removed, and the
terminology chosen by the Spanish community is preserved throughout: new
strings reuse the glossary already present in the file.

Filled strings:
- Reuse existing renderings so each English term keeps one Spanish form
  (chunk -> Fragmento, knowledge base -> Base de Conocimiento,
  workspace -> Espacio de Trabajo, skill -> Habilidad).
- Include the settings sidebar group headers added in 006f95ee5
  (Basics, Services, Preferences, Data, AI, Quality, Experience), which
  were otherwise falling back to English in the settings modal.
- Add real grammatical plural forms for every one/many/other variant.
  Spanish requires the `many` CLDR category, which English does not have.
- Preserve every {{placeholder}}, backtick, URL, HTML tag and CLI flag.

Repaired strings:
- 4 broken interpolations, including "Deleted {{name}}", whose placeholder
  had itself been translated to {{nombre}} and so rendered literally.
- 6 plural groups whose values carried the i18next key suffix as visible
  text (e.g. "{{count}} seleccionados_únicos"), and one where "_one" had
  been translated as the adjective "únicas".
- "Reset" now reads Restablecer rather than Reiniciar ("restart"), which
  misdescribed destructive actions such as Reset Vector Storage/Knowledge.
- "Access" (Acceso) separated from "Permissions" (Permisos); the *Access
  family was previously split between the two.
- "Chats Public Sharing" and "Chats Open Sharing" no longer render
  identically in the group permissions panel.
- "You" corrected to "Tú"; unaccented "Tu" is the possessive "your", and
  this string labels every user message bubble.
- "Channels" corrected from the singular "Canal".
- ~49 misspellings and missing accents (Conexxión, Interprete, Publicamente,
  Busqueda, Incrustración, Añador, Wev, actualiada, mantentrá).

Verified: 0 empty strings, key set identical to dev, 0 placeholder
mismatches, prettier --check passes, and all strings render correctly
through i18next 23.16.8 using the application's own options.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 20:23:24 -04:00
Timothy Jaeryang Baek
f677fdbf50 refac 2026-09-03 20:22:32 -04:00
spoofy
ca26632a2a
i18n: translate settings sidebar group headers (ru-RU, uk-UA) (#29591)
The settings navigation group labels added in 006f95ee5 (Basics,
Services, Preferences, Data, Experience, AI, Quality) were left
empty in every locale, so ru-RU and uk-UA fell back to the English
key text. Fill them for both locales.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-03 16:06:37 -05:00
Timothy Jaeryang Baek
006f95ee59 chore: i18n 2026-09-03 16:18:59 -04:00
Timothy Jaeryang Baek
cffd734a18 refac 2026-09-03 16:18:39 -04:00
Timothy Jaeryang Baek
894655f66b refac 2026-09-03 15:31:07 -04:00
spoofy
065002b49e
i18n: complete ru-RU and uk-UA translations (#29586)
Fill all remaining untranslated strings so both locales reach full
coverage:
- ru-RU: 967 strings translated
- uk-UA: 2110 strings translated

ru-RU also aligns terminology with the existing translations
(webhook -> вебхук, sub-agent -> субагент, reranker -> реранкер,
chunk -> чанк, parsing -> розбір/разбор) and standardizes "web search"
to «Поисковая система» across the whole file.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-03 15:27:03 -04:00
Classic298
8ba8786d6a
fix: pass file.meta through to vector DB chunks on file upload (#29499)
Custom metadata attached to a file upload (e.g. via the API's
metadata field) was stored on the file row but silently dropped
when the file was chunked and embedded for RAG, so it never reached
the LLM through retrieved sources.

process_file() builds chunk metadata in several branches; three of
them already merge the file's meta dict in, but the branch used for
a fresh upload processed through a document loader did not. Bring
it in line with the others so custom metadata flows through
consistently regardless of upload path.

Reported in open-webui/open-webui#29486.
2026-09-03 14:57:50 -04:00
Classic298
aa48106fca
fix: unregister the direct-connection stream listener on every exit (#29509)
A direct-connection streaming request registers a per-request socket.io handler before it asks the browser to start the completion, and that handler was only removed once the response had been fully streamed. Every other way the request could end left it behind for the life of the process: the call to the browser raising, a non-success status, an ack with no arguments or an ack that is not an object, cancellation while waiting for that ack and a response body closed or garbage-collected before it finished. A client that repeatedly hits a failing direct connection grows the server's handler table and the closures it holds without bound, and nothing ever cleans it up.

The removal now runs on every exit from the request, through one named helper that pops with a default so it is safe to run twice on the paths where both the generator's finally and the response's background task fire. The guard covers the exchange up to the status read and catches BaseException, because cancellation is not an Exception, and re-raises it unchanged.

Measured across thirteen exit paths: twelve leak a handler on dev and none of those leak here. The thirteenth, a response body that is never iterated at all, behaves the same on both. 200 failing requests leave 200 handlers on dev and none on this branch. Streamed bytes on the success path, exception types and cancellation behaviour are unchanged.
2026-09-03 14:57:39 -04:00
Classic298
8600b0564b
fix: use raw strings for the regex escapes in the knowledge filesystem (#29515)
The regex detection helpers wrote their backslash escapes in ordinary string
literals, so Python reported six invalid escape sequences on import. They work
today because Python leaves an unrecognised escape as its two characters, but
that behaviour is deprecated and becomes a syntax error in a future release, at
which point the knowledge filesystem tools stop importing at all.

The literals are now raw, which is the same two characters with no warning.
Pattern detection and normalisation are unchanged: verified identical output
over every string up to length five drawn from the characters these helpers
look at.
2026-09-03 14:57:25 -04:00
Anthony Chen
2831207d40
i18n: update Traditional Chinese entries - #29514 (#29549)
* add: new translations

* Update translations in zh-TW locale
2026-09-03 14:57:14 -04:00
Cyp
a9408ede42
i18n: update Korean translations - #29555 (#29556)
* i18n: update Korean translations

* Update Korean translation

---------

Co-authored-by: bglee <bglee@hct.co.kr>
2026-09-03 14:55:06 -04:00
Kylapaallikko
38cf156849
Update fi-FI translation.json (#29584)
Added missing translations.
2026-09-03 14:54:47 -04:00
G30
0cf48a04c1
fix: make slash command labels readable in light theme (#29512)
The `/` command menu is portalled into a tippy popup by
getSuggestionRenderer, and tippy's base stylesheet sets
.tippy-box { color: #fff }. The app's `transparent` theme neutralises
tippy's background, padding and margin but never its text colour, so
everything inside the popup inherits white.

DropdownMenu paints bg-white dark:bg-gray-850 dark:text-white, defining
a text colour only for dark mode, and the command rows in SlashCommands
set no colour of their own. In light theme that leaves Temporary, Model
and Settings as white text on a white menu.

Reset the colour on the transparent theme so the popup inherits from
body, which is #000 in light and #eee in dark. The secondary text and
the prompt rows were already unaffected because they set explicit
theme-aware colours.
2026-09-01 22:12:21 -04:00
Anthony Chen
dfc24dfb43
add: new translations (#29514) 2026-09-01 22:11:10 -04:00
Tim Baek
2a960a59fe
Merge pull request #29309 from open-webui/dev
0.11.3
2026-08-31 10:55:42 -04:00
Classic298
6ea321370e
chore: rebuild the 0.11.3 changelog section (#29307)
Keeps the entries already on dev and merges in the two this branch carried, then
puts the section back into the shape the format asks for.

Added now holds the contrast accessibility mode gives the dropdown menus, their
submenus and the model picker, followed by the general improvements placeholder
and the translation entry in their reserved positions. The translation entry
moves out of Changed, takes the standard wording and drops its link, which that
entry never carries; Changed is left out, having nothing else in it.

Each entry is condensed to a single sentence, the backticked column name becomes
a quoted one, and the interface font entry gains the commit it came from, having
carried no reference at all. Fixed is ordered by reach, from the conversation
branches and the failed upgrade down to the font and the disconnect control.
2026-08-31 09:48:50 -05:00
Timothy Jaeryang Baek
b299585025 chore: format 2026-08-31 10:48:28 -04:00
Timothy Jaeryang Baek
50413f3482 refac 2026-08-31 10:47:49 -04:00
Timothy Jaeryang Baek
59d3c5b064 refac 2026-08-31 10:37:46 -04:00
Timothy Jaeryang Baek
cbcbf7e326 refac 2026-08-31 10:36:14 -04:00
Badar Rahman
a3a72068c2
i18n: add missing Indonesian (id-ID) translations for v0.11.1 aligned with en-US catalog (#29294)
Co-authored-by: DarRahman <darrahman@users.noreply.github.com>
2026-08-31 10:27:38 -04:00
Timothy Jaeryang Baek
1457000ba6 refac 2026-08-31 10:27:01 -04:00
Timothy Jaeryang Baek
8c0c7b3b6c refac 2026-08-31 10:19:05 -04:00
Tim Baek
edcfddab0f
Merge pull request #29284 from open-webui/dev
refac
2026-08-31 01:51:41 -04:00
Timothy Jaeryang Baek
148391ded0 refac 2026-08-31 01:51:17 -04:00
Tim Baek
736707eeee
Merge pull request #29282 from open-webui/dev
0.11.2
2026-08-31 01:48:08 -04:00
Timothy Jaeryang Baek
471b5cbbb1 refac 2026-08-31 01:41:17 -04:00
Timothy Jaeryang Baek
a6f9751401 refac 2026-08-31 01:37:36 -04:00
Timothy Jaeryang Baek
6a2a131650 refac 2026-08-31 01:36:46 -04:00
Timothy Jaeryang Baek
2a4ef46ac8 refac 2026-08-31 01:29:36 -04:00
Timothy Jaeryang Baek
2daa610cba refac 2026-08-31 01:28:40 -04:00
Classic298
89716ea880
perf: stop scanning every socket.io payload for binary data (#28180)
* perf: stop scanning every socket.io payload for binary data

Every socket.io event the backend sends was first walked recursively to check whether any value was a bytes object needing binary attachment framing. Open WebUI never emits binary, so the walk always came back empty and the work was thrown away. It has no early exit and allocates at every level, so it scaled with the full size of the message, and the messages are the big ones: chat streaming re-emits the whole assistant message on every update, note collaboration sends document state as a JSON array with one entry per byte. With the Redis manager it ran once per instance per emit on top of that, since every instance builds its own copy of the packet.

The server now installs a Packet subclass with binary events off, through python-socketio's own serializer hook, the same mechanism its msgpack serializer uses. Inbound binary attachments are decoded to int lists rather than refused, so the one frontend path that sends a raw Uint8Array keeps working and handlers can still echo client data straight back out. One scan remains in multi-instance setups: python-socketio's Redis manager calls it on the base Packet class directly, where the serializer hook cannot reach.

Measured per encode:

| payload | before | after |
|---|---|---|
| chat completion re-emit (7.5 KB JSON) | 30 us | 13 us |
| collaborative document state (292 KB JSON) | 9.0 ms | 1.7 ms |

With ENABLE_ORJSON=true, where the scan is nearly the whole encode cost: 20 us to 2.3 us, and 7.8 ms to 0.14 ms.

Closes #28164

* fix: match the other Yjs emits and send the full state as an array

Collaboration.ts sent the initial full-document state as a raw Uint8Array while the other two Yjs emit sites convert with Array.from first. socket.io framed that one as a binary attachment, so with the JSON-only packet class the server turns it into a list of ints and re-broadcasts it as JSON: a 10240-byte state update becomes 36561 JSON characters. Converting at the emit site keeps the wire form uniform across all three sites.

Also trims the JSONOnlyPacket docstring, which claimed attachments already arrive as int lists when the override is what converts them, and annotates the new reconstruct_binary parameters.
2026-08-31 01:22:06 -04:00
Classic298
061f5e3a6d
perf: stop re-parsing the whole tool-argument buffer on every streamed chunk (#28858)
* perf: stop re-parsing the whole tool-argument buffer on every streamed chunk

Converting an OpenAI stream to Anthropic events buffers each tool call's arguments and, to find out when the JSON is complete, parsed the entire buffer again on every chunk. A tool call with large arguments pays that parse thousands of times, and the cost grows with the square of the argument size.

The parse now runs only when the buffer could actually be complete. A JSON object can only close on its final brace, so a chunk that does not end there cannot complete it. Arguments that are not an object, or that start with whitespace, keep parsing on every chunk exactly as before.

Measured on CPython 3.12 with 130 KB of tool arguments over 7648 chunks:

| | before | after |
|---|---|---|
| parses | 7648 | 1 |
| time | 382 ms | 0.82 ms |

The block closes on exactly the same chunk as before, verified by replaying randomized fragmentations of objects with braces inside strings, escaped characters, unicode escapes, arrays, bare scalars, leading and trailing whitespace and a buffer that never completes, against both JSON backends.

* refactor: read tool['arguments'] directly in the JSON completion guard

Restores the pre-existing comment above the guard to its original wording and drops the `buffered` local, so the guard and the parse call both read `tool['arguments']`, the name the rest of the file already uses for that buffer. Behaviour is unchanged: same three conditions in the same order, same short-circuit result.

* perf: strip whitespace in the tool-argument completion guard

The character guard only looked at the first and last byte of the buffer, so a
chunk that ended in a space still triggered a full parse and a tool argument
with leading whitespace fell back to parsing on every chunk. Stripping first
collapses both cases to a single parse at the end of the stream.

Measured on a streamed tool call, parses and wall time for the whole stream,
orjson on the left of the slash and stdlib json on the right:

| argument shape | before | after |
|---|---|---|
| 20 KB string, char-by-char deltas | 3678 parses, 56 / 28 ms | 1 parse, 3.1 / 3.1 ms |
| 200 KB, 20-char deltas | 1473 parses, 176 / 63 ms | 1 parse, 3.1 / 2.8 ms |
| 8 KB prose, leading whitespace | 715 parses, 5.5 / 2.5 ms | 1 parse, 0.18 ms |
| 8 KB of spaces inside a value | 713 parses, 6.0 / 2.9 ms | 1 parse, 0.83 / 0.72 ms |

The strip costs about 20 ns per delta on arguments that have no whitespace at
either end, which is where the old form was already optimal: a 20 KB compact
argument goes from 191 to 216 us over 1786 deltas. Soundness is unchanged, the
guard can still only skip a parse that would have failed: 2660892 buffers
(exhaustive to length 6 over a JSON-lexical alphabet, every prefix of 26 named
cases with a trailing byte appended, and every codepoint below U+3000 after a
complete document) with zero cases where a parse would have succeeded.
2026-08-31 01:18:05 -04:00
Classic298
d7674c5174
perf: bounded non-blocking session pool reaper, fewer blocking pool round trips (#28835)
* perf: stop the Socket.IO session pool blocking the websocket event loop

With WEBSOCKET_MANAGER=redis the session pool is a synchronous Redis client, so every call into it blocks the whole worker's event loop, not just the caller. Two paths did it constantly: the orphan reaper walked the pool one round trip per session with no await anywhere, freezing the loop for the entire sweep every cycle, and nearly every socket event re-read the sender's session back out of Redis. Other users' events and every in-flight generation on that pod wait behind both.

The reaper now walks the pool in HSCAN batches and deletes in bulk, yielding between batches, and no longer sleeps past half the lock TTL, which previously guaranteed a failed renew every cycle. The per-event reads are gone: Socket.IO events only reach the worker holding the connection, and that worker already saved the same session dict locally when the user authenticated, so it was asking Redis for its own data. The writes stay, since those are what other pods read.

Measured at 4000 users / 16 containers, Redis 1.1 ms away:

| | before | after |
|---|---|---|
| reaper sweep, 5k sessions | 5.6 s, loop frozen throughout | 62 ms, 4.2 ms worst block |
| same, crash recovery with every session expired | 11.5 s | 96 ms |
| heartbeat / usage ping / disconnect | 2 / 3 / 2 round trips | 1 / 2 / 1 |
| 50-member channel post | 50 round trips, 57.2 ms block | 0 round trips, 0.02 ms |
| loop time per wall second at rest | 103 ms (10.3%) | 67 ms (6.7%) |

The alternative, converting RedisDict to the async client, fixes the same paths with a far larger blast radius (every call site gains await, and `in`/`[]`/`del` cannot be awaited so the dict interface goes) and still round-trips for data already in memory. Two deliberate behaviour changes: a heartbeat re-adds a session the reaper already removed, so a tab that survives a stall recovers instead of staying out of the pool until it reconnects; and disconnect no longer skips Yjs document cleanup when the pool entry is already gone, which previously leaked that document's update log forever.

Closes #28172

* perf: cut disconnect and user session lookup pool round trips, harden the session reaper

Follow-up on top of the session pool reaper branch. With WEBSOCKET_MANAGER=redis two paths still blocked the worker's event loop on synchronous Redis calls. Every disconnect listed all models in use cluster-wide and fetched each one individually, one blocking round trip per model. Disconnecting all sessions of a user (admin role change or deletion) pulled the entire session pool in one HGETALL and decoded every entry in a single uninterrupted block.

Disconnect now fetches the usage pool once with items(), going from 2+N+M round trips to 2+M (N models in use cluster-wide, M models the session used), and its delete of an emptied model entry is KeyError-guarded because another node can remove the same key between snapshot and delete; unguarded, that race aborted the handler and skipped its Yjs document cleanup. The user session lookup reuses the reaper's HSCAN batches and yields to the loop between pages. The reaper previously died permanently on the first Redis connection error, on every node at once during an outage; it now logs, releases the lock and returns to retrying acquisition.

* refac: keep the socket pool perf work to the round trips

A review pass on this branch turned up four changes riding along with the round-trip work without belonging to it, so they are backed out here. The `disconnect` handler keeps its `if sid in SESSION_POOL:` guard, so USAGE_POOL and ydoc cleanup stay off the path for sockets that never authenticated. `RedisDict.set()` keeps its own inline HDEL. `get_session_ids_by_user_id` stays synchronous over one HGETALL, since it runs on user delete and role change rather than per message. The crash-resilience wrapper around the reaper loop is dropped; if that guard is worth having, it belongs in its own change.

What stays is the perf part. The reaper now sweeps the pool in bounded HSCAN batches and deletes expired sids with one HDEL per batch, down from HKEYS plus an HGET and a per-sid HDEL across the whole pool. The `disconnect` handler reads USAGE_POOL with a single HGETALL, down from HKEYS plus one HGET per model in use. Session lookups in the socket handlers come from the local Socket.IO store, which removes one Redis GET from every heartbeat, usage, channel and ydoc event.

Naming and annotations follow the file: `get_session_pool_batches` for the module's `get_` prefix, `RedisDict.pop_many` so both reaper branches use one word for removing keys, a named `SCAN_BATCH_SIZE`, and types on the new helpers.

* fix: invalidate the RedisDict write signature on batch delete

RedisDict.set() skips the write when the payload fingerprint matches the last one this process wrote, so a mutation that goes around set() has to clear that fingerprint. The new batch delete did not, leaving a stale fingerprint behind: the next refresh with identical content is treated as already written and silently skipped, so the hash stays empty.

Renamed pop_many to delete_many. In a dict emulation pop removes and returns; this returns nothing and cannot without an extra HMGET, so the name promised something it does not do. delete_many matches __delitem__ and the HDEL underneath. Its only call site is the session pool reaper, whose behaviour is unchanged: same fields deleted, same batching, same return.
2026-08-31 01:17:53 -04:00
Classic298
ac6a8c0082
chore: changelog (#29107)
* chore: add changelog entries for 0.11.2

Documents the commits landed on dev after the 0.11.1 changelog entry. Added covers the richer terminal file previews with page thumbnails, the reduced per-message overhead on deployments without pipelines, more room in file previews on touch screens, and the wider accessibility coverage. Fixed covers twelve user-facing corrections, among them stalled streaming on reasoning models, post-tool-call thinking leaking into replies, banners with underlined text failing to render, pinned models carrying the previous model's tools, disabled admin models, skill-mention text loss, and the workspace Knowledge list staying empty. Changed records the rename of High Contrast Mode to Accessibility Mode. Also records the Polish, Simplified Chinese, German, Catalan, and Portuguese (Brazil) translation updates. Issue template, pull request template, Docker workflow and locale catalog regeneration edits are omitted as they are not user-facing.

* chore: add the Redis Cluster stop and Valves overflow entries to 0.11.2

Documents the two user-facing commits landed on dev since the previous
changelog entry. Fixed gains the Redis Cluster stop signal, where the stop
button did not take effect when the request landed on a different instance
than the one streaming the reply, placed with the streaming and thinking
entries it shares a domain with; and the Valves dialog overflow, where a
valve with a long line of selected options stretched its input past the
edge of the dialog and over the page behind it, placed with the narrow
screen layout entry.

The section date moves to 2026-08-29 to cover the newer commits. The issue
and pull request template wording and the German locale catalog are omitted,
the former as contributor-facing rather than user-facing, the latter as
German is already named in the translation entry.

* chore: add the recurring calendar event entries to 0.11.2

Documents the calendar recurrence changes landed on dev after the previous
changelog commit. Fixed records repeating events working out their occurrences
from their own date and time rather than from a start date carried inside the
repeat rule, which could place them on the wrong weekday or hour. Changed
records the new limit refusing events that repeat more often than once a day.
The EXRULE handling and the timezone resolution rewrite are omitted: both reach
the same user-visible outcome as before, only by a clearer route.

* chore: add the security advisory notice to 0.11.2

Adds the standard advisory notice as the first item in the Fixed section. The
calendar recurrence work landed on dev under an unmarked commit message and
bounds the occurrences a single stored event can force the server to walk, so
the release carries a fix whose details are not spelled out in the entries
below it. The notice is the fixed wording and takes no reference links of its
own; the individual entries keep theirs.

* chore: add the structured output crash entry to 0.11.2

Fixed records the conversation that failed in the browser and stopped showing
the assistant reply until a reload, together with the recovery of chats already
saved in that state. It sits directly below the advisory notice as the most
disruptive correction in the section. The Irish catalog update joins the
translation entry; the Portuguese (Brazil) pass needs no change there, as that
language is already named.

* chore: add the dropdown, SQLite search and tool server entries to 0.11.2

Fixed gains the dropdown that opened past the edge of a narrow screen and the
dropdown list that ignored the interface theme, kept together as one group, plus
case-insensitive matching for accented and non-Latin text on SQLite installs,
placed beside the existing SQLite entry, and the tool server connection that was
sent an empty authorization header when saved without a key.

The advisory notice already stands at the top of the section, so the unmarked
backend commit needs no further flag there.

* chore: add the chat reload entry to 0.11.2

Fixed records the conversation that reloaded itself whenever any response in it
finished while an older unfinished reply sat in the history, now narrowed to the
reply the update concerns. It joins the response lifecycle group below the stop
entry. The commit carries no pull request or issue, so it is referenced by
commit.

* chore: add the interface font and touch resize entries to 0.11.2

Added gains the font family field in Interface settings, which applies a locally
installed font across the interface and falls back to the standard font when
cleared, and the side panel divider that can now be dragged by touch or stylus
while a mouse drag keeps tracking beyond the window edge. Both sit above the
reserved accessibility, general improvements and translation entries, with the
touch entry beside the existing touch screen one.

Each aspect of the font setting arrived in a single commit, so it is recorded as
one entry with no separate note for its configurability.

* chore: add the automation schedule and model registry entries to 0.11.2

Added records the model list refresh that no longer has every worker rewrite the
whole list to the shared cache when nothing changed, placed beside the existing
performance entry.

Fixed records the two automation schedule defects from the same pull request as
separate entries, because the symptoms differ: a counted schedule shown as a
single run and rewritten to one on save, and a schedule carrying a start date
losing its weekly or monthly setting and printing raw rule text in the list.
Both join the scheduling group below the recurring event entry.

* chore: extend the touch resize entry to the main sidebar

The sidebar divider received the same pointer handling the side panel dividers
got, so the existing entry now names the sidebar and carries both commits rather
than repeating itself as a second entry. The follow-up that moved the divider
border to the matching edge is listed with them, being a further correction to
the same divider and too small to record on its own.

* chore: record the preview focus and caveat changes in 0.11.2

Fixed gains the arrow keys that paged an open document or slide preview from
anywhere on the page, which also took those keys away from the field being typed
in, now confined to the focused preview.

The richer previews entry absorbs the removal of the notice warning that a
preview might differ from the download, the caveat having gone with the
approximation it described, and the accessibility entry absorbs the previews
becoming reachable by keyboard and announcing themselves. Neither warranted an
entry of its own, both being continuations of work already recorded.
2026-08-31 01:14:29 -04:00
Timothy Jaeryang Baek
5d74df95ed chore: format 2026-08-31 01:11:41 -04:00
Classic298
b75e2670b7
fix: keep a custom recurrence rule when the editor reopens it (#29260)
Loading an automation whose rule the visual controls cannot represent switched the schedule to Custom but left the bookkeeping the seeding block reads on the previous value, so that block immediately replaced the rule with a freshly built default. The rule was lost when the editor opened, before anything was saved, and cloning carried the default across as well. Recording the switch alongside it leaves the stored rule in place.

Verified in a browser against the same build without this line: a minutely rule and a yearly rule now survive reopen and save byte for byte, cloning keeps the original, and every schedule the editor itself produces, along with switching to Custom by hand, behaves exactly as before.
2026-08-31 00:08:28 -05:00
Classic298
1976387808
fix: stop labelling a counted schedule as a one-off (#29261)
The schedule label treated any rule whose text contained COUNT=1 as a single run, so counts such as 10, 12 and 14 were shown as "Once" together with the date of the first run, on the automations list and on the automation page alike. The label now matches a count of exactly one.

This covers the two places that render the label. The schedule editor reads the count the same way and changes separately. Rules that carry a start date still fall through to the raw rule text, exactly as they already did without a count; that parsing gap changes separately too.

Verified in a browser against the same build without these lines: ten ordinary schedules render identically in both places, and a genuine single-run schedule is still labelled as one.
2026-08-31 00:08:01 -05:00
Timothy Jaeryang Baek
fd679e1dac refac 2026-08-31 01:06:15 -04:00
Classic298
9f680bb80b
perf: skip the tool approval drain lookup for fresh chat messages (#29142)
Every chat completion request re-loads the target conversation's entire
message history from the database inside drain_approved_tool_calls() before
discovering there is nothing to drain: a fresh message always points at a
newly minted assistant message with no stored output, so the full-history
read (one SELECT of every chat_message row plus building the message map,
uncached, on top of the identical read process_chat_payload already did) is
pure overhead on every message.

Queued tool approvals can only ever be acted on by a resume or continue
request, and exactly those requests carry assistant_message_id in their
payload. The drain now returns early when the field is absent, removing one
O(conversation length) query per chat message while resume, continue, reject
and pause flows behave exactly as before, independent of the approval mode.
2026-08-30 23:57:11 -05:00
TOM
949876f9c0
fix(i18n): disable key/ns separator splitting at runtime to match i18next-parser config (#29161) 2026-08-31 00:54:08 -04:00
Timothy Jaeryang Baek
6609918bfe refac 2026-08-31 00:53:55 -04:00
Timothy Jaeryang Baek
9962d122c9 refac 2026-08-31 00:46:34 -04:00
Classic298
188fc83a79
fix: surface files the browser cannot read during a knowledge base directory sync (#29135)
* fix: surface files the browser cannot read during a knowledge base directory sync

Syncing a local folder into a knowledge base could fail with nothing but "Error accessing directory": no failing file name, no network request, no server log, and no console output either, because production builds strip console.error. On Windows this happens once the absolute path of a file passes the platform limit, at which point the browser refuses to open a file it just listed.

The directory scan now handles that per file. It names the first failing path and how many files are affected, and stops before anything is uploaded. Stopping is the point: a file missing from the manifest is treated as deleted by the sync, so continuing would remove the knowledge base copy of a file that still exists on disk.

Dragging a folder in hit the same failure and reported nothing at all, and the Firefox picker path returned its promise without awaiting it, so a rejection escaped the error handler and surfaced only as an unhandled rejection. Both report through the existing handler now, and production builds keep console.error so the underlying exception stays visible.

* fix: narrow the change to the silent drag-and-drop folder failure

Dropping a folder onto a knowledge base did nothing at all when the browser refused to open one of the files inside it: the rejection escaped the async drop listener, so the user got no toast, no upload and no clue why. The listener now routes that failure through the same error handler the directory picker already uses, so one path and one message cover both ways of adding a folder.

The rest of the branch is reverted. Dropping `console.error` from the esbuild `pure` list un-stripped 597 call sites across 103 files from every production bundle, which is a repo-wide logging policy change that needs its own argument. The picker-side collect-and-count machinery only reworded a toast the existing catch already showed, and the Firefox `return await` fix is a different bug in a different path.
2026-08-31 00:45:26 -04:00
Timothy Jaeryang Baek
8ed5487693 refac 2026-08-31 00:41:11 -04:00
Timothy Jaeryang Baek
1caf22b5a8 refac 2026-08-31 00:40:15 -04:00
Timothy Jaeryang Baek
873fb741c2 refac 2026-08-31 00:39:16 -04:00
Timothy Jaeryang Baek
b6d5055228 refac 2026-08-31 00:35:17 -04:00
Timothy Jaeryang Baek
e8bdbd716b refac 2026-08-31 00:33:43 -04:00
Timothy Jaeryang Baek
e96b6464b4 refac 2026-08-31 00:33:05 -04:00
Timothy Jaeryang Baek
7a11154182 refac 2026-08-31 00:32:46 -04:00
Classic298
be958d7b04
fix: a rejected ask_user call ending the turn with no reply (#29252)
* fix: a rejected ask_user call ending the turn with no reply

The documented behaviour of the built-in ask_user tool is that a call breaking its rules comes back to the model as an error. Instead the reply stopped there: the error was recorded as the tool result, the model was never asked again, and the user was left with a dead chat and no answer.

The rejection is now handed back like any other failed tool result, so the model sees it and can correct itself within the normal tool-call iteration limit. Any ordinary tool the model emitted in the same turn still runs.

A call rejected for arriving alongside other ask_user calls also left those siblings without a result, which the UI shows as a tool call stuck on "Executing..." forever. Every invalid call now gets its own result. Two ask_user calls on their own also reported the wrong reason, saying the call must be made by itself rather than that only one is allowed per turn.

Fixes #29077

* Keep the original ask_user validation order

Restores the pre-existing check order and the unchanged output id fallback, so this change only alters the return shape needed for staging, and trims a comment that narrated the lines below it.

* Correct the ask_user sibling-call error message

* Shorten the ask_user sibling-call error message

* Drop the untrue sibling-call claim from the ask_user error

The ask_user error text told the user and the model "The others ran.", but that sentence is written into the turn output before any sibling tool call has executed, so it can be plainly false. Under a saved chat with tool approval set to ask, the turn pauses right afterwards and the siblings sit at pending/queued, so the user reads "The others ran" directly above the approval prompt for tools that have not run, and reads it again beside the rejection result if they decline. When the model sends two ask_user calls and nothing else, nothing runs at all and the sentence is emitted twice.

The staging helper cannot see what happens to the sibling calls, so it no longer narrates it. The remaining two sentences hold in every flow: ask_user really is dropped from the executed calls whenever this error is set, and calling it on its own is always the right retry.
2026-08-31 00:17:21 -04:00
G30
95032b6c61
fix: let the model defaults capability and prompt suggestion sections scroll (#29235) 2026-08-31 00:11:34 -04:00
Timothy Jaeryang Baek
756241b34a refac 2026-08-31 00:11:13 -04:00
Timothy Jaeryang Baek
84d0940da1 refac 2026-08-31 00:11:04 -04:00
G30
9a669197c8
fix: let setting row controls shrink so long values do not squeeze the label (#29229) 2026-08-31 00:06:37 -04:00
Timothy Jaeryang Baek
09163ccc73 refac 2026-08-31 00:06:05 -04:00
Timothy Jaeryang Baek
2140c189e1 refac 2026-08-31 00:05:34 -04:00
Timothy Jaeryang Baek
81b9afb731 refac 2026-08-31 00:03:51 -04:00
Timothy Jaeryang Baek
64e6c9f010 refac 2026-08-30 23:56:01 -04:00
Timothy Jaeryang Baek
97c5f52bbc refac 2026-08-30 23:55:29 -04:00
Timothy Jaeryang Baek
e4dbfb1276 refac 2026-08-30 23:50:30 -04:00
Timothy Jaeryang Baek
aeb126b95d refac 2026-08-30 23:46:02 -04:00
Timothy Jaeryang Baek
df495a7945 refac 2026-08-30 23:42:35 -04:00
G30
039976ef24
fix: unregister a sidebar folder from the registry when it unmounts (#29121) 2026-08-30 23:35:34 -04:00
Timothy Jaeryang Baek
d5e35ea6f4 refac 2026-08-30 23:32:42 -04:00
Timothy Jaeryang Baek
6c2e0d3fe8 chore: format 2026-08-30 23:31:23 -04:00
Timothy Jaeryang Baek
7d694570aa refac 2026-08-30 21:47:28 -04:00
Timothy Jaeryang Baek
b3ba6823a9 refac
Co-Authored-By: G30 <50341825+silentoplayz@users.noreply.github.com>
2026-08-30 21:41:03 -04:00
Timothy Jaeryang Baek
ddc886fdc1 refac 2026-08-30 21:39:13 -04:00
Timothy Jaeryang Baek
49aab7451c refac 2026-08-30 21:36:17 -04:00
Classic298
d8133c905a
fix: serve module scripts and wasm assets with the correct MIME type (#29139)
On Windows hosts the built-in code interpreter fails immediately with "Failed to fetch dynamically imported module: .../pyodide/pyodide.asm.mjs", and the browser console shows the server answered with a MIME type of "text/plain". Code execution is unusable for those users.

Python's mimetypes module reads the Windows registry after loading its own table, so a stray registry entry silently replaces the correct type for an extension and Starlette then labels the file with it. Browsers enforce strict MIME checking for module scripts and streaming WASM compilation, so the pyodide loader gets refused. The same workaround already existed for .js; this extends it to the two other extensions pyodide ships, and moves it out of the frontend-build branch so the unconditionally mounted /static assets are covered as well.

Fixes #29133
2026-08-30 21:32:31 -04:00
Timothy Jaeryang Baek
120409ef01 refac 2026-08-30 17:45:50 -04:00
Timothy Jaeryang Baek
e4694f82eb refac 2026-08-30 17:45:18 -04:00
Timothy Jaeryang Baek
b356b80f8c refac 2026-08-30 17:44:39 -04:00
Timothy Jaeryang Baek
492ccf3ac0 refac 2026-08-30 17:39:56 -04:00
Timothy Jaeryang Baek
78d8c9166f refac 2026-08-30 17:38:45 -04:00
Timothy Jaeryang Baek
e250be48ee refac 2026-08-30 17:30:11 -04:00
Timothy Jaeryang Baek
f0ffa7508e refac 2026-08-30 17:26:03 -04:00
Classic298
0e65c65cc7
fix: stop counted schedules being rewritten as one-offs on save (#29263)
* fix: stop labelling a counted schedule as a one-off

The schedule label treated any rule whose text contained COUNT=1 as a single run, so counts such as 10, 12 and 14 were shown as "Once" together with the date of the first run, on the automations list and on the automation page alike. The label now matches a count of exactly one.

This covers the two places that render the label. The schedule editor reads the count the same way and changes separately. Rules that carry a start date still fall through to the raw rule text, exactly as they already did without a count; that parsing gap changes separately too.

Verified in a browser against the same build without these lines: ten ordinary schedules render identically in both places, and a genuine single-run schedule is still labelled as one.

* fix: stop counted schedules being rewritten as one-offs on save

The schedule editor decides that a rule is a one-off by looking for the text COUNT=1 anywhere in it. A rule that runs ten times carries COUNT=10, which contains that text, so opening such an automation shows it as a single run and saving writes a genuine one-off rule back. One open and save is enough to silently turn a ten run schedule into a one run schedule, with whatever date happened to sit in the rule. The check now requires that no further digit follows, the same test the two schedule labels already use.

The same screens also failed to read counted rules at all. Both label helpers and the editor parser split the stored rule on semicolons after stripping the RRULE prefix, so when the rule carries a DTSTART line the first piece is that whole line and the frequency is never found. The automations list then printed the raw rule text where a human label belongs, and the editor fell back to a plain daily schedule, quietly discarding the weekly or monthly settings on the next save. All three sites now drop the DTSTART part before splitting. They split on whitespace, so the newline form and the space separated form are both handled, matching how the backend already strips it.
2026-08-30 16:12:43 -04:00
Classic298
b8f279b8fb
perf: stabilize the model registry signature across workers (#29264)
The Redis-backed model registry skips its write when the content signature matches what is already stored. That skip has never worked across processes. Two of the values it hashes come out of Python sets, and set iteration order varies with each process's hash seed, so every worker computed a different signature for identical content and every worker rewrote the whole registry on every refresh.

Sorting both makes the signature depend on content alone. Measured on a 120 model registry, 522 KiB serialized: a refresh whose content already matches drops from GET, HKEYS, HSET and SET at 5.1 ms to a single GET at 2.2 ms per worker, and the 522 KiB write leaves the wire entirely.

Verified across 12 child processes with 12 distinct hash seeds: 12 different signatures before, 1 after. Filter execution order is unaffected, because the filter pipeline re-sorts by priority and id before running.
2026-08-30 16:12:31 -04:00
Timothy Jaeryang Baek
cfa2d25317 refac 2026-08-30 12:42:13 -04:00
Timothy Jaeryang Baek
b9765fe979 refac 2026-08-30 12:36:16 -04:00
Timothy Jaeryang Baek
22379ded1a refac 2026-08-30 12:22:32 -04:00
Timothy Jaeryang Baek
58a3fadbf3 refac 2026-08-30 12:13:06 -04:00
Timothy Jaeryang Baek
26f37426b7 refac 2026-08-30 12:09:34 -04:00
G30
6fa50a6558
fix: keep select dropdowns inside the viewport (#29226) 2026-08-30 12:06:27 -04:00
joaoback
f76fd904b3
i18n: add pt-BR translations for newly added UI items and consistency pass (#29217)
New **pt-BR** translations for items introduced in the latest releases, plus a consistency/quality pass across existing strings (grammar, tone, capitalization, pluralization). Placeholders and hotkeys preserved. No logic changes.
2026-08-30 11:58:03 -04:00
Aindriú Mac Giolla Eoin
1c59d2b1b0
i18n: Updated Irish translation (#29247) 2026-08-30 11:56:59 -04:00
Classic298
1f529c4eb3
fix: structured output renderer crashing on an empty output slot (#29250)
A chat could hard-fail in the browser with "TypeError: can't access property content" and stop rendering the assistant message until a reload.

Streamed response items and content parts were placed at the index the provider reports. That index is not bounded by the length of the array the client has built up, so an entry could land past the end and leave a gap behind. Spreading the array on the next event turned that gap into a real empty entry, and the renderer then dereferenced it while looking for message text.

Output items that would land past the end are now appended, since later events locate them by id anyway. Content and summary parts are padded up to the index instead, because a part carries no id and the text streamed for it is addressed by that same index. The renderer and the structured editor now skip an empty entry as well, so chats already saved in the broken state still display and edit.

Fixes #29244
2026-08-30 11:56:41 -04:00
Timothy Jaeryang Baek
a93c508038 refac 2026-08-29 16:14:32 -04:00
Timothy Jaeryang Baek
348751b6a5 refac 2026-08-29 15:07:48 -04:00
Classic298
a5ea8b0b8a
fix: stopping a response across instances on Redis Cluster (#29165)
On Redis Cluster deployments the stop button never stopped a running response when the request landed on a different instance than the one streaming it. The pub/sub listener that carries the stop signal between instances never managed to subscribe, so the command was published to a channel nobody was listening on.

The listener subscribes through a cluster client that connects lazily, and redis-py resolves the pub/sub node from a slot cache that is still empty at that point, which fails with a bare KeyError. Awaiting initialize() first fills that cache. It is a no-op on standalone and Sentinel clients, so nothing has to branch on the deployment type, and it stays inside the reconnect loop so a failover refreshes the cache instead of resubscribing against a stale one.

Before 0.11.1 the listener died on that first exception and cross-instance stop never worked at all. The reconnect loop added in 0.11.1 turned it into a startup window plus KeyError retry spam in the logs. Reported upstream as redis/redis-py#4296.

Fixes #19840
2026-08-29 15:00:52 -04:00
Classic298
26074e0a46
i18n: complete and correct German (de-DE) translations (#29179)
Fills the two remaining untranslated strings in the German catalog and corrects a number of existing entries.

The catalog addresses the user formally with "Sie" in over two hundred strings but had drifted to the informal "du" in around thirty, including "Wählen Sie ein Modell" sitting directly alongside "Wähle eine Option". Those now use "Sie", or the infinitive where the surrounding labels already use it. The "Du" chat bubble label and the two model-facing system prompts are deliberately left informal.

"Explored" carried a trailing ": " that the English source and every other translated locale lack. The summary text beside it only renders when there are tool calls or code interpreter runs, so a details group containing only reasoning items rendered a dangling "Untersucht:" in German.

"Persistent" was left in English beside its already translated sibling option "Ephemeral" ("Flüchtig"), leaving a half translated storage dropdown. The calendar tool description promised listing, searching, creating, updating and deleting calendars, when it acts on calendar events.

The remaining changes fix compounds written as two words ("Skill Beschreibung", "Datei upload", "Audio tag"), replace the non-word "managen" with "verwalten", correct a grammatical gender and a plural, and normalize the only two ellipsis characters in the file to the three periods used by every other entry.

Only de-DE is touched. No catalog regeneration, no other locale files.
2026-08-29 14:53:42 -04:00
G30
fe947f7e68
fix: prevent valve inputs from overflowing the valves modal (#29203) 2026-08-29 14:52:29 -04:00
TOM
b2050bcd4d
Polish translation update (#29184) 2026-08-28 20:31:09 -04:00
Classic298
83556188f5
fix: hide file preview zoom buttons on touch devices (#29176)
On a phone the file preview zoom bar shows plus and minus buttons that duplicate pinch-to-zoom and sit on top of an already small preview. They are hidden on coarse pointers now, so touch users get the preview area back and still zoom the way they expect.

The check is the pointer type rather than the viewport width, because what matters is whether the person can pinch, not how narrow their window is. A narrow desktop window keeps the buttons, a tablet does not.

The zoom percentage doubles as the reset control and has no gesture equivalent, so it stays visible, along with all page and slide navigation. Keyboard zoom is unaffected. Word document previews are left alone: they have no pinch support at all, so hiding their buttons would remove zooming entirely.

Fixes #29152
2026-08-28 18:46:53 -04:00
Timothy Jaeryang Baek
a235bf076c refac 2026-08-28 17:50:37 -04:00
spoofy
afb9f3c2d1
fix: name the Folder modal close button (WCAG 4.1.2) (#29160) 2026-08-28 17:49:22 -04:00
Timothy Jaeryang Baek
233681464f refac 2026-08-28 13:45:56 -04:00
Timothy Jaeryang Baek
797b4c51d1 refac 2026-08-28 13:40:45 -04:00
Timothy Jaeryang Baek
99cb8257f7 refac 2026-08-28 13:34:10 -04:00
G30
12d4b4ac59
fix: give the underline marked extension a renderer (#29118) 2026-08-28 12:23:06 -04:00
Classic298
88bbe4e1d7
perf: skip pipeline filter session setup when no filters exist (#29146)
process_pipeline_inlet_filter() and its outlet counterpart construct and tear
down an aiohttp ClientSession, with its own connector and cookie jar, on
every chat completion and every task generation request just to iterate an
empty filter list. On deployments without pipelines, which is the default,
that is wasted setup on every message.

Both functions now return the payload untouched before the session is
created when there is nothing to call. The per-call saving is small, a few
microseconds of object construction per request on the pinned aiohttp; the
point is that requests stop paying setup for a feature that is not
configured.
2026-08-28 12:22:23 -04:00
Timothy Jaeryang Baek
e6031ea6da refac 2026-08-28 12:18:37 -04:00
Timothy Jaeryang Baek
62e5eab60b refac
Co-Authored-By: G30 <50341825+silentoplayz@users.noreply.github.com>
2026-08-28 12:18:23 -04:00
fengyufeiyang-dot
95492cc0b4
i18n: complete zh-CN translation for Open WebUI v0.11.1 (#29151)
Co-authored-by: fengyufeiyang-dot <fengyufeiyang-dot@users.noreply.github.com>
2026-08-28 11:28:55 -04:00
Classic298
b3591e60b1
fix: apply the selected model's tools and skills when starting a chat from the sidebar (#29058)
Some checks failed
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
Clicking a pinned model in the sidebar started a new chat with that model but sent the previous model's tool_ids and skill_ids, and they stayed until the page was reloaded. The model dropdown was unaffected, and so was temporary chat.

A pinned entry links to /?model=<id>, which runs the new-chat path. That path applied the model's defaults and then restored the composer draft over the top, and the draft still held the selection from whichever model was active when it was written. The draft save is debounced, so the stale value was reliably the one read back.

The draft is now restored before the defaults are applied, so the model always decides which tools and skills are active while the unsent prompt, files and approval mode are still kept. Starting a new chat with several models selected now clears the selection instead of carrying the draft's over, since there are no per-model defaults to apply in that case.

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-27 19:24:19 -04:00
Timothy Jaeryang Baek
17cc566707 refac 2026-08-27 19:23:44 -04:00
Classic298
9277879bcf
i18n: complete the German (de-DE) translation (#29108)
173 keys in the de-DE catalog still had empty values, so German users saw those strings rendered in English: the whole terminal file browser, the tool-call approval prompts, chat variables, the model manager, calendar navigation, the accessibility labels for zoom, camera and call controls, and several admin settings panels.

Each string was translated against its actual call site rather than in isolation, so the grammatical form fits the widget it renders in (imperatives on buttons, nouns on labels and select options, participles on toasts). Terminology and the formal "Sie" register follow what the catalog already uses elsewhere, and technical literals were left alone on purpose: the CSV header hint mirrors the file that "Download CSV Template" actually produces, and the MIME pattern, the snake_case variable placeholder and the product names stay verbatim.

Only value strings changed; key order and formatting are untouched.
2026-08-27 18:26:04 -04:00
Classic298
3749e7dc74
fix: stop streaming responses breaking on a duplicate output key (#29053)
* fix: stop streaming responses breaking on a duplicate output key

With reasoning-capable models the chat froze mid-stream: the first chunk of the answer appeared, nothing followed, and the whole message only showed up once generation finished. The browser console showed a Svelte each_key_duplicate error.

When a stream event addresses an output slot past the end of the array, the missing slots were filled with the event's own item, id included, so a gap of two left two entries claiming the same id. The next chunk for that item was matched by id, landed in the first of the two, and the rendered list ended up with two items sharing a key, which Svelte refuses to update.

Only the addressed slot now takes the event's item, and the slots before it are anonymous placeholders. Replayed the reported event sequence against the real code: keys are unique again and the chunks stay in order instead of being split across the copies.

* fix: stream reasoning deltas when the provider also sends reasoning_details

Providers such as OpenRouter emit reasoning_details alongside the reasoning
text on the same delta. Merging those details cleared the pending event
unconditionally, discarding the response.reasoning_text.delta that had just
been built, so the client received no reasoning until the response completed
and the thinking block only appeared after generation finished.

The event is now only dropped when the details were all there was to report.
Details persistence is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014uuEg4AXPs9zE3vVUfN1Fj

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-27 18:25:49 -04:00
Classic298
87bed3f0b3
fix: stream post-tool-call thinking into the Thoughts section (#29052)
After a tool call, the model's thinking was streamed into the chat as if it were the main response, and only jumped into the collapsed Thoughts section once the turn finished. Every further tool call repeated it.

Each tool round appended an empty placeholder message item to the output and sent it to the browser, then dropped it again from the copy used to offset the next round's item indices. The browser therefore held one item more than the backend counted, so the first thinking chunk of the next round was written into that leftover message item and rendered as normal text until the finished output replaced it.

The placeholder is removed. It was never needed: a message item is already created when actual content arrives, and dropping it also stops an empty assistant message being sent back to the model on the follow-up request.

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-27 16:56:00 -04:00
G30
fa7de50c60
refac: align the admin Users tab counts with the workspace tab count behavior (#29080) 2026-08-27 16:45:34 -04:00
G30
1c5128c4aa
fix: keep a calendar event visible after its date is moved (#29085) 2026-08-27 16:44:58 -04:00
Classic298
0afe69e1a7
fix: stop deleting user text that looks like a skill mention (#29051)
Any `<$...>` run in a chat message was treated as an inline skill mention and removed before the request reached the model, so text like `<$(=MonthStart($(vMaxMonthEndINC)))"}, [Registration day] >` silently vanished mid-message and the model only saw the part before it.

The mention regexes accepted any character except `|` and `>` as the skill id, so they matched far more than real mentions. Skill ids are already validated as `[a-z0-9_-]+` when a skill is created, so both regexes now require that charset. Ordinary text passes through untouched while `<$id>`, `<$id|Label>` and `</id|Label>` still resolve and strip as before.

Verified against the reported message (now preserved verbatim) and the three mention forms.

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-27 16:44:33 -04:00
Aleix Dorca
f7db201ff5
i18n: Update catalan translation.json (new strings and typos) (#29043) 2026-08-27 16:43:58 -04:00
G30
98c0ac9355
fix: stop the workspace model id and timestamp overlaying the model name on narrow screens (#29084) 2026-08-27 16:43:36 -04:00
G30
aad1072874
refac: reset the shared chats modal search when the modal closes (#29082) 2026-08-27 16:43:23 -04:00
Timothy Jaeryang Baek
ae549d3e4a refac 2026-08-27 16:41:43 -04:00
Timothy Jaeryang Baek
9d95a0148b refac 2026-08-27 16:06:25 -04:00
Timothy Jaeryang Baek
061fb43432 refac 2026-08-27 15:34:20 -04:00
Timothy Jaeryang Baek
240c795efe refac 2026-08-27 13:12:55 -04:00
Timothy Jaeryang Baek
56d296ef1b refac
Some checks failed
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
2026-08-25 21:45:34 -04:00
joaoback
2bc122e05e
i18n: add pt-BR translations for newly added UI items and consistency pass (#29029)
New **pt-BR** translations for items introduced in the latest releases, plus a consistency/quality pass across existing strings (grammar, tone, capitalization, pluralization). Placeholders and hotkeys preserved. No logic changes.
2026-08-25 21:44:12 -04:00
G30
673938fd49
Merge pull request #29037 from silentoplayz/fix/admin-models-disabled-hidden
fix: restore visibility of disabled models visible in the admin Models list
2026-08-25 21:43:51 -04:00
Tim Baek
d3e8bf3405
Merge pull request #29027 from open-webui/dev
Some checks failed
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Release to PyPI / release (push) Has been cancelled
Release / publish (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
0.11.1
2026-08-25 17:17:36 -04:00
Timothy Jaeryang Baek
db22d89cab refac 2026-08-25 16:57:22 -04:00
Timothy Jaeryang Baek
5c62cc0517 chore: format 2026-08-25 16:53:53 -04:00
Timothy Jaeryang Baek
f71e9570c0 refac 2026-08-25 16:52:31 -04:00
Classic298
0366f5d3d8
chore: Update CHANGELOG.md (#27839)
* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* doc: changelog entries for the terminal preview same-origin setting and the automations bulk toggle
2026-08-25 16:43:59 -04:00
Timothy Jaeryang Baek
54d7a22370 refac 2026-08-25 16:34:49 -04:00
Timothy Jaeryang Baek
f4a0d3c973 refac 2026-08-25 16:33:03 -04:00
Timothy Jaeryang Baek
1d6d4e6e66 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-25 16:27:17 -04:00
Timothy Jaeryang Baek
067114c280 refac 2026-08-25 16:04:43 -04:00
Classic298
b1bfc18762
perf: cache the serialized builtin tool spec instead of deep-copying it per request (#28860)
Every chat request hands each builtin tool a fresh copy of its cached spec, because callers mutate what they get. That copy was a full deepcopy of a nested dict, repeated per tool per message.

The builder now caches the spec already serialized, so a request only parses it back. Parsing is what produces the independent tree callers mutate, and the cached value becomes an immutable string, so a request can no longer reach the cached object at all.

Measured on CPython 3.12 with a 1.1 KB spec and 20 builtin tools per request:

| | before | after |
|---|---|---|
| stdlib json, the default | 276.2 us | 66.4 us |
| orjson | 279.5 us | 37.5 us |

Builtin specs are plain JSON by construction: pydantic normalizes every default before it reaches the schema, so a tuple, set, enum or datetime cannot appear in one, and an unserializable default is dropped rather than embedded.
2026-08-25 16:00:34 -04:00
Classic298
45f4a87e85
fix: bound extracted document metadata by the upload size limit by default (#29025)
"RAG_METADATA_MAX_VALUE_CHARS" ships unset, and unset means no bound at all, so the limit only protects the deployments that already knew to configure it. A small Office document is a zip archive, and one crafted to expand enormously during extraction can turn a few hundred kilobytes into gigabytes of metadata held in memory; uploading it a handful of times is enough to exhaust a server and take Open WebUI down with it.

When no explicit limit is configured, the bound now follows "RAG_FILE_MAX_SIZE" instead of being absent, on the reasoning that a document cannot legitimately carry more metadata than the file itself is allowed to be. That keeps the number from being an arbitrary guess: it is whatever the administrator already decided an upload may weigh. Setting "RAG_METADATA_MAX_VALUE_CHARS" explicitly still wins, and a deployment that leaves both unset is unchanged, which is the same posture the upload limit itself takes.

"RAG_FILE_MAX_SIZE" is in MB and is treated as unset when it is zero, matching how the document loader already reads it.
2026-08-25 15:50:27 -04:00
Timothy Jaeryang Baek
6dcc2d5269 refac 2026-08-25 15:48:30 -04:00
Timothy Jaeryang Baek
140d2cf4b5 refac 2026-08-25 15:47:55 -04:00
Classic298
d198d950c6
perf: stop re-copying the response text on every stream save (#28821)
Every streamed delta saves a snapshot of the in-progress response so a reconnecting client can resume it, and each save rebuilt the assistant text from scratch. On the Chat Completions path that re-joined every accumulated chunk, including on saves carrying no new text, so a long answer followed by a large tool call re-joined the whole answer once per argument chunk. The Responses API path never collects those chunks and reads the text back out of the output items instead, where the blank check copied it in full every time.

The joined string is now kept and reused until another chunk arrives, since content_parts is only ever appended to; the nonlocal declaration that suggested otherwise was already dead and is dropped, and inlining the single-use helper removes an unreachable branch with it. The blank check in get_output_text now tests the text rather than allocating a stripped copy of it, which is equivalent for all twelve of its callers. Text streaming on the Chat Completions path is unchanged, since a text delta always appends before it saves.

| stream | before | after |
| --- | --- | --- |
| 20k-char answer, 2000 tool-argument chunks | 21.4 ms | 0.06 ms |
| Responses API, 40k deltas, 200k chars | 80.7 ms | 50.5 ms |

Without Redis nothing extra is retained, since the snapshot store already held that string; with Redis one copy of the response text stays alive while the stream runs.
2026-08-25 15:41:54 -04:00
Classic298
ac85b0f2a2
refac: gate code interpreter tag detection to legacy tool-calling mode (#29024)
Tag detection for the code interpreter ran regardless of the tool-calling mode, so a model in Native (Agentic) Mode that emitted <code_interpreter> blocks in ordinary reply text had that code sent to the executor. Native mode never teaches the tag format and exposes execute_code as a builtin tool, so the parser had nothing legitimate to pick up there.

Gates detection on the legacy mode, matching the condition that already decides whether the tag prompt is injected at all. The five authorization checks are unchanged, and native mode keeps executing through the tool.

Deployments on native mode whose models emit the tags unprompted will now see them rendered as text.
2026-08-25 15:33:55 -04:00
Timothy Jaeryang Baek
28f2965934 refac 2026-08-25 15:26:45 -04:00
Timothy Jaeryang Baek
e3a7a64d82 refac 2026-08-25 15:25:07 -04:00
Timothy Jaeryang Baek
278e97589e refac 2026-08-25 15:22:19 -04:00
Timothy Jaeryang Baek
02400c7a50 refac 2026-08-25 15:15:18 -04:00
Timothy Jaeryang Baek
f158f892f4 refac 2026-08-25 15:14:11 -04:00
Timothy Jaeryang Baek
bc4c91e6d6 refac 2026-08-25 15:11:06 -04:00
Timothy Jaeryang Baek
a610d77137 refac
Co-Authored-By: Fares <26122914+faqeel@users.noreply.github.com>
2026-08-25 15:05:02 -04:00
Timothy Jaeryang Baek
35fbde0a3f refac 2026-08-25 15:00:53 -04:00
Timothy Jaeryang Baek
684111715f refac 2026-08-25 14:56:23 -04:00
Timothy Jaeryang Baek
a21c8d15ee refac 2026-08-25 14:48:38 -04:00
Timothy Jaeryang Baek
20fe43d9da refac 2026-08-25 14:48:01 -04:00
Timothy Jaeryang Baek
3374b21a7d refac 2026-08-25 14:39:25 -04:00
Timothy Jaeryang Baek
8be4c5fa6a refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-25 14:35:22 -04:00
Timothy Jaeryang Baek
92a1502126 refac 2026-08-25 14:22:20 -04:00
Timothy Jaeryang Baek
5ce198b1d7 refac 2026-08-25 14:20:28 -04:00
Timothy Jaeryang Baek
21366fa1b1 refac 2026-08-25 14:19:58 -04:00
Timothy Jaeryang Baek
20f35d157b refac 2026-08-25 14:14:13 -04:00
Timothy Jaeryang Baek
c4b3e6840f refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-25 14:07:50 -04:00
Timothy Jaeryang Baek
7198e9d0df refac 2026-08-25 14:06:25 -04:00
Timothy Jaeryang Baek
2e6d61e6d0 refac 2026-08-25 13:51:58 -04:00
Timothy Jaeryang Baek
d241df8d79 refac 2026-08-25 13:39:54 -04:00
Timothy Jaeryang Baek
1c13fedb16 refac 2026-08-25 13:18:38 -04:00
Timothy Jaeryang Baek
a4738e0459 refac 2026-08-25 12:45:12 -04:00
Timothy Jaeryang Baek
3c1017f6c3 refac 2026-08-25 12:22:19 -04:00
Classic298
2d2bcb5332
fix: long streamed lines no longer abort the response (#28114)
Some providers send one very large piece of a streamed answer in a single go: a long reasoning trace, a code execution result, a turn with many tool calls, or a response echo carrying a big tool list. Anything past 128 KB in one line killed the chat mid-answer with a misleading `400, message: Got more than 131072 bytes when reading`. Nothing was rejected upstream, that is our own reader giving up on an oversized line.

Open WebUI already had code that assembles lines itself with no such limit, but it only ran when CHAT_STREAM_RESPONSE_CHUNK_MAX_BUFFER_SIZE was set. Unset is the default, and in that case the raw capped reader was used instead, so a default install always broke. That path now always assembles lines, and the setting goes back to being what its name says: an optional cap, off by default. It applies to the Ollama stream as well, since both now share the same reader.

The assembly loop only splits once a line actually completes, because the old one re-concatenated and re-split the whole buffer on every network chunk. Without that, allowing long lines would have traded an error for multi-second event loop stalls.

| | 20 MB in one line | 200k small lines |
| --- | --- | --- |
| before | 4249 ms | 27.3 ms |
| after | 37 ms | 25.2 ms |
2026-08-25 12:16:37 -04:00
Timothy Jaeryang Baek
beb3c114d3 chore: python-docx dep 2026-08-25 11:59:37 -04:00
Classic298
e3e4bd87df
refac: consolidate the web fetch address checks onto the request path (#27823)
* fix: apply the SSRF checks to redirect targets on every web fetch path

Two guards protect server-side fetches: a private-IP check and the operator's `WEB_FETCH_FILTER_LIST`. Neither reached a redirect hop on the aiohttp paths, and the filter list never reached one on the requests paths either.

aiohttp answers IP-literal hosts itself without consulting a resolver, so `_SSRFSafeResolver` was never invoked for a hop such as `http://169.254.169.254/` and the private-IP check simply did not run. With redirect following enabled, a submitted public URL that redirects to an IP literal reached loopback, RFC1918 and cloud-metadata addresses, and the response body was returned to the caller. The filter list was consulted only in `validate_url`, on the originally submitted URL, so a redirect to a filter-listed host was fetched without it ever being applied.

`_SSRFSafeResolver` is replaced by `_SSRFSafeConnector`, which hooks `_resolve_host` so the IP check also covers the IP-literal shortcut and both DNS cache paths. The filter list moves to a per-request hook on each transport, `connect()` for aiohttp and `send()` for the requests adapter, because those see the request destination: at the connection layer a proxied request presents the proxy's host, and a pooled connection skips resolution entirely. This covers every hop, including redirects, on all five aiohttp call sites and both requests sessions. The Playwright loader already validated each hop and is unchanged.

Both gaps required `AIOHTTP_CLIENT_ALLOW_REDIRECTS=true`, which is not the default.

Two behaviour changes for operators. The filter list now applies to redirect targets rather than only to submitted URLs. Under a forward proxy it is evaluated against the request destination instead of the proxy, which also fixes allowlist entries rejecting every fetch in proxied deployments.

* refac: match the web fetch filter list against resolved addresses

The filter list is now evaluated against the hostname together with the addresses it resolves to, at URL validation and on each connection, on both transports. An IPv6 address is also matched by the IPv4 address it carries.

* refac: screen outbound fetch addresses against reserved ranges ipaddress misses

`ipaddress.is_global` was the only test behind the web-fetch address check, and it answers a narrower question than "may we fetch this". Several special-purpose ranges are globally routable by registry while nothing on them is a legitimate destination, so they passed. Classification now screens those ranges on top of `is_global`, and applies the same screen to the IPv4 address embedded in an IPv6 transition encoding rather than only to the literal. All three checkpoints share the predicate, so they all inherit it.

The range list is the exact complement of what CPython's `ipaddress` already models, checked entry by entry against both IANA special-purpose registries. Prefixes IANA marks globally reachable are deliberately left out, so no real destination changes behaviour. Verified against 31 addresses covering every entry, their transition-encoded forms, and public controls in both families: 31/31 expected after, 18/31 before.

* refac: match web fetch filter entries that name an address or a range

A filter entry that parses as an address or a CIDR range is matched by containment rather than by DNS label suffix, so a range covers the addresses inside it and an address matches however it is spelled. A range entry previously matched nothing at all, silently.

The built-in list gains the special-purpose networks that ipaddress.is_global reports as reachable while nothing on them is a legitimate destination, so taking an address out of reach is a WEB_FETCH_FILTER_LIST change rather than a release. Those entries hold whether or not local web fetch is enabled; the private-address rule still follows the toggle.
2026-08-25 11:15:48 -04:00
Timothy Jaeryang Baek
ca4e07a40b refac 2026-08-25 11:07:54 -04:00
Timothy Jaeryang Baek
6330350a40 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-25 10:39:49 -04:00
Timothy Jaeryang Baek
176fa46212 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-25 10:03:53 -04:00
Timothy Jaeryang Baek
2f97c9fce3 refac 2026-08-25 10:00:43 -04:00
Timothy Jaeryang Baek
aa3d569610 refac 2026-08-25 09:59:07 -04:00
Timothy Jaeryang Baek
eadce55e34 refac 2026-08-25 09:55:18 -04:00
Timothy Jaeryang Baek
5078d987f8 refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
2026-08-24 21:15:54 -04:00
Timothy Jaeryang Baek
2a0274a0a0 refac 2026-08-24 20:34:59 -04:00
Timothy Jaeryang Baek
f3f7659da7 refac 2026-08-24 19:56:07 -04:00
Timothy Jaeryang Baek
170ad0595d refac 2026-08-24 19:48:49 -04:00
Timothy Jaeryang Baek
a9a802d95f refac 2026-08-24 19:48:12 -04:00
Timothy Jaeryang Baek
18bf0ade7b refac 2026-08-24 19:25:03 -04:00
Timothy Jaeryang Baek
11db926a7b refac 2026-08-24 19:20:24 -04:00
Timothy Jaeryang Baek
0169b10979 refac 2026-08-24 19:15:10 -04:00
Classic298
dc03e7e595
refac: keep external connections until their last knowledge base is removed (#28113)
Deleting an external knowledge base now clears its connection only when an admin removes the last knowledge base referencing it, matching the connection delete route.
2026-08-24 19:11:40 -04:00
Classic298
6b438f1a79
refac: stream web-fetched documents to disk instead of buffering them (#28945)
The binary branch of the web fetch read the entire response body into memory
before writing it out. It now streams in blocks, applies the configured file
size limit the same way the sibling URL endpoint already does, and removes the
temporary file when a download fails partway instead of leaving it behind.
2026-08-24 19:06:13 -04:00
Timothy Jaeryang Baek
06e7aac219 refac 2026-08-24 19:05:21 -04:00
G30
2978a03c68
fix: restore the model avatar reset option in the model editor (#29007) 2026-08-24 19:02:22 -04:00
Timothy Jaeryang Baek
d5b66533e7 refac 2026-08-24 18:56:39 -04:00
G30
e6648aefb5
fix: honor the terminal file browser filesystem root instead of clamping to home (#29006) 2026-08-24 17:55:51 -05:00
Sven Horvath
5735123f50
fix: keep the pending note save when an update carries no content snapshot (#28669)
stop_item_tasks() ran unconditionally while create_task() only ran when the
update carried data, so an update without a content snapshot cancelled the
pending save without scheduling a replacement and the edits were never
written.
2026-08-24 18:47:43 -04:00
Classic298
bf08835d6f
perf: allow disabling websocket per-message-deflate (#28613)
With delta streaming most websocket frames are tiny per-token deltas, and
per-message-deflate pays zlib work on every outgoing frame per subscriber for
near-zero gain there; under heavy streaming that shows up as measurable server
CPU. The frames that still benefit are the rare large ones (final message,
sources), and even a 100k token message is only a few hundred KB uncompressed,
which any network delivers without noticeable delay.

UVICORN_WS_PER_MESSAGE_DEFLATE=false (default true, current behavior) disables
the extension in every entry point: open-webui serve and dev, start.sh both
invocations, start_windows.bat and dev.sh. Verified against a running
instance: with the flag off the server declines the client-offered
permessage-deflate extension, with defaults it still negotiates it.
2026-08-24 18:46:07 -04:00
Timothy Jaeryang Baek
97466deea1 refac 2026-08-24 18:38:29 -04:00
Timothy Jaeryang Baek
9dff5e9327 refac 2026-08-24 18:37:39 -04:00
Timothy Jaeryang Baek
536b9edec0 refac 2026-08-24 18:35:04 -04:00
Timothy Jaeryang Baek
b96d2b12da refac 2026-08-24 18:29:36 -04:00
Timothy Jaeryang Baek
e3e82b1471 refac 2026-08-24 18:29:19 -04:00
Timothy Jaeryang Baek
e9efb95a9c refac 2026-08-24 18:17:11 -04:00
Timothy Jaeryang Baek
7a533d0d5b refac 2026-08-24 18:14:30 -04:00
Timothy Jaeryang Baek
6cb2449ab7 refac 2026-08-24 18:10:08 -04:00
Timothy Jaeryang Baek
cf4ac9c8db refac 2026-08-24 18:07:03 -04:00
Timothy Jaeryang Baek
fd8cc2ba4a refac 2026-08-24 18:06:59 -04:00
Timothy Jaeryang Baek
98ee2bdfd3 refac 2026-08-24 17:56:11 -04:00
Classic298
043cf330d2
perf: throttle last_active_at writes by default (#28177)
Presence tracking writes each user's last_active_at on every authenticated request, every API key request and every websocket heartbeat. The throttle for it already exists but ships unset, and unset means no throttle at all, so a stock deployment pays one UPDATE plus COMMIT per user per request. The 30 second frontend heartbeat alone is 2 write transactions per minute per open tab, before any actual UI traffic.

Defaulting the throttle to 60 seconds collapses that to at most one write per user per worker per minute. Presence is only ever read at minute granularity, so nothing visible changes.

60 rather than the 300 to 500 the docs currently suggest, because a user counts as active for 3 minutes after their last write and that window is hardcoded in the backend and again in the frontend. Any interval at or above 180 seconds makes people who are actively using the instance drop out of the active user count. Letting the window follow the interval instead would need the value shipped to the client, so that is a separate change.

0 still disables the throttle, and now costs nothing at all: the decorator returns the undecorated function instead of a wrapper that re-checks a constant on every call.

Closes #28165
2026-08-24 17:53:31 -04:00
G30
da9245626e
fix: keep the viewport in place when older messages load above it (#28657) 2026-08-24 17:51:20 -04:00
G30
683c92d064
fix: fetch sidebar folders once per refresh instead of three times (#28662) 2026-08-24 17:50:46 -04:00
Timothy Jaeryang Baek
a6834f089b refac 2026-08-24 17:47:10 -04:00
Timothy Jaeryang Baek
9e7c9360b7 refac 2026-08-24 17:43:43 -04:00
Timothy Jaeryang Baek
8c1f3d3824 refac 2026-08-24 17:39:30 -04:00
G30
f2313d0c72
fix: keep the workspace tab counts in sync with each section's list (#28983) 2026-08-24 17:36:12 -04:00
Timothy Jaeryang Baek
a914868e3c refac 2026-08-24 17:25:33 -04:00
G30
83d049a465
fix: skip the sidebar refresh when a chat is dropped back where it already is (#28664) 2026-08-24 17:20:44 -04:00
Sebastian
4d5084025f
fix: download nltk data somewhere a non-root UID can read (#28866)
nltk.download picks the first entry of nltk.data.path that already exists and is
writable. None do here, so punkt_tab lands in /root/nltk_data, and /root is mode
0700. The corpus is then unreachable whenever the container does not run as
root:

    nltk.data.find('tokenizers/punkt_tab')
    LookupError: Resource punkt_tab not found.

/usr/local/share/nltk_data is already on nltk.data.path, so nothing changes at
the read side.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 17:19:15 -04:00
Timothy Jaeryang Baek
ecad20b77f refac 2026-08-24 17:17:53 -04:00
Timothy Jaeryang Baek
91917b2395 refac 2026-08-24 17:16:01 -04:00
Timothy Jaeryang Baek
363ad352fe refac 2026-08-24 17:12:56 -04:00
Classic298
23b3a69bc2
fix: keep folder parent references acyclic (#28748)
Moving a folder under one of its own subfolders was accepted. A folder in a parent loop is never a root, so it and everything under it silently disappeared from the sidebar, and there was no way to get it back from the UI.

The move is now rejected with a 400, folders whose parent chain loops are put back at the root on the next folder list, and the folder tree traversals skip ids they have already visited so existing data in that state stays workable.
2026-08-24 17:06:19 -04:00
G30
18edfff2d6
fix: discard unsaved group settings and key the groups list (#28076)
* fix: stop unsaved group settings from persisting into the group list state

* fix: key the groups list so an open modal cannot rebind to another group
2026-08-24 16:31:39 -04:00
Timothy Jaeryang Baek
aeda6ff13a refac 2026-08-24 16:29:57 -04:00
G30
336d8841f4
fix: make embedded message content inert in the sidebar chat hover preview (#27770) 2026-08-24 07:36:44 -04:00
G30
fd7024f198
fix: log tool server connectivity failures without a traceback (#27757) 2026-08-24 07:35:57 -04:00
G30
6b4131d1d7
fix: log terminal proxy connectivity failures as single lines and handle client disconnects (#27755) 2026-08-24 07:35:38 -04:00
Classic298
baeb2dfb83
fix: apply RDS IAM token auth to the pgvector engine (#27754)
With `DATABASE_ENABLE_IAM_TOKEN_AUTH=true` and `VECTOR_DB=pgvector`, startup failed at vector store initialisation with `fe_sendauth: no password supplied`, so the two features could not be used together.

`PgvectorClient` builds its own engine and never got the `do_connect` listener that refreshes the RDS IAM token, and the `ScopedSession` branch that would have reused the instrumented main engine is unreachable because `PGVECTOR_DB_URL` defaults to `DATABASE_URL` and is therefore never falsy.

The pgvector engine now goes through `enable_iam_token_auth()` like the main and Alembic engines. Since a token authenticates exactly one host/port/user, that function now attaches the listener only to engines pointing at the same target, so a `PGVECTOR_DB_URL` aimed at a separate database keeps the password from its own URL instead of having it overwritten; the skip is logged with both identities.

Fixes #27752
2026-08-24 07:35:19 -04:00
G30
16b20c651d
fix: restore per-item retrieval settings in the model editor knowledge section (#27686) 2026-08-24 07:34:30 -04:00
G30
01452ff62f
fix: mount the code editor in its own container so duplicated message renders don't collide (#27740) 2026-08-24 07:34:07 -04:00
Classic298
9cf1a07960
fix: use the pooled client timeout for the Anthropic Messages passthrough (#27675)
* fix: use the pooled client timeout for the Anthropic Messages passthrough

The native `/api/v1/messages` passthrough still referenced `openai.AIOHTTP_CLIENT_TIMEOUT`, which stopped existing when `routers/openai.py` moved onto `session_pool.get_client_timeout()`. Every passthrough request therefore raised `AttributeError: module 'open_webui.routers.openai' has no attribute 'AIOHTTP_CLIENT_TIMEOUT'` before it was sent, and the surrounding handler turned that into a 502 "Open WebUI: Server Connection Error", so Anthropic-format clients such as Cline could not reach any model at all.

Use `get_client_timeout(stream=...)` like the OpenAI and Ollama proxies do, so the configured `AIOHTTP_CLIENT_TIMEOUT` applies and streaming requests additionally get the idle-read timeout.

Fixes #27595

* fix: authenticate native Anthropic requests with x-api-key

The Anthropic Messages passthrough and the token-count forwarding both build their upstream request through `get_anthropic_request_target`, which sends the connection key as `Authorization: Bearer <key>`. Anthropic's OpenAI-compatible `/chat/completions` endpoint accepts that, which is why the model works in the chat UI, but the native `/v1/messages` and `/v1/messages/count_tokens` endpoints do not: they require the key in `x-api-key` and reject a bearer token with 401 `Invalid bearer token` (and `jwt auth is not yet supported on count_tokens`). They also require an `anthropic-version` header, which was never sent.

For `api.anthropic.com` connections, send `anthropic-version` and move the key into `x-api-key`, dropping the bearer header. Connections using session, OAuth or Entra ID auth keep their token untouched, LiteLLM passthrough connections are unaffected, and admin-configured custom headers still win over both defaults.

Fixes #27695
2026-08-24 07:33:33 -04:00
Timothy Jaeryang Baek
8a170897ba refac 2026-08-24 07:28:09 -04:00
Timothy Jaeryang Baek
495296346e refac 2026-08-24 06:26:49 -04:00
Timothy Jaeryang Baek
ef455fcef9 refac 2026-08-24 06:25:09 -04:00
Classic298
091c44c621
perf: stop rescanning the whole response for tag boundaries on every streamed chunk (#28861)
Streamed responses are scanned for reasoning and code interpreter tags. To work out where the last complete tag ended, the scanner searched backwards from the start of the accumulated text on every chunk, once per tag set. Ordinary prose contains no angle bracket, so that search never stopped early and read the entire response back every time. The cost grows with the square of the response length, and this scanning is on unless a model turns it off.

The two positions are now carried forward as the text grows, so each chunk only scans the characters it added.

Measured on CPython 3.12, a 270 KB response streamed in 27000 chunks:

| response text | before | after |
|---|---|---|
| no newlines | 7690 ms | 40.6 ms |
| with newlines | 5695 ms | 41.7 ms |

The carried positions match a full rescan at every step of 36282 randomized replays, covering text with no markers, newlines only, dense markers, real tags and truncation part way through.
2026-08-24 05:11:32 -05:00
Timothy Jaeryang Baek
978d257214 refac 2026-08-24 05:40:53 -04:00
Classic298
16c2a9eda4
fix: index the chat queries that make large SQLite instances unusable (#27663)
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
The timer scheduler polls once a second and cancels on every message send and chat open, the sidebar lists chats ordered by `updated_at`, and the folder badges count unread chats per folder. None of those could be served by an index, so each call read most of the `chat` table, and because `meta` sits after the chat payload column SQLite had to walk every row's overflow pages to get there. On a large history that stalls the sidebar, every chat switch and every send, and the idle poll alone burns about a quarter of a CPU core.

Timers now keep their due time in a dedicated `chat.timer_at` column behind a partial index, and the chat list, unread and unfinished-reply queries each get an index matching their filter and ordering. Existing pending timers are backfilled from their meta by the migration. Dropping the `internal` and `type` checks also makes a forked timer chat inert, where a fork used to copy `meta` verbatim and become a second claim target that could fire a duplicate timer.

Measured on SQLite, same rows returned:

| query | before | after |
|---|---|---|
| idle timer poll (2000 chats, 0.43 GB) | 170 ms | 0.04 ms |
| cancel on send and chat open (4000 chats, 377 MB) | 200 ms | 0.04 ms |
| sidebar chat list (15000 chats, 1.26 GB) | 157 ms | 1.8 ms |
| folder unread badges (15000 chats, 1.4 GB) | 54 ms | 0.2 ms |

PostgreSQL 17 serves all of them as index-only scans with no sort node. Exercised through fresh install, upgrade with seeded data, downgrade and re-upgrade on SQLite and PostgreSQL 17.

Fixes #27622
2026-08-23 16:11:54 -05:00
enixCode
29e8d7db67
i18n: fill missing fr-FR translations (#28951)
75 entries in the French locale had an empty value. i18next falls back to the
key when a value is empty, so French users were shown raw English strings
across permission settings, empty states, form placeholders and error toasts.

- fill those 75 entries; no key is added, removed or reordered
- follow the conventions already present in the file: "Entrez ..." for input
  hints, "Échec de ..." for failures, Chat -> Conversation, Token left as is
- interpolation placeholders preserved; no existing translation modified
2026-08-23 16:02:27 -04:00
Classic298
f73f09a3e0
refac: drop the redundant .keys() from two dict membership tests (#28859)
`x in d` and `x in d.keys()` are identical for a plain dict, so the `.keys()` call builds a throwaway view and reads as if it were doing something. Both sites operate on a plain dict: `combined` in `merge_and_sort_query_results` is a local `dict()`, and `ui_settings` comes from `UserSettings.model_dump()` where `ui` is annotated `dict | None` and is already guarded against None on the preceding line.

No behaviour change, and no measurable speedup either, so this is a readability cleanup rather than a performance one.

Sites where `.keys()` is load-bearing are left alone: the `list(d.keys())` snapshots taken before mutating during iteration, and the places where `.keys()` is the iteration or comprehension source rather than a membership test.
2026-08-23 16:02:09 -04:00
Timothy Jaeryang Baek
b30b11d4c9 refac 2026-08-23 16:01:36 -04:00
Classic298
ac091273b7
fix: keep streamed text when a filter or provider sends non-string content (#28840)
A stream filter function, or a provider that puts something other than a string in a delta, makes the streaming handler concatenate a string with a non-string. That raises TypeError, and the broad handler wrapped around the whole per-chunk block swallows it at debug level and moves on. The chunk's text never reaches the message the user sees, and nothing above debug level says why.

The content and reasoning fields are now coerced to text once, where they are read off the delta, ahead of every consumer. The coercion is guarded on truthiness, so falsy values such as an empty list still skip the block exactly as before, and the accumulated content receives byte for byte what it received previously.

Checked against 14 delta shapes covering strings, empty values, numbers, booleans, None, lists, dicts and a content array: the truthiness gate and the accumulated content are identical before and after.
2026-08-23 15:35:53 -04:00
G30
603e85c569
fix: apply connection prefix id to the model display name as well as the id (#28950) 2026-08-23 15:27:09 -04:00
Timothy Jaeryang Baek
f64c0c87e8 refac 2026-08-23 15:14:43 -04:00
Timothy Jaeryang Baek
e623c02acc refac 2026-08-23 15:06:30 -04:00
Timothy Jaeryang Baek
78f48a21ee refac 2026-08-23 14:40:48 -04:00
Timothy Jaeryang Baek
fb4f476316 refac 2026-08-23 13:49:50 -04:00
Timothy Jaeryang Baek
2578174637 refac 2026-08-23 13:47:02 -04:00
Solaris-star
5ee7140b4c
fix: render code editor drawer above settings modal (#27648)
The ComfyUI workflow.json Edit drawer (CodeEditorModal -> Drawer) used
z-999, while the admin Settings dialog (Modal) uses z-9999. Since both
are appended to <body>, the drawer rendered behind the settings dialog
and appeared to open 'in the background' (#27647).

Add an optional zIndexClass prop to Drawer (default z-999, preserving
existing behaviour for all other callers) and pass z-99999 from
CodeEditorModal so the editor surfaces above any enclosing modal.
2026-08-23 13:44:41 -04:00
Timothy Jaeryang Baek
bf3a58dbcd refac 2026-08-23 13:40:13 -04:00
Timothy Jaeryang Baek
5093a99389 refac 2026-08-23 13:36:53 -04:00
Timothy Jaeryang Baek
886248de36 refac 2026-08-23 13:33:51 -04:00
Classic298
d16d62d1f1
fix: duplicate checkbox markers when serializing note task lists (#27671)
Task lists in Notes serialized to markdown as `- [ ] [ ]` with the item text pushed onto a separate line after a blank line, so previewing or downloading a note produced a broken checklist, and checking an item left the second `[ ]` behind as plain text.

TipTap renders each task item as a checkbox inside a label plus a block-wrapped body. The GFM turndown plugin matches that checkbox and emits its own `[ ]`, which landed next to the marker the task item rule already writes, and the block wrapper left blank lines around the text that the old leading-whitespace strip could not remove.

Register a rule that drops the checkbox so the task item rule is the only source of the marker, and trim the block wrapper while indenting continuation lines so nested lists and code fences stay inside the item.

Fixes #26067
2026-08-23 13:31:46 -04:00
Classic298
945c521ed2
refac: make the web search error message a plain constant (#28948)
The web search error message was a lambda with a passthrough branch that returned whatever it was handed. Since #28942 both call sites pass no arguments, so that branch is unreachable, and it is the trap that let a caller drop a raw exception object into an HTTP response body and turn an intended 400 into an unserialisable 500.

A plain string constant removes the trap and lines the message up with every other fixed message in that file. Behaviour is unchanged: the response detail comes out byte for byte identical, because the enum already overrides __str__ to render members as their value. Verified on Python 3.11 and 3.12, both producing the same string and the same JSON body.
2026-08-23 13:27:02 -04:00
Timothy Jaeryang Baek
842c1f9d67 refac 2026-08-23 13:10:07 -04:00
Timothy Jaeryang Baek
0fb542b376 refac 2026-08-23 12:59:59 -04:00
Classic298
fca3be5416
fix: web search failures return HTTP 500 with an empty body instead of 400 (#28942)
Any failure during a web search comes back to the client as a bare HTTP 500 with nothing in it. The handler tries to build a 400 whose detail is the caught exception object itself, FastAPI cannot serialise that into a response body, so rendering the error response fails and the request falls through to the generic 500 handler. In chat this surfaces as a web search that fails with no explanation at all, and the most common trigger is simply selecting a search engine without configuring its API key.

This routes the failure through the standard error formatter, which is what the sibling handler for content loading failures in the same function already does. Web search failures now return 400 with a readable message, and the exception itself keeps going to the server log exactly as before.

Passing str(e) into the response was the other option and was rejected: the rest of the backend deliberately keeps provider exception text out of client responses and in the log, and provider exceptions here can carry request details that should not be echoed back.
2026-08-23 12:59:25 -04:00
Classic298
069f49fcd2
refac: remove unreachable rate limit handler from DuckDuckGo web search (#28943)
The DuckDuckGo search path catches RatelimitException from the ddgs library. That exception is defined by the library but never raised anywhere in it, checked against the pinned 9.14.4 and against 9.11.3, so the handler could never run. The two fallbacks around it were dead for the same reason: ddgs.text() returns a non-empty list or raises, so None and an empty list are not outcomes it can produce.

Removing all three leaves one call and changes nothing observable. A refused or rate limited search already came out as a failed search, with the error shown to the user and the traceback in the log, and it still does.

The backend argument is now passed as backend or 'auto' rather than conditionally omitted, because 'auto' is the library's own default for that parameter, so every configured value including unset and empty resolves exactly as before. Verified by running the old and the new function side by side against a stubbed library covering normal results, the domain filter, all four backend settings and a failing search, with identical results in every case.
2026-08-23 12:59:05 -04:00
Timothy Jaeryang Baek
3c66d639e3 refac 2026-08-23 12:33:34 -04:00
Timothy Jaeryang Baek
c1c81f8127 refac 2026-08-23 03:40:40 -04:00
Timothy Jaeryang Baek
f3f76095d1 refac 2026-08-23 02:34:08 -04:00
Timothy Jaeryang Baek
4807866a1c refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-23 01:31:30 -04:00
Timothy Jaeryang Baek
7abe11346a refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
2026-08-22 09:13:13 -04:00
Timothy Jaeryang Baek
d17f06a235 refac 2026-08-22 08:43:46 -04:00
Timothy Jaeryang Baek
1b3b9375bb refac 2026-08-22 08:21:53 -04:00
Timothy Jaeryang Baek
d7d935275a refac 2026-08-22 08:18:38 -04:00
G30
302ffc8b7e
fix: skip the autocompletion debounce when the document shrank past the captured position (#28824) 2026-08-22 08:12:40 -04:00
G30
d1c207d091
fix: give docx and pptx previews in the file modal a definite frame height (#28878) 2026-08-22 08:11:29 -04:00
Timothy Jaeryang Baek
ccbb3303f2 refac 2026-08-22 08:05:52 -04:00
G30
883c7434fb
fix: tolerate reasoning items without started_at when closing them at stream end (#28872)
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
2026-08-21 15:51:21 -07:00
Timothy Jaeryang Baek
dcff244f9e refac 2026-08-21 15:50:50 -07:00
Aleix Dorca
6a999f357b
i18n: Update catalan translation.json (#28826) 2026-08-21 12:08:53 -07:00
G30
d2bc98eaeb
fix: let the composer's model selector shrink so narrow containers keep every control visible (#28912) 2026-08-21 12:04:59 -07:00
Classic298
7229fac0c4
refactor: remove the orphaned admin settings components (#28855)
Some checks are pending
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Two components under the admin settings area are not imported by any route or component. The admin settings shell went dead when the admin settings route became a redirect into the settings modal, which imports every admin tab directly and carries its own tab list and search. The model selector beside it lost its last importer in a separate models refactor. Every child component the shell used is still imported by the settings modal, so nothing goes with them.

This removes around 590 lines that still turn up in every search across the admin area.
2026-08-20 13:37:53 -07:00
Classic298
c7f306031d
refactor: remove the unreachable async half of the Mistral loader and an unused Datalab helper (#28839)
The Mistral OCR loader has a full async pipeline beside its synchronous one: an async load, its own upload, signed URL, OCR, delete and retry helpers, a pooled session and a batch loader on top. The only way in was the batch loader, which nothing calls, so the entire async half was unreachable. Everything that loads documents goes through the synchronous path, and the shared loader entry point runs it in a worker thread. The Datalab loader carries a public request status poller with no caller either, since its own load inlines the polling it needs.

With the async half gone, the retry classifier's two aiohttp branches can no longer be reached, since the only retried calls are synchronous, so those go with it along with the aiohttp import that existed solely to feed them, and a timeout attribute that nothing reads any more. The class docstring loses the three bullets that only described the removed pipeline, and four docstrings stop calling themselves the sync version of something that no longer has an async counterpart.

This removes around 350 lines and leaves one code path per loader instead of one live path and one that cannot be entered.
2026-08-20 13:14:14 -07:00
Timothy Jaeryang Baek
7d4747dfd7 refac 2026-08-20 13:13:51 -07:00
Classic298
18c604baa9
fix: voice mode produces no audio when the task model returns empty content (#28724)
With "show emoji in call" enabled, voice mode stayed completely silent and no request ever reached the configured TTS server. Reasoning models served with a reasoning parser return `message.content` as null and put the text in `reasoning_content`, and the emoji helper called `.replace()` on that null value and threw.

The call overlay ran the emoji request first, inside the same `try` block as speech synthesis, so that error skipped the entire TTS section. The audio cache was never filled, and the playback loop kept re-queueing the same content every 200 ms without ever playing it. Read aloud was unaffected because it synthesizes speech directly, which is why the failure looked specific to voice mode.

Fixed on both sides: the optional chain in `generateEmoji` now covers `content`, and the emoji request in the call overlay gets its own catch, matching the speech synthesis call directly below it. An emoji failure now costs the emoji instead of the whole reply.
2026-08-20 13:03:55 -07:00
Classic298
5586964bb2
fix: keep the usage pool cleanup task alive across lock loss and Redis errors (#28834)
With WEBSOCKET_MANAGER=redis on a multi-node deployment, the usage pool cleanup task could stop permanently for the whole cluster. Nodes that lost the startup lock race gave up for good after three attempts, and the winner died on a single failed renew or on any Redis connection error, releasing the lock with nobody left to take it over. From then on expired entries accumulated in the usage pool until a node restarted, so /api/usage over-reported models in use and every disconnect handler walked an ever-growing pool.

The task now retries lock acquisition forever like the session pool cleanup does, and any error is logged and answered by releasing the lock and returning to acquisition, so a transient failure costs one cleanup cycle and every node stays a takeover candidate. The delete of an emptied model entry is KeyError-guarded because a disconnect handler on another node can remove the same key between the sweep's snapshot and its delete; unguarded, that race was a permanent task killer that needed nothing rarer than a chat finishing while its tab closed.
2026-08-20 12:59:38 -07:00
Classic298
b0fdc00452
perf: write task payloads to Redis as bytes (#28833)
Saving a streaming response serialized the payload with orjson, decoded it to
str, scanned it for the three Unicode line separators and let redis-py encode
it straight back to UTF-8: on an 8 MB non-ASCII chat that is 6.9 ms and ~22 MB
of transient buffers per write, synchronously on the event loop.

json_codec now exposes dumps_bytes, which returns the serialized payload as
UTF-8 bytes without the line-separator escaping, and the two Redis writes in
tasks.py use it. That escaping only protects line-framed protocols such as
SSE; every reader of these Redis values re-parses them before anything is
served, and the escaped and raw forms parse identically, so mixed versions
during a rolling deploy interoperate both ways. The same write drops to
0.9 ms and one 8 MB buffer (7.5x), with 31-66% saved on KB-sized writes.
With ENABLE_ORJSON off, dumps_bytes wraps stdlib json, behaviour unchanged.

The str path keeps the escaping but applies it with chained str.replace
instead of a translate table, cutting a separator-containing 8 MB payload
from 312 ms to 5.7 ms with byte-identical output.
2026-08-20 12:58:52 -07:00
G30
a0e7d0e3a4
fix: send the model id when ejecting from the model selector (#28766) 2026-08-20 12:58:14 -07:00
G30
528259695c
fix: key selected item lists in model editor selectors to stop checkbox state reuse (#28837) 2026-08-20 12:57:39 -07:00
Timothy Jaeryang Baek
f822605b35 refac 2026-08-20 12:56:45 -07:00
Timothy Jaeryang Baek
8a42aa53e8 refac 2026-08-19 22:48:32 -07:00
Timothy Jaeryang Baek
46dce79eb0 refac 2026-08-19 22:28:41 -07:00
Classic298
4df2d9a7aa
perf: filter workspace models by access in SQL instead of loading every model (#28795)
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
Exporting workspace models loaded every model row, built a full response object with its owner for each, and only then dropped the ones the caller may not see. On a large model table that made the export endpoint slow in proportion to models the user cannot even access.

The owner-or-grant check now happens in the query itself, reusing the permission filter this file already applies to the paginated list endpoint, so only visible rows are ever hydrated. The by-user wrapper had one caller left and is gone with it.

Measured with 500 workspace models of which 3 are visible to the caller: 5 queries and ~12.7 ms before, 4 queries and ~2.8 ms after. The resulting set is unchanged for owner, public, direct-user, group and multi-grant entries, and base model entries stay excluded as before.
2026-08-19 18:15:01 -07:00
Classic298
ce22e0bb15
perf: stop rescanning the whole chat JSON on every streamed event write (#28820)
Every streamed event that persists to a chat (status updates, citations,
file attachments, message content) serialized the entire conversation JSON
three times: a null-byte check of the stored row, a second sanitize of the
whole blob after merging in the event payload and the flush of the UPDATE
itself. The middle pass rescans megabytes of already-clean history for null
bytes that can only come from the small incoming payload, so long chats pay
for their full history on every single event.

The write paths now sanitize just the incoming message, message id and
status dict and keep the row-level sanitize, so legacy rows with null bytes
still self-heal as before. Median per-event write time (sqlite, orjson):
1 MB chat 15.0 ms to 11.1 ms, 4 MB 65.2 ms to 52.4 ms, 10 MB 159.9 ms to
124.5 ms, roughly 20 percent less per event. As a side effect the
chat_message dual write now receives the sanitized message; previously null
bytes in non-content fields were cleaned in the blob but written raw to
chat_message, which failed that insert on PostgreSQL. Verified byte-identical
rows against the previous implementation across nine scenarios covering null
bytes in every input, legacy dirty rows, a missing title and a NULL chat
column.
2026-08-19 17:52:05 -07:00
Timothy Jaeryang Baek
33dff414e8 refac 2026-08-19 16:35:49 -07:00
G30
dd7158b45f
fix: clear the integrations search when leaving a section (#28812) 2026-08-19 16:14:53 -07:00
Timothy Jaeryang Baek
ea55d38793 refac 2026-08-19 16:14:27 -07:00
Classic298
7cf6051a74
perf: resolve group membership once per folder listing instead of once per entry (#28810)
Listing a user's folders re-checks which entries they may still see, and it resolved their group membership again for every folder, then again inside the collection and note branches for every entry. A comment in that helper claims one membership fetch for the whole listing, but the caller invokes it once per folder, so the claim never held.

The listing now resolves membership once, and only when some folder actually carries entries, then threads it through the file, collection and note checks. Callers that do not supply it are unchanged and still resolve for themselves.

Measured with twenty folders holding six files, two knowledge bases and two notes each: 245 queries and ~145 ms before, 186 and ~117 ms after. The folders returned, and the entries the integrity pass writes back, are unchanged. That was checked against entries the caller owns, entries shared through a group, entries shared with nobody, another user's files, and an unrecognised entry type.
2026-08-19 11:16:29 -07:00
Classic298
767a1157f1
refac: bind tool server cookies per connection (#28707)
The tool callable now takes its connection's cookie jar as a parameter, matching how its headers are already passed and how the terminal tool factory in the same module builds its callables.
2026-08-19 11:16:16 -07:00
Classic298
abc69000b3
fix: treat a non-numeric calendar alert_minutes as unset (#28790)
A calendar event's `meta` is a free-form dict, so `meta.alert_minutes` can hold any JSON type, while the upcoming-events lookup assumed it was a number and compared it directly. It now ignores a value that is not numeric and falls back to the default alert window for that event.

Handled on the read side rather than on the write path so events already stored with a non-numeric value are covered too. Numeric values are untouched, including the negative "no alert" sentinel.
2026-08-19 11:16:00 -07:00
Classic298
2e7df54673
fix: surface attached chat references in <attached_files> (#28788)
A chat attached via the "+" menu or dropped from the sidebar references an
existing chat by id and carries no url. add_file_context() filtered on
`file.get('url')`, so the reference was dropped from <attached_files>
entirely and the model was never told it existed.

When the RAG file-context path is enabled the chat content still reaches
the model as <source> context, which masked this. With file_context
disabled that path is skipped, and get_attached_knowledge() only promotes
collection/note items into <attached_knowledge> - so an attached chat was
visible in the UI but invisible to the model, which then reported having
no chat attachments despite having a view_chat tool available.

Keep chat references and emit their id so the model can resolve them with
view_chat. The url attribute is now conditional, since a chat has none;
the id guard it replaces was dead once the filter guarantees a url or a
chat id.

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-19 11:15:22 -07:00
Classic298
ebd4d9c6cc
fix: honor bypass_system_prompt on the pipe route (#28739)
The tool-call continuation re-submits with bypass_system_prompt=True, but only
routers/openai.py and routers/ollama.py checked it, so pipe and manifold models
had the system prompt applied again on every continuation. Since
add_or_update_system_message() prepends rather than replaces, N tool-call rounds
left N+1 copies of the system prompt in the payload.
2026-08-19 11:08:02 -07:00
Classic298
4f98a5184f
perf: resolve model-attached file access with a targeted query (#28802)
Checking whether a user may reach a file loaded and validated every workspace model that user can access, then scanned each model's knowledge list in Python for one file id. Folder listings run that check once per file, so opening a folder of twenty files rebuilt the whole accessible-model set twenty times, and the same check sits on every retrieval and download path.

The lookup now runs the other way round: the database returns the models that attach the file, and only those are access-checked. The text match on the metadata column is a prefilter and the knowledge entries still decide, so a file id that merely appears in a description grants nothing; file ids are server-generated uuids, so the match can only be too wide, never too narrow.

Measured with 500 accessible workspace models: a single check drops from 9 queries and ~20 ms to 6 and ~2.6 ms, and a twenty-file folder listing from 180 queries and ~680 ms to 120 and ~56 ms. A 72-case matrix over owner, public, direct-user and group grants, for both read and write, returns exactly what it returned before, and write still requires the model owner to own the file. The check also no longer writes to the database while answering a read-only question.
2026-08-19 11:07:43 -07:00
Classic298
dbf715cb63
perf: stop scanning every skill on each listing and chat turn (#28798)
Listing skills ran one database query per skill in the instance. A non-admin opening the list on a workspace with 500 skills issued over 500 queries, the paginated list re-resolved the caller's group membership once per row, and every chat message carrying a skill loaded every skill the user can read, full body and owner included, to use the two or three it actually referenced.

Skills now arrive already filtered: the owner-or-grant check runs in the query as an EXISTS subquery, the same way prompts and the search endpoints already do it, the per-item write flag uses the existing batch grant lookup, and the chat path asks only for the skill ids the request names.

Measured with 500 skills of which 3 are visible to the caller: 504 queries and ~300 ms before, 4 queries and ~2.6 ms after. The resulting set is unchanged for owner, public, direct-user, group and multi-grant entries, for both read and write.
2026-08-19 11:07:33 -07:00
Classic298
9bdb072690
refac: remove unused knowledge base accessors (#28794)
Three methods on KnowledgeTable have no callers anywhere in the repository. get_knowledge_bases_by_user_id loaded every knowledge base and filtered them in Python, which search_knowledge_bases already does in SQL with pagination. get_knowledge_by_id_and_user_id duplicates check_access_by_user_id with the permission hardcoded to write. update_knowledge_data_by_id writes a data column that a migration dropped, so it could only ever raise and return None through its own except block.

What remains is one per-entry access helper and one SQL-filtered list path, so nobody reaches for the slower or the broken variant by accident.

No behaviour change.
2026-08-19 11:07:21 -07:00
Classic298
284da2ae49
perf: reuse the already loaded chat when assembling builtin tools (#28809)
Assembling the builtin tools for a chat message fetched the chat row a second time to answer one question: whether this is a note chat. The caller had loaded that same row a few lines earlier, from the same id in the same metadata dict, and had already evaluated the same predicate for its own note handling. So every message with builtin tools enabled read the whole conversation blob twice.

The caller now works the flag out once and passes it down. Tool assembly no longer touches a chat model at all, so the two files cannot drift apart when the shape of that metadata changes.

Measured with a stub request across five chat shapes, a note chat, a plain chat, an internal chat that is not a note, a chat id with no row behind it, and an unsaved chat id: the returned tool set is identical in every case and the query count drops from six to five. The note tools are still enabled for a note chat with the notes feature switched off, which is the only thing that predicate decides.
2026-08-19 11:06:49 -07:00
Classic298
6db64c4855
perf: batch the shared folder listing instead of fetching one folder at a time (#28804)
Opening the shared folder list fetched every shared folder in its own query, fetched a chunk of them a second time to walk their children, and looked up each distinct owner separately. With forty folders shared with a user that is over a hundred queries before any subtree work starts.

The folders and their owners now come back in one query each, and the inheritance pass reuses the rows already in hand. Both folder listings also gained an explicit order: the sidebar merges shared subfolders in response order without sorting them, and neither query had an ORDER BY, so on Postgres a folder rename could reshuffle its siblings.

Measured with forty shared folders and no subtrees: 181 queries and ~105 ms before, 92 and ~66 ms after. With subtrees attached, 203 folders in total, it is 341 queries before against 252 after; the remainder is the recursive child walk, which this change deliberately leaves alone. The returned set, permissions and owner names are unchanged, including for a grant pointing at a deleted folder row, a folder the caller owns that is also shared with them, a folder whose owner record is gone, and a child folder that is itself directly shared.
2026-08-19 11:06:37 -07:00
Classic298
81fe43f210
perf: write a chat's messages in one transaction instead of one per message (#28806)
Saving a chat rewrote its message rows one at a time. Each message took its own session out of the pool and committed on its own, and the save endpoint hands over the entire merged history rather than only what changed, so a two hundred message chat cost two hundred sessions and two hundred commits on every save.

The messages now go through a single select and a single commit. The field mapping for the insert and the update branch moved into two small helpers, so the batch and the single-message path cannot drift apart.

Measured on a two hundred message chat with one message edited: 201 queries and 200 transactions before, 2 queries and 1 transaction after, ~149 ms against ~6 ms. Re-saving an unchanged history now costs one select and no writes at all.

One behaviour change worth stating: a message the database cannot store used to be skipped on its own, and now costs the rest of that same save. This table is a rebuildable fast path, so the reader falls back to the history on the chat row and re-triggers the backfill, and the next save reconciles everything still present. A per-message retry was tried and dropped, because a commit that lands but still raises would re-apply the usage merge and double the recorded token counts.
2026-08-19 12:46:49 -05:00
G30
bcb50fe7b0
fix: derive integrations toggle state from the selected ids (#28807) 2026-08-19 12:46:17 -05:00
Timothy Jaeryang Baek
5ea9ff3ed9 refac 2026-08-18 22:09:42 -07:00
Timothy Jaeryang Baek
3fdfbd5138 refac
Some checks are pending
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
2026-08-18 19:37:20 -07:00
Timothy Jaeryang Baek
b838860dc0 refac 2026-08-18 19:30:42 -07:00
Classic298
21e390561d
fix: revoke existing sessions when a password changes (#28725)
Some checks failed
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
Changing a password left every other logged-in device working until the JWT expired on its own, up to four weeks with the default settings. The hardening docs already promise the opposite: with Redis configured a password change is supposed to put the user's tokens on the revocation list, but only sign-out and OIDC back-channel logout ever wrote to it.

Both password-change paths, self-service and an admin resetting someone's password, now stamp the per-user revocation marker that token validation already checks, so every session issued before the change stops working. The acting device is signed out as well and asked to sign in again, which is the safer default when the password is being changed precisely because the old one may be compromised. Without Redis nothing can be revoked, as before, and the backend now logs a warning saying so.

The marker is written through one shared helper, so its lifetime follows the configured JWT lifetime instead of a fixed 30 days and never expires at all when JWT_EXPIRES_IN disables expiry. Back-channel logout picks that up too, where a long or disabled JWT lifetime previously let the marker expire while the tokens it revoked were still valid. API keys keep working, they are separate credentials with their own lifecycle.

Discussed in #28647.
2026-08-17 13:56:29 -07:00
Classic298
3fc491d22f
chore: drop test-only dependencies from the Docker image and the published package (#28726)
The Python test suite was deleted in 4527c747b but its dependencies stayed behind, so pytest, pytest-docker and the docker SDK still install into every image variant, and moto joins them for anyone running pip install open-webui[all]. No Python test file remains in the repository, nothing imports these packages, and no CI job runs pytest. They are removed from backend/requirements.txt and from the all extra, which are the only two channels they ship through.

netcat-openbsd goes for the same reason. It was added in January 2024 without a consumer and nc has never been invoked anywhere in the repository, in any script, workflow or compose file. Both the readiness wait and the healthcheck use curl, and the Ollama install script does not ask for it either.

uv.lock is regenerated output, not hand-edited. It drops three of the four packages plus three transitives that nothing else needs, with no version changes and no additions. pytest stays locked because pytest-asyncio in the dev group still requires it. The dependency markers it adds on the CUDA and numpy entries are inert: each one is a superset of the condition its parent already installs under, and the resolved default install set is identical before and after.

This saves roughly 2 MB uncompressed, which is nothing next to the image as a whole. The point is that a production image stops shipping a test framework and a Docker socket client it never uses.

Everything else stays and is load-bearing. The container installs pip packages at runtime for user-authored tools and functions, so it needs git and a working compiler for anything that is not a prebuilt wheel, and libmariadb-dev for the manual MariaDB install. zstd is required for updating Ollama inside the bundled image. black looks dev-only but backs the code formatting endpoint.

Ref: https://github.com/open-webui/open-webui/discussions/28716
2026-08-17 13:56:08 -07:00
Timothy Jaeryang Baek
ffda8aea80 refac 2026-08-17 01:34:37 -07:00
Timothy Jaeryang Baek
f67875ec57 refac 2026-08-17 01:30:21 -07:00
G30
88c55b86b1
feat: emit auth.login on SSO logins and attribute SSO logouts (#27619)
* feat: emit the auth.login event on SSO logins

* feat: attribute SSO logouts in the auth.logout event payload
2026-08-17 02:22:25 -06:00
Timothy Jaeryang Baek
a3a81fee03 refac 2026-08-17 01:21:58 -07:00
Classic298
646a568ae6
fix: enforce global web search and image generation switches on the legacy function-calling path (#27669)
The legacy function-calling path acted on the client-supplied `features` dict after checking only the per-user permission, so a user who still held `features.web_search` or `features.image_generation` could keep triggering web searches and image generation after an administrator had switched those off instance-wide. The native function-calling path already gates the equivalent builtin tools on `web.search.enable` and `image_generation.enable` in `get_builtin_tools`, so the two paths disagreed and the admin-level switch did not actually stop the outbound provider calls it was turned off to stop.

Gate the legacy web search handler on `web.search.enable` at its call site, and gate `chat_image_generation_handler` on the two image switches internally. The image handler needs the check inside it because `image_generation.enable` and `images.edit.enable` are independent: editing stays available when generation is disabled, matching the `/images/generations` and `/images/edit` routes and the native `generate_image`/`edit_image` tools. The handler calls `image_generations`/`image_edits` directly and so bypasses the route guards, which is why the check has to live at the caller.

The "Creating image" status event moves below the new guard so a disabled configuration returns without leaving an unresolved progress indicator in the chat.
2026-08-17 02:15:05 -06:00
G30
7ea46a37d0
fix: clip settings modal contents to its rounded corners (#27617) 2026-08-17 02:12:26 -06:00
Timothy Jaeryang Baek
b1dc945bd6 refac 2026-08-17 00:59:13 -07:00
Timothy Jaeryang Baek
0b27fa5e87 refac 2026-08-17 00:57:57 -07:00
Classic298
017075a2d7
perf: drop unused database session dependencies from seven endpoints (#28178)
Seven route handlers declare a request-scoped database session as a FastAPI dependency and then never touch it. Three of them are `GET /api/v1/users/user/settings`, `/user/status` and `/user/info`, which the frontend hits on every page load, and all three carry a comment saying the user object is already available, so the parameter is leftover from the refactor that removed the refetch. The other four are admin-only external-knowledge connection endpoints that read their data from the config store.

Measured on a route with and without the dependency, 20k requests, best of 5:

| | µs per request |
| --- | --- |
| no dependency | 16.18 |
| unused session dependency | 62.85 |

The dependency costs about three times as much as everything else the request does put together. It is worth being precise about why, because the obvious guess is wrong: this is not database I/O and not connection pool pressure. SQLAlchemy connects lazily, so a session that is never used checks out zero connections, verified by watching the pool's counter stay at zero across the request. The cost is FastAPI resolving an extra async-generator dependency onto the request's exit stack, plus constructing and closing the session object.

Deleting the seven parameters is the whole change. An AST scan over the backend finds exactly these seven handlers before and none after.
2026-08-17 01:53:00 -06:00
Timothy Jaeryang Baek
87d9b7e84e refac 2026-08-17 00:51:04 -07:00
Timothy Jaeryang Baek
6e468c5b95 refac 2026-08-17 00:50:54 -07:00
Timothy Jaeryang Baek
4ec6ee1441 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-17 00:47:32 -07:00
Classic298
ba0c4b3932
fix: don't hold a database connection for the lifetime of an SSE stream (#28183)
With database session sharing enabled, which the docs recommend for PostgreSQL and for multi-replica deployments, the knowledge pending-files and file process-status endpoints each pinned one pooled connection for as long as their SSE stream stayed open, up to one and two hours respectively. A file wedged in processing keeps a stream open for the full duration, so a handful of users sitting on that page can consume every connection in the pool, and the held transactions sit idle and block autovacuum on those tables.

Both handlers took a request-scoped session for their access checks, and FastAPI only releases a yield dependency once the response body has finished streaming, so the session outlived the handler by the whole life of the stream. Neither generator ever used it. They no longer take that dependency, and the queries they run already open their own short-lived sessions when none is passed. This is the approach the chat completion endpoints already use for the same long-response problem.

Measured against a pool with capacity 11: before, at most 11 concurrent streams could ever be open and every further attempt failed, deterministically across repeat runs. After, 25 of 25 opened. Non-stream latency is unchanged, within run-to-run noise, and behaviour is identical whether session sharing is on or off.
2026-08-17 01:46:54 -06:00
Timothy Jaeryang Baek
e968445812 refac 2026-08-17 00:43:47 -07:00
Timothy Jaeryang Baek
d799e81edb refac 2026-08-17 00:42:16 -07:00
G30
c0d09a5de9
fix: list publicly shared read-only notes in the Read Only view (#27637) 2026-08-17 01:40:56 -06:00
James Kerrane
44f4d5b94f
chore: refresh outdated version examples in bug report template (#28188)
* refactor: remove unused optional assignees key

According to the GitHub docs (https://docs.github.com/en/communities/using-templates-to-encourage-useful-issues-and-pull-requests/syntax-for-issue-forms#top-level-syntax) this key is optional. Since it's unused, it is fine to remove.

* chore: bump version examples for software

Older versions might confuse people filing new issues, so newer versions of mentioned software are used as examples.
2026-08-17 01:39:13 -06:00
Timothy Jaeryang Baek
f1a64ccfc2 refac 2026-08-17 00:35:12 -07:00
Timothy Jaeryang Baek
76d0160295 refac 2026-08-17 00:31:01 -07:00
Timothy Jaeryang Baek
9550731cc1 refac 2026-08-17 00:24:47 -07:00
Classic298
189c14fc4d
fix: match both JSON text spellings when searching serialised JSON columns (#28399)
Three searches LIKE against cast(json_col AS text), which means they have to match
bytes a JSON encoder wrote. Encoders disagree on non-ASCII: stdlib escapes it to
\uXXXX, orjson writes it raw. Which one produced a row depends on the codec in force
when it was written, so any single pattern finds only half the table.

models.py hard-codes the stdlib spelling, with a comment asserting SQLite stores
JSON via json.dumps(ensure_ascii=True). Model.meta is a JSONField, which has
serialised through JSONCodec since ENABLE_ORJSON was introduced, so on that setting
it stores raw UTF-8 and the escaped pattern matches nothing: non-ASCII workspace
model tag search is broken today. prompts.py and automations.py hard-code the
opposite spelling and miss rows written the other way.

json_text_variants returns both spellings a string can take inside serialised JSON,
collapsing to one for ASCII, and the three call sites OR over them. Rows written
under either setting are now found under either setting, which also covers a
database holding a mix of the two.

Case handling is unchanged. models.py keeps matching non-ASCII tags case-sensitively
on SQLite, whose LOWER() is ASCII-only and would not fold the stored text the way
str.lower() folds the tag. ASCII tags collapse to a single variant and take exactly
the query they took before.

Verified on SQLite across every combination of codec-that-wrote-the-row and
codec-the-app-is-running, for an ASCII and a CJK tag, over all three call sites: 24
of 24 match, against 12 of 24 before. Quoting still bounds whole-tag matches, so
searching "weather" does not match a row tagged "weathervane".

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-17 01:24:05 -06:00
G30
695d33aa7c
fix: let the chat column shrink so the collapsed sidebar rail is not pushed off screen (#28501) 2026-08-17 01:23:25 -06:00
G30
f5a5a434b9
fix: record an error state when a timer's chat completion raises (#27785) 2026-08-17 01:22:47 -06:00
G30
31897b7e34
fix: stop channel message hover actions overlapping code and table toolbars (#27737) 2026-08-17 01:20:54 -06:00
Timothy Jaeryang Baek
ad8c79f686 refac 2026-08-17 00:18:35 -07:00
Lin Junrong
211906d799
fix: keep references to lifespan background tasks (#28053)
periodic_usage_pool_cleanup, periodic_session_pool_cleanup and
scheduler_worker_loop were started with asyncio.create_task and their
handles discarded. The event loop keeps only a weak reference to a task,
so a task with no other referent can be garbage collected while it is
suspended at an await. All three are while True loops meant to run for
the process lifetime, and if one is collected the failure is silent:
pool entries stop being cleaned up, or automations and calendar alerts
stop firing, with nothing logged.

Six lines above, redis_task_command_listener is already stored on
app.state and cancelled on shutdown. This applies the same treatment to
the other three.

Closes #28052
2026-08-17 01:18:22 -06:00
Timothy Jaeryang Baek
0007369f4e refac 2026-08-17 00:16:11 -07:00
Timothy Jaeryang Baek
b6dc70c93b refac 2026-08-17 00:16:07 -07:00
Timothy Jaeryang Baek
d2af19ae3c refac 2026-08-17 00:15:36 -07:00
Classic298
3d630491c6
fix: re-syncing an existing model no longer fails silently (#28036)
POST /api/v1/models/sync only worked when every model in the payload was new. As soon as one id already existed, the whole call blew up and the endpoint still answered HTTP 200 with an empty list, so nothing was updated and well-behaved clients saw a success. Only a first-ever sync into an empty catalogue went through.

The update branch splatted the model dump (which already carries user_id and updated_at) and then passed both again as explicit keyword arguments, which is a duplicate-keyword TypeError before SQLAlchemy ever sees it. The insert branch right below merged the same values into a dict first, so it never collided.

Fixed by building that dict once and using it for both branches, matching how sync_functions already does it. Left the broad exception handler alone: it is the reason the failure was silent, but changing the error contract of sync_models is a separate call.

Fixes #28033
2026-08-17 01:14:11 -06:00
Classic298
b933292d63
refactor: track visited ids when resolving a chat's current message (#28035)
`delete_message_from_history` follows `childrenIds` down to the deepest leaf without recording where it has been. Record it.
2026-08-17 01:13:51 -06:00
Timothy Jaeryang Baek
f100edb708 refac 2026-08-17 00:12:13 -07:00
Classic298
27402ff210
refac: issue Playwright web loader requests from the shared HTTP clients (#28634)
The Playwright loader's route interceptor now performs each intercepted request with the same requests/aiohttp clients the other web loader paths already use and fulfills the page with that response, rather than having the browser issue it. Redirect handling, header forwarding and cookie delivery to the browser are unchanged.

Two consequences worth knowing. Page requests now leave from the backend instead of the browser, so with PLAYWRIGHT_WS_URL set they originate from a different host, and TLS is verified against certifi plus AIOHTTP_CLIENT_SSL_CERT_FILE rather than the browser's own trust store. And because the synchronous interceptor blocks, sub-resources on that path fetch one at a time: 30 assets at 40ms went from 2.01s to 3.01s, and 8 assets at 500ms from 1.05s to 4.50s. The asynchronous path is unaffected, at 0.65s and 1.05s respectively.
2026-08-17 01:06:12 -06:00
Classic298
73c1f5806a
refactor: match provider identity lookups via JSON subscript (#28624)
Both the OAuth and SCIM user lookups now compare the nested JSON value with SQLAlchemy's subscript operator, which emits the correct SQL for each supported database on its own. This replaces the hand-written sqlite and postgresql branches and the column-level contains() call they used.
2026-08-17 01:05:16 -06:00
G30
90724cdee0
fix: drop the white backdrop behind model icons in the admin Models list (#27612) 2026-08-17 01:03:37 -06:00
G30
8fc5ffe26e
fix: persist the Open Sharing permission in default user permissions (#27609) 2026-08-17 01:03:07 -06:00
Timothy Jaeryang Baek
0480ca9653 refac
Co-Authored-By: Solaris-star <67425364+solaris-star@users.noreply.github.com>
2026-08-17 00:01:38 -07:00
Timothy Jaeryang Baek
927ce0eae6 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-16 23:58:56 -07:00
Damien S
54cefd2b99
fix: preserve complete user context in agentic retrieval (#27642)
* fix: preserve user info in agentic RAG tools

* fix: preserve user info in file access checks

---------

Co-authored-by: Damien SPINELLI <damien.spinelli@external.list.lu>
2026-08-17 00:56:48 -06:00
xyonium
686d8dc54c
fix: strip prefix id from model name in /responses endpoint (#28575)
The /openai/responses endpoint forwarded the prefixed model id (e.g.
"myprovider.gpt-4o") to the upstream provider instead of the stripped
native name, causing "model not found" errors when a connection has a
Prefix ID configured.

generate_chat_completion() already strips the prefix before forwarding;
apply the same strip_provider_model_prefix() call in responses() after
the urlIdx routing (which needs the prefixed id) and re-serialize the
body afterwards.

Also fixes the Azure non-v1 deployment path, which built the deployment
URL from the prefixed model name.

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-17 00:53:00 -06:00
Classic298
3df485582d
fix: reject skill IDs that are not URL path safe (#27660)
A skill ID goes straight into the path of every mutating skill endpoint (/api/v1/skills/id/{id}/...), but create only replaced spaces with hyphens. An ID containing a "/" was stored verbatim as the primary key, so the route never matched, the request fell through to the SPA static mount and the client got 405 Method Not Allowed. The skill could not be opened, edited, toggled or deleted, by admins either, and since skill.name is UNIQUE it could not be recreated under a corrected ID. Percent-encoding does not help: uvicorn decodes the path before Starlette routes it, so the only remaining fix was a direct database write.

Create now rejects any ID outside [a-z0-9_-] with 400 instead of silently storing an unreachable one. Two frontend paths that fed unsanitized IDs into it are fixed as well: the manual "Skill ID" field, which was bound with no sanitization at all and is the path that reproduces on every version, and the markdown import, which put the raw frontmatter name into the ID before opening the editor in clone mode, where the reactive slugify is disabled.

Existing rows with an unreachable ID are not repaired here; rewriting a primary key would also have to re-point the access grants keyed on it.

Fixes #27655
2026-08-17 00:52:16 -06:00
Timothy Jaeryang Baek
954613944b refac 2026-08-16 23:51:38 -07:00
Classic298
805bfca5af
feat: emit group events on OAuth group sync (#27657)
With ENABLE_OAUTH_GROUP_MANAGEMENT enabled, every SSO login reconciles the user's group membership against the IdP claims, adding and removing them from groups and, with ENABLE_OAUTH_GROUP_CREATION, creating groups that do not exist yet. None of it emitted an event, so the same membership change was observable when an admin made it through the UI or when it arrived over SCIM, but invisible when the IdP drove it. That is the path that changes membership most often.

Emits group.member_added and group.member_removed per membership transition and group.created for each auto-created group, using the same payload keys as the groups router. The member events are published only when the write returned a group, so a failed or no-op write emits nothing, and both loops already run only on an actual transition. update_user_groups takes the request so the events can be published; it has a single caller.
2026-08-17 00:50:47 -06:00
Timothy Jaeryang Baek
d5d50169f4 refac 2026-08-16 23:48:05 -07:00
Timothy Jaeryang Baek
75df30c0ea refac 2026-08-16 23:47:02 -07:00
Timothy Jaeryang Baek
3e186abdd9 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-16 23:41:58 -07:00
G30
2813eb44f2
fix: stop the What's New modal drawing two bullets per changelog entry (#28676) 2026-08-17 00:40:57 -06:00
G30
ad72dc6658
fix: scroll to the top of a chat on the first click of Scroll to Top (#28659) 2026-08-17 00:40:14 -06:00
Timothy Jaeryang Baek
16f118d77a refac 2026-08-16 23:38:34 -07:00
G30
884388cac3
fix(search): wire up mark as unread in the search chats modal (#28136) 2026-08-17 00:34:49 -06:00
Timothy Jaeryang Baek
736e38338e refac 2026-08-16 23:32:23 -07:00
Timothy Jaeryang Baek
a40f6f2860 refac 2026-08-16 23:28:48 -07:00
Timothy Jaeryang Baek
b5da50f3df refac 2026-08-16 23:26:24 -07:00
Classic298
cd9db21c52
refac: bind tool server cookies per connection (#28630)
The tool callable now takes its connection's cookie jar as a parameter, matching how its headers are already passed and how the terminal tool factory in the same module builds its callables.
2026-08-17 00:25:07 -06:00
Classic298
7d392bedc9
refac: align the channel completion gate with the message update route (#28631)
The gate now applies the same authorship condition the channel message update route already uses, so both paths agree on which messages a caller may modify.
2026-08-17 00:24:49 -06:00
Classic298
1756c9d5d2
fix: default pinned models stop applying after a user's first page load (#28069)
Changing "Default Pinned Models" in admin settings had no effect for anyone who had already opened Open WebUI once. The sidebar copied the admin default into that user's own settings the first time it rendered and saved it to the server, which marked them as having customized their pins, so every later change to the default was ignored for them. Merely loading the page was enough, the user never had to touch a pin.

The default is now resolved for display only, through a shared store that falls back to the admin list while the user has no pins of their own, the same way default models already work. Nothing is written to the user's settings until they actually pin, unpin or reorder something, at which point their choice takes over for good. Unpinning everything still persists an empty list rather than snapping back to the default.

Users whose settings were already overwritten by the old behaviour keep that copy, since a stored pin list cannot be told apart from a deliberate one.

Fixes a drag-reorder path that mixed sidebar positions with stored ones, and stops the sidebar section reopening itself after any unrelated settings change.
2026-08-17 00:22:51 -06:00
Timothy Jaeryang Baek
1a376ac17f refac 2026-08-16 23:21:00 -07:00
Timothy Jaeryang Baek
b211c407f4 refac 2026-08-16 23:14:08 -07:00
G30
1d1c14bd77
fix: compare folder names for uniqueness with lower() equality instead of a LIKE pattern (#28695) 2026-08-17 00:00:02 -06:00
Timothy Jaeryang Baek
e4694d6c3f refac 2026-08-16 22:59:10 -07:00
Timothy Jaeryang Baek
f7a533cda6 refac 2026-08-16 22:57:01 -07:00
Timothy Jaeryang Baek
3258330729 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-16 22:56:19 -07:00
Timothy Jaeryang Baek
62fc436999 refac 2026-08-16 22:52:24 -07:00
Timothy Jaeryang Baek
31d08d592c refac
Some checks failed
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
2026-08-15 01:10:14 -06:00
Timothy Jaeryang Baek
467be93e6d refac 2026-08-15 01:09:39 -06:00
Timothy Jaeryang Baek
7fc5fa1ff3 refac 2026-08-15 01:02:06 -06:00
Timothy Jaeryang Baek
3eb65f4715 refac 2026-08-15 00:06:38 -06:00
Timothy Jaeryang Baek
30f82788bc refac 2026-08-15 00:05:45 -06:00
Timothy Jaeryang Baek
bbfdbd59f2 refac 2026-08-15 00:05:37 -06:00
Timothy Jaeryang Baek
98b9df0398 refac 2026-08-15 00:05:11 -06:00
Timothy Jaeryang Baek
0f821398ca refac 2026-08-14 23:52:13 -06:00
Timothy Jaeryang Baek
1b39ff352a refac 2026-08-14 23:50:34 -06:00
Timothy Jaeryang Baek
a5ea732c1e refac 2026-08-14 23:47:00 -06:00
Timothy Jaeryang Baek
516cf1a9a6 refac 2026-08-14 21:44:13 -06:00
Timothy Jaeryang Baek
d02b6a21fc refac
Some checks failed
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
2026-08-14 00:55:21 -06:00
Timothy Jaeryang Baek
c755ef60c6 refac 2026-08-14 00:54:33 -06:00
Timothy Jaeryang Baek
a1579a01ff refac 2026-08-14 00:22:17 -06:00
Timothy Jaeryang Baek
f7767d6be7 refac 2026-08-13 22:08:42 -06:00
Timothy Jaeryang Baek
d9014b3483 refac 2026-08-13 22:01:33 -06:00
Timothy Jaeryang Baek
57bd08304e refac 2026-08-13 22:01:17 -06:00
Timothy Jaeryang Baek
14e4d72d9a refac 2026-08-13 22:01:15 -06:00
Timothy Jaeryang Baek
256cce505b refac 2026-08-13 21:51:04 -06:00
Timothy Jaeryang Baek
55c202e841 refac 2026-08-13 21:48:01 -06:00
Timothy Jaeryang Baek
fa94a5ab24 refac 2026-08-13 21:36:41 -06:00
Timothy Jaeryang Baek
b4738d1a2e refac 2026-08-13 21:31:49 -06:00
Timothy Jaeryang Baek
653562d660 refac 2026-08-13 21:15:26 -06:00
Timothy Jaeryang Baek
a32a17965c refac 2026-08-13 21:06:51 -06:00
Timothy Jaeryang Baek
083e351441 refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
2026-08-13 21:02:17 -06:00
Timothy Jaeryang Baek
ec36972c2b refac 2026-08-13 20:02:08 -06:00
Timothy Jaeryang Baek
7d99b2716a refac 2026-08-13 19:59:11 -06:00
Timothy Jaeryang Baek
b018feb741 refac 2026-08-13 18:33:51 -06:00
Timothy Jaeryang Baek
133549a87e refac 2026-08-13 18:19:25 -06:00
Timothy Jaeryang Baek
4465f52a3e refac 2026-08-13 18:13:43 -06:00
Timothy Jaeryang Baek
2c01d59335 refac 2026-08-13 17:26:38 -06:00
Timothy Jaeryang Baek
7593698080 refac 2026-08-13 16:42:51 -06:00
Timothy Jaeryang Baek
2649e3305c refac 2026-08-13 16:42:10 -06:00
Timothy Jaeryang Baek
794671a988 refac 2026-08-13 16:16:10 -06:00
Timothy Jaeryang Baek
b7292890cc refac 2026-08-13 16:13:47 -06:00
Timothy Jaeryang Baek
f96b717566 refac 2026-08-13 16:04:53 -06:00
Timothy Jaeryang Baek
b3a5fd3875 refac 2026-08-13 15:57:42 -06:00
Timothy Jaeryang Baek
76583749ed refac 2026-08-13 15:52:44 -06:00
Timothy Jaeryang Baek
ec9bf5a64f refac 2026-08-13 15:49:50 -06:00
Timothy Jaeryang Baek
52c5e3b20d refac 2026-08-13 15:43:27 -06:00
Timothy Jaeryang Baek
c93c6d6fc4 refac 2026-08-13 15:42:54 -06:00
Timothy Jaeryang Baek
c1f914a626 refac 2026-08-13 15:38:25 -06:00
Timothy Jaeryang Baek
bfb68feea7 refac 2026-08-13 15:33:42 -06:00
Timothy Jaeryang Baek
2befa8f796 refac 2026-08-13 15:26:19 -06:00
Timothy Jaeryang Baek
7dfbdd221a refac 2026-08-13 15:20:59 -06:00
Timothy Jaeryang Baek
5c05608e3a refac 2026-08-13 15:02:18 -06:00
Timothy Jaeryang Baek
8d25ad00e2 refac 2026-08-13 14:52:30 -06:00
Timothy Jaeryang Baek
60feca71a6 refac 2026-08-13 14:50:43 -06:00
Timothy Jaeryang Baek
a17cb174ad refac 2026-08-13 14:46:08 -06:00
Classic298
5cb87e6630
fix: announce toggle state of integrations menu rows to screen readers (#27667)
Every toggle row in the chat integrations menu (filters, Web Search, Image, Code Interpreter, Tools, Skills) is a button whose on/off state was carried only by the decorative Switch inside it. Screen readers announced the row name and nothing else, so there was no way to tell whether a tool or feature was active without looking at it.

Each row button now carries aria-pressed, and the Switch wrapper is marked inert so the nested role=switch stops competing with the row for the announcement and stops adding a nameless tab stop. Hit testing skips inert content, so clicking the switch still toggles the row.

Tool rows that are not yet authenticated omit aria-pressed: activating those starts an OAuth redirect rather than toggling, so announcing them as an unpressed toggle would be wrong.

The Web Search, Image and Code Interpreter rows also had a state-flipping aria-label ("Disable Web Search") on top of aria-pressed, which announces as "Disable Web Search, pressed" and reads as the opposite of the truth. Removed: the visible row text already names each control.

Fixes #17150
2026-08-13 13:04:13 -06:00
Timothy Jaeryang Baek
ba885d0026 refac 2026-08-13 13:03:56 -06:00
G30
5c8d9e69c4
ci: auto-label feature requests with the enhancement label (#28531) 2026-08-13 12:55:03 -06:00
Timothy Jaeryang Baek
17c190bf99 refac 2026-08-13 02:07:02 -06:00
Timothy Jaeryang Baek
f8c5fda283 refac 2026-08-13 02:03:13 -06:00
Timothy Jaeryang Baek
31c1ffd55a refac 2026-08-13 00:45:02 -06:00
Timothy Jaeryang Baek
e1acd7e7ca refac 2026-08-13 00:38:32 -06:00
Timothy Jaeryang Baek
25802c048e refac 2026-08-13 00:34:10 -06:00
Timothy Jaeryang Baek
8260d527ee refac 2026-08-13 00:22:58 -06:00
Timothy Jaeryang Baek
85c3d0ae2f refac 2026-08-13 00:07:31 -06:00
Timothy Jaeryang Baek
3813fd4cdc refac 2026-08-12 23:20:19 -06:00
Timothy Jaeryang Baek
4e03d89414 refac 2026-08-12 23:17:37 -06:00
Timothy Jaeryang Baek
90bb94abf9 refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
2026-08-12 18:13:47 -06:00
Timothy Jaeryang Baek
86b7bf1f7e refac 2026-08-12 18:10:51 -06:00
Timothy Jaeryang Baek
2e93987490 refac 2026-08-12 17:54:10 -06:00
Timothy Jaeryang Baek
939bcdb79e refac
Co-Authored-By: G30 <50341825+silentoplayz@users.noreply.github.com>
2026-08-12 01:13:38 -06:00
G30
1deeaf71da
fix: stop the calendar update tool from nulling every omitted field (#27777) 2026-08-12 01:09:07 -06:00
G30
2ab0311b99
fix: reject COUNT rules without DTSTART so limited automations cannot run forever (#27781) 2026-08-12 01:08:55 -06:00
G30
783e87c0c5
fix: dock the thread reply input below the scroll area instead of inside it (#27768) 2026-08-12 02:06:20 -05:00
G30
104a0f2f11
fix: remove vestigial api_base_url param that broke the Tavily web loader (#27636) 2026-08-12 02:05:26 -05:00
Timothy Jaeryang Baek
e44e16cb59 refac 2026-08-12 01:02:02 -06:00
Timothy Jaeryang Baek
c086b80313 refac 2026-08-12 00:51:12 -06:00
Timothy Jaeryang Baek
ab41dcc487 refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
2026-08-11 17:44:34 -06:00
Timothy Jaeryang Baek
9c21d4ed3b refac 2026-08-11 17:42:25 -06:00
G30
e17dfae72e
fix: enforce the image generation flag on the legacy chat feature path (#27759)
* fix: enforce the image generation flag on the legacy chat feature path

* fix: refresh the config store when image generation is disabled in admin settings

* fix: hide active feature pills when the feature is no longer available
2026-08-11 17:38:10 -06:00
Timothy Jaeryang Baek
1744b63f16 refac 2026-08-11 17:35:24 -06:00
Timothy Jaeryang Baek
4f9a0ebf71 refac 2026-08-11 17:35:05 -06:00
Timothy Jaeryang Baek
65c3539666 refac 2026-08-11 15:13:59 -06:00
Timothy Jaeryang Baek
0a598e6ff7 refac 2026-08-11 15:09:48 -06:00
Timothy Jaeryang Baek
1674e5a9ef refac 2026-08-11 12:50:50 -06:00
Timothy Jaeryang Baek
724d2ebbf1 refac 2026-08-11 11:46:12 -06:00
Timothy Jaeryang Baek
ac0368b4ab refac 2026-08-11 11:40:29 -06:00
Timothy Jaeryang Baek
f79b443c22 refac 2026-08-11 11:40:24 -06:00
G30
e58245d5b4
fix: place model editor capability checkboxes directly before their labels (#27788) 2026-08-11 11:13:04 -06:00
G30
25455e943d
fix(ui): resolve owner avatars against the API base url and add a fallback (#28272) 2026-08-11 11:05:35 -06:00
Timothy Jaeryang Baek
29541cbb52 refac 2026-08-11 01:27:50 -06:00
Timothy Jaeryang Baek
c8f8fa451a a11y
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-08-11 01:25:20 -06:00
Timothy Jaeryang Baek
865c80c160 refac 2026-08-11 01:17:19 -06:00
Timothy Jaeryang Baek
f0bfcd4097 refac 2026-08-11 01:15:05 -06:00
Timothy Jaeryang Baek
be4afd7545 refac 2026-08-11 01:11:56 -06:00
Timothy Jaeryang Baek
e963d36e39 refac 2026-08-11 01:08:58 -06:00
KastenMonster
2ac01a42be
feat: Add openserp as a provider in the ui (#27594)
* add OpenSerp Option

* fix: change input label to the docs label

---------

Co-authored-by: KastenMonster <kastenmonster@groundonline.de>
2026-08-11 00:21:05 -06:00
Timothy Jaeryang Baek
9122c24ea2 refac 2026-08-11 00:06:21 -06:00
G30
79382d1b19
fix: stop profile preview avatars from squashing in narrow flex rows (#28000) 2026-08-10 23:49:45 -06:00
Timothy Jaeryang Baek
68c67a1e0f refac
Co-Authored-By: G30 <50341825+silentoplayz@users.noreply.github.com>
2026-08-10 23:48:02 -06:00
Timothy Jaeryang Baek
5cd9a39534 refac 2026-08-10 23:40:31 -06:00
Timothy Jaeryang Baek
bd250a0e24 refac 2026-08-10 23:37:45 -06:00
Classic298
934802e186
fix: enforce features.memories permission on the legacy memory context path (#27668)
Revoking a user's `features.memories` permission removed their access to the memories API and to the native function-calling memory tools, but their stored memories were still injected into the system context on the legacy function-calling path.

The branch in `process_chat_payload` only checked the client-supplied `features['memory']` flag plus the global `memories.system_context.enable` switch, with no user-permission check. `add_memory_context` did not compensate: it only checks `model_allows_memory`, which is a model capability rather than a permission, and the one call inside it that does check the permission (`query_memory`) has its 403 swallowed by a `try/except`, so `Memories.get_memories_by_user_id` and the neighbourhood scan still fed the system prompt.

Gate the branch with the same permission check the native path already performs in `get_builtin_tools`, matching the neighbouring `web_search` and `image_generation` branches.

Only the caller's own memories were injected into the caller's own context, so there was no cross-user exposure. The practical effect was that the permission toggle did not do what its name implies: an admin who revoked it still got memory content injected for that user.
2026-08-10 23:36:27 -06:00
G30
d959e3e312
fix: scale the settings modal height with tall viewports instead of capping at 54rem (#27615) 2026-08-10 23:34:04 -06:00
Classic298
4eb0394511
fix: merged response receiving empty model responses after reload (#27673)
Clicking "Merged Response" in a multi-model chat often made the merging model answer with "It appears that the responses provided from the other models were empty".

The merge handler collected each model's answer via `history.messages[id].content`. Assistant messages are persisted by the backend with `output` only (`upsert_message_to_chat_by_id_and_message_id` writes `done`/`role`/`output`, never `content`), so `content` is only populated in the browser session that generated the responses, where the streaming handler mirrors it. Once the chat is reloaded from the database, every assistant message has `content: ''` and the merge request is sent with a list of empty strings, which is exactly what the merging model then reports. That is why the failure looks random: merging works right after generating, and fails after a refresh or when reopening the chat.

Read the responses through `getOutputText(message.output) || message.content`, the same fallback already used by every other read site (`ResponseMessage`, `Overview/Node`, `SearchModal`, `ChatItem`, `ChatMenu`, `Navbar/Menu`).

Fixes #26962
2026-08-10 23:31:21 -06:00
Classic298
fcc130c9bb
fix: send stream_options.include_usage for backend-initiated chats (#27661)
Only the frontend added `stream_options: {include_usage: true}` to the completion payload, gated on the model's `usage` capability. Every backend-initiated run builds its own payload (automations, timers, subagents, channels) and omitted it, so those responses came back without token counts and never rendered the usage block, even with the capability enabled on the model.

Set it in `chat_completion` instead, the single handler all of those callers go through, and drop the two duplicate copies (the Anthropic-compat handler and the frontend). Capabilities are read from the resolved model before the custom-model fallback can rebind it, and the flag is applied after the model's `stream_response` override so a non-streaming model is unaffected.

Fixes #27653
2026-08-10 23:30:46 -06:00
G30
085f1b5ca4
feat: add a user setting to toggle sidebar chat hover previews (#27632)
Co-authored-by: Tim Baek <tim@openwebui.com>
2026-08-10 23:29:51 -06:00
G30
8e83d83324
fix: left-edge clipping of the terminal cloud icon and account profile avatar (#27691) 2026-08-10 23:28:17 -06:00
Solaris-star
2387ce63cb
fix(chat): keep regenerate button visible in High Contrast mode (#27644)
Every action button on a response message gates its visibility on
'isLastMessage || ($settings?.highContrastMode ?? false)' so that buttons
stay visible (not hover-only) when High Contrast mode is on. The two
Regenerate buttons — the one inside RegenerateMenu and the fallback in the
{:else} branch — were missed and still gated on isLastMessage alone, so on
any non-last message the Regenerate icon was invisible until mouse-over even
with High Contrast enabled (#27638).

Add the same highContrastMode clause to both buttons, matching the other
action buttons in this file.
2026-08-10 23:27:42 -06:00
Classic298
be9af1653e
fix: anchor chat input expand button to the input row (#27676)
The "expand input" button was positioned with `fixed top-0 right-0`. That only kept it near the composer by accident: `#message-input-container` sets `backdrop-blur-sm`, and a backdrop-filter makes an element the containing block for fixed descendants, so the button resolved to the top-right corner of the entire composer instead of the text area it belongs to.

That corner is already taken. The `@`-tagged model chip renders as the first row of the same container with its dismiss button at the right end, so with a multi-line prompt and a tagged model the two controls are drawn on top of each other. The attached-files row has the same problem: the button paints over the first thumbnail and its remove button.

Anchor the button to the wrapper that holds the text area instead, using `relative`/`absolute`, so it always sits at the top-right of the input row and below whatever rows precede it. With no chip and no files the position is unchanged. As a side effect the button is no longer a child of the `overflow-auto` scroller, so it can no longer be clipped or scrolled out of view on long prompts.

Fixes #26736
2026-08-10 23:26:53 -06:00
G30
b91bb67b55
fix: hide chat fork actions when chat import permission is disabled (#27711)
Co-authored-by: Nameless-Monster-Nerd <Nameless-Monster-Nerd@users.noreply.github.com>
2026-08-10 23:24:59 -06:00
Classic298
5c79ccc9e5
refactor: walk chat message history by map key (#28034)
`get_message_list` moves through `messages_map` by key but tracked each message's own `id` field, which the message body does not have to carry. Track the key instead.
2026-08-10 23:24:41 -06:00
G30
80d2f4154a
fix: align public_tools and public_notes sharing defaults with config (#27716)
The SharingPermissions model defaulted public_tools and public_notes to
True while the config defaults (USER_PERMISSIONS_WORKSPACE_TOOLS_ALLOW_PUBLIC_SHARING
and USER_PERMISSIONS_NOTES_ALLOW_PUBLIC_SHARING) are both False.

On an instance whose stored user.permissions config predates these keys,
GET /api/v1/users/default/permissions fills the gap from the model and
reports both as enabled, while has_permission fills it from
DEFAULT_USER_PERMISSIONS and denies. Saving any unrelated permission then
persists the model's True, granting public tool and note sharing the admin
never enabled.
2026-08-10 23:24:18 -06:00
G30
a3d33b4cf3
fix: hide chat delete actions when the chat delete permission is disabled (#27714)
The chat delete endpoints gate on chat.delete for non-admins, but none of
the UI affordances that reach them were conditioned on it, so a user
without the permission is offered controls that always fail with
"Access prohibited".

Gates every entry point on the same condition the backend enforces,
matching the existing chat.share / chat.export / chat.import gates in the
same components:

- Sidebar chat menu (ChatMenu) Delete item — also covers the Search
  modal's menu, which reuses ChatMenu
- Sidebar Shift+hover inline Delete button (ChatItem)
- Sidebar hidden #delete-chat-button keyboard-shortcut target (ChatItem)
- Chat navbar menu Delete item (Navbar/Menu)
- Search modal Shift-held inline Delete button (SearchModal)
- Settings -> Data Controls -> Delete All Chats
- Settings -> Archived Chats per-row Delete
2026-08-10 23:24:02 -06:00
G30
fe62934be7
fix: derive overview node spacing from rem so branches never overlap (#27995) 2026-08-10 23:23:31 -06:00
Classic298
2207876ae7
fix: generate valid WEBUI_SECRET_KEY in start_windows.bat (#28061)
The key generation loop redirected input from a non-existent file
(`SET /p WEBUI_SECRET_KEY=<!random!>>%KEY_FILE%`), printing "The system
cannot find the file specified." once per iteration and leaving the key
file empty, so startup failed with "WEBUI_SECRET_KEY is not set".

Build a fixed-length alphanumeric key by indexing into a charset with
%RANDOM% and write it once with `<nul set /p`. Also quote the key file
path and use delayed expansion so paths with spaces work.


Claude-Session: https://claude.ai/code/session_01CmgBivWjad68mX4yBVWMi2

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-10 23:22:43 -06:00
Timothy Jaeryang Baek
bd8378f643 refac 2026-08-10 23:22:08 -06:00
Classic298
1f22cccd22
perf: stop formatting every exported log record twice under OTEL log export (#27840)
With ENABLE_OTEL and ENABLE_OTEL_LOGS set, InterceptHandler builds the message once for loguru and then hands the same LogRecord to the OpenTelemetry handler, whose _translate calls record.getMessage() a second time. That used to be free, because the message was already a finished f-string with nothing to substitute. Now that log calls pass lazy %-args, the second call re-runs the whole interpolation, so every exported record is formatted twice.

The two getMessage() calls on a 78 kB retrieval record:

    before  373.0 us
    after     0.1 us

Stamping the built message back onto the record makes the second call a plain string return. msg and args are both in OpenTelemetry's _RESERVED_ATTRS, so neither ever reaches the exported attributes. The isinstance guard matters: _translate exports a non-str msg such as the dicts routers/audio.py logs as a typed body rather than a string, so those records are left untouched, and they have no %-args to format twice anyway. Body, attributes and severity were compared against LoggingHandler._translate for str, dict, list, int, None, exception and exc_info records.
2026-08-10 23:13:52 -06:00
Timothy Jaeryang Baek
ce3c175e26 refac 2026-08-10 23:13:10 -06:00
Timothy Jaeryang Baek
de289eb1aa refac 2026-08-10 23:07:47 -06:00
Timothy Jaeryang Baek
d8ae7ed405 refac 2026-08-10 23:01:41 -06:00
Cyp
bee1ded5ab
i18n: update Korean translations (#27681)
Co-authored-by: bglee <bglee@hct.co.kr>
2026-08-10 23:01:24 -06:00
Timothy Jaeryang Baek
629cdcb530 refac 2026-08-10 22:57:21 -06:00
Timothy Jaeryang Baek
89922cc9d5 refac 2026-08-10 22:53:37 -06:00
Timothy Jaeryang Baek
2a6e671f54 refac 2026-08-10 22:47:39 -06:00
Timothy Jaeryang Baek
c2107e5bb3 refac 2026-08-10 22:36:42 -06:00
Kylapaallikko
b7de04da14
Update fi-FI translation.json (#27700)
Added missing translations and improved existing ones.
2026-08-10 22:27:44 -06:00
G30
38a03830c2
fix: show a placeholder when an image cannot be loaded (#27730)
A message referencing a file that no longer exists rendered a broken
image, because nothing anywhere noticed the failed load. What the user saw
depended on the caller's alt text: in a response the alt is the message
content, so the entire reply was rendered inside the image frame, and the
preview still opened full-screen on a dead image.

Image.svelte now tracks a failed load and renders an 'Image unavailable'
placeholder instead, suppressing the preview for a source that cannot be
shown. The failed state resets when the source changes, so a replaced or
corrected URL is retried. An onError callback lets a caller react without
changing behaviour for the eight existing call sites, which are untouched.

Responses now pass the file name (or 'Generated Image') as alt rather than
the message content, which was never a description of the image.
2026-08-10 22:27:13 -06:00
Timothy Jaeryang Baek
0b4b7ae5ff refac 2026-08-10 22:20:53 -06:00
G30
04c22f0c41
fix: use the local date for the event modal's default start and end (#27779) 2026-08-10 22:19:54 -06:00
Classic298
3dbb4078b3
fix: repair two broken logging calls, one of which makes VECTOR_DB=opengauss unusable (#27838)
SRC_LOG_LEVELS became an empty dict when per-module log levels were dropped, and env.py keeps it only as a legacy name. opengauss.py is the last thing in the tree that still indexes it, at module scope, so importing the module raises KeyError: 'RAG' and any deployment on VECTOR_DB=opengauss dies the first time it touches the vector store. The factory imports it lazily, which is why nothing else trips over it. Deleting the line is the whole fix: every other vector backend takes getLogger(__name__) and inherits the root level.

colbert.py passes an argument to a message with no placeholder to consume it:

    log.info('ColBERT: Loading model', name)

At INFO, which is the default, logging evaluates 'ColBERT: Loading model' % ('colbert-ir/colbertv2.0',) and raises TypeError: not all arguments converted during string formatting. The record is swallowed by handleError, so loading a ColBERT reranker prints '--- Logging error ---' plus a traceback to stderr instead of the model name. Adding %s prints the name and drops the traceback.
2026-08-10 22:19:41 -06:00
Timothy Jaeryang Baek
3793b0c886 refac 2026-08-10 22:15:54 -06:00
Timothy Jaeryang Baek
3e9b075954 refac 2026-08-10 22:11:05 -06:00
Timothy Jaeryang Baek
2a45fa04cb refac 2026-08-10 21:58:12 -06:00
Teitur Bendtsen
b74c144c96
i18n: add Faroese translation (#28159)
Co-authored-by: T209211 <tbe@betri.fo>
2026-08-10 21:51:45 -06:00
Adam I. Horvath
cbd47eaa67
i18n: improve hungarian translations (#28214)
adds missing hungarian translations and improves existing hungarian ui labels.
2026-08-10 21:44:40 -06:00
Classic298
38fcee7f21
Update pull_request_template.md (#28246) 2026-08-10 21:43:33 -06:00
developersorli
f44647e251
feat(i18n): add complete Slovenian (sl-SI) translations (#28268) 2026-08-10 21:43:02 -06:00
Classic298
a680f21e12
feat: make OAuth admin settings read-only when ENABLE_OAUTH_PERSISTENT_CONFIG is off (#28276)
When ENABLE_OAUTH_PERSISTENT_CONFIG is off (the default), oauth.* config is
never persisted and is read from environment variables, but the admin panel
still let admins edit the OAuth/OIDC fields and silently dropped every save on
restart, which kept confusing users who missed the docs warning
(open-webui/open-webui#28247).

The OAuth/OIDC section is now read-only in that case: the admin oauth config
endpoint reports the flag and the UI wraps the section in a disabled fieldset,
slightly dimmed with every control inert but all values still visible, plus a
note naming the env var. Saving skips the OAuth POST since nothing can change.
With the flag enabled the section behaves exactly as before.

Known limits: the guard is UI-side only (the POST endpoint keeps accepting
writes, unchanged), and disabled fields mean values cannot be selected and the
masked client secret cannot be revealed while read-only. Switch.svelte gains a
disabled:cursor-not-allowed style that applies to any disabled switch app-wide.

Co-authored-by: Tim Baek <tim@openwebui.com>
2026-08-10 21:41:07 -06:00
G30
a3e5d0b362
fix(chat): attribute shared chat messages to their author, not the viewer (#28274) 2026-08-10 21:40:54 -06:00
Timothy Jaeryang Baek
8edab5020e refac 2026-08-10 21:39:27 -06:00
Timothy Jaeryang Baek
3c010951db refac 2026-08-10 21:21:52 -06:00
Timothy Jaeryang Baek
48a5696042 refac 2026-08-10 21:06:40 -06:00
Timothy Jaeryang Baek
d6679082e5 refac 2026-08-10 21:00:01 -06:00
Timothy Jaeryang Baek
943294df9a refac 2026-08-10 20:46:38 -06:00
G30
d661fb49b8
ci: relabel bug reports when the title is corrected after opening (#28333)
The labeler only listened for issues.opened, so it judged a title exactly
once. Reporters who omit the prefix are asked by the triage bot to add one,
and the title they then fix is never looked at again, leaving a genuine bug
report unlabelled until somebody notices it by hand.

It now also runs on issues.edited. Body only edits return immediately, so
the extra runs are limited to titles actually changing, and a bug label a
maintainer has already removed is not restored: on an edit the issue events
are checked for a previous removal first. The opened path is unchanged and
makes no additional API call.

The title pattern also required a delimiter after the bracket, so the
bracketed form the comment advertises, "[Bug] something is broken", never
matched unless it happened to be written "[Bug]: something is broken".
Both forms match now, while "[bug/perf]" keeps matching and titles that
merely mention the word, such as "[UI Bug]" or "fixed a typo", still do not.
2026-08-10 20:45:02 -06:00
G30
d10d552117
fix(chats): surface the error message instead of [object Object] (#28260) 2026-08-10 20:44:32 -06:00
Timothy Jaeryang Baek
178ccb30e1 refac 2026-08-10 20:37:53 -06:00
Timothy Jaeryang Baek
5cecb7dbfa refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
2026-08-10 20:30:56 -06:00
G30
148283f974
fix: release a queued message once an attached URL finishes (#28381)
A message sent while an attachment is still uploading is held in the chat
queue until the file settles. For files that release works, because the
message input calls onUpdate once the upload completes, which refreshes the
queued entry and asks the queue to run.

Attaching a URL takes a different path. uploadWeb sets the item to uploaded
but calls neither onUpdate nor processNextInQueue, and the queue is
otherwise only revisited when a generation finishes or when the chat is
mounted while idle. In a new chat with nothing generating, none of those
happen, so a message queued behind a URL waits with no event able to
release it until the chat is reloaded.

uploadWeb now asks the queue to run when it is done, the same way the file
path already does.
2026-08-10 20:24:35 -06:00
Timothy Jaeryang Baek
b20bcdbba7 refac 2026-08-10 20:22:29 -06:00
Timothy Jaeryang Baek
5ec16e76e6 refac 2026-08-10 20:13:03 -06:00
Classic298
5462c02af0
fix: OIDC login fails when the provider adds a private JOSE header (#28065)
Logging in through CyberArk Identity dies at the callback with "Unsupported {'app_id'} in header" and the user sees "The email or password provided is incorrect". Any provider that puts a vendor-specific parameter in the ID token header hits this; CAS was already patched by name, CyberArk is the next one.

Authlib 1.7 verifies ID tokens with joserfc, which rejects header parameters it does not recognise. The old fix registered `client_id` so CAS would work, which only ever fixes one provider at a time. This turns off the unknown-header rejection instead, so any private header parameter is ignored rather than fatal. Signature verification, the algorithm allowlist, `crit` handling and value validation of registered headers all still run, so nothing that actually protects the token is relaxed.

Fixes #28062
2026-08-10 20:06:31 -06:00
Timothy Jaeryang Baek
a41faa3c22 refac 2026-08-10 20:00:43 -06:00
Classic298
c5ec01b1f9
fix: make the aiodns resolver opt-in and pin aiodns to 3.6.1 (#28242)
Since v0.11.0 shipped aiodns, aiohttp silently switched every outbound request from the OS resolver to c-ares. On some Windows hosts the bundled c-ares 1.34.6 (pycares 5) discovers only 127.0.0.1:53 as nameserver, so every external provider lookup fails (#28013). In Docker the long-lived c-ares channel intermittently stops resolving container names while Docker's embedded DNS keeps answering, which wipes the Ollama model list and fails all in-flight chats with a misleading "Model not found" (#28215).

This restores the pre-0.11 ThreadedResolver (OS resolver) by default and gates the c-ares path behind a new env var, AIOHTTP_CLIENT_ASYNC_DNS_RESOLVER, off by default. The event-loop DNS perf improvement is now opt-in for deployments whose resolver setup is known to work with c-ares, instead of a process-wide side effect of the package being installed.

aiodns is also downgraded and pinned to 3.6.1 (pycares<5), the last release before the broken c-ares 1.34.6 build, so opting in does not hit the Windows regression. The hardcoded AsyncResolver in the Mistral OCR loader now follows the same switch. Simply removing aiodns instead was not an option because opting in would then be impossible, and #28215 showed the Docker failure is c-ares itself, not aiodns 4.x.
2026-08-10 19:52:29 -06:00
Timothy Jaeryang Baek
8d1c205d8e refac 2026-08-10 19:46:46 -06:00
Timothy Jaeryang Baek
b606e13da3 refac 2026-08-10 19:44:56 -06:00
Classic298
eff5c4a2d9
feat: parse :::writing block metadata and use the subject as the block title (#28280)
Newer OpenAI chat models put metadata on the opening line of a colon fence block, like :::writing{variant="email" id="48173" subject="Short question" recipient="mail@example.com"}. The tokenizer matched that line and discarded it, so every block rendered under the same generic "Writing" heading no matter what it contained.

The opening line is now parsed into an attributes map on the token and the header uses it: the subject becomes the title, the recipient follows it and the full string is reachable on hover when the row is too narrow for it. Blocks without metadata render exactly as before, and the other fence types get the parsed attributes for free.

Attributes are read only from inside the {...} braces, not from the whole opening line. Scanning the whole line turned ordinary prose containing key="value" into metadata, and it backtracked quadratically: a 40k character opening line took 586ms to parse, and that runs again on every re-lex while the message streams. Anchored to the braces it is 0.0ms.

Nothing here turns the recipient into a link or a send action. That metadata is model output and can be steered by whatever is in the context, so a prefilled mail action is a separate decision rather than a side effect of parsing.
2026-08-10 19:44:06 -06:00
Timothy Jaeryang Baek
385d08bea5 refac 2026-08-10 19:42:05 -06:00
Timothy Jaeryang Baek
11739a2de8 refac 2026-08-10 19:41:05 -06:00
Classic298
92f9f36c69
Update CODE_OF_CONDUCT.md (#28349) 2026-08-10 19:34:33 -06:00
Timothy Jaeryang Baek
e4dd6c4bf1 refac 2026-08-10 19:33:48 -06:00
Timothy Jaeryang Baek
d22bb6703f refac 2026-08-10 19:28:21 -06:00
Classic298
e5b24a22d0
fix: model ID whitelists accepting duplicate entries (#28251)
Adding a model ID that was already on the list in the connection settings modal simply appended it again, so the same model could sit in the whitelist any number of times. The arena model modal had the same flaw, its dropdown kept offering models that were already selected.

The connection modal now rejects a duplicate with a toast and trims the input first; surrounding whitespace renders invisibly in the list, so an untrimmed ID would slip past the duplicate check and still show up as a visually identical row. The arena modal instead filters already-added models out of the dropdown, matching the existing model selector in the admin settings, so a duplicate can no longer be picked at all. Both modals also drop duplicates when loading a stored list, so configs that already contain them are cleaned on their next save.

Until such a config is re-saved, one residual effect of old data remains: a duplicated ID in an arena model's stored list keeps double weight in the random model draw. New duplicates can no longer be created through the UI.

Fixes #28249
2026-08-10 19:27:12 -06:00
Timothy Jaeryang Baek
ff74bfa6a1 refac 2026-08-10 19:25:26 -06:00
Timothy Jaeryang Baek
f8ac75d188 refac 2026-08-10 19:21:52 -06:00
G30
8836dcb59f
fix(ui): fall back to the default avatar when a profile image fails to load (#28270) 2026-08-10 19:18:32 -06:00
G30
121f2404ee
fix(retrieval): report why a URL could not be read instead of blaming the knowledge base (#28362)
Fetching a URL and saving it were reported as one thing. Everything from
reading the URL to writing the vector database sat inside a single try,
whose handler blamed the knowledge base, so a page that could not be
fetched, parsed or resolved was reported as a knowledge base error even
though nothing had reached the knowledge base yet. Reading the URL now has
its own handler that names the URL, and the knowledge base message is left
to the step that actually touches it.

When YouTube refused a transcript the reason was discarded earlier still:
the loader caught the error, logged it, and returned an empty document
list, so the empty result failed downstream and even the salvageable
explanation was gone before a message was produced. The loader now raises
YoutubeTranscriptError carrying a readable reason, mapped from the
transcript library's own exception types. Blocked requests mention that a
proxy can be configured, and disabled, age restricted, unavailable and
missing language cases each say what actually happened.

URLs that attach successfully are unaffected.
2026-08-10 19:18:04 -06:00
Classic298
d9e23b90c1
refac: share one folder write-access check across chat folder_id paths (#28366)
Chat creation and chat moves each carried their own copy of the same folder_id validation, resolving the folder and checking ownership and shared write access in slightly different ways. Both now call a single has_folder_write_access helper, which the chat-completions creation path uses as well, so ownership, inherited write grants and nonexistent or malformed ids behave identically everywhere a chat folder_id is set. The owner case also costs one query fewer than before.
2026-08-10 19:17:18 -06:00
Timothy Jaeryang Baek
72a909fd2f refac 2026-08-10 19:16:21 -06:00
Timothy Jaeryang Baek
ec03e88144 refac 2026-08-10 19:08:46 -06:00
Timothy Jaeryang Baek
b5f86e6a43 refac 2026-08-10 19:08:06 -06:00
Timothy Jaeryang Baek
30d08a42f8 refac 2026-08-10 18:52:58 -06:00
Timothy Jaeryang Baek
407c40f72c refac 2026-08-10 18:52:34 -06:00
Timothy Jaeryang Baek
b4d3b27caf refac 2026-08-10 18:52:27 -06:00
Timothy Jaeryang Baek
d4461bd6f3 refac 2026-08-10 18:52:18 -06:00
Timothy Jaeryang Baek
eeaf1a1df0 refac 2026-08-10 18:51:38 -06:00
Timothy Jaeryang Baek
90a0e61cef refac 2026-08-10 18:50:43 -06:00
Timothy Jaeryang Baek
37f2548155 refac 2026-08-10 18:46:36 -06:00
Timothy Jaeryang Baek
13346c5f16 refac 2026-08-10 18:43:55 -06:00
Classic298
8d6a7c8308
perf: route native JSON columns through JSONCodec instead of stdlib json (#28396)
JSONField serializes with JSONCodec, but columns declared as SQLAlchemy's own JSON
type go through the engine's serializer instead, and no engine set one. That left
Chat.chat - the largest blob the app stores - on stdlib json.dumps/loads no matter
what ENABLE_ORJSON was set to, while the rest of the app used the codec. SQLAlchemy
invokes it once per write and once per read, so every chat read and write paid a
full stdlib pass over the whole conversation on top of whatever the caller did.

Both engine constructors are now wrapped so the codec is wired in by default and
cannot be missed by a call site that forgets it; an explicit json_serializer still
wins. The 10 create_engine/create_async_engine calls in this module go through the
wrappers. Vector-store engines (pgvector, mariadb, opengauss) are separate databases
and are left alone.

Serializing and deserializing chat-shaped blobs, median of 11 runs:

| chat blob | write | read |
| --- | --- | --- |
| 600 msgs (2.8 MB) | 10.1 -> 1.7 ms | 8.4 -> 3.7 ms |
| 3000 msgs (14.2 MB) | 51.9 -> 8.0 ms | 48.5 -> 27.8 ms |
| 6000 msgs (28.5 MB) | 105.9 -> 29.5 ms | 112.7 -> 80.5 ms |

With ENABLE_ORJSON off JSONCodec is stdlib json, so this is a no-op until the flag
is set - the change cannot regress a default deployment.

With it on, a round-trip probe through a native JSON column returns objects equal to
the stdlib ones on all 12 shapes tried: ASCII, CJK, emoji, astral-plane, unicode
keys, null bytes, lone surrogates, floats, ints above 2**63 and 2**64, line
separators, empty and deeply nested. Stored text changes for non-ASCII, which is
written as raw UTF-8 rather than backslash-uXXXX escapes and is correspondingly
smaller. Nothing queries that text by escape except two Postgres safety filters in
chats.py, and both still hold: a null byte is escaped identically by both codecs,
and the title filter reads a text column rather than JSON. The ->> and json_extract
searches decode the string before matching, so escaping cannot reach them.

Two differences are inherent to JSONCodec and already apply to every JSONField
column: ints beyond 2**64-1 come back as float, and NaN/Infinity serialize to null
rather than the bare literals stdlib emits - the latter being invalid JSON that a
Postgres json column rejects today. Neither shape occurs in chat blobs. Alembic
builds its own engine and stays on stdlib, which is fine in both directions since
each codec reads the other's output.


Claude-Session: https://claude.ai/code/session_014BXoM6QiFJKisxcxKAXii8

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-10 19:39:49 -05:00
Timothy Jaeryang Baek
8fbfd14a8b refac 2026-08-10 01:38:32 -06:00
Timothy Jaeryang Baek
060648f939 refac 2026-08-10 01:36:09 -06:00
Timothy Jaeryang Baek
61110677d4 refac 2026-08-10 01:32:34 -06:00
Timothy Jaeryang Baek
f9cd49443c refac 2026-08-10 01:32:29 -06:00
Timothy Jaeryang Baek
4e69166017 refac 2026-08-10 01:32:11 -06:00
Timothy Jaeryang Baek
048c063993 refac 2026-08-10 01:26:04 -06:00
Timothy Jaeryang Baek
ff7467b4c5 refac 2026-08-10 01:14:53 -06:00
Timothy Jaeryang Baek
8dd23f74c9 refac 2026-08-10 00:52:14 -06:00
Timothy Jaeryang Baek
1b72899f24 refac 2026-08-10 00:26:44 -06:00
Timothy Jaeryang Baek
a33fa05adc refac 2026-08-10 00:19:52 -06:00
Timothy Jaeryang Baek
5b8975b7da refac 2026-08-10 00:05:55 -06:00
Timothy Jaeryang Baek
2dadc5435a refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
2026-08-09 13:22:46 -06:00
Timothy Jaeryang Baek
5caa91a493 refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Waiting to run
Frontend Build / Unit Tests (push) Waiting to run
2026-08-08 18:47:57 -06:00
Classic298
74a7902821
fix: apply response.output_item.done instead of ignoring it (#28310)
The Responses API handler had a branch for response.output_item.done whose own comment said it was handled specifically below, but it never ran. The generic branch matching any response.*.done event came first in the chain and matched this event too, so it fell through and returned the accumulated output unchanged, leaving the dedicated branch below unreachable since the feature was added.

Moving the dedicated branch above the generic one makes the event apply. On a compliant stream this changes nothing, since response.completed replaces the whole output with the same data straight afterwards. It matters when a provider is less tidy: one that never sends response.content_part.added leaves the assistant's own reply unextractable from the next turn's context, and one that omits response.content_part.done drops the annotations that only arrive with the finished item. Both are repaired by honouring the event.

Worth knowing: the item replaces whatever the deltas accumulated, with no guard against a provider sending back less than it streamed. A reasoning item arriving without its content would therefore lose the reasoning body, which is the same shape of provider brokenness that #27800 already needed a guard for.
2026-08-08 18:33:39 -06:00
Classic298
a39126c27c
fix: catch the socket.io timeout in the event caller (#28311)
An interactive prompt raised by __event_call__ was meant to come back as an error dictionary when it timed out. It never did: sio.call raises socketio.exceptions.TimeoutError, which does not inherit from the builtin TimeoutError the handler was catching, so the exception escaped into plugin code instead. Because that exception carries no message, the call sites that wrap plugin calls in except Exception as e turned it into an empty string, so a timed-out prompt looked like an empty answer rather than a failure, and the error branches written for it were dead.

The handler now catches socketio's class alongside the builtin, so a timeout returns the intended error dictionary and a plugin can tell the two apart.

The session eviction that sat inside that handler is removed rather than switched on. It had never executed, and it is wrong in both directions: it compares the pool entry by value, which the heartbeat rewrites every thirty seconds, so it would usually not fire, and when it did fire on a short timeout it would evict a live tab whose user had simply not answered yet, with nothing to restore the entry short of a reload. Genuinely dead sessions are already reaped on missed heartbeats by periodic_session_pool_cleanup.

WEBSOCKET_EVENT_CALLER_TIMEOUT is unset by default, which means no timeout at all, so this only affects deployments that set it.
2026-08-08 18:33:12 -06:00
Classic298
fc8a9b8ed6
fix: stop the Responses delta handler falling through to a crash (#28312)
The generic response.*.delta branch could leave the streaming handler in two states that crash the caller. It bound its result only inside the guard that checks the target item exists, but returned that result outside the guard, so a delta arriving before its output item, or carrying an index past the end, raised UnboundLocalError. Separately, an event name with only two dot-separated parts failed the length check and fell off the end of the branch, so the function returned None and both call sites raised TypeError unpacking it.

Where the response is streamed to a browser both crashes were swallowed at debug level and cost a chunk. On the direct API path there is no handler between here and the server, so the caller kept its 200 while the body was cut short with no [DONE], and the outlet filters never ran.

The return now sits inside the guard with a branch-level fallback that hands back the accumulated output untouched, which is what the sibling done branch and every other skip path in this function already do. Deltas whose item exists behave exactly as before.

Dropping an orphan delta is deliberate rather than synthesizing the missing item: response.output_item.added appends without regard to output_index, so a placeholder would be duplicated when the real item arrives, and a fabricated function_call would have no name or call id.
2026-08-08 18:33:00 -06:00
Timothy Jaeryang Baek
009999f363 refac 2026-08-08 15:47:10 -06:00
Timothy Jaeryang Baek
8faaf2cd1e refac
Some checks failed
Python CI / Ruff Format (3.11) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
2026-08-05 10:37:38 -05:00
Timothy Jaeryang Baek
c1c07cbe0f refac 2026-08-05 07:57:55 -05:00
Timothy Jaeryang Baek
6c4d0ace16 refac 2026-08-05 07:44:06 -05:00
Timothy Jaeryang Baek
29eeda9f9a refac 2026-08-05 07:12:28 -05:00
Timothy Jaeryang Baek
9c7ce154e7 refac 2026-08-05 07:04:30 -05:00
Timothy Jaeryang Baek
cbb3aade2b refac 2026-08-05 06:41:30 -05:00
Timothy Jaeryang Baek
0800c21c64 refac 2026-08-05 00:47:49 -05:00
Timothy Jaeryang Baek
8dbbc206c5 refac
Some checks are pending
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
2026-08-04 00:41:26 -05:00
Classic298
2d18727ab8
perf: build info log messages lazily so raising the log level actually saves work (#27837)
Some checks failed
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
Raising GLOBAL_LOG_LEVEL to WARNING buys quieter output but not less work: 241 INFO call sites interpolate their payload into an f-string before the logging call gets to drop it. The heaviest is get_doc, which logs every chunk id and metadata dict in a collection, so on the full-context retrieval path that is the entire knowledge base, once per chat request.

That one line at WARNING, CPython 3.12:

| knowledge base | payload | before   | after   |
| -------------- | ------- | -------- | ------- |
| top-k of 3     | 1.2 kB  | 3.8 us   | 0.07 us |
| 500 chunks     | 201 kB  | 583.6 us | 0.08 us |
| 5000 chunks    | 2.0 MB  | 5.8 ms   | 0.15 us |

The lazy form log.info('query_doc:result %s %s', result.ids, result.metadatas) hands the payload to record.getMessage(), which the InterceptHandler only reaches once a record has passed the level check. Output at INFO is byte-identical. Two sites that already built their message eagerly, one str concat and one % operator, move to the same lazy form.
2026-08-02 15:39:10 -05:00
Classic298
52145eede9
perf: take the orjson fast path for ensure_ascii=False callers (#27841)
Some checks are pending
Python CI / Ruff Format (3.12) (push) Waiting to run
Python CI / Ruff Format (3.11) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
2026-07-31 20:39:10 -05:00
Classic298
615807ad0b
chore: remove unused json import in models/chats.py (#27796)
Some checks failed
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Python CI / Ruff Format (3.11) (push) Waiting to run
Python CI / Ruff Format (3.12) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Blocked by required conditions
Create and publish Docker images with specific build args / notify-helm-charts (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Blocked by required conditions
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Blocked by required conditions
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
`backend/open_webui/models/chats.py` imports `json`, but the module contains no `json.` references. The only remaining matches in the file are SQL function names such as `json_each` and `json_typeof` inside `text()` strings, which are unrelated to the import.

Noticed while profiling the chat read/write path: the import made it look as though the module serialized locally, when all of that happens in `utils/misc.py:sanitize_data_for_db`.

One-line deletion, no behaviour change.
2026-07-31 19:09:41 -05:00
Classic298
798f3935ae
fix: keep streamed Responses output when response.completed reports an empty output array (#27800)
The `response.completed` handler replaced the accumulated output with the terminal event's `output` whenever that key was present, guarded only by `is not None`. An empty array satisfies that guard, so a provider that finishes the stream with `"output": []` wiped everything collected from `response.output_item.added`, `response.output_text.delta` and `response.output_item.done`.

The assistant message was then persisted with `output: []` and empty content, which shows up as a reply that renders correctly while streaming and disappears the moment the stream ends.

Fall back to the accumulated output when the terminal array is empty. A spec-compliant `response.completed` still wins, since a populated array is truthy, and when nothing was streamed the accumulated output is empty too, so the fallback cannot invent content.

Fixes #27789
2026-07-31 19:09:32 -05:00
Classic298
ac8af4996c
perf: index group_member on (user_id, group_id) (#27822)
Permission checks are the most repeated database work in a request, and every one of them asks the same question: which groups is this user in. Today that question cannot use an index.

`group_member` has only its primary key and a `(group_id, user_id)` unique constraint. That constraint leads on `group_id`, so a lookup by `user_id` has to walk the entire membership table, every time. `Groups.get_groups_by_member_id` sits under `has_permission`, `has_access`, `check_model_access` and the `AccessGrants` fallbacks, so an ordinary chat completion pays that walk several times before the model is even called, and the admin user list pays it once per row.

The cost scales with total memberships across all users rather than with the size of any one user's, so it stays invisible on a small instance and then arrives all at once on a large one.

Measured on SQLite, timing the real join from `get_groups_by_member_id`:

| memberships | before | after |
|---|---|---|
| 5,000 | 0.04 ms | 0.03 ms |
| 50,000 | 0.10 ms | 0.04 ms |
| 200,000 | 1.33 ms | 0.04 ms |
| 500,000 | 2.94 ms | 0.04 ms |

The after column is flat because the lookup becomes a seek instead of a scan. Concretely: on a deployment with 500k memberships, say 10,000 users in 50 groups each, one chat completion currently spends roughly 15 ms of database time answering the same question over and over. Afterwards it is under 0.2 ms. On a small install you will not be able to measure the difference, and that is fine, the point is that the curve stops bending.

The index is `(user_id, group_id)`. The trailing column makes those lookups index-only, since `group_id` is the column they select. Queries that lead on `group_id`, such as `get_group_user_ids_by_id` and the `chat_messages` subqueries, are already served by the existing unique constraint and are unaffected.

What to expect when the migration runs: on PostgreSQL this is a plain `CREATE INDEX`, which takes a SHARE lock, so reads continue while writes to `group_member` block until it completes. The table holds one row per membership, so expect sub-second even on the numbers above. `CONCURRENTLY` cannot be used here because the migration runner wraps the upgrade in a transaction, and it is not warranted at this table size.
2026-07-31 19:09:15 -05:00
Classic298
52cfb02c72
perf: build debug log messages lazily so disabled debug logs cost nothing (#27834)
GLOBAL_LOG_LEVEL defaults to INFO, so every log.debug(...) in the backend is discarded, but the message is built first: 187 call sites interpolate their payload into an f-string before the logging call runs, so the work happens on every request and the result is thrown away. The worst one sits in process_chat_payload and stringifies the whole request body, full conversation history included, once per chat completion.

That one line with DEBUG disabled, CPython 3.12:

| conversation | payload | before   | after   |
| ------------ | ------- | -------- | ------- |
| 4 messages   | 1.2 kB  | 3.4 us   | 0.07 us |
| 20 messages  | 17 kB   | 24.8 us  | 0.07 us |
| 60 messages  | 123 kB  | 216.6 us | 0.07 us |

The lazy form log.debug('form_data: %s', form_data) hands the payload to record.getMessage(), which the InterceptHandler only reaches once a record has passed the level check. With DEBUG enabled the emitted lines are byte-identical, f'{x=}' sites included: those map to %r. MistralLoader._debug_log callers get the same treatment, since that wrapper already forwards *args.
2026-07-31 19:09:01 -05:00
G30
b67804f2b6
fix: do not auto-open artifacts from the sidebar chat hover preview (#27773) 2026-07-31 17:58:02 -04:00
Timothy Jaeryang Baek
78ed5a0235 refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-07-31 17:45:07 -04:00
Timothy Jaeryang Baek
bb0f898b43 refac 2026-07-31 17:41:14 -04:00
Timothy Jaeryang Baek
3becec6ccf refac 2026-07-31 17:35:34 -04:00
Timothy Jaeryang Baek
5b333d75c6 refac 2026-07-31 17:34:39 -04:00
Classic298
ec0e60033b
perf: use the orjson codec for the permission deep copy (#27807)
`get_permissions` deep-copies the default permission tree with a `json.loads(json.dumps(...))` round trip before merging group permissions into it. It runs on signin, signup, the permissions endpoint, OAuth, and the chat-completion middleware.

It now goes through `JSONCodec`, which selects orjson when `ENABLE_ORJSON` is set. The intermediate string never leaves the expression, so neither the escaping nor the separator differences between the two backends are observable; only the resulting object is used.

`default_permissions` always originates from `Config.get('user.permissions')`, a SQLAlchemy `JSON` column, so the tree is JSON-native by construction and the round trip is exact.

Note for anyone tempted to simplify this to `copy.deepcopy`: measured on the real `DEFAULT_USER_PERMISSIONS` shape over 200k iterations, `deepcopy` takes 3.51s against 1.45s for the stdlib round trip and 0.40s for orjson. The round trip is the fast option, not a workaround.

With `ENABLE_ORJSON` unset, which is the default, `JSONCodec` is stdlib `json` and this call site behaves exactly as before.
2026-07-31 17:32:14 -04:00
Classic298
e2221fb662
perf: use the orjson codec to parse Jupyter kernel messages (#27812)
The code interpreter parses every message from the Jupyter kernel websocket with stdlib `json`, in a loop that runs for the duration of an execution. Messages carrying large stdout or a base64 image payload are the expensive ones.

It now goes through `JSONCodec`, which selects orjson when `ENABLE_ORJSON` is set. Every consumer of the parsed message reads strings only: `content.text`, `content.data['text/plain']` and `['image/png']`, `content.traceback`, and `content.execution_state`. Jupyter renders large integers into `text/plain` as strings rather than JSON numbers, so no numeric round trip is involved.

The one-shot `execute_request` message this module sends keeps stdlib `json`; it is a small fixed-shape dict sent once per execution.

With `ENABLE_ORJSON` unset, which is the default, `JSONCodec` is stdlib `json` and this call site behaves exactly as before.
2026-07-31 17:32:05 -04:00
Classic298
d03e9af0b7
perf: use the orjson codec to parse Oracle vector metadata (#27813)
`_json_to_metadata` parses the metadata of every result row returned by search and get. It now goes through `JSONCodec`, which selects orjson when `ENABLE_ORJSON` is set.

The text it parses is produced by Oracle's own `JSON_SERIALIZE`, and the column is a native `JSON` type, so the database normalises whatever was written and the reader never depends on the writer's escaping.

The matching `_metadata_to_json` write deliberately keeps stdlib `json`: it passes `default=self._decimal_handler`, orjson accepts none of stdlib's keyword arguments, and dropping the handler would turn a currently successful insert of a `Decimal` into a hard failure. The read side has no such constraint.

With `ENABLE_ORJSON` unset, which is the default, `JSONCodec` is stdlib `json` and this call site behaves exactly as before.
2026-07-31 17:31:14 -04:00
Timothy Jaeryang Baek
d721b0d196 refac 2026-07-31 17:30:47 -04:00
Classic298
ace84b4ae9
perf: use the orjson codec for outbound Ollama request bodies (#27811)
`routers/ollama.py` serializes the outbound body with stdlib `json` on six inference paths: `/api/chat`, the OpenAI-compatible completions and chat completions proxies, embeddings, the Anthropic messages proxy, and responses. All six carry a full conversation or an embedding batch.

They now go through `JSONCodec`, which selects orjson when `ENABLE_ORJSON` is set. Every one is passed to `send_request`, which hands it to aiohttp as `data=`; aiohttp encodes `str` as UTF-8 and derives `Content-Length` from the encoded bytes. None is hashed, cached, length-measured or persisted.

Admin model management keeps stdlib: `/api/unload`, `/api/pull`, `/api/delete` and `/api/show` serialize fixed one- or two-key dicts, as do the blob download and upload progress events and the error frames. Codec dispatch on those costs about what it saves.

With `ENABLE_ORJSON` unset, which is the default, `JSONCodec` is stdlib `json` and this call site behaves exactly as before.
2026-07-31 17:26:38 -04:00
Classic298
006a63e641
perf: use the orjson codec for the Anthropic passthrough request body (#27810)
`passthrough_anthropic_messages` in `main.py` serializes the full request payload with stdlib `json` before sending it upstream. It is the largest single serialization on that path, since the body carries the whole conversation.

It now goes through `JSONCodec`, which selects orjson when `ENABLE_ORJSON` is set. The result is passed to aiohttp as `data=`, which encodes `str` as UTF-8 and sets `Content-Length` from the encoded bytes. The payload originates from a parsed request dict, so it holds only JSON-native types, and the serialized string is never hashed, compared or persisted.

The remaining stdlib `json` calls in this module are left alone: two are a debug log line and a fixed Ollama unload payload, and one parses an upstream error body.

With `ENABLE_ORJSON` unset, which is the default, `JSONCodec` is stdlib `json` and this call site behaves exactly as before.
2026-07-31 17:26:07 -04:00
Classic298
243a39dc9d
perf: read the model pool with one HGETALL instead of one HGET per model (#27821)
`request.app.state.MODELS` is a `RedisDict` when Redis is configured. Unpacking it with `{**pool}` makes Python call `keys()` and then `__getitem__` once per key, which is one HKEYS plus one HGET per model, issued sequentially through a synchronous client. At 200 models that is 201 blocking Redis round trips per call.

`RedisDict.items()` is a single HGETALL, so `dict(pool.items())` fetches the same data in one round trip. `utils/chat.py:184` already does exactly this and carries a comment explaining why; these ten call sites were missed.

They are on the direct-connection branch of the task endpoints (title, tags, follow-up, autocomplete, query generation and the rest), of `chat_completed`, and of context compaction, so they run for background tasks fired on ordinary chat turns.

Behaviour is unchanged. The merged mapping is identical, the explicitly added direct model still overrides any pool entry with the same id, and when Redis is not configured the pool is a plain dict where `dict(d.items())` and `{**d}` are equivalent.

It also closes a race. `RedisDict.set` writes with HSET and then HDELs the stale keys, so a key returned by HKEYS could be deleted before its HGET arrived, raising `KeyError` out of the dict literal and failing the request mid model refresh. The old path could likewise observe a mix of pre- and post-refresh entries. HGETALL is atomic, so the caller now always sees one coherent snapshot.
2026-07-31 17:25:53 -04:00
Classic298
6be11d4fc9
chore: remove dead json imports (#27815)
Fourteen modules import `json` without using it. Ruff flags every one with F401, and a word-boundary search for `json` in each file matches only the import line itself, including inside strings, comments and annotations.

Two exclusions, both deliberate. Migration files are left alone: the import is equally dead there, but those files are frozen history and not worth the churn. `models/chats.py` has the same dead import and is handled in its own change, so it is skipped here to avoid two changes touching the same line.

No behaviour change.
2026-07-31 17:25:40 -04:00
Classic298
1d6735ff0b
fix: escape line separators in orjson output (#27819)
orjson emits U+2028, U+2029 and U+0085 raw, where stdlib `json.dumps` escapes them under its default `ensure_ascii=True`. Python treats all three as line boundaries, so with `ENABLE_ORJSON` set, one of them inside model output splits a `data: {...}` SSE frame in half. Both halves then fail to parse and the delta is dropped with no error.

`utils/middleware.py` reassembles frames with `splitlines()`, so an affected response silently loses content on the direct API path. External clients are exposed as well: httpx's `LineDecoder` reimplements the same line-boundary semantics, so any SDK reading the OpenAI-compatible stream through `aiter_lines` breaks on a raw separator.

The three characters are escaped on the way out of `ORJSONCodec.dumps`. That restores parity with stdlib and fixes every reader at once, rather than patching one consumer and leaving external clients broken. They are the complete set: of the ten code points `splitlines()` treats as boundaries, the other seven are below U+0020, where JSON already forces an escape.

The membership guard is load bearing. Calling `translate` unconditionally costs roughly 1.5 us on a typical SSE chunk against 0.115 us for the serialization it wraps, so it would spend more than orjson saves. The three scans cost about 0.04 us.

Payloads containing none of the three are returned unchanged, byte for byte. With `ENABLE_ORJSON` unset, which is the default, none of this code runs.

U+2028 and U+2029 are common in text extracted from PDFs and word processor documents, so the realistic trigger is a model quoting an uploaded file back to the user.
2026-07-31 17:25:30 -04:00
joaoback
8d333335b9
i18n: add pt-BR translations for newly added UI items and consistency pass (#27814)
New **pt-BR** translations for items introduced in the latest releases, plus a consistency/quality pass across existing strings (grammar, tone, capitalization, pluralization). Placeholders and hotkeys preserved. No logic changes.
2026-07-31 17:25:10 -04:00
Classic298
f4c6a76651
perf: use the orjson codec for Valkey vector metadata (#27805)
The Valkey backend serializes chunk metadata on every insert and parses it back on every result row in `get` and `query`. Both directions now go through `JSONCodec`, which selects orjson when `ENABLE_ORJSON` is set.

The stored `metadata_json` field is never matched against as text. `_build_filter_expression` only emits TAG predicates, and the TAG fields are `id`, `hash`, `file_id`, `source` and `knowledge_base_id`; `metadata_json` appears only as a return field that is immediately re-parsed. So rows written with escaped non-ASCII and rows written raw are indistinguishable to every reader, and no migration is needed.

`process_metadata` already stringifies datetimes and strips null bytes and lone surrogates before the write, so the two backends cannot disagree about what is serializable here.

Both read `except` clauses widen from `(json.JSONDecodeError, TypeError)` to `(ValueError, TypeError)`. The codec falls back to engineio's codec, which installs `parse_int=_safe_int` and raises a bare `ValueError` for integer literals longer than 100 characters; the narrower clause would have let that escape and abort a search instead of yielding empty metadata. `json.JSONDecodeError` is a `ValueError` subclass, so this is a strict superset. That removes the module's last use of stdlib `json`, so the import goes with it.

With `ENABLE_ORJSON` unset, which is the default, `JSONCodec` is stdlib `json` and this call site behaves exactly as before.
2026-07-31 17:24:57 -04:00
Classic298
466e05801b
perf: stop query_collection blocking the event loop (#27824)
RAG vector search runs in a thread pool, but then calls `future.result()` on the event loop thread, so the whole worker freezes until every collection answers. Every other user's token stream stops for that long. It's the default retrieval path.

Now `asyncio.gather` over `asyncio.to_thread`, matching what `routers/retrieval.py:2779` already does for the same call.

Measured with 3 queries across 4 collections, 60 ms search, and a second request wanting a turn every 5 ms:

| | before | after |
|---|---|---|
| RAG call | 62.0 ms | 61.2 ms |
| other request's turns | 0 | 7 |
| its worst stall | 62.5 ms | 16.0 ms |

Same results, same order, same `(result, error)` contract. Cancellation now lands mid-search instead of after every thread finishes. Threads move from an unbounded per-call pool to the loop's bounded shared one.
2026-07-31 17:24:31 -04:00
Timothy Jaeryang Baek
810378c0b8 refac
Some checks failed
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
2026-07-27 19:39:36 -04:00
Timothy Jaeryang Baek
b6b16d5871 refac 2026-07-27 19:24:03 -04:00
Tim Baek
01f4282f1f
Merge pull request #27590 from open-webui/dev
Some checks failed
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Frontend Build / Unit Tests (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Has been cancelled
Frontend Build / Format & Build (push) Has been cancelled
Release to PyPI / release (push) Has been cancelled
Release / publish (push) Has been cancelled
Python CI / Ruff Format (3.11) (push) Has been cancelled
Python CI / Ruff Format (3.12) (push) Has been cancelled
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Has been cancelled
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Has been cancelled
Create and publish Docker images with specific build args / notify-helm-charts (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Has been cancelled
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Has been cancelled
refac
2026-07-27 05:47:11 -05:00
Timothy Jaeryang Baek
483adf7040 refac 2026-07-27 06:46:42 -04:00
672 changed files with 152639 additions and 95601 deletions

View file

@ -18,3 +18,6 @@ uploads
**/*.db **/*.db
_test _test
backend/data/* backend/data/*
.venv
.git

View file

@ -16,6 +16,15 @@ CORS_ALLOW_ORIGIN='*'
# Set to false to keep memory tools enabled without adding memory context to the system context. # Set to false to keep memory tools enabled without adding memory context to the system context.
ENABLE_MEMORY_SYSTEM_CONTEXT=true ENABLE_MEMORY_SYSTEM_CONTEXT=true
# Set to true to add compact row/column stats to parsed CSV retrieval context.
ENABLE_RAG_CSV_SUMMARY=false
# Set to true to preserve backing file records, storage blobs, and per-file vectors when files are removed from knowledge bases.
ENABLE_KNOWLEDGE_FILE_RETENTION=false
# Comma-separated chunk metadata keys to expose to the model alongside retrieved content.
RAG_SOURCE_METADATA_KEYS=''
# Set to false to disable workspace Tools and Functions. # Set to false to disable workspace Tools and Functions.
ENABLE_PLUGINS=true ENABLE_PLUGINS=true

View file

@ -1,41 +1,35 @@
name: Bug Report name: Bug Report
description: Create a detailed bug report to help us improve Open WebUI. description: Tell us what broke in Open WebUI.
title: 'issue: ' title: 'issue: '
labels: ['bug', 'triage'] labels: ['bug', 'triage']
assignees: []
body: body:
- type: markdown - type: markdown
attributes: attributes:
value: | value: |
# Bug Report # Bug Report
## Important Notes Use this for real, reproducible bugs. A clear issue is the most useful contribution: include the affected workflow, the expected result, the actual result, and the details needed for someone else to reproduce it.
- **Before submitting a bug report**: Please check the [Issues](https://github.com/open-webui/open-webui/issues) and [Discussions](https://github.com/open-webui/open-webui/discussions) sections to see if a similar issue has already been reported. If unsure, start a discussion first, as this helps us efficiently focus on improving the project. Duplicates may be closed without notice. **Please search for existing issues AND discussions. No matter open or closed.** Before submitting, search open and closed [Issues](https://github.com/open-webui/open-webui/issues) and [Discussions](https://github.com/open-webui/open-webui/discussions). The issue may already be reported or fixed on `dev`.
- Check for opened, **but also for (recently) CLOSED issues** as the issue you are trying to report **might already have been fixed on the dev branch!** **Test on the latest release AND on `dev`, right before you submit this report, not last week.** A huge share of reports are for bugs already fixed on `dev`, sometimes weeks earlier, because the reporter only tested an old version and never rechecked. Reports that don't reproduce on latest and on current `dev` at submission time will be closed without further discussion, no exceptions.
- **Respectful collaboration**: Open WebUI is a volunteer-driven project with a single maintainer and contributors who also have full-time jobs. Please be constructive and respectful in your communication. Please do not open a code pull request for this report unless a maintainer asks for one, or the change is only i18n/localization. If you want to share code as reference, include it here as a local diff or patch. Actionable reproduction details are the most useful next step.
- **Contributing**: If you encounter an issue, consider submitting a pull request or forking the project. We prioritize preventing contributor burnout to maintain Open WebUI's quality. Security vulnerabilities must not be reported publicly. Use the [GitHub security page](https://github.com/open-webui/open-webui/security) instead.
- **Bug Reproducibility**: If a bug cannot be reproduced using a `:main` or `:dev` Docker setup or with `pip install` on Python 3.11, community assistance may be required. In such cases, we will move it to the "[Issues](https://github.com/open-webui/open-webui/discussions/categories/issues)" Discussions section. Your help is appreciated!
- **Scope**: If you want to report a SECURITY VULNERABILITY, then do so through our [GitHub security page](https://github.com/open-webui/open-webui/security).
- type: checkboxes - type: checkboxes
id: issue-check id: issue-check
attributes: attributes:
label: Check Existing Issues label: Before Submitting
description: Confirm that you’ve checked for existing reports before submitting a new one.
options: options:
- label: I have searched for any existing and/or related issues. - label: I searched open and closed issues and discussions for an existing report.
required: true required: true
- label: I have searched for any existing and/or related discussions. - label: I reproduced this bug on the latest release AND on the current `dev` branch, right before submitting this report. I did not just check an old version or rely on a check from days ago.
required: true required: true
- label: I have also searched in the CLOSED issues AND CLOSED discussions and found no related items (your issue might already be addressed on the development branch!). - label: I understand that maintainers want a well-written issue before any code pull request.
required: true required: true
- label: I am using the latest version of Open WebUI. - label: This is not a security vulnerability.
required: true required: true
- type: dropdown - type: dropdown
@ -44,9 +38,9 @@ body:
label: Installation Method label: Installation Method
description: How did you install Open WebUI? description: How did you install Open WebUI?
options: options:
- Git Clone
- Pip Install
- Docker - Docker
- Pip Install
- Git Clone
- Other - Other
validations: validations:
required: true required: true
@ -55,67 +49,51 @@ body:
id: open-webui-version id: open-webui-version
attributes: attributes:
label: Open WebUI Version label: Open WebUI Version
description: Specify the version (e.g., v0.6.26) description: Specify the version, commit, or image tag.
placeholder: v0.11.0, dev commit SHA, or Docker tag
validations: validations:
required: true required: true
- type: input
id: ollama-version
attributes:
label: Ollama Version (if applicable)
description: Specify the version (e.g., v0.2.0, or v0.1.32-rc1)
validations:
required: false
- type: input - type: input
id: operating-system id: operating-system
attributes: attributes:
label: Operating System label: Operating System
description: Specify the OS (e.g., Windows 10, macOS Sonoma, Ubuntu 22.04, Debian 12) description: Specify the OS and version.
placeholder: Windows 11, macOS Tahoe, Ubuntu 26.04, Debian 13
validations: validations:
required: true required: true
- type: input - type: input
id: browser id: browser
attributes: attributes:
label: Browser (if applicable) label: Browser
description: Specify the browser/version (e.g., Chrome 100.0, Firefox 98.0) description: If the bug appears in the browser, include browser and version.
placeholder: Chrome 151.0, Firefox 153.0.3
validations: validations:
required: false required: false
- type: checkboxes - type: input
id: confirmation id: ollama-version
attributes: attributes:
label: Confirmation label: Ollama Version
description: Ensure the following prerequisites have been met. description: Include this if Ollama is involved.
options: placeholder: v0.32.5
- label: I have read and followed all instructions in `README.md`. validations:
required: true required: false
- label: I am using the latest version of **both** Open WebUI and Ollama.
required: true - type: textarea
- label: I have included the browser console logs. id: summary
required: true attributes:
- label: I have included the Docker container logs. label: Summary
required: true description: What is wrong, in a few sentences?
- label: I have **provided every relevant configuration, setting, and environment variable used in my setup.** validations:
required: true required: true
- label: I have clearly **listed every relevant configuration, custom setting, environment variable, and command-line option that influences my setup** (such as Docker Compose overrides, .env values, browser settings, authentication configurations, etc).
required: true
- label: |
I have documented **step-by-step reproduction instructions that are precise, sequential, and leave nothing to interpretation**. My steps:
- Start with the initial platform/version/OS and dependencies used,
- Specify exact install/launch/configure commands,
- List URLs visited, user input (incl. example values/emails/passwords if needed),
- Describe all options and toggles enabled or changed,
- Include any files or environmental changes,
- Identify the expected and actual result at each stage,
- Ensure any reasonably skilled user can follow and hit the same issue.
required: true
- type: textarea - type: textarea
id: expected-behavior id: expected-behavior
attributes: attributes:
label: Expected Behavior label: Expected Behavior
description: Describe what should have happened. description: What should have happened?
validations: validations:
required: true required: true
@ -123,7 +101,7 @@ body:
id: actual-behavior id: actual-behavior
attributes: attributes:
label: Actual Behavior label: Actual Behavior
description: Describe what actually happened. description: What actually happened?
validations: validations:
required: true required: true
@ -131,32 +109,21 @@ body:
id: reproduction-steps id: reproduction-steps
attributes: attributes:
label: Steps to Reproduce label: Steps to Reproduce
description: | description: Include the exact commands, settings, URLs, model/provider setup, and user actions needed to hit the bug.
Please provide a **very detailed, step-by-step guide** to reproduce the issue. Your instructions should be so clear and precise that anyone can follow them without guesswork. Include every relevant detail—settings, configuration options, exact commands used, values entered, and any prerequisites or environment variables.
**If full reproduction steps and all relevant settings are not provided, your issue may not be addressed.**
**If your steps to reproduction are incomplete, lacking detail or not reproducible, your issue can not be addressed.**
placeholder: | placeholder: |
Example (include every detail): 1. Start Open WebUI with ...
1. Start with a clean Ubuntu 22.04 install. 2. Configure ...
2. Install Docker v24.0.5 and start the service. 3. Open ...
3. Clone the Open WebUI repo (git clone ...). 4. Click ...
4. Use the Docker Compose file without modifications. 5. See ...
5. Open browser Chrome 115.0 in incognito mode.
6. Go to http://localhost:8080 and log in with user "test@example.com".
7. Set the language to "English" and theme to "Dark".
8. Attempt to connect to Ollama at "http://localhost:11434".
9. Observe that the error message "Connection refused" appears at the top right.
Please list each step carefully and include all relevant configuration, settings, and options.
validations: validations:
required: true required: true
- type: textarea - type: textarea
id: logs-screenshots id: logs-screenshots
attributes: attributes:
label: Logs & Screenshots label: Logs, Screenshots, and Config
description: Include relevant logs, errors, or screenshots to help diagnose the issue. description: Include relevant browser console logs, server/container logs, screenshots, and configuration. If something does not apply, say so.
placeholder: 'Attach logs from the browser console, Docker logs, or error messages.'
validations: validations:
required: true required: true
@ -164,13 +131,6 @@ body:
id: additional-info id: additional-info
attributes: attributes:
label: Additional Information label: Additional Information
description: Provide any extra details that may assist in understanding the issue. description: Anything else that might help us understand the report.
validations: validations:
required: false required: false
- type: markdown
attributes:
value: |
## Note
**If the bug report is incomplete, does not follow instructions or is lacking details it may not be addressed.** Ensure that you've followed all the **README.md** and **troubleshooting.md** guidelines, and provide all necessary information for us to reproduce the issue.
Thank you for contributing to Open WebUI!

View file

@ -1,76 +1,73 @@
name: Feature Request name: Feature Request
description: Suggest a new feature or improvement description: Describe what you would like Open WebUI to support.
title: 'feat: ' title: 'feat: '
labels: ['triage'] labels: ['triage']
body: body:
- type: markdown - type: markdown
attributes: attributes:
value: | value: |
## Before Submitting # Feature Request
Please check **open AND closed** [Issues](https://github.com/open-webui/open-webui/issues) and [Discussions](https://github.com/open-webui/open-webui/discussions) for similar requests. If you find one, add your input there instead. Describe the requested behavior, the problem it solves, and any examples, mockups, screenshots, or workflows that clarify the request. A clear issue or discussion is the most useful contribution.
### Scope Guidelines Search open and closed [Issues](https://github.com/open-webui/open-webui/issues) and [Discussions](https://github.com/open-webui/open-webui/discussions) before submitting. If the request needs broad product, UX, architecture, compatibility, or maintenance discussion, please start in [Discussions](https://github.com/open-webui/open-webui/discussions) so the community can weigh in.
Feature requests that require significant implementation effort should be posted in the **Ideas** section of [Discussions](https://github.com/open-webui/open-webui/discussions) instead. We will move oversized feature requests to Discussions to keep the Issues tab focused on actionable items. Please do not open a code pull request for this request unless a maintainer asks for one, or the change is only i18n/localization. If you want to share code as reference, include it here as a local diff or patch. Clear product context is the most useful next step.
If your request might impact the broader community, please open a Discussion first so others can weigh in on the design. Security vulnerabilities must not be reported publicly. Use the [GitHub security page](https://github.com/open-webui/open-webui/security) instead.
### Be Respectful
Open WebUI is a volunteer-driven project maintained by a small team. We value constructive, positive communication. Please be mindful of maintainers' time and energy.
### Contributing
If you encounter an issue, we encourage you to submit a pull request or fork the project. We actively work to prevent contributor burnout and maintain project quality.
### Reproducibility
If a bug cannot be reproduced with a `:main` or `:dev` Docker setup, or a `pip install` with Python 3.11, it may be moved to the "Issues" section in Discussions for community assistance.
- type: checkboxes - type: checkboxes
id: existing-issue id: existing-request
attributes: attributes:
label: Check Existing Issues label: Before Submitting
description: Confirm you have searched for similar requests.
options: options:
- label: I have searched all existing **open AND closed** issues and discussions and found none comparable to my request. - label: I searched open and closed issues and discussions for an existing request.
required: true required: true
- label: I checked whether this already exists on the `dev` branch or latest source.
- type: checkboxes required: true
id: feature-scope - label: I understand that maintainers want a well-written issue or discussion before any code pull request.
attributes: required: true
label: Verify Feature Scope - label: This request is not a security vulnerability.
description: Confirm this request belongs in Issues rather than Discussions.
options:
- label: I believe this feature request is appropriately scoped for the Issues section as described above.
required: true required: true
- type: textarea - type: textarea
id: problem-description id: problem-description
attributes: attributes:
label: Problem Description label: Problem
description: Is this related to a problem? Describe the pain point clearly. description: What is missing, frustrating, confusing, or unnecessarily hard today?
placeholder: "e.g., I'm frustrated when..." placeholder: "I'm trying to..., but..."
validations: validations:
required: true required: true
- type: textarea - type: textarea
id: solution-description id: desired-behavior
attributes: attributes:
label: Proposed Solution label: Desired Behavior
description: Describe what you would like to happen. description: What would you like to happen instead?
placeholder: "I would like Open WebUI to..."
validations: validations:
required: true required: true
- type: textarea
id: why-it-matters
attributes:
label: Why This Matters
description: Who benefits, and what workflow does this unlock or improve?
validations:
required: true
- type: textarea
id: examples
attributes:
label: Examples or References
description: Add mockups, screenshots, links, prompts, workflows, or examples from other tools.
validations:
required: false
- type: textarea - type: textarea
id: alternatives-considered id: alternatives-considered
attributes: attributes:
label: Alternatives Considered label: Alternatives or Workarounds
description: Describe any alternative solutions or workarounds you have considered. description: What have you tried instead, if anything?
validations:
- type: textarea required: false
id: additional-context
attributes:
label: Additional Context
description: Add any other context, mockups, or screenshots about the feature request.

View file

@ -1,104 +1,95 @@
<!-- <!--
⚠️ CRITICAL CHECKS FOR CONTRIBUTORS (READ, DON'T DELETE) ⚠️ Important checks for contributors:
1. Target the `dev` branch. PRs targeting `main` will be automatically closed. 1. DO NOT OPEN A CODE PULL REQUEST unless a maintainer explicitly asked you to, or the change is strictly limited to i18n/localization.
2. Do NOT delete the CLA section at the bottom. It is required for the bot to accept your PR. 2. Target the `dev` branch. PRs targeting `main` will be closed.
3. Do not delete the Contributor License Agreement section at the bottom. The CLA bot requires it.
--> -->
# Pull Request Checklist # Pull Request
### Note to first-time contributors: Please open a discussion post in [Discussions](https://github.com/open-webui/open-webui/discussions) to discuss your idea/fix with the community before creating a pull request, and describe your changes before submitting a pull request. **Do not open a code pull request unless a maintainer has explicitly requested it or the change is limited to i18n/localization.**
This is to ensure large feature PRs are discussed with the community first, before starting work on it. If the community does not want this feature or it is not relevant for Open WebUI as a project, it can be identified in the discussion before working on the feature and submitting the PR. The most useful way to help is to give us a clear understanding of the problem: report reproducible bugs in [Issues](https://github.com/open-webui/open-webui/issues) and share proposals in [Discussions](https://github.com/open-webui/open-webui/discussions). We use that context to evaluate solutions and refine the implementation internally, accounting for the broader codebase and ongoing work. External implementations usually require substantial reworking to fit the project's standards, and coordinating those revisions usually takes more effort than developing the solution internally. Please follow this process before investing time in a pull request. PRs opened outside these guidelines are generally closed without review.
<!-- ## Maintainer Request
### ⚠️ Important: Your PR is a contribution, not a guarantee of merge.
The most impactful way to contribute to Open WebUI is through well-written bug reports, detailed feature discussions, and thoughtful ideas. These directly shape the project. If you do open a pull request, please know that Open WebUI is held to the highest standard of code quality, consistency, and architectural coherence, and every line merged becomes something the core team must own, maintain, and support indefinitely. Submitted code may be refactored, rewritten, or used as inspiration for a different implementation. This is not a reflection of your work's quality. It is how we ensure that a small team can deeply understand and evolve every part of the codebase. Link the maintainer's request for this PR, or state that the change is limited to i18n/localization.
-->
**Before submitting, make sure you've checked the following:** ## Checklist
- [ ] **Linked Issue/Discussion:** This PR references an existing [Issue](https://github.com/open-webui/open-webui/issues) or [Discussion](https://github.com/open-webui/open-webui/discussions) — `Closes #___` / `Relates to #___`. If one does not exist, create one first. PRs without a linked issue or discussion may be closed without review. - [ ] I have read and I understand the [contribution policy](https://docs.openwebui.com/contributing/#submit-code).
- [ ] **Target branch:** The pull request targets the `dev` branch. **PRs targeting `main` will be immediately closed.** - [ ] This PR targets the `dev` branch.
- [ ] **Description:** A concise description of the changes is provided below. - [ ] This PR links to a well-described, confirmed Issue or active Discussion: `Closes #___` / `Relates to #___`.
- [ ] **Changelog:** A changelog entry following [Keep a Changelog](https://keepachangelog.com/) format is included at the bottom. - [ ] A maintainer explicitly asked me to open this PR, or this PR only updates i18n/localization.
- [ ] **Documentation:** Relevant documentation has been added or updated in the [Open WebUI Docs Repository](https://github.com/open-webui/docs). - [ ] The change is one logical unit with no unrelated commits.
- [ ] **Dependencies:** Any new or updated dependencies are explained, tested, and documented. - [ ] I matched nearby code patterns and avoided unnecessary new settings, abstractions, or dependencies.
- [ ] **Testing:** Manual tests have been performed to verify the fix/feature works correctly and does not introduce regressions. Screenshots or recordings are included where applicable. - [ ] I manually tested the changed workflow and any nearby behavior that could be affected.
- [ ] **No Unchecked AI Code:** This PR is either human-written or has undergone thorough human review AND manual testing. Unreviewed AI-generated PRs may be closed immediately. - [ ] I have not added or rewritten automated tests, fixtures, snapshots, or testing infrastructure unless a maintainer explicitly requested them.
- [ ] **Self-Review:** A self-review of the code has been performed, ensuring adherence to project coding standards. - [ ] I updated relevant docs, including the [Open WebUI Docs Repository](https://github.com/open-webui/docs), if needed.
- [ ] **Architecture:** Smart defaults are preferred over new settings. Local state is used for ephemeral UI logic. Major architectural or UX changes have been discussed first. - [ ] I added screenshots for UI changes, and a recording when motion or interaction matters.
- [ ] **Git Hygiene:** The PR is atomic (one logical change), rebased on `dev`, and contains no unrelated commits. - [ ] I reviewed any AI-generated code before submitting it.
- [ ] **Title Prefix:** The PR title uses one of the following prefixes: - [ ] The PR title uses one of the prefixes listed below.
- **BREAKING CHANGE**: Changes affecting backward compatibility
- **build**: Build system or dependency changes
- **ci**: CI/CD workflow changes
- **chore**: Refactoring, cleanup, or non-functional changes
- **docs**: Documentation additions or updates
- **feat**: New features or enhancements
- **fix**: Bug fixes or corrections
- **i18n**: Internationalization or localization changes
- **perf**: Performance improvements
- **refactor**: Code restructuring
- **style**: Formatting changes (whitespace, semicolons, etc.)
- **test**: Test additions or corrections
- **WIP**: Work in progress
# Changelog Entry ## Title Prefix
### Description Use one of the following prefixes:
- [Describe the changes, including motivation and impact] - **BREAKING CHANGE**: Changes affecting backward compatibility
- **build**: Build system or dependency changes
- **ci**: CI/CD workflow changes
- **chore**: Refactoring, cleanup, or non-functional changes
- **docs**: Documentation additions or updates
- **feat**: New features or enhancements
- **fix**: Bug fixes or corrections
- **i18n**: Internationalization or localization changes
- **perf**: Performance improvements
- **refactor**: Code restructuring
## Summary
Describe the change, the problem it solves, and the impact on users.
## Verification
Describe how you reproduced the problem and manually checked the behavior before and after the change. Include exact steps, setup details, and relevant logs, screenshots, or recordings. Report results from relevant existing checks and anything you could not verify.
Do not add or rewrite automated tests unless a maintainer explicitly requests them. Tests that repeat an implementation's assumptions can pass while preserving the same mistake; maintainers determine the regression coverage needed. Do not remove, disable, or weaken existing tests to make the change pass.
## Changelog Entry
### Added ### Added
- [New features, functionalities, or additions] -
### Changed ### Changed
- [Changes, updates, refactorings, or optimizations] -
### Deprecated
- [Deprecated functionality or features]
### Removed
- [Removed features, files, or functionalities]
### Fixed ### Fixed
- [Bug fixes or corrections] -
### Removed
-
### Security ### Security
- [Security-related changes or vulnerability fixes] -
### Breaking Changes ### Breaking Changes
- **BREAKING CHANGE**: [Changes affecting compatibility or functionality] -
--- ## Additional Context
### Additional Information Add anything maintainers should know before review.
- [Any additional context, notes, or references to related issues/commits] ## Contributor License Agreement
### Screenshots or Videos
- [Attach relevant screenshots or videos demonstrating the changes]
### Contributor License Agreement
<!-- <!--
🚨 DO NOT DELETE THE TEXT BELOW 🚨 DO NOT DELETE THIS SECTION.
Keep the "Contributor License Agreement" confirmation text intact. Your PR will not be reviewed or merged until you check the box below confirming that you have read and agree to the CLA.
Deleting it will trigger the CLA-Bot to INVALIDATE your PR.
Your PR will NOT be reviewed or merged until you check the box below confirming that you have read and agree to the terms of the CLA.
--> -->
- [ ] By submitting this pull request, I confirm that I have read and fully agree to the [Contributor License Agreement (CLA)](https://github.com/open-webui/open-webui/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT), and I am providing my contributions under its terms. - [ ] By submitting this pull request, I confirm that I have read and fully agree to the [Contributor License Agreement (CLA)](https://github.com/open-webui/open-webui/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT), and I am providing my contributions under its terms.
> [!NOTE]
> Deleting the CLA section will lead to immediate closure of your PR and it will not be merged in.

View file

@ -311,7 +311,6 @@ jobs:
runs-on: ubuntu-latest runs-on: ubuntu-latest
if: ${{ !cancelled() && (github.ref == 'refs/heads/main' || startsWith(github.ref, 'refs/tags/v')) }} if: ${{ !cancelled() && (github.ref == 'refs/heads/main' || startsWith(github.ref, 'refs/tags/v')) }}
needs: [merge] needs: [merge]
continue-on-error: true
strategy: strategy:
fail-fast: false fail-fast: false
matrix: matrix:
@ -403,5 +402,16 @@ jobs:
[ -z "$TAG" ] && continue [ -z "$TAG" ] && continue
DEST="${DOCKERHUB_IMAGE}:${TAG}" DEST="${DOCKERHUB_IMAGE}:${TAG}"
echo " -> ${DEST}" echo " -> ${DEST}"
docker buildx imagetools create -t "${DEST}" "${SOURCE}" for ATTEMPT in 1 2 3; do
if docker buildx imagetools create -t "${DEST}" "${SOURCE}" && \
docker buildx imagetools inspect "${DEST}"; then
break
fi
if [ "${ATTEMPT}" = "3" ]; then
echo "Failed to copy ${DEST} after ${ATTEMPT} attempts"
exit 1
fi
echo "Copy attempt ${ATTEMPT} for ${DEST} failed, retrying in 15s..."
sleep 15
done
done <<< "${{ steps.tags.outputs.tags }}" done <<< "${{ steps.tags.outputs.tags }}"

View file

@ -2,7 +2,7 @@ name: Issue Labeler
on: on:
issues: issues:
types: [opened] types: [opened, edited]
permissions: permissions:
issues: write issues: write
@ -22,22 +22,118 @@ jobs:
return; return;
} }
const isEdit = context.payload.action === 'edited';
const titleWasEdited = Boolean(context.payload.changes?.title);
if (isEdit && !titleWasEdited) {
return;
}
const title = issue.title ?? ''; const title = issue.title ?? '';
const body = issue.body ?? ''; const body = issue.body ?? '';
// Freeform bug reports: "issue: ...", "bug: ...", "fix: ...", "[Bug] ...", "issue/UX: ..." // Freeform bug reports: "issue: ...", "bug: ...", "fix: ...", "[Bug] ...", "issue/UX: ..."
const bugLikeTitle = /^\s*\[?(bug|issue|fix)\]?\s*[:/\-]/i.test(title); const bugLikeTitle = /^\s*(\[\s*(bug|issue|fix)\b[^\]]*\]|(bug|issue|fix)\s*[:/\-])/i.test(title);
// API/CLI-created issues that reproduce the bug report form structure. // API/CLI-created issues that reproduce the bug report form structure.
// Only headings distinctive to the bug form (both are required fields there) — // Only headings distinctive to the bug form (both are required fields there) —
// generic headings like "Expected Behavior" also appear in freeform feature requests. // generic headings like "Expected Behavior" also appear in freeform feature requests.
const bugFormBody = /###\s*(Installation Method|Open WebUI Version)/i.test(body); const bugFormBody = /###\s*(Installation Method|Open WebUI Version)/i.test(body);
if (bugLikeTitle || bugFormBody) { if (!bugLikeTitle && !bugFormBody) {
await github.rest.issues.addLabels({ return;
}
if (isEdit) {
const events = await github.paginate(github.rest.issues.listEvents, {
owner: context.repo.owner, owner: context.repo.owner,
repo: context.repo.repo, repo: context.repo.repo,
issue_number: issue.number, issue_number: issue.number,
labels: ['bug'] per_page: 100
}); });
const bugLabelWasRemoved = events.some(
(event) => event.event === 'unlabeled' && event.label?.name === 'bug'
);
if (bugLabelWasRemoved) {
return;
}
} }
await github.rest.issues.addLabels({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: issue.number,
labels: ['bug']
});
label-feature-requests:
runs-on: ubuntu-latest
steps:
- name: Add "enhancement" label to unlabeled feature requests
uses: actions/github-script@v7
with:
script: |
const issue = context.payload.issue;
if (issue.labels.some((label) => label.name === 'enhancement')) {
return;
}
// A human (or the bug form) already classified this as a bug;
// do not stack a second, contradictory classification on it.
if (issue.labels.some((label) => label.name === 'bug')) {
return;
}
const isEdit = context.payload.action === 'edited';
const titleWasEdited = Boolean(context.payload.changes?.title);
if (isEdit && !titleWasEdited) {
return;
}
const title = issue.title ?? '';
const body = issue.body ?? '';
// Feature requests: "feat: ...", "feature: ...", "feature request: ...",
// "enhancement: ...", "enh: ...", "[Feature Request] ..." — the feature
// request form titles every submission "feat: ", so form submissions are
// covered by the same pattern.
const featureLikeTitle =
/^\s*(\[\s*(feat|feature|enhancement|enh)\b[^\]]*\]|(feat|feature( request)?|enhancement|enh)\s*[:/\-])/i.test(
title
);
// API/CLI-created issues that reproduce the feature request form structure.
// Only headings distinctive to that form.
const featureFormBody = /###\s*(Proposed Solution|Alternatives Considered)/i.test(body);
if (!featureLikeTitle && !featureFormBody) {
return;
}
if (isEdit) {
const events = await github.paginate(github.rest.issues.listEvents, {
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: issue.number,
per_page: 100
});
const enhancementLabelWasRemoved = events.some(
(event) => event.event === 'unlabeled' && event.label?.name === 'enhancement'
);
if (enhancementLabelWasRemoved) {
return;
}
}
await github.rest.issues.addLabels({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: issue.number,
labels: ['enhancement']
});

45
.github/workflows/regression.yaml vendored Normal file
View file

@ -0,0 +1,45 @@
# ─────────────────────────────────────────────────────────────────────────────
# Tests — run the open-webui/tests unit suite against a release candidate
# Release pull requests go from dev into main and are titled with the version
# ─────────────────────────────────────────────────────────────────────────────
name: Tests
on:
pull_request:
branches: [main, dev]
types: [opened, synchronize, reopened, edited]
# An edit must not cancel a running suite: the replacement run would skip it and still report green.
concurrency:
group: regression-${{ github.ref }}
cancel-in-progress: ${{ github.event.action != 'edited' }}
jobs:
# Into main only the release branch counts; into dev any version title does, so the
# suite can be exercised outside a release. Expressions have no regex, hence the literal prefixes.
regression:
name: Suite
if: >-
(github.event.pull_request.base.ref == 'dev' ||
github.event.pull_request.head.ref == 'dev') &&
(github.event.action != 'edited' || github.event.changes.title != null) &&
(startsWith(github.event.pull_request.title, '0.') ||
startsWith(github.event.pull_request.title, '1.'))
permissions:
contents: read
uses: open-webui/tests/.github/workflows/regression.yml@main
with:
open-webui-ref: ${{ github.event.pull_request.head.sha }}
# Single check to require in branch protection.
result:
name: Result
needs: [regression]
if: always()
runs-on: ubuntu-latest
timeout-minutes: 5
permissions: {}
steps:
- name: Fail unless the suite passed or was not required
if: needs.regression.result != 'success' && needs.regression.result != 'skipped'
run: exit 1

View file

@ -5,6 +5,641 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [0.11.4] - 2026-09-21
### Added
- 📉 **Far smaller slim image.** A slim build now comes down at around 175 MB, near enough 89% smaller than the last release, the local models, the packages around them and the tools that installed them all gone from it; what that changes about the way an instance behaves is set out under Changed below and in the documentation. [Commit](https://github.com/open-webui/open-webui/commit/cb942bb94c8dc7941336088fb3392e2398ff56c1), [Commit](https://github.com/open-webui/open-webui/commit/d27aa72ab4a7b5632b4ad49e8467081ad3d7ebb4)
- 📦 **Smaller standard image.** The image no longer carries a second copy of Python, two sets of fonts nothing ever loaded, packages nothing imports, or the tool that installed them, taking about 170 MB off a standard build. [#29731](https://github.com/open-webui/open-webui/pull/29731), [#29723](https://github.com/open-webui/open-webui/pull/29723), [#29725](https://github.com/open-webui/open-webui/pull/29725), [#29726](https://github.com/open-webui/open-webui/pull/29726), [#29728](https://github.com/open-webui/open-webui/pull/29728), [Commit](https://github.com/open-webui/open-webui/commit/91f8775b28b52c9ae7f2ab990cf2490bda8055d6), [Commit](https://github.com/open-webui/open-webui/commit/98fcb844e1b19f7dd6289af26273cdec5447dc52), [Commit](https://github.com/open-webui/open-webui/commit/508de20779168003e538bb936b49a31d4a8fb8bb), [Commit](https://github.com/open-webui/open-webui/commit/a1c02098aa2687c72482a59117efe643b785df51)
- 🧑‍💻 **Skills from a terminal.** Skills a connected terminal server offers now sit beside workspace skills everywhere skills are picked — the "$" and "/" menus, the integrations menu and the skills panel, each marked Terminal — and are used the same way: picking one puts its instructions, its folder and the files it ships with in front of the model, and a model that was only told a skill exists can open it itself. They follow whichever terminal is selected and clear when that changes. [Commit](https://github.com/open-webui/open-webui/commit/e69236bccb1e12d14045b09b76ebac0490014851)
- 📔 **Terminal instructions file.** A model working with a terminal is now handed the AGENTS.md sitting in that terminal's home directory, read afresh at the start of every turn, so the instructions you keep beside your work reach the model without being pasted in. [Commit](https://github.com/open-webui/open-webui/commit/946be432375057dc18dbcc7c3513feb322a83a4b), [Commit](https://github.com/open-webui/open-webui/commit/f6922a4c4293449805938c80fbe54174994a1fc6)
- 📇 **Automatic skill discovery.** Every skill you can reach is now listed to the model by name and description, and the full text of one is loaded only when it decides to use it; before, a skill you had not selected in the message box was invisible to it, and this applies to models with built-in tools on. [Commit](https://github.com/open-webui/open-webui/commit/e69236bccb1e12d14045b09b76ebac0490014851)
- 🌄 **Model background images.** A workspace model can now carry a background image, uploaded in its editor and drawn behind the chat whenever that model is selected, sitting below a folder's own background and above your personal one, and it travels with the model through export and import. [Commit](https://github.com/open-webui/open-webui/commit/66addbd6b47bfb25cf9aa37ec70d24db5b28b6b6)
- 📓 **Skill creation from a chat.** Typing "/skills:create" in a chat that already has content, with a terminal selected, turns the workflow you just went through into a reusable skill written to ".agents/skills" under that terminal's root working directory rather than whatever folder the shell happens to sit in, and authored to Open WebUI's skill standards. [Commit](https://github.com/open-webui/open-webui/commit/113c56fc8c1986359556107a20e206159ca32f59), [Commit](https://github.com/open-webui/open-webui/commit/924a4a10fbd0be508a69faf66bf08ac9761bc24d), [Commit](https://github.com/open-webui/open-webui/commit/58078ab3045cc7aee409d9215f68adf0e02c9165), [Commit](https://github.com/open-webui/open-webui/commit/d25f6c7135e0aee93dc2c840ce320d97617a6e47), [Commit](https://github.com/open-webui/open-webui/commit/a096961a31499be23890b98a040ecf2917405568)
- 🖥️ **Terminal tabs per command.** The terminal pane now carries a tab for each command a model is running alongside your own shell, so you can watch them as they go, move between them and take the shell yourself, each tab opening with the command that produced it. Opening the pane puts you in a tab, starting your shell where nothing else is running, and closing the last tab folds the pane away again. [Commit](https://github.com/open-webui/open-webui/commit/54a7a7a7ce22725074c29c7e827446f5dce1421c), [Commit](https://github.com/open-webui/open-webui/commit/f3eade42aead0a3b96ada3a300d6c294c89f74a8), [Commit](https://github.com/open-webui/open-webui/commit/de1203f9b5d3d8b2a84bf5c89d1c61542469367a), [Commit](https://github.com/open-webui/open-webui/commit/675b9839f17df94f866c871cbdb9881323a18a8a)
- 📂 **Folder uploads to terminals.** The file browser beside a terminal takes a folder now, dropped onto it or picked from the Upload Folder entry in its menus, and rebuilds what is inside it as it goes, subfolders and all, where before a drop uploaded only the files sitting loose at the top. [Commit](https://github.com/open-webui/open-webui/commit/f80ef8bd001d7ef650fb278d6fbd05bea4afad0f), [Commit](https://github.com/open-webui/open-webui/commit/4cc0d48b4d83503199bcc5f2322881722971a499), [Commit](https://github.com/open-webui/open-webui/commit/674760bfc1122e0a19fe299e05a86d1cbb528e6f)
- ⚖️ **Side-by-side file comparison.** Picking two files in the file browser lights up a Compare button that lays them side by side or one above the other, numbering the lines, marking what was added and removed down to the part of the line that changed, counting both, and letting you swap which is which or leave whitespace out of it, a choice it remembers; a terminal too old to offer the comparison says so rather than failing quietly. [Commit](https://github.com/open-webui/open-webui/commit/8556033c6b64fa53e44152ebab1f889a578529af), [Commit](https://github.com/open-webui/open-webui/commit/3808eace6c2beb1904f0f04fdd443ec544f94acb)
- 🌐 **Suggested prompts in your language.** A fresh install now offers its suggested starter prompts in the language the interface is set to, rather than the same four in English for everyone, and the model defaults panel carries a way back to those defaults once its own suggestions have been edited. [Commit](https://github.com/open-webui/open-webui/commit/30881dbcc966206cc15fb0b2276b41d88f7c0f09), [Commit](https://github.com/open-webui/open-webui/commit/9fba2843b16a111b035949f5865e29ec6cf5ac9c)
- 📏 **Exa result length cap.** Web search through Exa can be held to a number of characters per result, set beside its key in the admin web search settings or as "EXA_MAX_CONTENT_LENGTH". [Commit](https://github.com/open-webui/open-webui/commit/12b14124b9376eccf2ba7cf3dba92bc145493703)
- 🇪🇺 **European web search option.** Staan can now be picked as the web search provider, giving deployments that need search inside the EU an option they do not have to host themselves, configured from the admin web search settings or through "STAAN_API_KEY", "STAAN_MARKET" and "STAAN_MAX_SNIPPETS". [#30138](https://github.com/open-webui/open-webui/pull/30138), [#26006](https://github.com/open-webui/open-webui/discussions/26006), [#30303](https://github.com/open-webui/open-webui/pull/30303), [#30301](https://github.com/open-webui/open-webui/issues/30301)
- 🗣️ **Per-language names and descriptions.** A model, tool, skill, function, banner or arena entry can now hold its name, description, starter prompts and valve labels once per language, written in a searchable table in its own editor or brought in as a JSON file, which is refused where a translation drops one of the placeholders the original fills in. The interface shows the wording for the language it is set to, and falls back to the plain text where that language has none. [Commit](https://github.com/open-webui/open-webui/commit/7b6562e3358956ec71eed05352538a646d047734), [Commit](https://github.com/open-webui/open-webui/commit/858ab727d2d80bb3c9ef40f01a32de7a12ca75c9), [Commit](https://github.com/open-webui/open-webui/commit/3facfa61d413945d7144d26b827fcce1f3aef7c5), [Commit](https://github.com/open-webui/open-webui/commit/85b11a4f3504434bd9cd0748629b8a84817fb4f2)
- ✏️ **Interface text you can reword.** An administrator can now replace the interface's own wording language by language from a panel in the admin settings, with a replacement refused where it drops one of the placeholders the original fills in. [Commit](https://github.com/open-webui/open-webui/commit/67ac1a4e937271a02bfb02a330aa99dc8192cbf4)
- 🔓 **Turning off the sign-in form.** The box asking for an email address and a password can now be taken off the sign-in page from the authentication settings, where until now it could only be set before the server started, leaving single sign-on or a directory to sign people in. [Commit](https://github.com/open-webui/open-webui/commit/c4a349651e34bf5d0920bb5e26707ae9961ac39f)
- 🏷️ **Custom file metadata.** Metadata attached to an uploaded file now travels with the pieces that file is split into and arrives with the retrieved sources, the oversized and internal fields left out, and an operator can name in "RAG_SOURCE_METADATA_KEYS" which of those fields the model itself gets to see alongside the text. [#29499](https://github.com/open-webui/open-webui/pull/29499), [#29486](https://github.com/open-webui/open-webui/issues/29486), [Commit](https://github.com/open-webui/open-webui/commit/894655f66b9563890e63c76090ddecc90a311eb2), [#29502](https://github.com/open-webui/open-webui/pull/29502), [#29696](https://github.com/open-webui/open-webui/pull/29696)
- 🗑️ **Quick delete shortcut.** Holding Shift over a note in the list or grid, or over a row on the automations page, turns its trailing controls into a delete button, removing the entry in one click rather than the three the menu and its confirmation ask for. [#29635](https://github.com/open-webui/open-webui/pull/29635), [#29633](https://github.com/open-webui/open-webui/issues/29633), [#29640](https://github.com/open-webui/open-webui/pull/29640), [#29637](https://github.com/open-webui/open-webui/issues/29637)
- 🔀 **Diffs are drawn as diffs.** A diff or patch block in a reply is now laid out as one, with the file and hunk headings, the old and new line numbers, and the added and removed lines picked out down to the part of the line that changed, and a button to switch to the plain text and back. [Commit](https://github.com/open-webui/open-webui/commit/254e29b9af634145dde0450451a36f0dfbd8f421)
- ✒️ **A formatting switch for notes.** A note you can edit carries a Formatting switch in its menu: turned off, Markdown you type or paste stays as the characters you wrote, a paste keeps its plain text, and a web address is left as text, while formatting already in the note is untouched. [Commit](https://github.com/open-webui/open-webui/commit/d8f27e745bd31c89efa25cbe0f41bb0f70576fc5)
- 🎛️ **Admin model list filters.** The models page in the admin settings can now be narrowed to the base models a connection offers or to the ones built in the workspace, alongside the filters for enabled, disabled, visible and hidden. [Commit](https://github.com/open-webui/open-webui/commit/9e634c0c56e0a060717f1e24a11b082daf664036)
- 🔍 **Settings search by name.** The search box in settings now matches each setting's own name and description rather than a keyword list kept per page, in whichever language you are using and regardless of accents, it leaves out anything your permissions do not let you change, and opening a result no longer clears what you typed. [Commit](https://github.com/open-webui/open-webui/commit/7cbaabe02fc3c7c040008e82cf85f590bd0c346a), [Commit](https://github.com/open-webui/open-webui/commit/9d98ffcddf792be8d567ea39370d39fcac649acd), [Commit](https://github.com/open-webui/open-webui/commit/c82634b9d01adadbf9780ff84a0a8168fdc4fdea)
- 🆘 **A model for every chat.** Opening a conversation whose model has since been retired no longer leaves it stranded with nothing selected; it falls back to your default model, then the configured default, then the first one available, and a chat that still has a live model keeps it. [#29757](https://github.com/open-webui/open-webui/pull/29757)
- ⚡ **Non-blocking search.** Searching the text of chats and knowledge, and the grep a model runs over a knowledge base, now run beside the rest of the server rather than in front of it, so a long search no longer keeps other requests waiting. [#29621](https://github.com/open-webui/open-webui/pull/29621), [Commit](https://github.com/open-webui/open-webui/commit/d9c8de9c39fca7c4756e0f9bd599b09ebe435208)
- ✴️ **Lighter shared note editing.** Notes written by several people at once no longer echo every keystroke back to the server once per watcher; the traffic between editors falls by half with two of them and keeps falling as more join, so shared notes stay smooth as the room grows. [#28185](https://github.com/open-webui/open-webui/pull/28185)
- 🔢 **Sorted knowledge listings.** The "ls", "tree" and "find" commands a model runs over a knowledge base now sort by name and take "-t", "-S" and "-r" for newest first, largest first and reversed. [#29840](https://github.com/open-webui/open-webui/pull/29840)
- 🪪 **Authentication type header.** A request to an OpenAI or Ollama connection now carries "X-OpenWebUI-Auth-Type", saying whether the person behind it signed in through the browser or called with an API key; its name is set with "FORWARD_USER_INFO_HEADER_AUTH_TYPE" and "{{AUTH_TYPE}}" works in custom headers. [Commit](https://github.com/open-webui/open-webui/commit/ee46e2664ab23bacc536fed21c06ab19067d695b)
- 💡 **Follow-up ghost text.** Once a reply has finished, the first of the follow-up questions it suggests now sits greyed inside the empty message box as well as under the reply: Tab writes it out, and typing anything of your own clears it away. [Commit](https://github.com/open-webui/open-webui/commit/7aaa4a692e949724913834efc47c57572042703b)
- 🪵 **Model refusal logging.** A request turned away with "Model not found" now writes a warning naming the account, the model and why it was refused, while the message the caller sees stays as vague as before. [Commit](https://github.com/open-webui/open-webui/commit/f5fcf4c89fb27ed39825875d573446fa84aa2a5b)
- 🤲 **Cheaper chat requests behind Redis.** Where several servers share their websocket traffic through Redis, a chat request no longer pulls the whole shared model list from Redis to ask about a couple of models; each worker keeps its own copy and refetches only when the list actually changed, so the cost stops growing with the number of models configured. [#28176](https://github.com/open-webui/open-webui/pull/28176), [#28167](https://github.com/open-webui/open-webui/issues/28167)
- 💽 **Long conversations stream cheaper.** Writing or reading a single message no longer loads, checks and rewrites the entire conversation to change a few hundred bytes, so streaming into a long chat stops getting more expensive as it grows, and the whole round trip a message event costs falls to a fraction on conversations of thousands of messages. [#28184](https://github.com/open-webui/open-webui/pull/28184), [#28169](https://github.com/open-webui/open-webui/issues/28169)
- 🧱 **In-place streaming appends.** Streamed replies can be assembled with CPython's in-place string append, turned on with "ENABLE_CHAT_RESPONSE_STREAM_INPLACE_APPEND" and left off by default while it is rolled out gradually. [Commit](https://github.com/open-webui/open-webui/commit/113c56fc8c1986359556107a20e206159ca32f59), [Commit](https://github.com/open-webui/open-webui/commit/924a4a10fbd0be508a69faf66bf08ac9761bc24d), [#30066](https://github.com/open-webui/open-webui/pull/30066)
- 📘 **Direct connection guidance.** The connections and integrations pages in the personal settings now say plainly that a direct connection leans on your browser session to keep requests running, which suits testing and temporary use rather than everyday work. [Commit](https://github.com/open-webui/open-webui/commit/5d66dd6ad291db17894953714c01fd4303093e7e), [Commit](https://github.com/open-webui/open-webui/commit/15e2259a2ae1b7dec1f93a52078304dd3a79a14c)
- 🔄 **General improvements.** Various improvements were implemented across the application to enhance performance, stability, and security.
- 🌐 **Translation updates.** Translations for Japanese, Traditional Chinese, Korean, Finnish, Russian, Ukrainian, German, Spanish, Portuguese (Brazil), Arabic, Arabic (Bahrain), Azerbaijani, Bulgarian, Bengali, Tibetan, Bosnian, Catalan, Cebuano, Czech, Danish, Basque, Croatian, Dutch, Italian, Turkish, Simplified Chinese, Bosnian, Catalan, Danish, Estonian, Finnish, Galician, Croatian, Kabyle, Norwegian Bokmål, Portuguese (Portugal), Romanian and Turkmen were enhanced and expanded.
### Fixed
- 🛡️ **Security Advisory**: This release includes security and access-control fixes. We recommend updating production deployments at your earliest convenience. Not all security fixes in this version may be enumerated in the fixed section. Some may be withheld for a short time to give administrators time to upgrade. [Advisories](https://github.com/open-webui/open-webui/security)
- 🔑 **Tokens stay out of logs.** A failure part way through signing in with an identity provider no longer writes the credentials it was handed into the application log, recording the provider and the error it reported instead. [#29709](https://github.com/open-webui/open-webui/pull/29709)
- 🔐 **Sign-in role mapping.** Roles sent by an identity provider were in some setups not applied, leaving an account at the default role, and a sign-in whose roles cannot be read is now refused rather than let through. [Commit](https://github.com/open-webui/open-webui/commit/10d1cfe6375f207acaa531e857edb575ded2cfc3)
- 🚫 **Blocked sign-in groups saved.** Comma-separated group names typed into the blocked groups field now take effect once saved, where the save stored the text as written and the sign-in check then read it as nothing; a group name carrying a comma of its own survives a save too. [Commit](https://github.com/open-webui/open-webui/commit/3fc1146c13d7b4b7a67c7fd058014436f80e0dfd)
- 🚪 **Signing out ends the session.** Signing out or having every token revoked left any live connection that account already had in place, and those are now cut at the same moment; a token already revoked can no longer be replayed to cut someone else's newer session. [Commit](https://github.com/open-webui/open-webui/commit/3a6d0fd203050401481bdb8621b827089d393817)
- 🛡 **Diagrams and SVG stay on this origin.** Mermaid diagrams and SVG previews, attached or uploaded, no longer follow references pointing at another origin, a picture, style or configuration line among them, and a diagram that reaches outside is refused with an error; a file whose drawing leans on an external sprite now shows that part blank. [#30271](https://github.com/open-webui/open-webui/pull/30271)
- 📲 **Terminal sessions follow access changes.** A terminal connection now re-checks your access to it every ten seconds for as long as it stays open, so removing someone's permission or deactivating their account ends their live terminal on the next check, on every worker and not only the one they signed in through. [Commit](https://github.com/open-webui/open-webui/commit/a1189a2d757407530a0c1c85a41b46cd6a4f7da5), [Commit](https://github.com/open-webui/open-webui/commit/e8bd0661d3774248c318286849041b94bb861909)
- 🔒 **Connection model listing access.** The endpoints that list the models on a single Ollama or OpenAI connection could be reached by a signed-in account of any role, and now check the caller's role. [#29619](https://github.com/open-webui/open-webui/pull/29619)
- 🔏 **Knowledge file access.** A file's access through a knowledge base now follows its actual attachment alone, rather than also a collection name left behind on the file record. [#29937](https://github.com/open-webui/open-webui/pull/29937)
- 🌳 **Knowledge directories stay inside their base.** A directory id supplied to the knowledge file routes is now refused where it belongs to a different knowledge base, where a caller-supplied id could reach into a directory tree outside the one being addressed. [Commit](https://github.com/open-webui/open-webui/commit/a9541c18ca2a056f4ccfe56b9c733cba74b86b49), [#29887](https://github.com/open-webui/open-webui/pull/29887)
- 🍪 **Cookie forwarding scope.** A connection authenticating as the signed-in person is handed this instance's browser cookies only where its new Forward cookies switch is on, which an instance relying on it has to turn on after upgrading. [Commit](https://github.com/open-webui/open-webui/commit/b71744b17823f7a3a9a64cb363e79a7164ed5b23), [Commit](https://github.com/open-webui/open-webui/commit/3cda47cdb449aac970426f0929c9e5f8bb8bd807), [Commit](https://github.com/open-webui/open-webui/commit/f9f815c86220f809fad5ee1300c44f55dea74c79)
- 🗝️ **Revocation list fallback.** Where the revocation list is kept in Redis and Redis cannot be reached, a token is now accepted rather than the request failing, so signing out may not take effect until Redis is back. [Commit](https://github.com/open-webui/open-webui/commit/c1615bec2f5f143084b5fcc04f42ebfa0dade9df), [Commit](https://github.com/open-webui/open-webui/commit/6a85abb3f5e002f3069007213e16e4ef60776890)
- 🗒 **Home page chat drafts.** A message typed into a chat started from the home page now comes back after a reload, text, uploads and all, where the draft was stored under a key the page never read again. [#29762](https://github.com/open-webui/open-webui/pull/29762), [#29760](https://github.com/open-webui/open-webui/issues/29760)
- 🏞️ **S3-hosted chat images.** A chat image whose host echoes a Content-Encoding nothing asked for, an S3 or MinIO object uploaded with that metadata among them, no longer breaks the reply it sits in; it is sent as its link the way an unreachable image already was. [#29623](https://github.com/open-webui/open-webui/pull/29623)
- 🪆 **Branch descent guard.** Stepping between reply branches walked a chat's children without tracking where it had already been, so a history that pointed back at itself spun forever and locked the tab, and a child id with no message behind it threw; every one of the eleven places that walk now shares a single guarded helper. [#30070](https://github.com/open-webui/open-webui/pull/30070)
- 🗒️ **Note save on exit.** A note's title and its attachments save a moment after they change, and leaving the note inside that moment dropped the save, so the notes list kept showing the old title until the page was reloaded; the pending save now finishes before the next page opens. [#29745](https://github.com/open-webui/open-webui/pull/29745), [#29744](https://github.com/open-webui/open-webui/issues/29744)
- 📷 **Attachments sent without text.** Sending an image or a file with nothing typed to a model that has skills attached replaced the empty message with the list of skill names, so the model answered with its own catalogue instead of looking at what you sent. [#30045](https://github.com/open-webui/open-webui/pull/30045), [#30040](https://github.com/open-webui/open-webui/issues/30040)
- 🖍️ **SVG attachments read as text.** An SVG attached to a chat went up as a picture and came back rejected by every model that tried to decode it, and is now read as its source text instead, which needs a model that accepts file uploads. [#30102](https://github.com/open-webui/open-webui/pull/30102), [#30100](https://github.com/open-webui/open-webui/issues/30100)
- 🩹 **Notes survive a small edit.** Asked to add a section or change a few lines, a model would send the whole note back and lose the rest of it with nothing to undo, because the editing tool never said when to edit a range instead; it now spells out the range rules the handler already enforced. [#30048](https://github.com/open-webui/open-webui/pull/30048)
- 📋 **Note paste placement.** A plain paste replaced more of the note than was selected, and a paste into a code block broke out of the block instead of going inside it; both now land exactly where they were dropped. [Commit](https://github.com/open-webui/open-webui/commit/d8f27e745bd31c89efa25cbe0f41bb0f70576fc5)
- 🧿 **Insert into note on read-only notes.** The Insert into note action no longer appears on a note you may only read, which offered it where it could not land. [#30223](https://github.com/open-webui/open-webui/pull/30223), [#30174](https://github.com/open-webui/open-webui/issues/30174)
- 🤝 **Note sharing survives a save.** Saving a note you had been given access to took that sharing away, so the note went private and everyone else lost it. [#30175](https://github.com/open-webui/open-webui/pull/30175)
- 🌍 **Collaborative link handling.** In a note two people are writing at once, a plain web address arriving from the other editor was turned into a link on your screen but not on theirs; it is now left as written. [Commit](https://github.com/open-webui/open-webui/commit/d8f27e745bd31c89efa25cbe0f41bb0f70576fc5)
- 📑 **Tika 4 document formatting.** Where document extraction runs against a Tika server on version 4, the text now comes back as Markdown rather than a flat run of characters, so headings, lists and tables survive into what the model reads. [Commit](https://github.com/open-webui/open-webui/commit/ba34bee2d17e1c00238c85a8aab1a11a60db164e)
- ⏳ **Session expiry accuracy.** A session with an identity provider now expires with the token it actually calls with, rather than whichever of its two tokens ran out first, which renewed it early and dropped it whenever that failed. [Commit](https://github.com/open-webui/open-webui/commit/aaaf26fb8ede28854dd660b11b85ba6bafe64007)
- 🧩 **Tool steps in finished replies.** A reply that called tools showed each step as it ran, then lost them the moment the reply completed, because the provider's closing message replaced everything on screen rather than joining it; the closing message is now merged into what is already there, and a tool call is no longer mistaken for its own result. [Commit](https://github.com/open-webui/open-webui/commit/e1bfefdf9f8f2012cecfc7d81079bd6fff147932), [Commit](https://github.com/open-webui/open-webui/commit/31b272d3c93b87636c920f2aa68d6caaa07ae79a)
- 🧭 **Streamed replies follow along.** A reply arriving over the response events now keeps the view at the bottom as it is written, as replies on the older path already did, unless you have scrolled up yourself. [Commit](https://github.com/open-webui/open-webui/commit/dfde08aa7391d924359b27f5768411d7a533a6ec), [Commit](https://github.com/open-webui/open-webui/commit/88e78b7819b28ffe91f3dff24d8c5d992004197e)
- 🚿 **Tool results arrive when the tools are done.** The results of a model's tool calls now reach the reply as soon as the round that produced them finishes, instead of being held back with the rest of the streaming until the throttle let them through, and a channel reply or a continuation no longer carries the picture data a tool handed the model. [Commit](https://github.com/open-webui/open-webui/commit/478d1785fd27f400c614a5700f5b38ec9de412f9)
- 🧗 **Tagged blocks while streaming.** A reply using reasoning, solution or code interpreter tags briefly showed the raw tag text in the message as it arrived, because the chunk carrying it reached the browser before the cleaned output did; the cleaned output is now sent the moment a tag is taken out, and the code an interpreter block writes fills that block rather than the message body. [Commit](https://github.com/open-webui/open-webui/commit/58b36765a7c20f5943a3180bd289de48876d0878)
- 📗 **Excel in the code interpreter.** Reading or writing a spreadsheet in code the browser runs failed outright because the library that handles them never reached the browser, and it is now shipped alongside the rest. [#30140](https://github.com/open-webui/open-webui/pull/30140), [#30130](https://github.com/open-webui/open-webui/issues/30130)
- 📀 **Bundled charts and formatting.** Drawing a chart with seaborn in code the browser runs fetched the library over the internet at that moment rather than taking it from what ships, and the code editor's Format button failed outright for anyone who is not an administrator; both now work from what comes with the application, offline included. [#30148](https://github.com/open-webui/open-webui/pull/30148), [#30145](https://github.com/open-webui/open-webui/issues/30145)
- 🙋 **Typed answer submission.** In the card a model puts up to ask you a question, choosing one of its options on the last question sends your answers straight away, but typing your own into the Other box left Submit answers greyed out with no way to send it, and it now turns on as soon as that box has text. [#29494](https://github.com/open-webui/open-webui/pull/29494), [#29311](https://github.com/open-webui/open-webui/issues/29311)
- 🥇 **A truthful Recommended badge.** The card a model puts up to ask you a question marks its first option Recommended, but nothing told the model that, so the badge fell on whichever option happened to be listed first; models are now asked to put the option they recommend there. [#30196](https://github.com/open-webui/open-webui/pull/30196), [#30195](https://github.com/open-webui/open-webui/issues/30195)
- 🎚️ **Partial settings permissions.** An account barred from changing the interface settings can now save its system prompt, notifications, audio, keyboard shortcuts and pinned models, which were refused along with them. [Commit](https://github.com/open-webui/open-webui/commit/98a920168e2eea435ac15e1ad3d679946631e41d)
- 🎧 **Speech file types kept.** Saving the audio settings emptied the list of file types accepted for speech recognition, so transcription then turned away the recordings it had taken before. [#30208](https://github.com/open-webui/open-webui/pull/30208)
- 🚻 **Group picker without permission.** The picker for sharing a chat, a note or a knowledge base with a group was offered to accounts not allowed to share with groups, and is now hidden from them. [#30189](https://github.com/open-webui/open-webui/pull/30189)
- 🗜️ **Per-field settings saves.** Only the settings actually changed are now stored, rather than the whole object, so another tab's older copy no longer overwrites them and a default an administrator changes still reaches everyone, and pinning a model, reordering the list or picking a default from outside the settings window now saves the same way. [Commit](https://github.com/open-webui/open-webui/commit/98a920168e2eea435ac15e1ad3d679946631e41d), [#30183](https://github.com/open-webui/open-webui/pull/30183)
- ⚠️ **Failed settings feedback.** Settings that could not be saved were shown as saved anyway until the page was reloaded, because the interface stored them locally without waiting on the server; a failure now raises an error and leaves the panel as it was. [Commit](https://github.com/open-webui/open-webui/commit/98a920168e2eea435ac15e1ad3d679946631e41d)
- 🔠 **Menu text scaling.** The entries in the menus that drop down across the interface now scale with the rest of it, rather than staying at a fixed size while the menu around them grew. [#29493](https://github.com/open-webui/open-webui/pull/29493), [#29488](https://github.com/open-webui/open-webui/issues/29488)
- 🪞 **Account menu highlighting.** An account menu entry carrying a pin button beside it now lights up across the whole row in the shared colour, and a long label no longer pushes the pin out of the menu. [Commit](https://github.com/open-webui/open-webui/commit/e723dcda58f638a5da2398743d22d3e5854042bf)
- 🔲 **Shift-click file selection.** Shift-clicking a file in the terminal's file browser now adds that range to what is already selected, and clicking through a selected file removes its range. [Commit](https://github.com/open-webui/open-webui/commit/f80ef8bd001d7ef650fb278d6fbd05bea4afad0f), [Commit](https://github.com/open-webui/open-webui/commit/4cc0d48b4d83503199bcc5f2322881722971a499), [Commit](https://github.com/open-webui/open-webui/commit/674760bfc1122e0a19fe299e05a86d1cbb528e6f)
- 🛟 **Attachment name collisions.** Attaching a file to a message on an instance that puts attachments in a terminal's working directory replaced whatever file of that name was sitting there; the upload now takes the next free name, "report (1).pdf" beside "report.pdf", and the attachment shows the name it was saved under, though two uploads arriving at the same moment from different browsers can still land on the same name. [Commit](https://github.com/open-webui/open-webui/commit/d70053e44993f271d534fb87d2b40724b028fca2)
- 📛 **Failed upload feedback.** A file that could not be written to the terminal now says so, rather than passing in silence while the browser refreshed as though it had arrived. [Commit](https://github.com/open-webui/open-webui/commit/f80ef8bd001d7ef650fb278d6fbd05bea4afad0f), [Commit](https://github.com/open-webui/open-webui/commit/4cc0d48b4d83503199bcc5f2322881722971a499), [Commit](https://github.com/open-webui/open-webui/commit/674760bfc1122e0a19fe299e05a86d1cbb528e6f)
- ☑️ **Unchecked checkbox defaults.** A prompt variable written as a checkbox with a default of false opened the form already ticked, as did False, "false", 0 and "0", because any non-empty default counted as ticked; it is now ticked only where the value really is true. [#30037](https://github.com/open-webui/open-webui/pull/30037), [#30036](https://github.com/open-webui/open-webui/issues/30036)
- 🎯 **Prefilled question timing.** Opening a chat from a link holding a question sent it before the box had it, so a question naming a variable went off with the variable unfilled; the send now waits for the text to be in place and filled in. [Commit](https://github.com/open-webui/open-webui/commit/3808eace6c2beb1904f0f04fdd443ec544f94acb), [Commit](https://github.com/open-webui/open-webui/commit/a23b579233276e40159eb615917b9aeb7d5ed5c9)
- 📜 **Task list height.** The list of steps a model works through ran as long as it needed and pushed the rest of the reply down the page; it now stops at a quarter of the window's height and scrolls within itself. [Commit](https://github.com/open-webui/open-webui/commit/307b9b9133f0c7ad899f4b76226059da6f3177eb)
- ⎋ **Escape key targeting.** The shortcut for closing a dialog always shut the settings window, whichever dialog was actually in front, so a dialog opened from within settings took both away at once; each dialog now answers the shortcut for itself, as it already did for the escape key. [#29830](https://github.com/open-webui/open-webui/pull/29830), [#29817](https://github.com/open-webui/open-webui/issues/29817)
- 🎹 **Message pair shortcut.** The shortcut that adds an empty question and answer to a chat also sent whatever was typed in the message box, so the pair arrived alongside a message you had not meant to send yet. [#30167](https://github.com/open-webui/open-webui/pull/30167)
- 🥁 **Collapsing the task list.** Folding away the list of tasks under a reply also sent whatever you had typed in the message box. [#30202](https://github.com/open-webui/open-webui/pull/30202)
- 🕰️ **Temporary chat links.** A link carrying the temporary chat marker opened an ordinary chat, because the marker was written into the address but never read back when the page loaded, and such a link now opens the temporary chat it promises. [Commit](https://github.com/open-webui/open-webui/commit/67adde31936d1d2e656140f9edd898b1fbb18a1c)
- ⌨️ **Recording a new shortcut.** Pressing a key combination to record it as a shortcut also ran whatever that combination was already bound to, so setting one up did the thing you were trying to rebind. [#30160](https://github.com/open-webui/open-webui/pull/30160)
- 🔃 **Stale tab reload.** A tab still running the previous build met an error page after the server was updated instead of loading the new one, because every image built from Docker carried the same version stamp; the stamp now follows the build, and a stale tab reloads as it was meant to. [#29832](https://github.com/open-webui/open-webui/pull/29832), [#29831](https://github.com/open-webui/open-webui/issues/29831)
- 📌 **Channel code headers.** The bar naming a piece of code in a channel thread or its pinned messages now sits flush at the top of the panel, clipped to the block, rather than floating over the code as it scrolls. [#29836](https://github.com/open-webui/open-webui/pull/29836), [#29835](https://github.com/open-webui/open-webui/issues/29835)
- 📨 **Duplicated proxy headers.** A reply proxied from a terminal server or an OpenAI or Ollama connection no longer carries that server's own "Server" and "Date" headers, which had a reverse proxy in front logging a duplicate line for every one. [#29841](https://github.com/open-webui/open-webui/pull/29841), [#29824](https://github.com/open-webui/open-webui/issues/29824), [#29843](https://github.com/open-webui/open-webui/pull/29843)
- 🎣 **Testing an image connection.** The Verify button beside an image generation connection saved the whole image configuration first and then tested whichever engine was active rather than the connection beside it, and it now tests exactly that connection and changes nothing. [Commit](https://github.com/open-webui/open-webui/commit/64bbdf7a73724986fac8bcf6e880fe32ef9ac495)
- 🪝 **Responses API tool strictness.** A workspace tool, MCP server or OpenAPI server reaching a model on the Responses API was turned into a strict schema where it had never asked to be, so the model filled every optional field with empty strings, zeros and empty arrays, and search and filter tools were handed values where leaving them out was meant. [#30046](https://github.com/open-webui/open-webui/pull/30046), [#27750](https://github.com/open-webui/open-webui/issues/27750)
- 🪃 **Responses API tool calling.** A forced tool choice sent to a connection on the Responses API went out in the wrong shape and was refused by providers that check it, and a tool call in a reply that was not streamed came back as empty text, so nothing reading the API ever saw it. [#30095](https://github.com/open-webui/open-webui/pull/30095), [#30085](https://github.com/open-webui/open-webui/issues/30085)
- ⛓️ **Missing tools in links.** Opening a chat from a link whose "tools" or "tool-ids" parameter names a tool that no longer exists, or that the account cannot see, kept that id in the selection and sent it with the message; ids matching no tool the account has are now dropped and the rest of the link works as before. [#29803](https://github.com/open-webui/open-webui/pull/29803)
- 🚫 **Model editor error messages.** A workspace model that cannot be loaded for editing now says so, instead of sending you back with "You do not have permission to edit this model" whatever the real reason. [#29694](https://github.com/open-webui/open-webui/pull/29694), [#29629](https://github.com/open-webui/open-webui/issues/29629)
- 👥 **Directory updates that were dropped.** A change your identity provider sent as an add or a remove, or without naming the attribute it was changing, was accepted and then quietly discarded, so a rename or a deactivation never reached the account; those now take effect. [Commit](https://github.com/open-webui/open-webui/commit/ad9da981680c0fd9151c803a366a01545ea907c1)
- 🗃️ **Directory request validation.** A provisioning request carrying the wrong kind of value, or an attribute Open WebUI does not support, now comes back as an error instead of passing in silence, and a sync that changes nothing no longer marks the account as touched. [Commit](https://github.com/open-webui/open-webui/commit/ad9da981680c0fd9151c803a366a01545ea907c1)
- 🔘 **Model editor save button.** Saving a workspace model whose model list could not be refreshed afterwards now reports the error and frees the Save button, rather than leaving it disabled and spinning though the model had been saved. [Commit](https://github.com/open-webui/open-webui/commit/dc98e3023fc6e0113dbad76545cdb15e014db6ad), [Commit](https://github.com/open-webui/open-webui/commit/cc5479d16d5caf28ec19a16ffa1ee4f3dbc0f63f)
- 🫱 **Rating a reply again.** Switching a reply's rating from one thumb to the other kept the score and the reason given the first time, and rating a reply whose feedback had since been deleted failed outright; both now record the rating you just gave. [#30075](https://github.com/open-webui/open-webui/pull/30075), [#30077](https://github.com/open-webui/open-webui/pull/30077)
- 📶 **Sorting feedback by user.** The User column in the admin feedback history did nothing when clicked, and now sorts by who left the feedback. [#30191](https://github.com/open-webui/open-webui/pull/30191)
- 🪂 **Closing the Edit User dialog.** Changes typed into the admin Edit User dialog and then abandoned showed on the row in the user list until the page was reloaded, and closing the dialog now drops them. [#30193](https://github.com/open-webui/open-webui/pull/30193)
- 🪙 **Model defaults save.** Saving the model defaults in the admin settings put the selected, pinned and ordered models back as they stood when the page was opened, undoing anything changed in between. [#30206](https://github.com/open-webui/open-webui/pull/30206)
- 👤 **A custom gender shown back.** An account whose gender is a wording of its own came back to an empty dropdown in the account form, and the form now shows Custom with that wording beside it. [#30215](https://github.com/open-webui/open-webui/pull/30215)
- 🗓️ **Clearing a calendar event.** Emptying the repeat, the description or the location of a calendar event did not take and the old wording came back, and those fields can be cleared again. [#30204](https://github.com/open-webui/open-webui/pull/30204)
- 📕 **Required prompt dropdowns.** A prompt's form could be sent with a required dropdown left unchosen, and now asks you to pick something first. [#30078](https://github.com/open-webui/open-webui/pull/30078)
- 🎫 **Account form required fields.** The account form now refuses to save with a required field left empty, the way the admin user dialog does, and a date of birth is no longer demanded of accounts that never set one. [#30296](https://github.com/open-webui/open-webui/pull/30296), [#30295](https://github.com/open-webui/open-webui/issues/30295)
- 🐑 **Model clones keep their base.** Cloning a model from the admin Models settings now carries the model it was built on, where the clone came out detached from it, and an arena model no longer offers Clone at all. [#30080](https://github.com/open-webui/open-webui/pull/30080), [#30079](https://github.com/open-webui/open-webui/issues/30079)
- 🧤 **Model sharing survives a save.** Saving a model you cannot fully share no longer strips the access entries you cannot re-create, where an editor resending every stored entry had each one re-checked against what its author may grant, and a save from someone with narrow rights quietly took the model private for everyone else. [Commit](https://github.com/open-webui/open-webui/commit/754c4b5762ed0d79954ff28c4e06631004dc31b3), [#30093](https://github.com/open-webui/open-webui/pull/30093), [#30087](https://github.com/open-webui/open-webui/issues/30087)
- 📠 **Prompt version saves.** Saving a new version of a prompt without Set as Production leaves the live prompt exactly as it was, where the draft quietly took its place, and the editor now shows the production text the moment a version is set. [#30231](https://github.com/open-webui/open-webui/pull/30231), [#30230](https://github.com/open-webui/open-webui/issues/30230), [#30233](https://github.com/open-webui/open-webui/pull/30233), [#30232](https://github.com/open-webui/open-webui/issues/30232)
- 💬 **Model description round trips.** A workspace model whose description is switched from the default back to custom saves again, where the switch read the field as empty and the description was dropped on save. [#30249](https://github.com/open-webui/open-webui/pull/30249), [#30247](https://github.com/open-webui/open-webui/issues/30247)
- 🎚 **Compaction threshold saves.** The context compaction threshold set in general settings now reaches the model parameters, where the value never left the page. [#30270](https://github.com/open-webui/open-webui/pull/30270), [#30269](https://github.com/open-webui/open-webui/issues/30269)
- 📢 **Speech engine defaults.** Switching the text to speech engine now applies that engine's own default voice and model, where the change kept the previous engine's settings in place. [#30289](https://github.com/open-webui/open-webui/pull/30289), [#30288](https://github.com/open-webui/open-webui/issues/30288)
- 🗞 **MinerU key in local mode.** The document settings save again in local mode with the MinerU key left empty, where the form demanded a key it did not need. [#30299](https://github.com/open-webui/open-webui/pull/30299), [#30298](https://github.com/open-webui/open-webui/issues/30298)
- 🫂 **Group dialog reset.** The new-group dialog opens empty after a group is created, where the next one came up holding the group just made. [#30280](https://github.com/open-webui/open-webui/pull/30280), [#30279](https://github.com/open-webui/open-webui/issues/30279)
- 🧵 **Thread reply notifications.** Clicking the notification for a reply written inside a thread dropped you at the bottom of the channel with the thread still shut and the reply nowhere in sight; it now opens the thread the reply belongs to. [#29856](https://github.com/open-webui/open-webui/pull/29856), [#29855](https://github.com/open-webui/open-webui/issues/29855)
- 🔽 **Dropdown arrow spacing.** The arrow in the dropdowns drawn no wider than their contents, among them the provider on a new connection, no longer overlaps the last characters of the longest choice. [#29866](https://github.com/open-webui/open-webui/pull/29866), [#29865](https://github.com/open-webui/open-webui/issues/29865)
- 🧊 **Lowercase header handling.** Headers from an upstream that writes them in lower case, as anything served by uvicorn does, are now matched without regard to case, so "Content-Encoding" is stripped and clients stop failing to decompress a body the server had already decoded. [#29843](https://github.com/open-webui/open-webui/pull/29843)
- 🛑 **Complete chat stop.** Where a chat had more than one task in flight, stopping it, deleting it, or closing a note being worked on could stop the first and leave the rest running to the end, both because a task that had already finished ended the round early and because the list being worked through was rewritten underneath it as each one was cleared away; every task is now stopped in turn, so a reply that was calling tools stops calling them rather than running on to its own limit, though instances sharing their state through Redis were not affected. [Commit](https://github.com/open-webui/open-webui/commit/e35b907f737e625b900b3a03d61890f67ab0c4b0), [#29844](https://github.com/open-webui/open-webui/pull/29844), [#29816](https://github.com/open-webui/open-webui/issues/29816)
- 🖊️ **Channel code blocks.** Where a model answers in a channel with structured output, the code inside it was drawn as plain highlighted text rather than in the editor every other message uses, so it could not be edited in place and a diff in it could not be opened for editing; it now renders the same way as everywhere else. [#29861](https://github.com/open-webui/open-webui/pull/29861)
- 📐 **Code block edits on collapse.** An unsaved edit inside a code block stays on screen when the block is folded away and opened again, where collapsing it showed the saved text though the edit was still pending. [#30284](https://github.com/open-webui/open-webui/pull/30284), [#30283](https://github.com/open-webui/open-webui/issues/30283)
- 🚨 **Code run errors with output.** A code run that ends with an error now shows the error alongside what it printed, where a run that printed anything at all showed its error nowhere. [#30286](https://github.com/open-webui/open-webui/pull/30286), [#30285](https://github.com/open-webui/open-webui/issues/30285)
- 🔕 **Reactions on read-only channels.** The reaction picker no longer appears on a channel open to you as a viewer alone, matching the reply and menu options already hidden there. [#30241](https://github.com/open-webui/open-webui/pull/30241), [#30240](https://github.com/open-webui/open-webui/issues/30240)
- 🎯 **Knowledge search accuracy on PostgreSQL.** A fresh install using PostgreSQL built its search index before a single piece of text existed to organise it around, so the index was never fit for the content that arrived afterwards and every search quietly returned the wrong passages; the index now waits until there is enough text to build on, and until then searches read everything exactly. An install already carrying such an index can restore it by rebuilding that one index. [#30143](https://github.com/open-webui/open-webui/pull/30143), [#30134](https://github.com/open-webui/open-webui/issues/30134)
- 🗂 **Knowledge file filter on first click.** The File content filter in a knowledge base now narrows the listing the first time it is clicked, where the first click only armed the checkbox and everything stayed listed until it was clicked again. [#30211](https://github.com/open-webui/open-webui/pull/30211), [#30210](https://github.com/open-webui/open-webui/issues/30210)
- 🔭 **Knowledge base search recall.** Where many knowledge bases are stored together, a search of one holding a small share of what is stored found only a small share of its matches, and the shortfall grew as the store did; the search now keeps looking until it has enough from the knowledge base you asked for, and "PGVECTOR_ITERATIVE_SCAN" switches that off or makes it strict. [#30142](https://github.com/open-webui/open-webui/pull/30142), [#30135](https://github.com/open-webui/open-webui/issues/30135)
- 🪣 **Searches that found nothing.** A knowledge search that came back empty never handed its database connection back, so enough of them left knowledge search failing outright until the server was restarted; connections are returned now on PostgreSQL and on openGauss alike. [#30142](https://github.com/open-webui/open-webui/pull/30142), [#30133](https://github.com/open-webui/open-webui/issues/30133), [#30144](https://github.com/open-webui/open-webui/pull/30144)
- 🧽 **Knowledge folder deletion.** Deleting a folder without moving what was in it up a level dropped the files from the listing but left their text in the search index and the files themselves in storage, so a model went on retrieving and citing pages from a folder that was no longer there; the text is now removed with the folder, and a file no other knowledge base holds is deleted with it unless file retention is switched on. [Commit](https://github.com/open-webui/open-webui/commit/17dbc6f001aeea25ae1df528cb79bb272eca4a77)
- 🦀 **Regex searches over chat and knowledge files.** A pattern search now runs in time proportional to the text whatever the pattern, where a nasty expression could stall the search for good, and the patterns it accepts follow RE2's rules, with no lookarounds or backreferences and character classes matching ASCII only. [Commit](https://github.com/open-webui/open-webui/commit/97e013a66169c4fcd7f26ab9e41a4ff260ffd4da)
- ⏰ **Automation schedule parsing.** An automation whose rule puts the time in "DTSTART" is now read at that hour, rather than listed at midnight and opened at nine, which moved when it ran as soon as it was saved again. [Commit](https://github.com/open-webui/open-webui/commit/540467b90a430e47e536c20710878b342de659fb)
- ⏲️ **Recurrence counts and start dates.** A rule repeating a set number of times, ten or a hundred, is read as the limited repeat it is, where any count beginning with a one was treated as one-shot, and a rule carrying its start date on the same line as its repeat text now follows the start you picked, on schedules and calendar events alike. [#29262](https://github.com/open-webui/open-webui/pull/29262)
- ⌛ **Stalled schedules and stuck tasks.** A schedule whose rule takes too long to work out no longer holds up the round that evaluates it and is skipped with a warning, and background tasks shared through Redis now expire once their worker falls silent, after "REDIS_TASK_TTL" seconds. [Commit](https://github.com/open-webui/open-webui/commit/5fb869db221d9599a576e8e2eea9c3314947d25e)
- 📣 **Model mention chips.** A mention whose ID carried anything beyond letters, digits and a little punctuation, such as the brackets in some workspace model IDs, stayed on screen as the raw "<@…>" text both in the box you type in and in the message once sent; any ID without a space in it is now drawn as a chip. [#29864](https://github.com/open-webui/open-webui/pull/29864)
- 😀 **Multi-codepoint emoji.** An emoji whose shortcode is several codepoints, the flags among them, is inserted whole, where only its first part reached the message. [#30213](https://github.com/open-webui/open-webui/pull/30213), [#30212](https://github.com/open-webui/open-webui/issues/30212)
- 🔔 **Custom webhook names.** A webhook target named with a space or a slash is now tidied the way an automatic name always was, so it can still be edited, deleted, made the default or tested afterwards. [#29947](https://github.com/open-webui/open-webui/pull/29947)
- 🪛 **Paginated MCP tool lists.** A server that hands its tools back a page at a time had only the first page read, so the rest were never offered to a model; the whole list is now collected before the tools are built. [Commit](https://github.com/open-webui/open-webui/commit/ffae4116a8d58a820f1770c41c69dafe4d51c881)
- 💭 **Anthropic thinking blocks.** Open WebUI's own thinking blocks are no longer forwarded to an OpenAI-compatible backend, which since 0.11.0 made a strict server such as NVIDIA Dynamo refuse an Anthropic client's second turn; signed blocks still pass through. [#29849](https://github.com/open-webui/open-webui/pull/29849), [#29799](https://github.com/open-webui/open-webui/issues/29799)
- 🗨️ **Channel model terminals.** A model with a terminal chosen in the workspace had that choice honoured in a chat but dropped where it answered in a channel or ran as a channel automation, so it worked without one; it now carries the same terminal everywhere, alongside the tools, filters and features it already carried. [Commit](https://github.com/open-webui/open-webui/commit/c78ad89934095c4e42e3f059d400a24fe5681de2)
- 🧺 **Deleted channel messages leave no quotes.** Deleting a channel message now clears the quote of it sitting on every reply and drops it from the reply box, where a reply kept showing the deleted message and could still be sent addressed to it. [#30314](https://github.com/open-webui/open-webui/pull/30314), [#30313](https://github.com/open-webui/open-webui/issues/30313)
- 🔌 **Terminal tools need a terminal.** A model was offered the tools that read your terminal and type into it whether or not the chat had a terminal switched on and connected in your browser, so it could reach for one that was not there; those tools are now handed over only for the terminal the chat has open. [Commit](https://github.com/open-webui/open-webui/commit/ca1eefe2937081b010ffef3985997d17d3332fa3), [Commit](https://github.com/open-webui/open-webui/commit/1ddba7e2c6f625fbc2c131eb24d3b9775ad898ad)
- 🔤 **Custom header encoding.** The custom headers a connection sends are encoded once the values are filled in, so a person's name or group carrying anything beyond plain ASCII, a line break included, no longer breaks the request or reaches the other end as something else. [Commit](https://github.com/open-webui/open-webui/commit/7a4a4b93dca34f8ce0c481b180d3eea23797e984)
- 🆎 **Chromium spellcheck corrections.** Picking a suggestion from the browser's own spelling menu in the message box did nothing, or put the misspelling straight back, because a highlight meant for the notes editor was being drawn over the selection and rebuilding the text underneath it, taking the browser's spelling marks with it; that highlight is now kept out of the message box and only drawn where the editor is not in use. [#29952](https://github.com/open-webui/open-webui/pull/29952), [#29944](https://github.com/open-webui/open-webui/issues/29944)
- 🈁 **IME composition while renaming.** Confirming a chat rename with Enter or Escape part way through typing with an input method editor now finishes the composition without saving or cancelling the rename, where the key press acted at once. [Commit](https://github.com/open-webui/open-webui/commit/85146206f60a22385ed27dae30d4f00e7e2675eb)
- 🐳 **Dotless host addresses.** With local web fetching turned on, an address pointing at a container name on the same network, "http://apprise:8000" and the like, is now accepted rather than refused as invalid. [#29945](https://github.com/open-webui/open-webui/pull/29945), [#28161](https://github.com/open-webui/open-webui/issues/28161)
- 🪧 **False skill mentions.** Something written as "<$fh>" in a message, as Perl and other languages do, was taken for a mention of a skill and quietly removed before the model saw it, and drawn on screen as a chip; only mentions naming a skill that exists and is turned on are treated as mentions now. [Commit](https://github.com/open-webui/open-webui/commit/0edd731c7422870aa109fd0b758c6d68c06655da)
- 🎒 **Skill settings survive saving.** Saving a skill from its editor cleared the tags and translations it carried and switched it back on where it had been disabled, and all of that now survives the save. [#30185](https://github.com/open-webui/open-webui/pull/30185)
- 🫥 **Webhook avatar forwarding.** Turning "ENABLE_PROFILE_IMAGE_URL_FORWARDING" off stops browsers being sent on to outside picture addresses, and a channel webhook's picture was sent on regardless; it now serves the built-in picture like the rest, in the webhooks dialog as well as the message list, where the dialog had been sending every viewer's browser straight to the outside address. [#29889](https://github.com/open-webui/open-webui/pull/29889), [#29892](https://github.com/open-webui/open-webui/pull/29892)
- 🪟 **Statistics window origin.** The window that shares chat statistics with the community accepted requests from any page that opened it and answered to anywhere; it now reads and replies only where the community site is at the other end. [#29918](https://github.com/open-webui/open-webui/pull/29918)
- 🖼️ **Tool image rendering.** Where a tool answered with an image tucked inside an object or a list rather than on its own, the image was written into the conversation as its raw text, a single screenshot costing hundreds of thousands of tokens and crowding out everything else; such an image is now taken out wherever it sits and attached to the reply, so the model is handed the picture and you see it. An older fault that let every second image through untouched goes with it, and in a saved chat such an image is now kept as a file and referred to rather than written into the conversation itself, so the chat stays small. [#29665](https://github.com/open-webui/open-webui/pull/29665), [#29208](https://github.com/open-webui/open-webui/issues/29208), [Commit](https://github.com/open-webui/open-webui/commit/d372bec70427fe2d568e052ce5e1529e2ad41da9)
- ⏱️ **Stopped reply state.** Pressing stop saved the reply as finished while the parts inside it were still marked as running, so a block went on reading "Thinking..." or "Executing..." and came back that way after every reload; those parts are now closed off as the reply is stopped, and an open tab settles at once rather than only after a reload, while a tool call still waiting for your approval keeps its prompt. [#29495](https://github.com/open-webui/open-webui/pull/29495), [#29281](https://github.com/open-webui/open-webui/issues/29281)
- 🩺 **Knowledge sync errors.** A knowledge base sync that fails now names the file it happened on and what the browser said, rather than reporting nothing beyond "Error accessing directory". [#29507](https://github.com/open-webui/open-webui/pull/29507)
- 🚧 **Terminal proxy restrictions.** Requests passed through to a terminal server are now refused where they aim at that server's administrative endpoints, are not followed on to somewhere else, and are turned away where the path carries characters a parser would rewrite. [Commit](https://github.com/open-webui/open-webui/commit/51bb8cb142f72503e861eeee25ae4dc73c26c36b)
- ⛔ **Malformed tool calls.** Where a model asked for a tool with arguments that were not an object at all, a bare list or string, the reply stopped there; the model is now told what was wrong with the call and can try again. [Commit](https://github.com/open-webui/open-webui/commit/fed94c9f5af8a59660425d52df09e15fbedb25bc)
- 📡 **Broken stream reporting.** Where something failed part way through streaming an answer out of "/api/chat/completions", the stream simply stopped, leaving a client waiting on an answer that would never finish; it now closes with an error and a proper end of stream. [Commit](https://github.com/open-webui/open-webui/commit/c0fb36c9b833a85a3a7364e195cf54f3d6c7a787)
- 🧷 **Chat unblocked after an error.** A reply that failed, on a content filter or an exhausted quota, left the chat turning away everything you typed after it and stopped the message queue. Only the failed reply now ends, so the chat carries on and the other replies in a multi-model answer keep writing. [Commit](https://github.com/open-webui/open-webui/commit/dbb17a5725f9d7f844a6eee63ffca0bd077c7d94)
- 🫙 **Empty failed replies in history.** An assistant turn that ended in an error with nothing written is no longer handed back to the model as part of the conversation when you send your next message. [Commit](https://github.com/open-webui/open-webui/commit/dbb17a5725f9d7f844a6eee63ffca0bd077c7d94)
- 📭 **Empty page uploads.** Adding a web address to a knowledge base that came back without any text failed with a bare "Error uploading file" and, where the upload itself was refused, left the row sitting in the list; the reason now reaches you as it was given, and the row is taken away. [Commit](https://github.com/open-webui/open-webui/commit/6786ae1797eadaad7464a147213790e2d272822b)
- 🎞️ **Tool embed scope.** The frames a tool call can ask to have shown, which run scripts of their own, were drawn wherever a message was rendered, a channel among them; they are now drawn only in the replies of the chat you are in, and never in a channel. [#29985](https://github.com/open-webui/open-webui/pull/29985)
- 🚰 **Rejected picture addresses.** A model entry carrying a picture address the server refuses no longer leaves that address in memory, where anyone signed in could pile them up. [#29971](https://github.com/open-webui/open-webui/pull/29971)
- 🖌 **Editing a stored image.** Asking a model to edit an image this instance already holds now works whatever host its address names, where a container name, a default port or a self-composed host made the edit fail with a generic loading error. [#29691](https://github.com/open-webui/open-webui/pull/29691), [#29220](https://github.com/open-webui/open-webui/issues/29220)
- 🧪 **Memory replies carry less.** The memory tool no longer hands the model each memory's stored metadata, and a memory now records the model's id rather than the whole model entry. [Commit](https://github.com/open-webui/open-webui/commit/e9a0164690a8b1e190bdc8f4613e9d918b26d327)
- 🔇 **Memory fully off.** With memory switched off, stored memories are no longer folded into a reply's context and the memory tools are no longer offered to the model, where both went on reaching it behind the switch. [#30228](https://github.com/open-webui/open-webui/pull/30228), [#30227](https://github.com/open-webui/open-webui/issues/30227)
- 🧠 **Memory review behind the switch.** The background review that drafts new memories from a conversation no longer runs when memories are switched off or the account is barred from them, where it went on spending a task-model call every interval turn and failing at the write. [#30309](https://github.com/open-webui/open-webui/pull/30309)
- 🏗️ **Terminal server save button.** Saving a terminal server now waits for the save to finish before the dialog closes and cannot be set off twice by a second click. [Commit](https://github.com/open-webui/open-webui/commit/1cdd7aa459d6e96905324b452600ff56369d8a4e)
- 🗺️ **Terminal file panel paths.** The file panel beside a terminal now opens the file a model just wrote even when it is named with a relative path, where the panel could not match the name, jumped to the root and dragged the session's working directory with it. [#30282](https://github.com/open-webui/open-webui/pull/30282), [#30051](https://github.com/open-webui/open-webui/issues/30051)
- 🍴 **Forked chat folder.** Forking a chat put the copy in the original's folder even where you cannot write to that folder; it is now created outside any folder unless you can. [#30069](https://github.com/open-webui/open-webui/pull/30069)
- 🪢 **Dropping a folder in place.** Dragging a folder onto the folder it already sits in failed with "Folder already exists", and is now taken for the no-op it is. [#30169](https://github.com/open-webui/open-webui/pull/30169)
- 🗝️ **Read-only folders read-only everywhere.** A folder shared with you as a viewer no longer shows its edit controls on the empty-chat page, where they appeared and a save went through or failed depending on rights the page never checked. [Commit](https://github.com/open-webui/open-webui/commit/ee3ece1e2b8c9a38c94faf354ca020d66aba801d)
- 🚿 **Deleting a folder, keeping chats.** Removing a folder while keeping the chats inside it was refused for an account not allowed to delete chats, even though nothing was being deleted, and it now goes through. [#30163](https://github.com/open-webui/open-webui/pull/30163)
- 🏷️ **Tag cleanup after deletion.** An administrator deleting someone else's chat tidied unused tags out of their own account rather than the owner's, leaving the owner with tags nothing points at. [#30171](https://github.com/open-webui/open-webui/pull/30171)
- 🌱 **Forking past an unfinished reply.** A chat that held an interrupted reply anywhere in it refused every fork from then on and never recovered; forking now waits only on a reply actually being generated, and a turn paused for tool approval still forks with its prompt showing. [#30131](https://github.com/open-webui/open-webui/pull/30131), [#30128](https://github.com/open-webui/open-webui/issues/30128)
- 💾 **Deleted tool memory.** Deleting a tool or a function left the whole of its code in memory for as long as the server ran; it is now let go of along with the rest. [#29983](https://github.com/open-webui/open-webui/pull/29983)
- 🐌 **Sign-in rate limiting.** Counting sign-in attempts through Redis no longer stops the whole worker until Redis answers, so a slow Redis stops freezing every other request with it. [#29977](https://github.com/open-webui/open-webui/pull/29977)
- ♻️ **Session pool cleanup.** The task that clears out abandoned websocket sessions now carries on through an error from Redis instead of ending for good, and stops properly at shutdown. [#29976](https://github.com/open-webui/open-webui/pull/29976), [#29979](https://github.com/open-webui/open-webui/pull/29979)
- 🛰️ **Direct connections across workers.** A reply streamed over a direct connection no longer goes quiet part way through where several servers share their websocket traffic through Redis; the events it lives on now travel between workers the way the rest already did. [Commit](https://github.com/open-webui/open-webui/commit/0180efecf362d487e0c30f040f5948c325fbe337)
- 🫧 **Empty document ids.** A save arriving for a document with no id at all was filed against that empty id and never cleared, so anyone signed in could pile them up; nothing is filed for it now. [#29980](https://github.com/open-webui/open-webui/pull/29980)
- 🔬 **Page fetch CPU spin.** Fetching a page through the browser-driven loader never returned where that page opened a WebSocket, holding a worker thread at full CPU for the life of the process and costing another core on every further fetch, which left the whole instance slow. [#30050](https://github.com/open-webui/open-webui/pull/30050), [#30024](https://github.com/open-webui/open-webui/issues/30024)
- 🔁 **Duplicate search tracebacks.** A vector database outage wrote a full traceback twice for every collection and query pair, turning one outage into hundreds of identical stack traces per message on every replica; a single record now names every collection that failed. [#29981](https://github.com/open-webui/open-webui/pull/29981)
- 🧯 **Page fetch logging.** Fetching a page through the browser-driven loader filled the log with tracebacks where the page closed while it was still pulling pieces of itself, as sites behind Cloudflare and similar do; those requests are now let go of before the page closes. [#29325](https://github.com/open-webui/open-webui/pull/29325), [#28869](https://github.com/open-webui/open-webui/issues/28869)
- 📤 **Tool export scope.** Exporting all tools at once returned every tool the account could see, the source of a tool shared for reading included; it now returns only the tools it may edit, matching the single-tool export and the way models already export. [#29310](https://github.com/open-webui/open-webui/pull/29310)
- 🗳️ **Partial workspace exports.** Exporting the prompts or the models from the workspace wrote out only the page you happened to be looking at, so most of them were quietly left out of the file; both now export everything you are allowed to. [#30187](https://github.com/open-webui/open-webui/pull/30187)
- 🧳 **Imported chats keep more.** A chat brought back from an export arrived unpinned and unarchived and without the variables it was saved with, and all three now survive the round trip. [#30155](https://github.com/open-webui/open-webui/pull/30155), [#30151](https://github.com/open-webui/open-webui/pull/30151)
- 🗂️ **Unarchiving from search.** The menu on a search result offered to archive a chat that was already archived and said it had been archived when it had been brought back, and it now names and reports whichever of the two it did. [#30177](https://github.com/open-webui/open-webui/pull/30177)
- 🙈 **Folder filters with no match.** A chat search narrowed by a folder name that matches no folder now finds nothing, where the folder filter was quietly dropped and every chat came back. [#30273](https://github.com/open-webui/open-webui/pull/30273), [#29959](https://github.com/open-webui/open-webui/discussions/29959)
- 🧹 **Sidebar after bulk actions.** Archiving, deleting or unarchiving every chat at once, or importing a batch of them, left the folders and the pinned chats in the sidebar showing what was no longer there until the page was reloaded, and a bulk action that failed no longer reports success. [Commit](https://github.com/open-webui/open-webui/commit/8b3ee2827241ccc952a3073b2a6bbfad5df01827)
- 🔐 **Model pictures follow model access.** The picture belonging to a model is now shown only to people who can see that model, where anyone signed in could fetch it and tell an existing model from an unknown one by which picture came back. [#29700](https://github.com/open-webui/open-webui/pull/29700)
- 🚪 **Webhook pictures follow channel access.** The picture belonging to a channel webhook is now shown only to people with access to that channel, where anyone signed in could fetch it or be sent on to wherever it pointed, and it is refused outright where channels are turned off. [#29703](https://github.com/open-webui/open-webui/pull/29703)
- 📎 **Safer Word document previews.** Previewing a Word document no longer renders an HTML sub-document embedded inside it, and a link in one opens only where it points at a web address, a mail address or a telephone number. [#29699](https://github.com/open-webui/open-webui/pull/29699)
- 🚦 **Citation link schemes.** A source attached to a reply now opens only where it points at a web address, falling back to the panel that shows the source rather than following anything else. [#29701](https://github.com/open-webui/open-webui/pull/29701)
- 🈚 **Citation chips inside formatted text.** A citation inside bold, italic or linked text now renders its chip, where the formatting took it and the citation vanished from the sentence. [#30278](https://github.com/open-webui/open-webui/pull/30278), [#30277](https://github.com/open-webui/open-webui/issues/30277)
- 🐍 **Saving a tool or function.** Saving a tool or function in the admin pages no longer fails with a missing module error from the built-in code formatter, which was not installing everything it needed. [#29503](https://github.com/open-webui/open-webui/pull/29503)
- 🛠️ **Tool request duplication.** A tool that writes through an address carrying part of its input no longer has that part repeated in the body of the request as well, which servers checking their input strictly turned away, so those calls now go through. [#29717](https://github.com/open-webui/open-webui/pull/29717), [#29716](https://github.com/open-webui/open-webui/issues/29716)
- 🍎 **Answers survive on Apple Silicon.** Asking a question that searches a knowledge base with a locally run reranking model no longer takes the whole server down on a Mac, losing the answer and the connection with it. [#29735](https://github.com/open-webui/open-webui/pull/29735), [#29722](https://github.com/open-webui/open-webui/issues/29722)
- 📚 **Web results stop being cited.** Pages a web search only listed are no longer offered to the model as things to cite, which had it attaching a result id to text from a different result and the citations panel resolving that to a title that looked right. [#29631](https://github.com/open-webui/open-webui/pull/29631), [#29627](https://github.com/open-webui/open-webui/issues/29627)
- 🔎 **SearchApi errors, news and links.** Web search through searchapi.io now reports a bad key instead of coming back empty, reads the news results it returns alongside its ordinary ones, and hands the web loader the resolved destination link, so citations stop pointing at a redirect page. [#30308](https://github.com/open-webui/open-webui/pull/30308), [#30305](https://github.com/open-webui/open-webui/issues/30305)
- 🔼 **Honest version checks.** An instance that cannot reach the release listing now says the check failed, instead of reporting whatever it is running as the newest version and recording nothing about it. [#29626](https://github.com/open-webui/open-webui/pull/29626), [#29580](https://github.com/open-webui/open-webui/issues/29580)
- 🏟️ **Arena models report their errors.** A message to an arena model whose provider answers with an error now shows that error in the chat, where it used to fail on something unrelated and leave the real reason unsaid, and titles and tags no longer break the same way. [#29662](https://github.com/open-webui/open-webui/pull/29662), [#29658](https://github.com/open-webui/open-webui/issues/29658)
- 🏳️ **Nameless tool calls fail once.** A model endpoint that sends a tool call with no name at all now has that call fail on the spot, rather than the missing name being kept, stored with the message and sent back on the next turn for the endpoint to reject. [#29690](https://github.com/open-webui/open-webui/pull/29690), [#29686](https://github.com/open-webui/open-webui/issues/29686)
- 🧹 **Direct connections stop leaking listeners.** A server talking to a direct connection no longer leaves a listener behind for every request that ends any way but a clean finish, which grew without limit while a connection kept failing. [#29509](https://github.com/open-webui/open-webui/pull/29509)
- 🚀 **Cheaper model refreshes.** The model registry shared through Redis is now written only when the models themselves change, rather than on every refresh because of the countdown Ollama attaches to a model it holds in memory. [Commit](https://github.com/open-webui/open-webui/commit/649c012ecf308a994ea180127f7f8f94d0aec311)
- ✍️ **Continued reply text.** Asking for the rest of a cut-off reply in a temporary chat replaced what was on screen with only the new text, because the message being continued was read back from the saved chat it did not have. It is now taken from the request before the model is called, the continuation joins the same message instead of arriving as a second one, on a connection whose provider is set to llama.cpp the model is told to carry on from the text it is handed rather than repeat it back, opening the result in the message editor no longer shows a line break where the two halves meet, and, where haptic feedback is switched on, a continuation buzzes as it streams like any other reply. [Commit](https://github.com/open-webui/open-webui/commit/77d2000eb79e1cb6ae2004d40e8cca8c9e754cd0), [Commit](https://github.com/open-webui/open-webui/commit/6c7aa3543d21442241f6c53add5ff623ec816c44), [Commit](https://github.com/open-webui/open-webui/commit/57fc344873edc0db9e9f7ff3e9fb167cd80e3ef2), [Commit](https://github.com/open-webui/open-webui/commit/7eefeef4f17118f81c87acb464ad562803a6f26c), [Commit](https://github.com/open-webui/open-webui/commit/d418840aa9c4b77f613308cdaa4f2c6a062a8715), [Commit](https://github.com/open-webui/open-webui/commit/3795d5b29253d4b8d7a0adbab457c7317c40f3c6), [Commit](https://github.com/open-webui/open-webui/commit/57acc2b68f2f9e40b53aa7e609fdd52a4b0d15c4)
- 🔗 **Cancelled edits keep attachments.** Cancelling the edit of a message no longer strips the files attached to it, where dropping the edit took the attachments down with it. [#30281](https://github.com/open-webui/open-webui/pull/30281), [#30192](https://github.com/open-webui/open-webui/issues/30192)
- 🏎️ **Faster media page reads.** The browser-driven loader pulled every image, video and font a page referenced down through the server before any text was extracted, so a page carrying a few dozen audio players took ten seconds or timed out. Those requests are now dropped before they are made, and the same page comes back in under three seconds, having pulled 4 MB where it used to pull 55. [#29742](https://github.com/open-webui/open-webui/pull/29742), [#29741](https://github.com/open-webui/open-webui/issues/29741)
- 📝 **Starting a note from search.** Starting a note from the search box now works when you are already on the notes page, keeps the whole of what you typed including characters such as ampersands and hashes, and no longer makes a further note each time the browser back button is pressed. [#29645](https://github.com/open-webui/open-webui/pull/29645), [#29642](https://github.com/open-webui/open-webui/issues/29642)
- 📱 **Apple device replies.** An assistant reply no longer comes up blank in a home screen app, an in-app browser or a desktop-class window on Apple devices, where the check that avoided the drawing fault only recognised Safari itself. [#29734](https://github.com/open-webui/open-webui/pull/29734), [#29688](https://github.com/open-webui/open-webui/issues/29688), [#26712](https://github.com/open-webui/open-webui/issues/26712)
- 📊 **Single source relevance.** A reply drawing on a single source now shows how relevant that source is, where the figure appeared only once a second source joined it and so looked as though it came and went. [#29647](https://github.com/open-webui/open-webui/pull/29647), [#29646](https://github.com/open-webui/open-webui/issues/29646)
- 🔧 **Arduino sketches upload to knowledge.** A sketch file now reaches the plain text reader like the C++ and header files beside it, rather than being handed to a document extraction server that could make nothing of it and failing the upload. [#29673](https://github.com/open-webui/open-webui/pull/29673), [#29670](https://github.com/open-webui/open-webui/issues/29670)
- 📰 **Docling conversion failures.** A file that Docling refuses or fails to convert now fails the upload with the reason Docling gave, rather than breaking with a raw error or quietly filing a placeholder that was then indexed and handed to the model in place of the file. [#30107](https://github.com/open-webui/open-webui/pull/30107), [#29808](https://github.com/open-webui/open-webui/issues/29808)
- 📄 **Uploaded text kept as written.** A file whose text contains escape sequences such as the one standing for a non-breaking space is now stored and read by the model exactly as it was written, rather than having some of them rewritten depending on where in the file they sat. [#29736](https://github.com/open-webui/open-webui/pull/29736), [#29732](https://github.com/open-webui/open-webui/issues/29732)
- 🔦 **Readable slash command labels.** The entries in the slash command menu no longer show as white text on a white background in the light theme. [#29512](https://github.com/open-webui/open-webui/pull/29512), [#29510](https://github.com/open-webui/open-webui/issues/29510)
- ⌨️ **Literal arrow sequences.** A sequence such as three hyphens after a less-than sign is now shown as the characters it is made of rather than drawn as an arrow, which had text look changed when it never was. [#29595](https://github.com/open-webui/open-webui/pull/29595), [#29594](https://github.com/open-webui/open-webui/issues/29594)
- 🖌️ **Editing an image you uploaded.** An image already held by Open WebUI can now be used with image editing, where fetching its own link back over the network could fail on a private network or without a sign-in. [Commit](https://github.com/open-webui/open-webui/commit/50413f34824ea49d5b94d3a97f3fe4bb2e881e38)
- 🖼 **Playground image edits.** Editing an image in the Images playground now works, where every attempt came back rejected since the request carried its fields under a heading the endpoint never read. [Commit](https://github.com/open-webui/open-webui/commit/c07fa08b995e8d1a1fc2d94a88d8cea691bdc5ee)
- 🎨 **A tidier attach webpage dialog.** The row holding the Add button no longer carries a grey band of its own between the address box and the button, matching every other dialog. [#29664](https://github.com/open-webui/open-webui/pull/29664), [#29663](https://github.com/open-webui/open-webui/issues/29663)
- 🖱️ **A plain note date.** The date under a note title no longer shows a pointing hand or announces itself as something to press, having never done anything when clicked. [#29708](https://github.com/open-webui/open-webui/pull/29708)
- 🌇 **Folder backgrounds on creation.** The background image picked in the Create Folder dialog now arrives on the folder when it is created from the sidebar, where the image was discarded unless the dialog was opened from an existing folder. [#30218](https://github.com/open-webui/open-webui/pull/30218), [#30217](https://github.com/open-webui/open-webui/issues/30217)
- 🖇 **Delete Chat shortcut everywhere.** The keyboard shortcut that deletes the open chat now works wherever the chat was opened from, where it did nothing unless the chat's row happened to be on screen in the sidebar at that moment. [#30165](https://github.com/open-webui/open-webui/pull/30165), [#30164](https://github.com/open-webui/open-webui/issues/30164)
- 🚮 **Retired chat variables.** A model whose system prompt no longer declares a variable a previous prompt did stops asking for it, where every new chat kept opening the dialog and refusing to send until a value was entered. [#30173](https://github.com/open-webui/open-webui/pull/30173), [#30172](https://github.com/open-webui/open-webui/issues/30172)
- 👁️ **Compact hover previews in Safari.** Holding over a chat in the sidebar shows the small preview every other browser shows, rather than one laid out at the width and spacing of a full conversation. [#29734](https://github.com/open-webui/open-webui/pull/29734)
- 🔖 **Visible title generation faults.** When a new chat's automatic title cannot be generated, the reason now reaches the log at the default level, rather than only under debug logging where a broken feature looked the same as a switched-off one. [#30106](https://github.com/open-webui/open-webui/pull/30106), [#29533](https://github.com/open-webui/open-webui/issues/29533)
### Changed
- 🪶 **Slim starts with nothing configured.** The slim image leaves out the local models and the libraries around them, and still starts and holds a conversation on its defaults; the features that leaned on those models each need an external service of their own. [Commit](https://github.com/open-webui/open-webui/commit/cb942bb94c8dc7941336088fb3392e2398ff56c1)
- 🗄️ **Slim database support.** The slim image runs on SQLite, its default, or on PostgreSQL; pointed at MySQL, MariaDB or another engine, or started with AWS RDS IAM logins switched on, it stops with an error instead of starting, and those deployments need the standard image. [Commit](https://github.com/open-webui/open-webui/commit/d27aa72ab4a7b5632b4ad49e8467081ad3d7ebb4)
- 📁 **Slim file storage.** The slim image keeps files on local storage, its default; configured for an S3, Google Cloud or Azure bucket, it stops with an error instead of starting, and those deployments need the standard image. [Commit](https://github.com/open-webui/open-webui/commit/d27aa72ab4a7b5632b4ad49e8467081ad3d7ebb4)
- 🧮 **Slim searches only through pgvector.** Knowledge search on the slim image needs PostgreSQL with pgvector, and configured for another vector store the instance still starts, and the failure arrives the first time something is searched rather than at startup. [Commit](https://github.com/open-webui/open-webui/commit/cb942bb94c8dc7941336088fb3392e2398ff56c1), [Commit](https://github.com/open-webui/open-webui/commit/d27aa72ab4a7b5632b4ad49e8467081ad3d7ebb4)
- 🧠 **Slim embedding requirements.** The slim image carries no embedding or reranking model, so knowledge needs OpenAI, Ollama or Azure OpenAI embeddings and an external reranker, falling back to plain cosine scoring where none is set. [Commit](https://github.com/open-webui/open-webui/commit/cb942bb94c8dc7941336088fb3392e2398ff56c1)
- ✂️ **Slim document splitting.** Splitting a document along a downloaded tokenizer is unavailable on the slim image, which leaves splitting by character or by token count. [Commit](https://github.com/open-webui/open-webui/commit/cb942bb94c8dc7941336088fb3392e2398ff56c1)
- 📃 **Slim document readers.** The slim image reads text, Markdown, CSV, HTML and XML files as they are; uploading a PDF, a Word file or a presentation fails unless one of the external document extractors is configured. [Commit](https://github.com/open-webui/open-webui/commit/cb942bb94c8dc7941336088fb3392e2398ff56c1)
- 🎙️ **Slim speech requirements.** The slim image carries neither local Whisper nor local voices, so speech to text and text to speech need an external engine before they will work. [Commit](https://github.com/open-webui/open-webui/commit/cb942bb94c8dc7941336088fb3392e2398ff56c1)
- 🕸️ **Slim web page fetching.** The slim image carries no headless browser, so a web page is fetched over plain HTTP or through an external loader, and a page that draws itself with JavaScript comes back with less of its content than on the standard image. [Commit](https://github.com/open-webui/open-webui/commit/cb942bb94c8dc7941336088fb3392e2398ff56c1), [Commit](https://github.com/open-webui/open-webui/commit/d27aa72ab4a7b5632b4ad49e8467081ad3d7ebb4)
- 🔎 **Slim leaves out DDGS.** DDGS, the metasearch provider that needs no key of its own, is not carried in the slim image, so web search there needs a provider with a key. [Commit](https://github.com/open-webui/open-webui/commit/0fa4dea5ff64ea663f162d07c9a1175c39e4aad0)
- 📥 **Slim code interpreter packages.** The slim image leaves out the code interpreter's packages, so the browser fetches numpy, pandas, matplotlib, scikit-learn and the rest from "cdn.jsdelivr.net" and the interpreter stops working where that is blocked. [Commit](https://github.com/open-webui/open-webui/commit/98fcb844e1b19f7dd6289af26273cdec5447dc52)
- 🧰 **Slim git requirements.** The slim image no longer carries git, so a tool or function whose requirements point at a "git+https://" address fails to install and needs the standard image or a published package. [Commit](https://github.com/open-webui/open-webui/commit/30eed1251301f74e0dfeaad09c41e81320f790fc)
- 🧺 **LangChain community removal.** The readers for text, HTML, Word, CSV, PDF and Azure Document Intelligence are now written here rather than taken from "langchain-community", which is no longer installed; a tool or function importing it has to name it in its own requirements from now on. [Commit](https://github.com/open-webui/open-webui/commit/05484aa055a868a49842e1ddff169c19c98b755b)
- 🧾 **Undeclared package imports.** Packages that sat in the image only by accident, among them nltk, pymongo, the Google Drive client and the Gemini SDK, are no longer installed, so a tool or function importing one must name it in its own requirements. [#29725](https://github.com/open-webui/open-webui/pull/29725), [#29726](https://github.com/open-webui/open-webui/pull/29726), [Commit](https://github.com/open-webui/open-webui/commit/a1c02098aa2687c72482a59117efe643b785df51)
- 🆔 **Model ID whitespace.** A workspace model whose ID contains a space or a tab is now refused in the editor, through the API and on import; one already stored goes on answering but cannot be saved again until it is recreated. [Commit](https://github.com/open-webui/open-webui/commit/8a19e2f867063256bb2836649e2fe80af41ef748)
- 🖇️ **Link scheme rendering.** A link in a reply, a citation or a web search result is now rendered only where it points at a web address, a mail address, a telephone number or somewhere inside this instance; anything else, an "ftp://" address or an application link such as "obsidian://" or "vscode://" among them, is shown as the text it is. Two old oddities go with it: a source written as "HTTP://" now becomes a link, and a filename merely containing the letters http no longer becomes one that leads nowhere. [#29890](https://github.com/open-webui/open-webui/pull/29890)
- 🧲 **Integrations tab is opt-in.** The Integrations tab in personal settings, where tool and terminal connections of your own are managed, is now hidden until an administrator turns on Direct Integrations under Integrations or sets "ENABLE_DIRECT_INTEGRATIONS", and hiding it leaves existing connections working. [Commit](https://github.com/open-webui/open-webui/commit/6d8e63e3666e1b1aa5540ffda0816be5f40c5271), [Commit](https://github.com/open-webui/open-webui/commit/c82634b9d01adadbf9780ff84a0a8168fdc4fdea)
- 📡 **Image connection check endpoint.** The endpoint that checks an image generation connection has moved and now takes the connection to test in the request itself, so anything calling the old address needs updating. [Commit](https://github.com/open-webui/open-webui/commit/64bbdf7a73724986fac8bcf6e880fe32ef9ac495)
## [0.11.3] - 2026-08-31
### Added
- ♿ **Accessibility mode reaches the menus.** Accessibility mode now marks the menu entry you are pointing at and the model already chosen with a stronger background, across the dropdown menus, their submenus, and the model picker together with its filter and compare controls, so those cues carry the contrast the accessibility guidelines ask for in both themes. [Commit](https://github.com/open-webui/open-webui/commit/a6f9751401589ee73208295b6f6a7f6eae9c1b44), [Commit](https://github.com/open-webui/open-webui/commit/471b5cbbb16c3996c32e68808dde6f8898f64ecd)
- 🔄 **General improvements.** Various improvements were implemented across the application to enhance performance, stability, and security.
- 🌐 **Translation updates.** Translations for Indonesian were enhanced and expanded.
### Fixed
- 💥 **Chat branches stay connected after reloads.** A reply saved under an earlier message now stays listed under that message, so branch arrows, exports, reloads, and later edits keep the whole conversation in view, and chats already saved with that link missing are repaired when opened. [#29299](https://github.com/open-webui/open-webui/issues/29299)
- 🧱 **Upgrades fail clearly instead of starting half updated.** A failed database upgrade now stops at the migration error that caused it, instead of starting anyway and reporting a missing table or column such as 'chat.timer_at' later, which is the upgrade failure seen after moving from 0.11.0, 0.11.1, or 0.11.2. [#29280](https://github.com/open-webui/open-webui/issues/29280)
- 🔤 **Custom interface fonts reach more of the app.** The font chosen in interface settings now applies to dropdowns and other interface text that previously fell back to the standard font. [Commit](https://github.com/open-webui/open-webui/commit/1457000ba66547b24bd98012aa35ac16fd4bc696)
- 🔌 **Disconnect OAuth only where there is OAuth.** The disconnect control on a tool server reached over MCP now appears only where that server signs in through OAuth and an account is connected, rather than on servers that use no sign-in at all. [#29296](https://github.com/open-webui/open-webui/issues/29296)
## [0.11.2] - 2026-08-31
### Added
- 🖼️ **Richer previews for terminal files.** Word documents and slide decks produced in the terminal are now previewed as the finished document rather than an approximation, and every document preview gains a page strip down the side with numbered thumbnails you can click to jump straight to a page, and the notice warning that a preview might differ from the download is gone now that it does not. [Commit](https://github.com/open-webui/open-webui/commit/061fb434328ea8365cc4961519e8adbe5ac7b0c1), [Commit](https://github.com/open-webui/open-webui/commit/9d95a0148b3180ba48186bc6928713f969a69047), [Commit](https://github.com/open-webui/open-webui/commit/ae549d3e4af672a63a809e8fe2a6c91e0bfc1067), [Commit](https://github.com/open-webui/open-webui/commit/78d8c9166f1b6626b9aa559a812183ce216eef4f), [Commit](https://github.com/open-webui/open-webui/commit/492ccf3ac0d9aec1f3f2bc6825a3e6d3b32aa0d1)
- ⚡ **Less overhead on every message.** Deployments without pipelines configured, which is the default, no longer pay setup work for them on each chat message and background task. [#29146](https://github.com/open-webui/open-webui/pull/29146)
- 🏎️ **Lighter model list refreshes.** On deployments running several workers, a refresh that finds the model list unchanged no longer has every worker rewrite the whole list to the shared cache, cutting the work and the traffic each refresh costs. [#29264](https://github.com/open-webui/open-webui/pull/29264)
- 🧵 **Smoother busy websocket servers.** Large Redis-backed websocket deployments now spend far less time sweeping old sessions, reading their own session data back from Redis, and checking every outgoing websocket message for file attachments Open WebUI does not send, so channel posts, collaboration updates, heartbeats, and live chat updates put less load on busy servers. [#28835](https://github.com/open-webui/open-webui/pull/28835), [#28180](https://github.com/open-webui/open-webui/pull/28180)
- 🧰 **A new request filter step.** Filter authors can now use `request` to adjust the payload right before each model call, including follow-up calls after tool use, while existing `inlet`, `stream`, and `outlet` filters keep working as before. [Commit](https://github.com/open-webui/open-webui/commit/2daa610cba514dfe2eae008b1fc440d4ddb0e470), [Commit](https://github.com/open-webui/open-webui/commit/2a4ef46ac86ada3447eb6d6c50fd4c1241eebd4d)
- 🤏 **More room in file previews on touch screens.** File previews on phones and tablets no longer show zoom buttons that sit over an already small preview, leaving pinch to zoom to do the job. [#29176](https://github.com/open-webui/open-webui/pull/29176), [#29152](https://github.com/open-webui/open-webui/issues/29152)
- 👆 **Resizing panels by touch.** The divider beside the main sidebar or a side panel can now be dragged on a touchscreen or with a stylus, and a drag made with the mouse keeps following the pointer when it leaves the window. [Commit](https://github.com/open-webui/open-webui/commit/cfa2d25317b426c80dca43829afbf113eeccc365), [Commit](https://github.com/open-webui/open-webui/commit/f0ffa7508e75946434794941378675e0c7151eae), [Commit](https://github.com/open-webui/open-webui/commit/e250be48ee16d71202651e4109a175910a6519bb)
- 🔤 **Choice of interface font.** Interface settings now offer a font family field where you can name a font installed on your computer and watch it apply across the interface as you type, with an empty field returning to the standard font. [Commit](https://github.com/open-webui/open-webui/commit/b9765fe97943831e17a68490cd01c70a6ccf1c43)
- ♿ **Wider accessibility mode coverage.** Accessibility mode now lifts the contrast of more of the interface, including muted text, table text, icons, the colours elements change to on hover, and the highlight behind sidebar and menu items you point at, so low-contrast areas meet the level the accessibility guidelines ask for in both light and dark themes, and the close button on the folder dialog now announces itself to screen readers, as do document and slide previews, which can now be reached with the keyboard. [Commit](https://github.com/open-webui/open-webui/commit/e6031ea6daec897eddbcf0c8b03c2d6fad2b506e), [Commit](https://github.com/open-webui/open-webui/commit/797b4c51d1e31f614c8e6eba14ee347c61ad0c7a), [#29160](https://github.com/open-webui/open-webui/pull/29160), [Commit](https://github.com/open-webui/open-webui/commit/e4694f82ebd2e916c099e5e921c0ca8445ba88a7), [Commit](https://github.com/open-webui/open-webui/commit/120409ef016c1e866247e2aa1eba7606f797a20c)
- 🌐 **Translation updates.** Polish, Simplified Chinese, German, Irish, Catalan, and Portuguese (Brazil) translations were expanded, corrected, and brought up to date with newer interface text. [#29029](https://github.com/open-webui/open-webui/pull/29029), [#29043](https://github.com/open-webui/open-webui/pull/29043), [#29108](https://github.com/open-webui/open-webui/pull/29108), [#29151](https://github.com/open-webui/open-webui/pull/29151), [#29184](https://github.com/open-webui/open-webui/pull/29184), [#29179](https://github.com/open-webui/open-webui/pull/29179), [#29247](https://github.com/open-webui/open-webui/pull/29247), [#29217](https://github.com/open-webui/open-webui/pull/29217)
### Fixed
- 🛡️ **Security Advisory**: This release includes security and access-control fixes. We recommend updating production deployments at your earliest convenience. Not all security fixes in this version may be enumerated in the fixed section. Some may be withheld for a short time to give administrators time to upgrade. [Advisories](https://github.com/open-webui/open-webui/security)
- 💥 **Replies stop breaking mid-chat.** A conversation no longer fails in the browser and stops showing the assistant reply until you reload, and chats already saved in that state display and can be edited again. [#29250](https://github.com/open-webui/open-webui/pull/29250), [#29244](https://github.com/open-webui/open-webui/issues/29244)
- ⌨️ **Preview shortcuts stay in the preview.** Arrow keys now page a document or slide preview only while that preview has focus, instead of paging it from anywhere on the page and taking the arrow keys away from whatever you were typing in. [Commit](https://github.com/open-webui/open-webui/commit/b356b80f8c86d110951301b83607de245b4ef6cc), [Commit](https://github.com/open-webui/open-webui/commit/e4694f82ebd2e916c099e5e921c0ca8445ba88a7)
- 🧊 **Streaming no longer stalls on reasoning models.** A reply from a reasoning model now streams through to the end instead of showing its first few words and then freezing until generation finishes. [#29053](https://github.com/open-webui/open-webui/pull/29053), [#29035](https://github.com/open-webui/open-webui/issues/29035)
- 💭 **Thinking stays in Thoughts.** After a model uses a tool, its reasoning for the next step now appears in the collapsed Thoughts section instead of being written into the reply as ordinary text. [#29052](https://github.com/open-webui/open-webui/pull/29052), [#29040](https://github.com/open-webui/open-webui/issues/29040)
- 🧠 **Large tool calls stream faster through Anthropic-compatible clients.** When a streamed tool call sends a large block of arguments, Open WebUI now waits until that block can actually be complete before checking it, instead of doing the same expensive check again on every tiny piece. [#28858](https://github.com/open-webui/open-webui/pull/28858)
- 🛑 **Stop works across instances.** Stopping a response now takes effect on deployments that spread people across several instances backed by a Redis cluster, instead of the reply continuing to the end regardless. [#29165](https://github.com/open-webui/open-webui/pull/29165), [#19840](https://github.com/open-webui/open-webui/issues/19840)
- 🎯 **Fewer needless chat reloads.** A conversation holding an older unfinished reply no longer reloads itself each time some other response in it finishes, reloading only for the reply the update actually concerns. [Commit](https://github.com/open-webui/open-webui/commit/22379ded1aebf0537e1ee3f00a20809d47646bf6)
- 📖 **Banners with underlined text.** A banner containing underlined text now displays instead of failing to render. [#29118](https://github.com/open-webui/open-webui/pull/29118), [#29115](https://github.com/open-webui/open-webui/issues/29115)
- 🧰 **Pinned models start with their own tools.** Starting a chat from a pinned model in the sidebar now applies that model's tools and skills instead of carrying over the ones from the model you used last. [#29058](https://github.com/open-webui/open-webui/pull/29058), [#29050](https://github.com/open-webui/open-webui/issues/29050)
- ❓ **Rejected ask_user calls no longer kill the reply.** If a model asks the user a question in a way Open WebUI refuses, the model now receives that error and can continue or retry, instead of leaving the chat stopped with no answer. [#29252](https://github.com/open-webui/open-webui/pull/29252), [#29077](https://github.com/open-webui/open-webui/issues/29077)
- 🚫 **Disabled models no longer vanish.** Turning a model off in the admin Models list now keeps it in view so you can turn it back on, instead of it disappearing with no way to recover it. [#29037](https://github.com/open-webui/open-webui/pull/29037), [#29036](https://github.com/open-webui/open-webui/issues/29036)
- ✂️ **Message text kept intact.** Text containing angle brackets and a dollar sign is no longer mistaken for a skill mention and silently removed before your message reaches the model. [#29051](https://github.com/open-webui/open-webui/pull/29051), [#29041](https://github.com/open-webui/open-webui/issues/29041)
- 📆 **Moved calendar events stay visible.** Changing an event's date no longer makes it disappear from the calendar, and events already stuck in that state show up again. [#29085](https://github.com/open-webui/open-webui/pull/29085), [#29067](https://github.com/open-webui/open-webui/issues/29067)
- 🔁 **Repeats follow the event.** A repeating event now works out its occurrences from its own date and time, where a repeat rule carrying a start date of its own could place them on the wrong weekday or at the wrong hour. [Commit](https://github.com/open-webui/open-webui/commit/a93c5080380447e3ab9113cacb73b61c57cbc308)
- 🔟 **Repeating automations keep their count.** An automation set to run a fixed number of times is no longer shown as running once and quietly rewritten to a single run the next time you open and save it. [#29261](https://github.com/open-webui/open-webui/pull/29261), [#29263](https://github.com/open-webui/open-webui/pull/29263)
- 🗓️ **Schedules survive a save.** An automation carrying a start date now shows its real schedule in the list and keeps its weekly or monthly setting when saved, instead of showing raw rule text and falling back to daily. [#29263](https://github.com/open-webui/open-webui/pull/29263)
- 🧩 **Custom automation schedules stay custom.** Reopening an automation with a custom repeat rule no longer loses the custom rule from the editor. [#29260](https://github.com/open-webui/open-webui/pull/29260)
- 🔢 **Accurate admin user counts.** The counts on the admin Users tabs now follow your search and reset when you switch tabs, instead of showing stale or unfiltered numbers. [#29080](https://github.com/open-webui/open-webui/pull/29080), [#29079](https://github.com/open-webui/open-webui/issues/29079)
- 🧷 **Sidebar folders stop acting stale.** A folder that has been removed from the sidebar is now removed from the internal sidebar registry too, so later sidebar actions do not target a folder that is no longer on screen. [#29121](https://github.com/open-webui/open-webui/pull/29121)
- 📱 **Readable model list on small screens.** The model identifier and timestamp in the workspace model list no longer overlap the model name on narrow displays. [#29084](https://github.com/open-webui/open-webui/pull/29084), [#29083](https://github.com/open-webui/open-webui/issues/29083)
- 🧱 **Long setting values stop crushing labels.** Settings rows now let long controls shrink inside the available space instead of squeezing the label beside them. [#29229](https://github.com/open-webui/open-webui/pull/29229)
- 🧭 **Model defaults panels stay usable.** The model capabilities and prompt suggestion sections in admin settings now scroll inside their panels instead of growing past the available space. [#29235](https://github.com/open-webui/open-webui/pull/29235)
- 📲 **Dropdowns stay on screen.** A dropdown near the edge of a narrow screen now shifts inward to stay fully visible instead of opening past the edge with its option labels cut off. [#29226](https://github.com/open-webui/open-webui/pull/29226), [#29225](https://github.com/open-webui/open-webui/issues/29225)
- 🌗 **Dropdown lists follow the theme.** The list a dropdown opens now takes its light or dark colouring from the rest of the interface instead of the system default. [Commit](https://github.com/open-webui/open-webui/commit/58a3fadbf336029733d6e5fdde05fd5a24e656fa)
- 🎛️ **Valves dialog stays in bounds.** A valve whose selected options form a long line no longer stretches its input past the edge of the dialog and over the page behind it. [#29203](https://github.com/open-webui/open-webui/pull/29203), [#29202](https://github.com/open-webui/open-webui/issues/29202)
- 🧹 **Shared chats search resets.** Reopening the shared chats dialog now starts with an empty search box and the full list, rather than a leftover search term above unfiltered results. [#29082](https://github.com/open-webui/open-webui/pull/29082), [#29081](https://github.com/open-webui/open-webui/issues/29081)
- 🔎 **Case-insensitive search beyond English.** Searching and filtering by tag now ignore letter case for accented and non-Latin text on installations backed by SQLite, where only unaccented English letters were matched whatever their case. [Commit](https://github.com/open-webui/open-webui/commit/26f37426b7ef6c0c3b1363f413ac9a83ccaaeb71)
- 🏷️ **Labels with colons and example URLs translate correctly.** Interface text such as `Warning:`, `http://localhost:8000`, and `https://mineru.net/api/v4` now appears correctly instead of being misread by the translation system. [#29161](https://github.com/open-webui/open-webui/pull/29161), [#29154](https://github.com/open-webui/open-webui/issues/29154)
- 🔑 **Precise account matching on SQLite.** Signing in through an identity provider now matches your account on its exact identifier, so accounts whose identifier is an unusually long or zero-padded number are no longer at risk of being confused with another. [Commit](https://github.com/open-webui/open-webui/commit/17cc566707cd2ee78b460e39c7fad23b9c258c70)
- 🗃️ **Knowledge list loads reliably.** The Knowledge page in the workspace now fills in its list on opening instead of occasionally staying empty. [Commit](https://github.com/open-webui/open-webui/commit/e6031ea6daec897eddbcf0c8b03c2d6fad2b506e)
- 🎟️ **Tool servers without a key.** A tool server connection saved without a key is no longer sent an empty authorization header, which some servers refused outright. [Commit](https://github.com/open-webui/open-webui/commit/26f37426b7ef6c0c3b1363f413ac9a83ccaaeb71)
- 🐍 **Code interpreter starts on Windows.** The built-in code interpreter now serves its script and WebAssembly files with the browser-safe file types Windows hosts sometimes overwrite, so Pyodide can load instead of failing before code runs. [#29139](https://github.com/open-webui/open-webui/pull/29139), [#29133](https://github.com/open-webui/open-webui/issues/29133)
### Changed
- ⏱️ **Daily limit on repeats.** A calendar event can no longer be set to repeat more often than once a day, and saving one that repeats more often is now refused with a message explaining the limit. [Commit](https://github.com/open-webui/open-webui/commit/a93c5080380447e3ab9113cacb73b61c57cbc308)
- 🏷️ **High Contrast Mode is now Accessibility Mode.** The interface setting previously called 'High Contrast Mode' is now called 'Accessibility Mode', with the same switch in the same place. [Commit](https://github.com/open-webui/open-webui/commit/e6031ea6daec897eddbcf0c8b03c2d6fad2b506e)
## [0.11.1] - 2026-08-25
### Added
- 🚦 **Human in the loop tool approval.** Where an administrator has turned it on, you can switch a conversation from letting tools run freely to being asked first, so a model that wants to use a tool stops and waits for you to allow or deny it, one call at a time in a saved conversation, by button or by keyboard shortcut, with your choice remembered for this conversation and for future ones, switching back to running freely releasing anything already waiting, and automations, channel replies, and temporary chats unaffected. [Commit](https://github.com/open-webui/open-webui/commit/7d99b2716a0472b2100b3a71825d8eb3fcbbe877), [Commit](https://github.com/open-webui/open-webui/commit/ec36972c2b5a8d48713f1a240b0ed305e535b4cc), [Commit](https://github.com/open-webui/open-webui/commit/653562d660398c32a9a193450bbbee300d7195c3), [Commit](https://github.com/open-webui/open-webui/commit/fa94a5ab2431edba064150651a14dd4992ef29e8), [Commit](https://github.com/open-webui/open-webui/commit/55c202e841e76cd69679206cd0ecb4a65b039ea4), [Commit](https://github.com/open-webui/open-webui/commit/30f82788bc75e6965da3766e3bf43840dd971eed), [Commit](https://github.com/open-webui/open-webui/commit/bbfdbd59f29e246db1d8b5c2401bc7b11aa32c99), [Commit](https://github.com/open-webui/open-webui/commit/7fc5fa1ff3f6cd6f0efcbf847064ca8d21620ba6), [Commit](https://github.com/open-webui/open-webui/commit/3eb65f47151af4fb4ccfaf33d6e77154546e7259), [Commit](https://github.com/open-webui/open-webui/commit/62fc436999ad32e82d1405ac14d1f03e0f0341ec)
- 🙋‍♂️ **Models that can ask you a question.** A new built-in tool lets a model pause and put up to three multiple-choice questions to you before continuing, with room to type your own answer instead, and the question survives a reload in a saved conversation, so you can come back and answer it later rather than losing the conversation. [Commit](https://github.com/open-webui/open-webui/commit/4465f52a3eb521854cf190f91b0ea7cf3fa21830), [Commit](https://github.com/open-webui/open-webui/commit/133549a87ee371577453d897c2db2d8071223f27), [Commit](https://github.com/open-webui/open-webui/commit/b018feb7419e68314378b3cdd8b7b1b389400a98), [Commit](https://github.com/open-webui/open-webui/commit/083e35144152d6a301bed01aa6d898db3871ddcf), [Commit](https://github.com/open-webui/open-webui/commit/256cce505be8ddd4930e8cb2536faea718d3f789), [Commit](https://github.com/open-webui/open-webui/commit/14e4d72d9a21a10196e8e6efb04180cd9104b8e6), [Commit](https://github.com/open-webui/open-webui/commit/57bd08304e45f707768a898de9f50894929008dd), [Commit](https://github.com/open-webui/open-webui/commit/d9014b3483d4c6e8d99e50395afcdc538bcf05cd)
- 🖇 **Agents can now display terminal files directly.** A model can now show a file it made in a terminal directly in its reply, with a preview and a download button, instead of describing a path that led nowhere when clicked, and a new interface setting chooses whether these open in the reply or in the files pane. [Commit](https://github.com/open-webui/open-webui/commit/78f48a21eef330c0b78c33f2b5fc2169b084997c), [Commit](https://github.com/open-webui/open-webui/commit/e623c02acc70c4ee5d7f2eb2f32d9b7f39287663), [Commit](https://github.com/open-webui/open-webui/commit/f64c0c87e8d1bfdbe060ea5e5a3dee24c0323657), [#27650](https://github.com/open-webui/open-webui/issues/27650)
- 📶 **Streaming rebuilt from the ground up.** A reply now streams as small pieces of new text instead of resending the whole message so far with every update, so the data sent over a reply grows with its length rather than with its length squared, which on a server with many people chatting at once means far less processor time spent encoding, passing, and decoding those updates, far less load and memory on the shared cache that carries them between instances, and far less work in your browser, which no longer takes in the whole reply again and redraws the parts of it that have not changed on every update, cutting the data sent and the server work spent appending to a message by up to 1000x on a very long reply, and a reply still in progress is now kept aside on the server, so reopening the conversation after a refresh picks it up where it is rather than showing a blank message, on deployments backed by Redis. [Commit](https://github.com/open-webui/open-webui/commit/a1579a01ff43cacb357269707d36267ad35e01d6), [Commit](https://github.com/open-webui/open-webui/commit/c755ef60c6bd47ea25306bd898d9a6d1bd8e871d), [Commit](https://github.com/open-webui/open-webui/commit/d02b6a21fc02fb073782e25356968cfee3c45c36), [Commit](https://github.com/open-webui/open-webui/commit/3e186abdd91edee9e97e43c9b345714643a84cf5)
- 🪵 **Much faster throughout.** Hundreds of places across the application no longer assemble detailed log text that is switched off and thrown away unread, so sending messages, uploading and indexing files, running searches, signing in, and loading admin pages all get that time back, with the largest gains on busy servers, in long conversations, and on chats that draw from a large knowledge base. [#27834](https://github.com/open-webui/open-webui/pull/27834), [#27837](https://github.com/open-webui/open-webui/pull/27837)
- 🚀 **Faster model list lookups.** Title generation, tag suggestions, autocomplete, and other background steps of a chat turn now fetch the model list in one go, which keeps other people's responses flowing on busy Redis-backed instances with many models. [#27821](https://github.com/open-webui/open-webui/pull/27821)
- 🛰️ **Cheaper log export.** Deployments that export their logs to a telemetry collector, with "ENABLE_OTEL" and "ENABLE_OTEL_LOGS" both set, now prepare each exported line once instead of twice, which matters more than it used to now that log text is only assembled when something will actually read it. [#27840](https://github.com/open-webui/open-webui/pull/27840)
- 📇 **Faster permission checks on large instances.** Working out which groups you belong to is now a direct lookup rather than a scan of every membership on the server, so chats and the admin user list stay quick as an organization grows. [#27822](https://github.com/open-webui/open-webui/pull/27822)
- ⚙️ **Much faster JSON handling.** Saving and opening chats, reading settings, returning results from built-in tools, streaming replies, signing in and signing up, working out your permissions, and reading stored chunk details during knowledge base searches on Valkey and Oracle vector storage are all handled much faster across the application when the "ENABLE_ORJSON" option is turned on. [Commit](https://github.com/open-webui/open-webui/commit/bb0f898b431d5aa45efa7805956657ed9c3dd78d), [#28396](https://github.com/open-webui/open-webui/pull/28396), [#27841](https://github.com/open-webui/open-webui/pull/27841), [#27807](https://github.com/open-webui/open-webui/pull/27807), [#27805](https://github.com/open-webui/open-webui/pull/27805), [#27813](https://github.com/open-webui/open-webui/pull/27813)
- 📤 **Much faster outbound requests.** Conversations and embedding batches sent to Ollama and Anthropic models are packaged for delivery much faster, which is most noticeable in long chats when the "ENABLE_ORJSON" option is turned on. [#27811](https://github.com/open-webui/open-webui/pull/27811), [#27810](https://github.com/open-webui/open-webui/pull/27810)
- 🐍 **Much faster code interpreter output.** Printed output and generated images from code run in chat appear much faster when the "ENABLE_ORJSON" option is turned on. [#27812](https://github.com/open-webui/open-webui/pull/27812)
- 🪶 **Lighter page loads.** Several small requests the interface makes on every page load, along with a few administrative ones, no longer set up database access they never used, which took several times longer than the rest of the request put together. [#28178](https://github.com/open-webui/open-webui/pull/28178)
- ♻️ **One less read per message.** Sending a message no longer loads the whole conversation from the database twice over, which mattered most in long chats where that record is largest. [#28809](https://github.com/open-webui/open-webui/pull/28809)
- 🏁 **Faster skills on large instances.** Opening the skills list, or sending a message that uses one, no longer checks every skill on the instance one at a time, so both are far quicker where many skills exist and most of them are not yours. [#28798](https://github.com/open-webui/open-webui/pull/28798)
- 🩻 **Faster tools on large instances.** Listing or exporting tools no longer checks every tool on the instance one at a time, so the integrations menu and the tools workspace open faster where many exist. [Commit](https://github.com/open-webui/open-webui/commit/4807866a1cf47340f1b5ea76fded95f8114305f9)
- 🧊 **Faster file access checks.** Checking whether you may reach a file no longer walks every workspace model you can see looking for it, so opening a folder of files, downloading one, or retrieving from one is much quicker on instances with many models. [#28802](https://github.com/open-webui/open-webui/pull/28802)
- 🧱 **Faster folder listings.** Listing your folders now works out your group memberships once for the whole listing rather than again for every item in every folder. [#28810](https://github.com/open-webui/open-webui/pull/28810)
- 🧼 **Less work per update in a long chat.** Each update saved while a reply streams no longer re-examines the entire conversation, only the part being added, so the cost of an update stops growing with the length of the chat. [#28820](https://github.com/open-webui/open-webui/pull/28820)
- 🗃 **Cheaper attaching of sources and files to a reply.** Adding a source, file, or embedded item to a reply now reads just that one field rather than rebuilding the whole conversation to find it, which on a two hundred message chat is around 3.1 ms per item down to 0.65 ms, and no longer grows with the length of the conversation. [Commit](https://github.com/open-webui/open-webui/commit/536b9edec00547d5b84ef2e6ea0f929c054c1333), [Commit](https://github.com/open-webui/open-webui/commit/9dff5e93277aa1a2236d702e7e2f55dbb35ea7fa)
- 🚏 **Faster workspace model lookups.** Working out which workspace models you may edit no longer loads every model on the instance and discards most of them, which also speeds up exporting models and the file access checks that relied on it. [#28795](https://github.com/open-webui/open-webui/pull/28795)
- 📮 **Faster handing off a streaming reply.** Passing a reply in progress between instances now writes it once rather than converting it back and forth and scanning it for characters that only matter elsewhere, which on a large non-English conversation took most of the time spent on each write. [#28833](https://github.com/open-webui/open-webui/pull/28833)
- 🥵 **Constant load on an idle instance.** An instance sitting idle no longer works through every chat you have once a second looking for timers that are due, which on a large history kept about a quarter of a processor core busy doing nothing and could exhaust memory until the application was killed. [#27663](https://github.com/open-webui/open-webui/pull/27663), [#27622](https://github.com/open-webui/open-webui/issues/27622), [#27745](https://github.com/open-webui/open-webui/issues/27745)
- 📍 **Sidebar folders fetched once.** Refreshing the sidebar now asks for your folders once rather than three times, on page load and on every action that refreshes it. [#28662](https://github.com/open-webui/open-webui/pull/28662), [#28661](https://github.com/open-webui/open-webui/issues/28661)
- 💤 **Far fewer writes just from being signed in.** Recording that someone is online now writes at most once a minute for each person rather than on every single request, where an open tab alone caused two write transactions a minute before anyone touched anything. [#28177](https://github.com/open-webui/open-webui/pull/28177), [#28165](https://github.com/open-webui/open-webui/issues/28165)
- 🛰 **Less overhead on every request.** The layers each request passes through before it is handled are now one instead of five, which also removes a quarter of that cost from every piece of a streamed reply on instances that set security headers. [Commit](https://github.com/open-webui/open-webui/commit/b96d2b12dae5e953e520bec03f74e9b85b955dc7), [#28171](https://github.com/open-webui/open-webui/issues/28171)
- 🎏 **Turning off compression of live updates.** A new "UVICORN_WS_PER_MESSAGE_DEFLATE" setting stops the server compressing every live update it sends, which costs processor time on each one for almost no saving now that a reply streams as small pieces; compression stays on unless it is turned off. [#28613](https://github.com/open-webui/open-webui/pull/28613)
- 🌡 **Faster chat list and unread counts.** Opening the sidebar, and the unread markers on folders, no longer read through your whole chat history to produce a short list, which on an instance with 15000 chats took 2 to 4 seconds. [#27663](https://github.com/open-webui/open-webui/pull/27663), [#27622](https://github.com/open-webui/open-webui/issues/27622), [#27745](https://github.com/open-webui/open-webui/issues/27745)
- 🥁 **Long replies no longer slow as they grow.** A long reply is no longer re-examined from the beginning for reasoning and code blocks on every piece that arrives, so the work stops growing with the length of the reply, which on a long reply is around 190x less time spent on it. [#28861](https://github.com/open-webui/open-webui/pull/28861)
- 💽 **Faster saving of long chats.** A chat is now written to the database in one go rather than one message at a time, so saving a long conversation is much quicker and puts far less strain on the database, and saving one where nothing has changed writes nothing at all. [#28806](https://github.com/open-webui/open-webui/pull/28806)
- 📦 **Faster loading of shared folders.** Folders shared with you now load in a couple of queries rather than one for each folder and each owner, so the list appears sooner for anyone with many of them. [#28804](https://github.com/open-webui/open-webui/pull/28804)
- ⚡ **Uninterrupted chat during knowledge search.** Responses now keep streaming for everyone on the server while knowledge base searches run, instead of pausing until each search finishes. [#27824](https://github.com/open-webui/open-webui/pull/27824)
- 🔍 **Smarter chat search.** Searching your chats now finds conversations containing all of your words in any order rather than only the exact phrase you typed, with exact matches still listed first, and the preview snippet points at whichever word it found. [Commit](https://github.com/open-webui/open-webui/commit/0800c21c64c64810f24c2ec88cca5e36daebb10e)
- ⌨️ **Model switching from the message box.** Typing "/model" now tells you which model you are on, switching to another by name with "/model" followed by its id, or opening the model picker straight from the slash menu without reaching for the mouse. [Commit](https://github.com/open-webui/open-webui/commit/29eeda9f9abeaedf77176908cb036ace7f175e76), [Commit](https://github.com/open-webui/open-webui/commit/9c7ce154e79a3baea3a8223ab1e5b1371ce80dbb)
- 📎 **Sending while attachments upload.** Sending a message before its files have finished uploading now queues it and sends it automatically once they are ready, instead of refusing with an error, and each queued message shows the progress of its attachments. [Commit](https://github.com/open-webui/open-webui/commit/6c4d0ace163a89aba6e8cd2aa9fef185167ebe0f), [Commit](https://github.com/open-webui/open-webui/commit/c1c07cbe0f847e4dd04e21f2a6c5f12cdc0fc7ee), [#28381](https://github.com/open-webui/open-webui/pull/28381), [#28380](https://github.com/open-webui/open-webui/issues/28380)
- 📖 **Opening a document at the right page.** A model showing you a PDF, Word document, or slide deck from a terminal can now open it at a particular page or slide, so a reply that cites something on page 76 can put that page in front of you. [Commit](https://github.com/open-webui/open-webui/commit/fd8cc2ba4a226ccbf51e114c71f156676eabb2d9), [Commit](https://github.com/open-webui/open-webui/commit/cf4ac9c8db67031627f196fae894764367fba5b4), [Commit](https://github.com/open-webui/open-webui/commit/6cb2449ab70ed6aa5124fcd2b1291b48527af4cd)
- 💼 **Attachments that go straight to a terminal.** A terminal connection can now be set to receive files attached in chat into its own working directory rather than into the conversation, which also means files can be attached while using a model that cannot read them itself. [Commit](https://github.com/open-webui/open-webui/commit/8a42aa53e826d5f25d6443ca95e0173e3cf034cf), [Commit](https://github.com/open-webui/open-webui/commit/d7d935275a77fb185586a5130cd37acf1175ca0a), [Commit](https://github.com/open-webui/open-webui/commit/1b3b9375bb4de0f360782dd2f9d8c7c2903b8baa), [Commit](https://github.com/open-webui/open-webui/commit/d17f06a23501e5c39cf98d9c9642178da25343fb)
- 🔦 **Searching files in the terminal browser.** The file browser now has a search box that finds files by name and by what is inside them, and opening a result takes you to the matching line. [Commit](https://github.com/open-webui/open-webui/commit/7abe11346a1cc309da2c3fcec2f61faaeb47b96a)
- 🌲 **Browsing files as a tree.** The terminal file browser now expands folders in place rather than only navigating into them, remembers what you had open, offers a right-click menu, can show hidden files, sorts by size, expands a folder you hover over while dragging something onto it, and moves a whole selection in one go when you drop it. [Commit](https://github.com/open-webui/open-webui/commit/7dfbdd221ac0f7a839dd049e3e82b6a84c316860), [Commit](https://github.com/open-webui/open-webui/commit/516cf1a9a6154336182475b9df2d333b1579516d)
- 🧰 **Managing models on more servers.** Administrators can now download, load, and unload models on llama.cpp and LM Studio connections from the manage models dialog, remove them on llama.cpp, and start a download straight from the model picker's search box, alongside the Ollama support that was already there. [Commit](https://github.com/open-webui/open-webui/commit/85c3d0ae2fad58ffc2b92b1733a6c8bc7ab471db), [Commit](https://github.com/open-webui/open-webui/commit/8260d527ee97372a207ce9bd9c6dab4909061281), [Commit](https://github.com/open-webui/open-webui/commit/25802c048e6123fa602182949ad2a9f349e1e863), [Commit](https://github.com/open-webui/open-webui/commit/31c1ffd55a018a74c7e53a35fb9f9bbe33774b6f), [#28766](https://github.com/open-webui/open-webui/pull/28766)
- 📢 **Automations that post to a channel.** An automation can now be pointed at a channel instead of a chat, so its scheduled run appears as a message there for everyone to see, chosen from a new destination picker that also covers folders. [Commit](https://github.com/open-webui/open-webui/commit/2649e3305c49cb37101c112ae10ffd8beacb5885)
- 🙋 **Mentioning people in a channel.** Typing an at sign in a channel now lists that channel's own members first, before everyone else on the server, so the people you are likely to mean are at the top. [Commit](https://github.com/open-webui/open-webui/commit/ba885d0026ad3cce70c7cc36d1d6f59775264a12), [Commit](https://github.com/open-webui/open-webui/commit/dcff244f9e87e615ee0a3f345f7f801c5ebe370d), [#28883](https://github.com/open-webui/open-webui/issues/28883)
- 🔗 **Attaching any link.** Pasting a link into a chat or a knowledge base now works out what is behind it, downloading a document or image as a real attachment rather than treating everything as a web page to be read as text. [Commit](https://github.com/open-webui/open-webui/commit/8fbfd14a8b3fe7236883594fd18760076e4de0d5)
- 🔎 **Searching tools and skills in chat.** The integrations menu now has a search box for tools and for skills, so a long list can be narrowed by name instead of scrolled through. [#26709](https://github.com/open-webui/open-webui/issues/26709), [Commit](https://github.com/open-webui/open-webui/commit/954613944b317a78d5929456ae1388543aa052c8), [#28807](https://github.com/open-webui/open-webui/pull/28807), [#28812](https://github.com/open-webui/open-webui/pull/28812)
- 🧭 **More from the message box.** The slash menu now offers settings and, in a new chat, a toggle for temporary chat, alongside the commands that were already there. [Commit](https://github.com/open-webui/open-webui/commit/a5ea732c1e3d0d1f70a7e4b6a20c30c336114784)
- 🚨 **Being told when a file fails to process.** A file that cannot be processed for a knowledge base now raises a notification naming the file and what went wrong, and keeps that reason on the file, instead of quietly being marked as failed. [#27666](https://github.com/open-webui/open-webui/pull/27666), [#6311](https://github.com/open-webui/open-webui/issues/6311), [Commit](https://github.com/open-webui/open-webui/commit/1a376ac17fa0c3f957656a997b6f8ffacb1f6f30)
- 🔬 **Zoom controls on previewed images.** An image opened in the file browser now has zoom in, zoom out, and a reset button showing the current zoom, and pinching, scrolling, and holding a modifier key while scrolling now zoom and pan as they do elsewhere. [Commit](https://github.com/open-webui/open-webui/commit/2befa8f796266e92fa55861bb8eaa81639ff4053), [Commit](https://github.com/open-webui/open-webui/commit/bfb68feea766e8d5408fb6e278be56cbca0c4afe), [Commit](https://github.com/open-webui/open-webui/commit/ec9bf5a64f9f718e472123350ec83bce1b064884), [Commit](https://github.com/open-webui/open-webui/commit/467be93e6d7e4a31358f9c75ee67bfac1200c0a1)
- 🗂️ **Recognisable file icons.** The terminal file browser now marks each file with an icon for its type, so code, images, archives, documents, and configuration files can be told apart at a glance instead of sharing one generic page icon. [Commit](https://github.com/open-webui/open-webui/commit/60feca71a6a77b7d1eef1c172a61624d5fe48150), [Commit](https://github.com/open-webui/open-webui/commit/8d25ad00e2fc5f0a328a2098c2e73fb00f2ca934), [Commit](https://github.com/open-webui/open-webui/commit/c1f914a6268580f7722825390ac8a204743a4520)
- 📽️ **Truer PowerPoint previews.** Slide previews now render tables, charts, connectors, gradients, theme colours, bullets, fonts, and text alignment far closer to the original, and the viewer lets you move between slides with the arrow keys or the scroll wheel while the thumbnail strip follows along. [Commit](https://github.com/open-webui/open-webui/commit/8dd23f74c9a59fdfbdaefb4a403ef79bdd3dfed9), [Commit](https://github.com/open-webui/open-webui/commit/048c06399363baac4046717129f072172ee90805), [Commit](https://github.com/open-webui/open-webui/commit/c93c6d6fc48ee8870eec349cb4c5af88c9435cfb), [Commit](https://github.com/open-webui/open-webui/commit/76583749edb966c0239ce203a84c1ee8d8c7faee), [Commit](https://github.com/open-webui/open-webui/commit/b3a5fd3875dc3c29a11eb5435438e55685cad639), [Commit](https://github.com/open-webui/open-webui/commit/794671a9883d5e067c407067f324ffa25ed38e3f), [Commit](https://github.com/open-webui/open-webui/commit/f96b717566b2084bd5fe70fb622ab66b5a643c0d), [Commit](https://github.com/open-webui/open-webui/commit/b7292890ccea794750651abbcde01cda62551e1b), [Commit](https://github.com/open-webui/open-webui/commit/31d08d592c4ee46a7ce827d6ffdf8e3cf16bf2ac)
- 📄 **Faithful Word document previews.** Word documents now open as proper pages with headers, footers, footnotes, and embedded images intact, and can be zoomed, rather than being flattened into plain formatted text. [Commit](https://github.com/open-webui/open-webui/commit/ff7467b4c593f1e775311c088c63340d0cc1e4a2), [Commit](https://github.com/open-webui/open-webui/commit/060648f939447e2a9de390d05221ca896c50eeed)
- 🗝 **Deleting your API key.** An API key can now be revoked outright from your account settings, where the only way to retire one was to replace it with a new one. [#28874](https://github.com/open-webui/open-webui/issues/28874), [Commit](https://github.com/open-webui/open-webui/commit/b30b11d4c975b8ff9a6c6eb9e39fbb310aed1072)
- 🎛️ **Settings for the task model.** Administrators can now set the generation parameters used for background work such as titles, tags, follow-ups, search queries, and conversation summaries, either from the admin panel or through "TASK_MODEL_PARAMS", instead of those requests always using a fixed token limit that could cut a summary short. [#27604](https://github.com/open-webui/open-webui/issues/27604), [Commit](https://github.com/open-webui/open-webui/commit/f0bfcd40976dfb1e1876f86b39d6a659d894d15e), [Commit](https://github.com/open-webui/open-webui/commit/865c80c1600ecc2e606bef4e90540ce305b4b9de)
- 🎚️ **Default interface settings for everyone.** Administrators can now set system-wide defaults for the interface options in Settings, either from the admin panel or through "DEFAULT_INTERFACE_SETTINGS", with each person's own choices still taking precedence and anything left untouched shown as inherited and kept in step with later changes to the defaults. [Commit](https://github.com/open-webui/open-webui/commit/37f2548155efdd5cd8114ceb9b57c7864de4bef2), [Commit](https://github.com/open-webui/open-webui/commit/13346c5f1621b014e2fcaabd41375d6f7035dd0a), [Commit](https://github.com/open-webui/open-webui/commit/90a0e61cef119154e89d900d6921123b599a0389), [Commit](https://github.com/open-webui/open-webui/commit/eeaf1a1df01525e4b6d4eb73060b32a9a75ebfc6), [Commit](https://github.com/open-webui/open-webui/commit/d4461bd6f39936460a989e8bd50097de604b4b00), [Commit](https://github.com/open-webui/open-webui/commit/b4d3b27caf1587b783a46268c769132eb9211c80), [Commit](https://github.com/open-webui/open-webui/commit/407c40f72cdadcd9f57538d9ce38070662ce42c7), [Commit](https://github.com/open-webui/open-webui/commit/724d2ebbf1bd2fc90f27a2e5bc571d40afc3db4c), [Commit](https://github.com/open-webui/open-webui/commit/1674e5a9ef835696f0963cdbcf477c8b92ce0b8a)
- 🔠 **Interface scaling throughout.** The UI Scale setting and your browser's own text size now resize the whole interface consistently, including the sidebar, menus, dialogs, and file browser, rather than leaving parts of it fixed. [Commit](https://github.com/open-webui/open-webui/commit/30d08a42f8b0c32cc64dccd81203cc760651584a), [Commit](https://github.com/open-webui/open-webui/commit/b5f86e6a433b519e30a5a6f1b0a5bf985939c9a8), [Commit](https://github.com/open-webui/open-webui/commit/ec03e8814403422a6ab3dc62f4239fd7af439361), [Commit](https://github.com/open-webui/open-webui/commit/72a909fd2f3ac8d67554f6c85bc1723934454e2d)
- 🏷️ **Named writing blocks.** When a model wraps a draft such as an email in a writing block, the block is now titled with its subject and shows the recipient beside it, rather than every block reading simply as Writing. [#28280](https://github.com/open-webui/open-webui/pull/28280), [#28198](https://github.com/open-webui/open-webui/issues/28198)
- 🤝 **Files for delegated tasks.** A task handed to a sub-agent can now carry the attachments it needs, so an image or document from your conversation reaches the sub-agent instead of arriving as a file reference it cannot open and may answer about anyway. [Commit](https://github.com/open-webui/open-webui/commit/5ec16e76e6402980b39c26923b8ee26278f0b243), [#28213](https://github.com/open-webui/open-webui/issues/28213)
- 📟 **Terminal availability and scope.** Administrators can now decide for each managed terminal whether it appears in chats and in automations at all, and whether everyone shares a single workspace or each chat or automation gets its own, with per chat terminals waiting until the conversation has been saved. [Commit](https://github.com/open-webui/open-webui/commit/009999f3636b1a3451f8fe55232e6ab132a64e66)
- 🔒 **Read-only files in the terminal browser.** Files and folders you are not allowed to change are now labelled read-only, with uploading, editing, renaming, moving, and deleting turned off for them rather than failing at the moment you try. [Commit](https://github.com/open-webui/open-webui/commit/2dadc5435af77b0638af69200e3af1b9654417b6)
- 🔐 **Terminals that use your own login.** Managed terminals configured for session authentication now authenticate the terminal connection with your own token, where it previously sent no credentials at all. [Commit](https://github.com/open-webui/open-webui/commit/2dadc5435af77b0638af69200e3af1b9654417b6)
- 🎟 **Setting up a tool server that uses OAuth.** Adding one is now easier to get right: the connection dialog can authorize the account from the dialog itself, the check button tests the sign-in details rather than reporting a connection failure that was never going to succeed without them, and it is now labelled for what it does rather than suggesting it verifies the whole connection. [Commit](https://github.com/open-webui/open-webui/commit/f822605b3563c57030aa200492f68572cadcc4da), [#28552](https://github.com/open-webui/open-webui/issues/28552)
- 🪤 **Control over what embedded pages may do.** Two new interface settings decide whether pages shown inside a chat, such as an artifact or an HTML preview, may run scripts and start downloads. [Commit](https://github.com/open-webui/open-webui/commit/3c66d639e31ba8a7477d337b37d8d671ea2430cf), [#28924](https://github.com/open-webui/open-webui/issues/28924), [Commit](https://github.com/open-webui/open-webui/commit/842c1f9d677c0d9940cccdd18e34add13f746a68)
- 🗄 **Keeping files removed from a knowledge base.** A new "ENABLE_KNOWLEDGE_FILE_RETENTION" setting keeps the stored file and its search data when a file is taken out of a knowledge base, rather than deleting them. [Commit](https://github.com/open-webui/open-webui/commit/363ad352fec9553469852d111bc0506b896504a6)
- 🧾 **CSV shape in retrieval.** Turning on "ENABLE_RAG_CSV_SUMMARY" adds a short line naming the row count, data row count, column count, and column names of a CSV file to what the model sees, giving it the shape of the table alongside its contents. [Commit](https://github.com/open-webui/open-webui/commit/1b72899f246ba46ab8f7cd6aee2d56ef69217d82)
- 🔭 **OpenSERP in the search settings.** OpenSERP can now be picked as the web search engine in the admin panel, with a field for its address, rather than only being configurable through the environment. [#27594](https://github.com/open-webui/open-webui/pull/27594), [#27592](https://github.com/open-webui/open-webui/issues/27592)
- 🪧 **Profile changes from single sign-on.** A name, email address, or picture updated from an identity provider at sign-in now raises an event naming what changed, and the rest of the session uses the updated record rather than a stale copy. [Commit](https://github.com/open-webui/open-webui/commit/927ce0eae67af1b4a856d3e82fba0601c3b40d02)
- 📯 **Group changes from single sign-on.** Group memberships added or removed when someone signs in through an identity provider, and groups created automatically along the way, now raise the same events as the equivalent change made by an administrator or over directory sync. [#27657](https://github.com/open-webui/open-webui/pull/27657)
- 🔔 **Sign-in and sign-out events for single sign-on.** Signing in through an identity provider now raises the same login event that signing in with a password does, and signing out says which provider the session came from, so a function can set up or tidy up an account in another system when someone arrives or leaves. [#27619](https://github.com/open-webui/open-webui/pull/27619), [#27613](https://github.com/open-webui/open-webui/issues/27613)
- 🪛 **Naming background worker threads.** A new "THREAD_POOL_THREAD_NAME_PREFIX" setting labels the threads that background work runs in, so they can be told apart when reading a profile or a thread dump. [Commit](https://github.com/open-webui/open-webui/commit/4ec6ee14418edd04eaba9e34bd5868453f61df40)
- 📙 **OpenDocument files in a temporary chat.** A text document, spreadsheet, or presentation from an office suite that uses the OpenDocument format now has its text read out in the browser when attached to a temporary chat, where the model was handed the raw archive and answered that it could not read the file. [Commit](https://github.com/open-webui/open-webui/commit/9e7c9360b744c878ae0c38aa1caf3886f27b07ab), [#28906](https://github.com/open-webui/open-webui/discussions/28906)
- 🌍 **Pointing Tavily somewhere else.** A new "TAVILY_API_BASE_URL" setting sends Tavily searches and page fetches to a different address, for instances that reach the internet only through a gateway of their own or that use a compatible service. [Commit](https://github.com/open-webui/open-webui/commit/98ee2bdfd3e90faec6bfe8e7ebf5803159f379e3), [#28701](https://github.com/open-webui/open-webui/issues/28701)
- 🪟 **Honest OAuth settings.** When single sign-on settings come from the environment rather than being saved in the application, the admin panel now shows them as read-only with a note naming the setting that controls this, instead of accepting edits that were silently discarded on the next restart. [#28276](https://github.com/open-webui/open-webui/pull/28276)
- 📏 **Widening the chat controls pane.** The controls pane can now be dragged as wide as you like, where it stopped at a fixed limit regardless of screen size. [Commit](https://github.com/open-webui/open-webui/commit/0fb542b3764cefecf2366607a4685288051d1b46)
- 📱 **Smoother sidebar on mobile.** The sidebar now follows your finger as you swipe it open or closed, responds to a quick flick, dims the page behind it as it moves, and gives every chat row a menu button you can reach without a hover you cannot perform on a touchscreen. [Commit](https://github.com/open-webui/open-webui/commit/b20bcdbba72707e3b0cf2b9a6a5f3168b464ba3c), [Commit](https://github.com/open-webui/open-webui/commit/178ccb30e1253ea727fe5ccbb5cc3ca3832dc70b), [Commit](https://github.com/open-webui/open-webui/commit/943294df9a03d45e2708b330e876967af3463282), [Commit](https://github.com/open-webui/open-webui/commit/d6679082e5c6b0b54ca00a99a88e66a53a30c7d9), [Commit](https://github.com/open-webui/open-webui/commit/d8ae7ed40551362922925d9d6e47ba65d3658cfd), [Commit](https://github.com/open-webui/open-webui/commit/be4afd75452361ead376d9977cf2ca8cb93e55eb), [Commit](https://github.com/open-webui/open-webui/commit/ab41dcc487d1517f7c8c5d0b02a02cdaadedb03d)
- 🚪 **Sidebar that stays put.** Opening and closing the sidebar is now a smooth transition that keeps your chat list loaded, instead of rebuilding the list each time. [Commit](https://github.com/open-webui/open-webui/commit/3c010951db2cc349466658d3b52234cc22dad327), [Commit](https://github.com/open-webui/open-webui/commit/8edab5020eaa5d47b12a572eb632a891aa29440a), [Commit](https://github.com/open-webui/open-webui/commit/3e9b075954b7fc9d492a7ec832550b10b0bf80ed), [Commit](https://github.com/open-webui/open-webui/commit/3793b0c886f57630dc31320d3c0257c933c6eca1), [Commit](https://github.com/open-webui/open-webui/commit/0b4b7ae5ff3c0a17b58e8e85a5fddf190e3bda14)
- 👁️ **Turning off chat previews.** A new setting under Settings and Interface lets you switch off the preview card that appears when you hover a chat in the sidebar, useful for a quieter sidebar, for sharing your screen, or on a slow connection. [#27632](https://github.com/open-webui/open-webui/pull/27632), [#27639](https://github.com/open-webui/open-webui/issues/27639)
- ☑️ **Checkboxes beside their labels.** In the model editor and the admin model defaults, each capability, feature, and tool checkbox now sits directly in front of its own label instead of at the far edge of its column, where it could look like it belonged to the next one, and the label itself can be clicked to toggle it. [#27788](https://github.com/open-webui/open-webui/pull/27788), [#27771](https://github.com/open-webui/open-webui/issues/27771), [Commit](https://github.com/open-webui/open-webui/commit/1b39ff352a2fa3b57bd7815f44c50daf96937a14), [Commit](https://github.com/open-webui/open-webui/commit/0f821398ca9c9ddd8da521b5f9dcf103030202e8)
- ✍️ **Typing cursor while responding.** A blinking cursor now marks where the reply is being written, from the moment you send your message until generation finishes, in place of the previous loading placeholder. [Commit](https://github.com/open-webui/open-webui/commit/cbb3aade2b4e901c22aa9a30530634721c80b078)
- ✒️ **Underlined text.** Underlined text now appears underlined in a reply instead of showing the markup around it, and underlining is kept when you edit in a rich text box rather than being dropped. [Commit](https://github.com/open-webui/open-webui/commit/11db926a7b471e9477595786ef451a916174cc18), [#26904](https://github.com/open-webui/open-webui/issues/26904)
- 📥 **Adding group members from a file.** Administrators can now add many people to a group at once by uploading a CSV of names and email addresses, with a template to download and a message naming any row whose address does not match an account. [Commit](https://github.com/open-webui/open-webui/commit/f3f7659da754b51c216b17b2d6def4d59071d78f)
- 📑 **Apache Tika 4 support.** Administrators extracting document text with Tika can now choose which server version they run, from Admin Settings under Documents or through "TIKA_SERVER_VERSION", where only Tika 3 was understood before. [Commit](https://github.com/open-webui/open-webui/commit/170ad0595d9440113721eb06375a2fc99aefaad1), [#28939](https://github.com/open-webui/open-webui/issues/28939)
- 💓 **Tunable heartbeat for live updates.** A new "WEBSOCKET_HEARTBEAT_INTERVAL" setting controls how often each open tab checks in with the server, where it was fixed at 30 seconds, so a large deployment can cut background traffic that no one asked for. [Commit](https://github.com/open-webui/open-webui/commit/3c1017f6c3ffc7194f074f2cfd2ea5f49943c575), [#28166](https://github.com/open-webui/open-webui/issues/28166)
- ⌛ **Expiring abandoned reply state.** A new "REDIS_RESPONSE_STREAM_TTL" setting expires the saved state of a reply that never finished, so a server killed mid-answer no longer leaves that data behind for good. [Commit](https://github.com/open-webui/open-webui/commit/176fa462128d5298492db29c67080a4e1afc2642)
- 🖨️ **File and image detail parts on API requests.** A request sent to the OpenAI-compatible endpoint carrying an image detail level or a file part in its message content now forwards both to providers that use the Responses API, where they were dropped, while documents attached inside Open WebUI are unaffected because those still go through knowledge retrieval. [Commit](https://github.com/open-webui/open-webui/commit/ca4e07a40b4ff11989a25d84b517d3d0049f93ae)
- 🧺 **Leaner stored document metadata.** Bulky extraction details such as page layouts, tables and detected languages are no longer kept alongside a document in the vector store, and a new "RAG_METADATA_MAX_VALUE_CHARS" setting drops any remaining oversized value, falling back to the configured upload size limit so a document that expands enormously while being read cannot exhaust a server's memory. [Commit](https://github.com/open-webui/open-webui/commit/278e97589e71d119b887d5bca9d6ae32912d1dff), [Commit](https://github.com/open-webui/open-webui/commit/e3a7a64d82ab2dd06c681ce85027367ccc8234d4), [#29025](https://github.com/open-webui/open-webui/pull/29025)
- 💨 **No filter work on installs without filters.** A completed message on an install with no filter functions and no pipeline filters, which is the default, no longer rebuilds the whole conversation and ships it to the browser as an event nothing acts on. [Commit](https://github.com/open-webui/open-webui/commit/28f2965934f5b6fba0e0c38a15b2af6ee790819f)
- 📂 **Opening a file in a knowledge base.** A file listed in a knowledge base can now be opened and read straight from that list, where the name was shown but nothing happened when it was clicked. [Commit](https://github.com/open-webui/open-webui/commit/20f35d157bc7d535091ce6a90b517c23d26486c0), [#28086](https://github.com/open-webui/open-webui/issues/28086)
- 🧵 **Cheaper saving of a reply as it streams.** The resume snapshot taken on every piece of a streamed reply no longer rebuilds the whole answer each time, so the cost of a save stops growing with the length of the reply. [#28821](https://github.com/open-webui/open-webui/pull/28821)
- 📀 **Less repeated work setting up built-in tools.** Every chat request no longer rebuilds a fresh copy of each built-in tool's definition from scratch, which was paid once per tool on every message. [#28860](https://github.com/open-webui/open-webui/pull/28860)
- 🕹️ **Control over what a terminal port preview may reach.** A new interface setting decides whether a previewed port runs with access to same-origin browser APIs, so you can lock a preview out of them on installs where previews serve content you do not fully trust. [Commit](https://github.com/open-webui/open-webui/commit/54d7a223707f03172efbb9e754db6e69709956d0)
- ♿ **Improved UI accessibility.** A closed sidebar is no longer reachable by keyboard or announced by screen readers, the tool call blocks in a response can now be expanded with the keyboard, the buttons that normally appear on hover, such as message actions, file removal, and chat menus, now appear when you reach them with the keyboard as well, whatever you have tabbed to is marked with a clear outline throughout the application, and the rows in the integrations menu now tell a screen reader whether each tool or feature is switched on. [Commit](https://github.com/open-webui/open-webui/commit/48a5696042b414f6511911afa72b3289ad2797b9), [Commit](https://github.com/open-webui/open-webui/commit/bd250a0e2431f8c2c7e2f4a34c5327f21fc77d4b), [Commit](https://github.com/open-webui/open-webui/commit/c8f8fa451a60974dfa4ebaf7cd6163ef30c45fd6), [Commit](https://github.com/open-webui/open-webui/commit/29541cbb52659a8a6ee22d255f14e6d2f168a09b), [Commit](https://github.com/open-webui/open-webui/commit/ac0368b4abe4c880f715b236ee0024826b2e3e7a), [Commit](https://github.com/open-webui/open-webui/commit/c086b80313fe0f5ce60831a0d150f695646c66e4), [#27667](https://github.com/open-webui/open-webui/pull/27667), [#17150](https://github.com/open-webui/open-webui/issues/17150)
- 🔄 **General improvements.** Various improvements were implemented across the application to enhance performance, stability, and security.
- 🌐 **Translation updates.** Faroese was added, and translations for Slovenian, Hungarian, Finnish, Korean, Portuguese (Brazil), Catalan, and French were enhanced and expanded.
### Fixed
- 🛡️ **Security Advisory**: This release includes security and access-control fixes. We recommend updating production deployments at your earliest convenience. Not all security fixes in this version may be enumerated in the fixed section. Some may be withheld for a short time to give administrators time to upgrade. [Advisories](https://github.com/open-webui/open-webui/security)
- 🛂 **Knowledge search reaching past what you may read.** Searching knowledge bases now applies the list of collections you are allowed to open, where that restriction was handed to the vector store and silently discarded, so results could include material from knowledge bases you have no access to. [Commit](https://github.com/open-webui/open-webui/commit/1d6d4e6e6647e1d403438ede7bd9ba20bc4cc8f6)
- 💣 **Documents that unpack far beyond their size.** A Word, Excel, PowerPoint, OpenDocument or EPUB file that expands to far more than it stores is now rejected before it is read, where one could previously be used to exhaust a server's memory. [Commit](https://github.com/open-webui/open-webui/commit/2a0274a0a039dbe0a1ad4d24003b085aae7b896b)
- ✂️ **Long replies cut off partway.** A single oversized piece of a streamed reply, such as a long reasoning trace or a turn carrying many tool calls, no longer ends the answer early with a misleading error about byte counts, which affected every default installation. [#28114](https://github.com/open-webui/open-webui/pull/28114), [#25664](https://github.com/open-webui/open-webui/issues/25664)
- 🚧 **Web address checks that could be skipped.** The fetchable-address test and the operator's web fetch filter list now run on every outgoing request, where a proxied or already-open connection could bypass them and a filter entry written as an address range silently matched nothing at all. [#27823](https://github.com/open-webui/open-webui/pull/27823)
- 🧑‍💻 **Code execution reachable through a tag in a reply.** On installs using native function calling, the older path that runs code found inside a tag in the model's reply is no longer active alongside the built-in tool, so code execution happens only through an explicit tool call. [#29024](https://github.com/open-webui/open-webui/pull/29024)
- ⌚ **Recurrence rules that could tie up the server.** How often an automation repeats is now taken from the rule the scheduler actually parsed rather than from the text of the rule, so a crafted rule can no longer disagree with what gets scheduled and walk the server through an unbounded run of occurrences, and a rule carrying a time zone on its start date now schedules instead of erroring. [Commit](https://github.com/open-webui/open-webui/commit/067114c28038b77045a3a1983e2e2719b1a32ad8)
- 🗝️ **Changing a password now ends other sessions.** Changing your password, or an administrator resetting it for you, now stops every device that was already signed in, where they had stayed signed in on the old password until their session expired on its own, up to four weeks by default; the device making the change is signed out too and asked to sign in again, and this requires Redis, without which nothing can be revoked and a warning is now logged saying so. [#28725](https://github.com/open-webui/open-webui/pull/28725), [#28647](https://github.com/open-webui/open-webui/discussions/28647)
- 🧬 **Workspace models shadowing a real one.** Someone without administrator rights can no longer create, import, or edit a workspace model so that it takes over the identity of a model served by a connected provider, where doing so would have changed what everyone else got when they picked that model. [Commit](https://github.com/open-webui/open-webui/commit/ea55d38793014a4e3cd5a4046816fe22e69e9739)
- 🌳 **Folders disappearing when moved into themselves.** Moving a folder inside one of its own subfolders is now refused, where it was accepted and made that folder and everything in it vanish from the sidebar with no way to bring it back, while leaving the server walking the loop endlessly and querying the database as it went, which could exhaust a worker and its memory; any folder already in that state is returned to the top level. [#28748](https://github.com/open-webui/open-webui/pull/28748)
- 🧨 **Searching a knowledge base with a costly pattern.** A search pattern written so that it expands enormously before it even runs is now refused, where it could tie up the server; ordinary patterns are unaffected. [Commit](https://github.com/open-webui/open-webui/commit/d5b66533e7829654f6fb343abbaaeab988dbfde8), [#28284](https://github.com/open-webui/open-webui/pull/28284)
- 💧 **Attaching a very large file from a link.** A file fetched from a link is now written to disk as it arrives and stops at the configured size limit, where the whole thing was held in memory first with no limit applied, so a large enough file could exhaust the server; a download that fails partway no longer leaves the partial file behind. [#28945](https://github.com/open-webui/open-webui/pull/28945)
- ⛓ **Deleting one knowledge base removing a shared connection.** Deleting an external knowledge base now leaves its connection in place while other knowledge bases still use it, and only an administrator removing the last one clears it, where any user deleting theirs took the connection away from everyone. [#28113](https://github.com/open-webui/open-webui/pull/28113)
- 📡 **Intermittent connection failures.** Requests to model providers and to services on the same network no longer fail intermittently with name lookup errors, often surfacing as a misleading model not found message, because addresses are resolved through the system again by default, with the faster resolver still available through "AIOHTTP_CLIENT_ASYNC_DNS_RESOLVER". [#28242](https://github.com/open-webui/open-webui/pull/28242), [#28013](https://github.com/open-webui/open-webui/issues/28013), [#28215](https://github.com/open-webui/open-webui/issues/28215)
- 🗯️ **Losing the conversation with memory on.** With the memory tool enabled, the model can see the earlier messages in your conversation again, instead of answering the second message as though the first had never been sent. [#28400](https://github.com/open-webui/open-webui/issues/28400)
- 👻 **Vanishing responses.** Replies from Responses-API providers that report an empty output at the end of a stream no longer disappear the moment generation finishes, leaving an empty message in their place. [#27800](https://github.com/open-webui/open-webui/pull/27800), [#27789](https://github.com/open-webui/open-webui/discussions/27789)
- 📥 **Queued messages disappearing.** Messages waiting to be sent are put back in the queue if sending them fails, rather than vanishing without being sent. [Commit](https://github.com/open-webui/open-webui/commit/f79b443c226098564386d0a31708e71ec0146158)
- 🧵 **Replies cut short mid-stream.** A reply no longer breaks off part way through when a provider sends the pieces of its response in an unexpected order, which had left the answer truncated and skipped the filters that run once a message finishes. [#28312](https://github.com/open-webui/open-webui/pull/28312)
- 🧷 **Replies not carried into the next turn.** With providers that skip parts of the streaming sequence, the finished reply is now taken from the completed message, so it stays available as context for your next question and the citations that arrived with it are no longer dropped. [#28310](https://github.com/open-webui/open-webui/pull/28310)
- 🌊 **Replies arriving in oversized pieces.** Very large streamed pieces no longer break the response on default settings, where the reader that splits them safely only ran when a chunk size limit was configured. [Commit](https://github.com/open-webui/open-webui/commit/a33fa05adc6def8f3d098539a4786dc9c7bf61d2)
- 🩹 **Signing in after a long-delayed upgrade.** Accounts on instances that were upgraded from a version older than 0.6.41 to 0.9.6 or newer can sign in again, where an upgrade step had written their single sign-on identity in a form the application could not read afterwards, and a repair step corrects the affected accounts on startup. [#28107](https://github.com/open-webui/open-webui/pull/28107), [#28101](https://github.com/open-webui/open-webui/issues/28101), [Commit](https://github.com/open-webui/open-webui/commit/bd8378f643dd8a56366ad9d1843b7d40f67cf1f9)
- 🔑 **Signing in with some identity providers.** Logging in through a provider that adds its own vendor-specific information to the header of the sign-in token now completes, rather than failing at the final step with a message claiming the email or password was wrong. [#28065](https://github.com/open-webui/open-webui/pull/28065), [#28062](https://github.com/open-webui/open-webui/issues/28062)
- 🔌 **Role changes taking effect at once.** Changing someone's role now ends their live sessions no matter how the change was made, whether by a directory sync, an identity provider, a trusted header, or deleting the account, so permissions from their old role cannot linger, and their browser reconnects on its own. [Commit](https://github.com/open-webui/open-webui/commit/ce3c175e260709f359d7e6cbb3132f0572098b95)
- 🛑 **Memory permission being respected.** Taking away someone's memory permission now also stops their stored memories being added to the context of their conversations, which one path had continued doing regardless. [#27668](https://github.com/open-webui/open-webui/pull/27668)
- 🔍 **Listing a single connection's models.** Asking for the models or version of one particular connection is now restricted to administrators, and a request naming a specific backend is checked against the models that backend actually serves even where the access control bypass is turned on. [Commit](https://github.com/open-webui/open-webui/commit/16f118d77ad9d68c64116f94551ffb14bf8b2abd)
- ⚖️ **Sharing defaults matching what was configured.** On instances upgraded from older versions, public sharing of tools and notes no longer shows as switched on in the admin panel, and saving any unrelated permission no longer grants everyone a capability that was never enabled. [#27716](https://github.com/open-webui/open-webui/pull/27716), [#27715](https://github.com/open-webui/open-webui/issues/27715)
- 🗂️ **Folder permissions when starting a chat.** Starting a conversation filed into a folder now checks that you are allowed to write to that folder, a check the message sending path had been skipping, and every place a chat can be filed now treats ownership, shared access, and unknown folders the same way. [#28366](https://github.com/open-webui/open-webui/pull/28366)
- 🪧 **Clearer attachment failures.** A link that cannot be read now says so and names the link, and a YouTube video whose transcript is refused explains why and points at the proxy setting that exists for it, instead of both being reported as a knowledge base error. [#28362](https://github.com/open-webui/open-webui/pull/28362), [#28361](https://github.com/open-webui/open-webui/issues/28361)
- 🔎 **Chat search finding recent messages.** Searching your chats now looks inside the messages of current conversations on default installations, where it had only been reading an older storage format and missing their content entirely. [Commit](https://github.com/open-webui/open-webui/commit/0800c21c64c64810f24c2ec88cca5e36daebb10e)
- 🧭 **Your place in a compacted chat.** Opening a conversation whose history has been compacted now takes you to its most recent message instead of leaving you parked on the summary, and updates to an existing message no longer move your place in the conversation. [Commit](https://github.com/open-webui/open-webui/commit/5caa91a49304148696641d1cd39e43184ba8d748)
- 🖥️ **Chats with a personal terminal.** Sending a message with a terminal you added yourself under Settings selected no longer fails with a terminal unavailable error, which had blocked those chats since 0.11.0. [#27621](https://github.com/open-webui/open-webui/issues/27621), [Commit](https://github.com/open-webui/open-webui/commit/5b333d75c6adea4d8bde96439974d6e9c83d8198), [Commit](https://github.com/open-webui/open-webui/commit/3becec6ccfc7dc457270d4e9f0d8c957e8160ddc)
- 📆 **Default date for new events.** Creating a calendar event now starts on today's date rather than tomorrow's when you open the form in the evening, or yesterday's when you open it early in the morning. [#27779](https://github.com/open-webui/open-webui/pull/27779), [#27778](https://github.com/open-webui/open-webui/issues/27778)
- 🗓️ **Recurring event times.** Repeating calendar events now show at the time you set them for instead of being worked out in the server's time zone and shifted by the gap between the two. [#27774](https://github.com/open-webui/open-webui/issues/27774), [Commit](https://github.com/open-webui/open-webui/commit/d721b0d19621d01e35505105a778ce5681b2ffba)
- 🧩 **Chats during model list refreshes.** On direct connections, background work such as title and tag generation no longer fails or runs against a mix of old and new model entries while the model list is being refreshed. [#27821](https://github.com/open-webui/open-webui/pull/27821)
- 🎛️ **Chat Controls staying put.** Hovering a chat in the sidebar whose preview contains an artifact no longer forces the Chat Controls pane open and fills it with that artifact, over the chat you currently have open. [#27773](https://github.com/open-webui/open-webui/pull/27773), [#27772](https://github.com/open-webui/open-webui/issues/27772)
- 📨 **Reliable streaming with unusual characters.** Responses containing any of three rare invisible line break characters no longer arrive split or broken when the "ENABLE_ORJSON" option is turned on. [#27819](https://github.com/open-webui/open-webui/pull/27819)
- 🧮 **JSON options honoured again.** Options passed to the shared JSON helper are no longer silently ignored when the "ENABLE_ORJSON" option is turned on, falling back to the standard encoder that supports them. [Commit](https://github.com/open-webui/open-webui/commit/78ed5a0235c4de67828ba4b0d6147035067cec2c)
- 🚫 **Duplicate models in lists.** Adding a model that is already on a connection's allowed list is now rejected instead of quietly adding it a second time, the arena picker no longer offers models you have already chosen, and existing duplicates are cleaned up the next time the list is saved. [#28251](https://github.com/open-webui/open-webui/pull/28251), [#28249](https://github.com/open-webui/open-webui/issues/28249)
- 📁 **Dragging chats into shared folders.** A shared folder you can write to now highlights and accepts a dropped chat, while one you only have read access to no longer offers itself as a drop target for an action that could only fail. [Commit](https://github.com/open-webui/open-webui/commit/e4dd6c4bf14c7ea54c1effa6c40ad9e95185c413), [#28261](https://github.com/open-webui/open-webui/issues/28261)
- ⏰ **Listing automations through chat.** Asking a model to list your automations without naming a folder now returns every automation you have, instead of only the ones that sit outside a folder. [Commit](https://github.com/open-webui/open-webui/commit/f8ac75d188b7e07906aa1d5e648e3fc25a78ef2b)
- 🧠 **Faster follow-ups with memory enabled.** The memories handed to the model now appear in a stable order from one message to the next, so servers that reuse their work between turns no longer reprocess the whole conversation each time you reply. [#28292](https://github.com/open-webui/open-webui/issues/28292), [Commit](https://github.com/open-webui/open-webui/commit/ff74bfa6a117c6f03097034c74c2f9bdfb824203), [Commit](https://github.com/open-webui/open-webui/commit/d22bb6703f244e6769ac9b0f2e41d784ad6cf535)
- 🎯 **Custom model parameters combining.** Setting a custom parameter on a model no longer silently discards every custom parameter defined in the global defaults, and a value sent directly in an API request is no longer overwritten by the model's saved settings. [Commit](https://github.com/open-webui/open-webui/commit/11739a2de8eddf7ef26378356aff8a1bcab4a349), [#28241](https://github.com/open-webui/open-webui/issues/28241)
- 👥 **Sharing with people already added.** The access picker no longer offers people and groups that already have access, and the Users heading no longer appears above an empty list. [Commit](https://github.com/open-webui/open-webui/commit/385d08bea5899bcb160391db466662b0a0a6f801), [#28253](https://github.com/open-webui/open-webui/issues/28253)
- 🔧 **Full tool parameter descriptions.** A tool whose parameter description runs over several lines now passes the whole description to the model instead of only its first line. [Commit](https://github.com/open-webui/open-webui/commit/b606e13da3753027ead0e92e35e1399b60c96c8e)
- 📝 **Notes saved in an unexpected shape.** A note whose content was stored as structured data rather than text no longer breaks the notes page for everything else, and opens with that content shown as a formatted code block. [Commit](https://github.com/open-webui/open-webui/commit/8d1c205d8e7335ef7c292afb125a7270976f327b), [#28222](https://github.com/open-webui/open-webui/issues/28222)
- 💬 **Direct messages after an account is deleted.** A direct message conversation no longer counts a deleted account among its members, and opening a direct message with someone finds the existing conversation instead of starting a second one alongside it. [Commit](https://github.com/open-webui/open-webui/commit/a41faa3c226b20e7db7508c49c7d1d015ac9e973), [#28257](https://github.com/open-webui/open-webui/issues/28257)
- 🫥 **Deactivating a model.** Turning a model off no longer removes the wrong entry from the model list, or fails the list outright and leaves the model picker empty for everyone until the model is turned back on. [Commit](https://github.com/open-webui/open-webui/commit/5cecb7dbfad3994228ee53d19232650400462332), [#28202](https://github.com/open-webui/open-webui/issues/28202)
- 💭 **Readable errors on chat actions.** When moving, renaming, or otherwise changing a chat fails, the message explaining why now appears in place of an unhelpful object placeholder. [#28260](https://github.com/open-webui/open-webui/pull/28260), [#28259](https://github.com/open-webui/open-webui/issues/28259)
- 🪪 **Authorship in shared chats.** A chat shared with you now shows the name and picture of whoever wrote it, in the message list and in the overview panel, rather than crediting the messages to you. [#28274](https://github.com/open-webui/open-webui/pull/28274), [#28273](https://github.com/open-webui/open-webui/issues/28273)
- 🖇️ **Adding terminals over plain connections.** Saving a terminal connection now works when the interface is served without HTTPS, where the dialog would sit there doing nothing because the browser withholds the tool used to generate its identifier. [Commit](https://github.com/open-webui/open-webui/commit/2a45fa04cb1d07b258e0b880ab0c02577f0dc035), [#28148](https://github.com/open-webui/open-webui/issues/28148)
- 🗃️ **openGauss vector storage.** Deployments using openGauss for vector storage no longer fail the moment they touch it. [#27838](https://github.com/open-webui/open-webui/pull/27838)
- 🎚️ **ColBERT reranker startup.** Loading a ColBERT reranker now names the model in the log rather than printing a logging error and a traceback in its place. [#27838](https://github.com/open-webui/open-webui/pull/27838)
- 🖼️ **Images that no longer exist.** A message whose image file has been deleted now shows a small unavailable placeholder that cannot be opened, rather than a broken image that spilled the whole reply text into the picture frame and still opened full screen. [#27730](https://github.com/open-webui/open-webui/pull/27730), [#27728](https://github.com/open-webui/open-webui/issues/27728), [Commit](https://github.com/open-webui/open-webui/commit/f8c5fda283e6206d92e81ae3b83ffe5d7a144b03)
- 🛂 **Connecting external accounts.** Authorizing a tool's external account now completes only for the person who started it, rather than for whoever happens to return with the authorization, and signing out clears the session it relies on. [Commit](https://github.com/open-webui/open-webui/commit/c2107e5bb3689a69c170ca526925f4ed84bd00f5)
- 🪟 **Starting up on Windows.** The Windows start script now creates the secret key it needs on a fresh installation, instead of printing a run of file not found messages and then refusing to start, and it copes with an installation path that contains spaces. [#28061](https://github.com/open-webui/open-webui/pull/28061), [#28060](https://github.com/open-webui/open-webui/issues/28060)
- 🕸️ **Overlapping branches in the overview.** Branch nodes in a chat's overview keep a clear gap between them when the interface is scaled up, rather than sitting on top of one another. [#27995](https://github.com/open-webui/open-webui/pull/27995), [#27994](https://github.com/open-webui/open-webui/issues/27994)
- 🗑️ **Delete offered only when allowed.** The chat deletion controls in the sidebar, the chat menu, search, archived chats, and data controls no longer appear for people whose permissions do not allow deleting, where using them produced an access denied error. [#27714](https://github.com/open-webui/open-webui/pull/27714), [#27713](https://github.com/open-webui/open-webui/issues/27713)
- 🍴 **Fork offered only when allowed.** The fork action and the fork command no longer appear for people whose chat import permission is turned off, where using them produced an access denied error. [#27711](https://github.com/open-webui/open-webui/pull/27711), [#27692](https://github.com/open-webui/open-webui/issues/27692)
- ↕️ **Expand button in the message box.** The button that enlarges the message box no longer sits on top of a tagged model's dismiss button or the first attached file, and stays reachable in long prompts. [#27676](https://github.com/open-webui/open-webui/pull/27676), [#26736](https://github.com/open-webui/open-webui/issues/26736)
- 🔆 **Regenerate in high contrast mode.** With high contrast mode on, the regenerate button now stays visible on earlier replies instead of appearing only when you hover over them. [#27644](https://github.com/open-webui/open-webui/pull/27644), [#27638](https://github.com/open-webui/open-webui/issues/27638)
- ✂️ **Clipped icons and avatars.** The terminal icon beside the message box and the profile picture in account settings are no longer shaved flat along their left edge. [#27691](https://github.com/open-webui/open-webui/pull/27691), [#27690](https://github.com/open-webui/open-webui/issues/27690)
- 🪄 **Merged responses after a reload.** Merging the answers from several models now works on a conversation you have reopened, instead of the merging model reporting that the other responses were empty. [#27673](https://github.com/open-webui/open-webui/pull/27673), [#26962](https://github.com/open-webui/open-webui/issues/26962)
- 🔢 **Token counts for background chats.** Conversations started by automations, timers, sub-agents, and channels now report their token usage like any other chat, rather than arriving without it even when the model is set up to provide it. [#27661](https://github.com/open-webui/open-webui/pull/27661), [#27653](https://github.com/open-webui/open-webui/issues/27653)
- 📐 **Settings on tall screens.** The settings window now grows with the height of your display instead of stopping short and making you scroll inside it while space sits unused above and below. [#27615](https://github.com/open-webui/open-webui/pull/27615), [#27614](https://github.com/open-webui/open-webui/issues/27614)
- 🖱️ **Sections opening by accident.** Folders, collapsible sections, and tool call blocks now open and close only when you click them, rather than also reacting when you release the mouse over them after dragging or selecting text. [Commit](https://github.com/open-webui/open-webui/commit/bd250a0e2431f8c2c7e2f4a34c5327f21fc77d4b)
- 🔁 **Rebuilding knowledge base vectors.** Rebuilding the vectors for a knowledge base now also rebuilds them for each file it contains, so attaching a single file afterwards finds its content instead of quietly returning nothing and letting the model answer from thin air. [#28106](https://github.com/open-webui/open-webui/issues/28106), [Commit](https://github.com/open-webui/open-webui/commit/2a6e671f548970c8223692024e630b9936e9fa7c), [Commit](https://github.com/open-webui/open-webui/commit/89922cc9d585e10b026693681b44afdc4b874588)
- 📌 **Attaching a chat shared with you.** Attaching a conversation that was shared with you, directly or through a shared folder, now brings its content along instead of quietly attaching nothing. [Commit](https://github.com/open-webui/open-webui/commit/5cd9a395344e882b4384bec0899f503ad540aecc)
- 🧲 **The page staying still when typing.** Returning focus to the message box no longer scrolls the conversation, so switching chats, running a command, or picking something from a menu leaves your place on screen alone. [Commit](https://github.com/open-webui/open-webui/commit/9122c24ea2d506aad50384ca8aba1616d1bed626)
- 📷 **Round profile pictures on narrow screens.** Profile pictures in the admin user list and other lists no longer squash into ovals of differing widths when the window is narrow. [#28000](https://github.com/open-webui/open-webui/pull/28000), [#27999](https://github.com/open-webui/open-webui/issues/27999)
- 🎙️ **Voice mode in the notes editor.** The voice mode button is no longer offered in the chat embedded in a note, where it does not apply. [Commit](https://github.com/open-webui/open-webui/commit/9122c24ea2d506aad50384ca8aba1616d1bed626), [Commit](https://github.com/open-webui/open-webui/commit/e963d36e393aead56b941e99d4167fb0dc9d2d7a)
- ⏳ **Analytics stuck loading.** Choosing a custom date range in analytics without picking dates yet no longer leaves the tab spinning forever, including after leaving and coming back to it. [#28125](https://github.com/open-webui/open-webui/issues/28125), [Commit](https://github.com/open-webui/open-webui/commit/629cdcb5303958c5b99ee437030ab80cc418bc2e)
- 🧑‍🤝‍🧑 **Owner avatars on shared chats.** The picture beside a chat someone shared with you now loads, and falls back to the default image if it cannot, rather than leaving a blank gap when the interface and the server are on different addresses. [#28272](https://github.com/open-webui/open-webui/pull/28272), [#28271](https://github.com/open-webui/open-webui/issues/28271)
- 🎨 **Image generation and web search staying switched off.** Turning either off now takes effect at once on every path: sessions opened beforehand can no longer produce images or run searches, an image request no longer reaches the provider on a model using the older tool-calling method, the entry disappears from the integrations menu right away, and an active marker beside the message box no longer lingers after its feature is withdrawn. [#27759](https://github.com/open-webui/open-webui/pull/27759), [#27758](https://github.com/open-webui/open-webui/issues/27758), [#26842](https://github.com/open-webui/open-webui/issues/26842), [#27669](https://github.com/open-webui/open-webui/pull/27669)
- 📰 **Attached pages reaching the model.** The text pulled from an attached web page or YouTube video now actually reaches the model, rather than arriving empty so the reply had nothing to work from, and opening the source to check no longer fails. [#28378](https://github.com/open-webui/open-webui/issues/28378), [Commit](https://github.com/open-webui/open-webui/commit/9c21d4ed3ba9ba8e53def7e1a1366b74834cb42d)
- 🌐 **Tavily page fetching.** Reading a web page with Tavily selected as the loader works again, having failed on every attempt since 0.10.0. [#27636](https://github.com/open-webui/open-webui/pull/27636), [#27602](https://github.com/open-webui/open-webui/issues/27602)
- 🗒️ **Reply box in threads.** The reply box in a channel thread now stays at the bottom of the panel while you scroll back through the replies, instead of scrolling out of sight with them. [#27768](https://github.com/open-webui/open-webui/pull/27768), [#27767](https://github.com/open-webui/open-webui/issues/27767)
- 🎹 **Model picker shortcut.** The keyboard shortcut for opening the model picker works again, and a link to a chat naming a model you do not have still opens the picker with that name filled in. [Commit](https://github.com/open-webui/open-webui/commit/e1acd7e7ca5085e03babbea16b4770e51c38886e)
- ⌨️ **Reaching the download options by keyboard.** In the model picker, arrowing past the last result now moves through the options to fetch that model from each server that can supply it, so they can be chosen with the keyboard instead of only by clicking. [Commit](https://github.com/open-webui/open-webui/commit/4e03d89414a7be1203579f45225ca64f1221205f), [Commit](https://github.com/open-webui/open-webui/commit/25802c048e6123fa602182949ad2a9f349e1e863)
- 📂 **Opening a folder in the sidebar.** Selecting a folder now refreshes just that folder's chats rather than rebuilding the whole folder tree, and a folder that is empty or still loading says so instead of showing nothing. [Commit](https://github.com/open-webui/open-webui/commit/2e939874904e1ae54d02b8bc6d745931a8e1f673), [Commit](https://github.com/open-webui/open-webui/commit/86b7bf1f7e11109816b0d16ada708c637b50c520)
- ⏲️ **Changing an automation through chat.** Asking a model to change one thing about an automation no longer moves it out of its folder or drops its model when the model fills those fields in blank instead of omitting them. [Commit](https://github.com/open-webui/open-webui/commit/90bb94abf9a8c6fbb87c72952295ae39d2ad98e5)
- 🔂 **Automations that run a set number of times.** An automation asked to run a limited number of times is now rejected unless it says when to start counting from, rather than being accepted and then running indefinitely. [#27781](https://github.com/open-webui/open-webui/pull/27781), [#27780](https://github.com/open-webui/open-webui/issues/27780)
- 📅 **Editing a calendar event through chat.** Asking a model to change one thing about an event, such as its title, no longer fails or wipes the details you did not mention. [#27777](https://github.com/open-webui/open-webui/pull/27777), [#27776](https://github.com/open-webui/open-webui/issues/27776)
- 🧹 **Session cleanup on multi-instance setups.** The instance doing the periodic session cleanup now keeps its claim on that job alive between passes, so an idle deployment stops logging a renewal warning every two minutes and the claim no longer lapses for half of every cycle. [Commit](https://github.com/open-webui/open-webui/commit/939bcdb79e3dcad2278f1e7f4f517a3f2f36f3ec), [#27762](https://github.com/open-webui/open-webui/issues/27762)
- 🏠 **Starting folder in the file browser.** Reopening the file browser now keeps the folder you were in, instead of the breadcrumb losing its starting point and jumping you elsewhere. [Commit](https://github.com/open-webui/open-webui/commit/52c5e3b20d3cd8c3a15db50e2cc75bbd6e82068e)
- 🩺 **Repairing default model settings.** Instances whose stored default and pinned model settings had been written in the wrong shape are corrected on startup, so those defaults take effect again. [Commit](https://github.com/open-webui/open-webui/commit/5c05608e3ac0aefce8c1eed18c3e205c61f6b6fe)
- 🧽 **Cleaner conversation history for the model.** Internal bookkeeping attached to your messages, such as attachment records and token counts, is no longer sent to the model along with the conversation. [Commit](https://github.com/open-webui/open-webui/commit/a32a17965ca36f730ebf5ff53a2a8e17082b61a6)
- 🙈 **Needless request from the model picker.** Opening the model picker as a non-administrator no longer fires a request to an administrator-only settings endpoint that was always refused. [Commit](https://github.com/open-webui/open-webui/commit/b4738d1a2e6af1ce20adb322fc6d2aa2c9407901)
- ⏹️ **Stopping a reply that is waiting.** The stop button now ends a reply that is sitting waiting for you, such as one paused on a tool approval, rather than leaving the conversation stuck part way through. [Commit](https://github.com/open-webui/open-webui/commit/f7767d6be774720e16265c6d016bf6a24a1a8759)
- 🏷️ **Folder names with unusual characters.** Naming a folder is no longer refused because another folder's name happens to be similar, and a name ending in a backslash no longer fails outright on PostgreSQL, because names are now compared exactly rather than treated as a search pattern. [#28695](https://github.com/open-webui/open-webui/pull/28695), [#28694](https://github.com/open-webui/open-webui/issues/28694)
- 🤔 **Reasoning carried back to Ollama.** A model's earlier thinking is now passed back to Ollama in its own native field rather than pasted into the message as tagged text, so reasoning models keep their train of thought across turns. [Commit](https://github.com/open-webui/open-webui/commit/3258330729942b533dc5fe141e876cee7e5eb40d)
- 📎 **Default pinned models taking effect.** Changing the default pinned models now reaches people who have never chosen their own, where simply having opened the interface once was enough to freeze the list they first saw, and reordering a pin no longer moves the wrong one or reopens the sidebar section afterwards. [#28069](https://github.com/open-webui/open-webui/pull/28069), [#28067](https://github.com/open-webui/open-webui/discussions/28067)
- ✏️ **Editing other people's channel messages.** Asking a model to work on a message in a channel now only applies to your own messages, where write access to the channel had been enough to reach anyone's. [#28631](https://github.com/open-webui/open-webui/pull/28631)
- 🍪 **Signed-in tool servers.** A tool server that relies on your session now receives the credentials belonging to its own connection, rather than whichever were most recently prepared. [#28630](https://github.com/open-webui/open-webui/pull/28630)
- 🏗️ **Editing a folder from its page.** Renaming a folder, changing its icon, creating a subfolder, or deleting it from the folder's own page now updates the sidebar straight away, instead of leaving the old name and icon there, and the new subfolder missing, until a reload. [Commit](https://github.com/open-webui/open-webui/commit/a40f6f2860b49b4c7f11e369f11bd46e9451eabf), [Commit](https://github.com/open-webui/open-webui/commit/736e38338ede1fca5af0829e35585a72f8517d32), [#28692](https://github.com/open-webui/open-webui/pull/28692), [#28690](https://github.com/open-webui/open-webui/issues/28690)
- 📣 **Long channel names in the sidebar.** A channel with a long name no longer squeezes its own menu button out of the row. [Commit](https://github.com/open-webui/open-webui/commit/b5da50f3df51786972010c75c7736d5f95362e58), [#28671](https://github.com/open-webui/open-webui/pull/28671), [#28670](https://github.com/open-webui/open-webui/issues/28670)
- 📬 **Mark as unread in chat search.** Marking a chat unread from the search dialog now works and updates the sidebar, where the menu entry looked normal but did nothing at all. [#28136](https://github.com/open-webui/open-webui/pull/28136), [#28135](https://github.com/open-webui/open-webui/issues/28135)
- ⬆️ **Scroll to top on the first click.** In a long chat where older messages had not been loaded yet, one click of scroll to top now reaches the first message instead of stopping short and needing a second. [#28659](https://github.com/open-webui/open-webui/pull/28659), [#28658](https://github.com/open-webui/open-webui/issues/28658)
- 🔘 **Double bullets in the release notes.** Each entry in the what's new dialog shows a single bullet again, rather than two sitting at different heights. [#28676](https://github.com/open-webui/open-webui/pull/28676), [#28675](https://github.com/open-webui/open-webui/issues/28675)
- 📚 **Knowledge search and shared files in chat.** A model searching your knowledge bases or reading a file shared with you through a group now works, where it had failed since 0.11.0 and quietly answered as though the knowledge were empty, affecting instances that forward user details to their embedding service and, for shared files, every instance regardless of settings. [#27642](https://github.com/open-webui/open-webui/pull/27642), [#27641](https://github.com/open-webui/open-webui/issues/27641)
- 🔖 **Skill identifiers that cannot be reached.** Creating a skill whose identifier contains a character that is not allowed in a web address is now refused outright, rather than accepted and then permanently impossible to open, edit, turn off, delete, or recreate. [#27660](https://github.com/open-webui/open-webui/pull/27660), [#27655](https://github.com/open-webui/open-webui/issues/27655)
- 🔤 **Model names on connections with a prefix.** A connection that adds a prefix to its model names now strips it before sending a request through the responses endpoint, where the prefixed name was passed on and rejected as unknown. [#28575](https://github.com/open-webui/open-webui/pull/28575), [#28574](https://github.com/open-webui/open-webui/issues/28574)
- 🔓 **Turning on open sharing.** The open sharing permission can now be switched on in the default user permissions, where saving appeared to work but the setting was discarded and came back off. [#27609](https://github.com/open-webui/open-webui/pull/27609), [#27607](https://github.com/open-webui/open-webui/issues/27607)
- 🖌️ **White boxes behind model icons.** Model icons with a transparent background no longer sit on a white square in the admin models list, matching how they already appeared everywhere else. [#27612](https://github.com/open-webui/open-webui/pull/27612), [#27611](https://github.com/open-webui/open-webui/issues/27611)
- 🪞 **Matching the right account at sign-in.** Looking up an account by its identity provider details now matches the exact value, where the stored details were searched as loose text and a value contained within another's could be matched instead. [#28624](https://github.com/open-webui/open-webui/pull/28624)
- 🔄 **Syncing a model catalogue more than once.** Syncing models now updates the ones that already exist, where any repeat of a previous sync silently did nothing at all while still reporting success. [#28036](https://github.com/open-webui/open-webui/pull/28036), [#28033](https://github.com/open-webui/open-webui/issues/28033)
- 🕰️ **Saving a calendar event without a date.** Creating or editing an event with the date cleared now asks for one, where it was sent anyway, refused by the server, and reported as an unreadable error. [Commit](https://github.com/open-webui/open-webui/commit/f100edb70874808c93fab84eae1810595f9e9dc3), [#28133](https://github.com/open-webui/open-webui/issues/28133)
- 🔗 **Deleting a message in a looping chat.** Deleting a message no longer hangs when the conversation contains a cycle in its reply structure. [#28035](https://github.com/open-webui/open-webui/pull/28035)
- 🫀 **Scheduled work stopping without warning.** The routine that runs automations and calendar alerts can no longer be discarded while the application is running, which had silently stopped them firing with nothing reported, and it now stops cleanly on shutdown. [#28053](https://github.com/open-webui/open-webui/pull/28053), [#28052](https://github.com/open-webui/open-webui/issues/28052)
- 🔋 **Session and usage records left uncleared.** The routines that clear out stale sessions and finished model usage can no longer be discarded while the application is running, so those records stop accumulating unnoticed, and both now stop cleanly on shutdown. [#28053](https://github.com/open-webui/open-webui/pull/28053), [#28052](https://github.com/open-webui/open-webui/issues/28052)
- ⚗️ **Reasoning carried between turns.** A model's earlier thinking is now recognised from providers that report it in their own nested field, and reasoning that cannot be sent back without a signature is left out rather than being passed on and rejected. [Commit](https://github.com/open-webui/open-webui/commit/b6dc70c93b0d36e2659438e7a55aa8b21fee27f8)
- 🛎️ **Losing all your settings.** Your interface settings are no longer wiped by a session that failed to load them, which could happen with no action on your part and cleared everything from your theme to your model parameters; saving now changes only the settings you actually changed, and a session that cannot load them tells you instead of carrying on as though you had none. [#27766](https://github.com/open-webui/open-webui/issues/27766), [Commit](https://github.com/open-webui/open-webui/commit/ad8c79f68657bd3bcf5db6be650e498bb904b36b)
- 🖲️ **Losing the collapsed sidebar.** With the sidebar collapsed, opening a chat no longer pushes the narrow sidebar strip off the edge of the screen, which left no way to reopen the sidebar short of shrinking the window to phone size. [#28501](https://github.com/open-webui/open-webui/pull/28501), [#28500](https://github.com/open-webui/open-webui/issues/28500)
- 🧯 **Timers that fail without saying so.** A timer whose reply cannot be generated, such as one set against a model that has since been removed, is now recorded as failed with the reason, instead of being marked as completed while the reply never arrives. [#27785](https://github.com/open-webui/open-webui/pull/27785), [#27783](https://github.com/open-webui/open-webui/issues/27783)
- 🎲 **Timers stop for retired owners.** A timer whose owner has been deleted or demoted to pending is recorded as an error instead of running, so a retired account no longer answers through a timer it set while active. [#30220](https://github.com/open-webui/open-webui/pull/30220)
- 🖊️ **Message buttons in channels.** The buttons that appear when you hover a channel message now sit above the message rather than over its content, so they can be clicked on a message that starts with a code block or a table, and so the code and table controls stay clickable too. [#27737](https://github.com/open-webui/open-webui/pull/27737), [#27736](https://github.com/open-webui/open-webui/issues/27736)
- 🔡 **Searching for non-English tags and text.** Searching workspace models by tag, or prompts and automations by their contents, now finds entries containing characters outside the English alphabet, where roughly half were missed depending on which settings were in force when each one was saved. [#28399](https://github.com/open-webui/open-webui/pull/28399)
- 🔭 **Searching the calendar without an end date.** Asking a model to search your calendar without naming an end date now works on PostgreSQL, where the open-ended range was too large for the database to accept and the search failed outright. [Commit](https://github.com/open-webui/open-webui/commit/9550731cc17759f6862595b8cd849ae48695b5c1), [#27717](https://github.com/open-webui/open-webui/issues/27717)
- 🫧 **Attachments replaced by a loading dot.** Pinning a channel message, or otherwise updating one, no longer replaces its attachment with a loading indicator that never resolves until you reload or leave the channel. [Commit](https://github.com/open-webui/open-webui/commit/76d01602950a9e40823b53278557708d5bcd1036), [#27734](https://github.com/open-webui/open-webui/pull/27734), [#27731](https://github.com/open-webui/open-webui/issues/27731)
- 🛠️ **Rebuilding empty server lists on every request.** An instance with no tool servers or no terminal servers configured no longer rebuilds that empty list on every request that needs it. [Commit](https://github.com/open-webui/open-webui/commit/f1a64ccfc2eb2a58086c55fe413a455bb488c35a), [#28568](https://github.com/open-webui/open-webui/issues/28568)
- 🪫 **Errors logged for a cache that was simply empty.** A shared cache that has not been filled yet no longer logs an error suggesting its stored value is broken. [Commit](https://github.com/open-webui/open-webui/commit/f1a64ccfc2eb2a58086c55fe413a455bb488c35a), [#28568](https://github.com/open-webui/open-webui/issues/28568)
- ⚓ **Starting up as a non-root user.** Deployments that run the container as a non-root user, such as Kubernetes setups using runAsNonRoot, start again, where a bundled speech model file that only root could read had stopped them since 0.11.0, and a bundled text corpus is now stored somewhere a non-root user can reach. [#27651](https://github.com/open-webui/open-webui/issues/27651), [Commit](https://github.com/open-webui/open-webui/commit/0480ca9653f0d566eaedadf0af0d785a9938480b), [#28866](https://github.com/open-webui/open-webui/pull/28866)
- 📋 **Finding notes shared read only.** A note shared publicly for reading now appears in the read only view of your notes, where it was readable by anyone with the link but listed nowhere at all. [#27637](https://github.com/open-webui/open-webui/pull/27637), [#27487](https://github.com/open-webui/open-webui/issues/27487)
- 🫂 **Signing in from another application.** Signing in through an application that exchanges a token from your identity provider now applies your role and group memberships the same way signing in through the browser does, and an account whose provider sends no role keeps the one it has rather than being reset to the default. [Commit](https://github.com/open-webui/open-webui/commit/d799e81edbdc971c6deb096b6474cd95b93504bf), [Commit](https://github.com/open-webui/open-webui/commit/e9684458125c3202ebc3378aef32f0b6015171d8)
- 🗄️ **Empty models section in the sidebar.** The models section no longer appears with nothing in it when every pinned model has since been removed, renamed, or hidden. [#27634](https://github.com/open-webui/open-webui/pull/27634), [#27633](https://github.com/open-webui/open-webui/issues/27633)
- 🚧 **The files pane reopening by itself.** With a terminal selected, closing the files pane now keeps it closed, where saving any setting reopened it, including something as incidental as picking an emoji for a folder. [#28693](https://github.com/open-webui/open-webui/pull/28693), [#28691](https://github.com/open-webui/open-webui/issues/28691), [Commit](https://github.com/open-webui/open-webui/commit/c1c81f8127466a4cd59a16d8658155b1b8faa2d0)
- 🚰 **Watching file processing tying up the database.** Waiting for a file or a knowledge base to finish processing no longer holds a database connection open for as long as the page is watching, which on busy instances could use up every available connection and leave the rest of the application unable to reach the database. [#28183](https://github.com/open-webui/open-webui/pull/28183)
- 🗜️ **Embedding settings for other providers.** Saving your embedding settings now writes only the provider you have selected, where it also overwrote the address and key stored for the other two, losing them if their fields were not filled in. [Commit](https://github.com/open-webui/open-webui/commit/87d9b7e84e71b097eadf1df4f9852359104f17ed)
- ↔️ **Connectors in the side-by-side overview.** With the conversation overview laid out left to right, the lines between messages now join at the sides rather than the top and bottom, so they no longer cut across the boxes. [Commit](https://github.com/open-webui/open-webui/commit/6e468c5b9539d3b6057bab63d81e8d2974acff7b)
- 💫 **Thinking indicator with the fade turned off.** Turning off the fade effect for streaming text no longer hides the thinking indicator and its spinner, which had made a reasoning model look as though it had already finished from the first moment it started. [#28559](https://github.com/open-webui/open-webui/issues/28559), [Commit](https://github.com/open-webui/open-webui/commit/b1dc945bd6cf97a7067e603fc1b797301e626fb6)
- 🈳 **Conversations compacted too early on llama.cpp.** A conversation served by llama.cpp is no longer shortened at roughly half the size you configured, where its cached input was counted twice, so a chat showing 39,000 tokens was treated as 77,000 against a 70,000 limit. [#28590](https://github.com/open-webui/open-webui/issues/28590), [Commit](https://github.com/open-webui/open-webui/commit/0b27fa5e873c9f3cff7b14e4eeab8321bdd8970c)
- ◻️ **Settings tabs spilling past the corner.** Scrolling the list of tabs in settings no longer paints a tab or part of an icon across the dialog's rounded bottom corner, where it appeared to sit outside the dialog. [#27617](https://github.com/open-webui/open-webui/pull/27617), [#27616](https://github.com/open-webui/open-webui/issues/27616)
- 🎰 **Settings for a switched-off function.** A function that has been turned off no longer offers its per-user settings, and saving them is refused, where doing so loaded the function's code and stored settings that had no effect. [Commit](https://github.com/open-webui/open-webui/commit/a3a81fee03ba7ec0ceb4d1f456e80a8f4fc2320f)
- 🗃 **Shared folders reordering themselves.** The list of folders shared with you now keeps a consistent order, where on PostgreSQL renaming a folder could shuffle the ones beside it. [#28804](https://github.com/open-webui/open-webui/pull/28804)
- 📜 **System prompt repeated after a tool call.** On models served by a pipe or manifold, the system prompt is no longer added again each time a tool runs, where it built up one extra copy per round and was sent to the provider that way. [#28739](https://github.com/open-webui/open-webui/pull/28739), [#28736](https://github.com/open-webui/open-webui/issues/28736)
- 🔕 **Calendar reminders stopping for everyone.** A single event whose reminder time was stored as something other than a number no longer stops reminders being sent, for that event or for anyone else's, and falls back to the usual reminder window instead. [#28790](https://github.com/open-webui/open-webui/pull/28790)
- 🗨 **Attached conversations reaching the model.** A conversation attached to your message is now listed among its attachments, where it was left out entirely and the model was never told it was there. [#28788](https://github.com/open-webui/open-webui/pull/28788)
- ↔ **Dragging a side panel closed.** Dragging the chat controls, the note chat, or a channel thread panel closed by its edge no longer floods the browser console with errors and leaves stray handlers behind, and the divider can now be moved with the arrow keys once focused. [#28759](https://github.com/open-webui/open-webui/issues/28759), [Commit](https://github.com/open-webui/open-webui/commit/33dff414e829d10dc6691b1cac7457077fa6799c)
- 🎫 **Saving a message with unusual characters on PostgreSQL.** A message carrying characters PostgreSQL will not store outside its text no longer fails to save, where those characters were cleaned from the conversation but passed through raw to the separate message record. [#28820](https://github.com/open-webui/open-webui/pull/28820)
- 🔲 **Removing an item in the model editor.** Unticking a tool, skill, action, or filter no longer leaves the next one in the list looking unticked while it is still selected, needing two clicks to remove and passing the same confusion down the list each time. [#28837](https://github.com/open-webui/open-webui/pull/28837), [#28832](https://github.com/open-webui/open-webui/issues/28832)
- 📉 **Usage figures drifting upward on clustered setups.** The routine that clears out finished model usage no longer stops for good across the whole cluster after a brief interruption, which had left the usage figures counting models nobody was using and grew the work every disconnection had to do. [#28834](https://github.com/open-webui/open-webui/pull/28834)
- 🔇 **Voice mode staying silent with reasoning models.** Voice mode now speaks when the emoji option is on and the model behind it reports its answer as thinking rather than text, where the whole reply went unspoken and nothing reached the speech service at all. [#28724](https://github.com/open-webui/open-webui/pull/28724)
- 🏷 **Tags on a chat shared with you.** Opening a chat shared with you, or one in a shared folder, no longer fails to load its tags, and an administrator opening someone else's chat now sees the tags that chat actually carries. [Commit](https://github.com/open-webui/open-webui/commit/7d4747dfd73d7629227b10ae63c6854cc7543bee), [#28767](https://github.com/open-webui/open-webui/issues/28767)
- 🔻 **Message box controls in a narrow panel.** Narrowing the note chat panel, or squeezing the chat with a wide controls pane, no longer hides the attach and integrations buttons behind the model name or pushes the send button outside the box; the model name is shortened to make room instead. [#28912](https://github.com/open-webui/open-webui/pull/28912), [#28911](https://github.com/open-webui/open-webui/issues/28911)
- ⌛ **Replies that are all thinking and no answer.** A reply from a provider using the responses format that ends while the model is still in its reasoning, with no answer text after it, now finishes normally instead of failing the whole turn and leaving an unreadable error in place of the reply. [#28872](https://github.com/open-webui/open-webui/pull/28872), [#28871](https://github.com/open-webui/open-webui/issues/28871)
- 📃 **Word and PowerPoint previews overflowing.** Previewing one of these files now keeps the document inside its frame, with the zoom and slide controls staying put rather than scrolling away, and a presentation opens on its current slide instead of below the visible area. [#28878](https://github.com/open-webui/open-webui/pull/28878), [#28877](https://github.com/open-webui/open-webui/issues/28877)
- 🎞 **Workspace models in the admin models list.** Workspace models appear in the admin models list again, so they can be ordered, set as the default, and pinned for everyone; choosing one opens its own editor, where it opened the base model editor and could strip the model's base model, turning it into something else. [Commit](https://github.com/open-webui/open-webui/commit/ccbb3303f2ec5db0b16573bbb665a12f6764da66), [#27702](https://github.com/open-webui/open-webui/issues/27702)
- ⌨ **Errors after sending a long message.** With prompt autocompletion on, sending or clearing a message of several paragraphs within a second of typing no longer throws an error in the browser console. [#28824](https://github.com/open-webui/open-webui/pull/28824), [#28823](https://github.com/open-webui/open-webui/issues/28823)
- ❌ **Tool calls that failed looking successful.** A tool call that returned an error is now marked as failed with a red cross rather than a green tick, so a reply built on a failed call is easier to spot. [Commit](https://github.com/open-webui/open-webui/commit/f3f76095d18e07a3a944f22a4b25993bbffe90a3), [#28016](https://github.com/open-webui/open-webui/issues/28016)
- 🗜 **Download links in a cited source doing nothing.** A link in a citation shown as formatted content now downloads the file when clicked, where it silently did nothing at all. [Commit](https://github.com/open-webui/open-webui/commit/3c66d639e31ba8a7477d337b37d8d671ea2430cf), [#28924](https://github.com/open-webui/open-webui/issues/28924)
- 🗳 **Web searches failing without saying why.** A web search that fails now explains itself instead of returning nothing at all, which most often happens when a search engine has been selected without its key being configured. [#28942](https://github.com/open-webui/open-webui/pull/28942)
- ✅ **Checklists in notes.** A checklist in a note now previews and downloads as a proper checklist, where each item carried a stray second pair of brackets and its text began two lines below the box. [#27671](https://github.com/open-webui/open-webui/pull/27671), [#26067](https://github.com/open-webui/open-webui/issues/26067)
- 🧿 **Shortening a conversation with the wrong model.** Choosing to shorten long conversations with the model you are chatting with now does that, where it used the configured task model instead on any instance that has one. [Commit](https://github.com/open-webui/open-webui/commit/5093a9938937153e287db27a671f5ba1fb5d7592), [#27603](https://github.com/open-webui/open-webui/issues/27603)
- 🖥 **Stopping a reply after the shared cache restarts.** Stopping a reply now keeps working across a cluster after the shared cache restarts or its connection drops, where the part that carries a stop between instances gave up for good and silently, and only restarting the application brought it back. [Commit](https://github.com/open-webui/open-webui/commit/bf3a58dbcd18ddc2c7f130d8f9529477fe7cb042), [#28909](https://github.com/open-webui/open-webui/issues/28909)
- 📼 **Attached links to media and archives.** Attaching a link that leads to something other than a web page, such as a video or an archive, now reads it as the file it is rather than trying to treat it as text. [Commit](https://github.com/open-webui/open-webui/commit/886248de36e3a60c3687d1bee6af110e14110eca)
- 🗒 **Editing a workflow from settings.** Opening the code editor for a ComfyUI workflow from the images settings now brings it to the front, where it opened behind the settings dialog and could not be reached at all. [#27648](https://github.com/open-webui/open-webui/pull/27648), [#27647](https://github.com/open-webui/open-webui/issues/27647)
- 🖼 **Downloading a generated image.** Downloading an image from its preview now saves the image, where it could silently save a small file containing an authentication error instead, and a download that does fail now says so. [Commit](https://github.com/open-webui/open-webui/commit/2578174637e48cafa4bcb09adbb1b7f4b545a8d5), [#27723](https://github.com/open-webui/open-webui/pull/27723), [#27722](https://github.com/open-webui/open-webui/issues/27722)
- 🗂 **Directory sync listing local accounts.** A directory service syncing accounts over SCIM now sees only the accounts that came from a directory, where it also listed and could modify accounts created with a password in Open WebUI itself. [Commit](https://github.com/open-webui/open-webui/commit/fb4f476316a2f83e4d2914d535c19b3440cbc490)
- ↕ **Sorting a list of people.** Sorting the admin user list, or a channel's member list, now works when no search term has been entered, where the chosen order was ignored unless something was being searched for. [Commit](https://github.com/open-webui/open-webui/commit/fb4f476316a2f83e4d2914d535c19b3440cbc490)
- 🧶 **Text dropped from a reply by a filter.** A filter that rewrites a reply as it streams, or a provider that sends something other than plain text in a chunk, no longer causes that part of the reply to vanish without explanation. [#28840](https://github.com/open-webui/open-webui/pull/28840)
- 🏗 **Timers firing twice after a fork.** Branching a conversation that has a timer set no longer leaves the copy able to fire that timer as well. [#27663](https://github.com/open-webui/open-webui/pull/27663), [#27622](https://github.com/open-webui/open-webui/issues/27622), [#27745](https://github.com/open-webui/open-webui/issues/27745)
- 🗣 **Sentences skipped in voice mode.** Voice mode now speaks every sentence of a reply, where any sentence that completed in the same piece of the reply as another was silently never read out, which happened routinely with providers that send whole paragraphs at a time. [Commit](https://github.com/open-webui/open-webui/commit/495296346edfcd96a506a22fed4bb5b8faf3cd4d), [#28730](https://github.com/open-webui/open-webui/issues/28730), [#19861](https://github.com/open-webui/open-webui/issues/19861)
- 🖊 **Dragging a side panel wider than the window.** A side panel can no longer be dragged so wide that the chat beside it is squeezed away, and dragging one below its minimum width now closes it rather than sticking. [Commit](https://github.com/open-webui/open-webui/commit/ef455fcef9d6cb1136275c26e49bc3c5d6661795), [#28965](https://github.com/open-webui/open-webui/issues/28965)
- 🎙 **The wrong model selected when reopening a chat.** Reopening a conversation where a reply was regenerated with a different model now selects the model behind the reply you are looking at, where it picked the one used for the first attempt, which might be a model no longer available. [#27674](https://github.com/open-webui/open-webui/pull/27674), [#25052](https://github.com/open-webui/open-webui/issues/25052), [Commit](https://github.com/open-webui/open-webui/commit/8a170897bad569d93e069226a996942583ebde80)
- 🛜 **Coding tools that speak Anthropic's format.** A tool such as Cline pointed at Open WebUI using Anthropic's own message format can reach models again, where every request failed before it was even sent, and once that was corrected the request was rejected by Anthropic for being signed the wrong way. [#27675](https://github.com/open-webui/open-webui/pull/27675), [#27595](https://github.com/open-webui/open-webui/issues/27595), [#27695](https://github.com/open-webui/open-webui/issues/27695)
- 👯 **Code in pinned messages.** Opening the pinned messages of a channel now shows the code in those messages, where each block appeared empty and its contents were drawn into the channel behind the dialog instead, doubling them there. [#27740](https://github.com/open-webui/open-webui/pull/27740), [#27739](https://github.com/open-webui/open-webui/issues/27739)
- 🧻 **Logs flooded by an unreachable server.** A terminal or tool server that cannot be reached now records one line per attempt rather than a full stack trace, where a few minutes of downtime could fill the log with hundreds of megabytes and drown out everything else. [#27755](https://github.com/open-webui/open-webui/pull/27755), [#27751](https://github.com/open-webui/open-webui/issues/27751), [#27757](https://github.com/open-webui/open-webui/pull/27757), [#27756](https://github.com/open-webui/open-webui/issues/27756)
- 🖱 **Clicking inside a chat preview.** Clicking an image or a source in the preview that appears when you hover a chat in the sidebar no longer flashes a viewer open and shut, since the preview is meant only to be read. [#27770](https://github.com/open-webui/open-webui/pull/27770), [#27769](https://github.com/open-webui/open-webui/issues/27769)
- 🛢 **Using pgvector with database access by role.** An instance on Amazon RDS that signs in to its database with a temporary credential rather than a stored password now starts when pgvector is the vector store, where the two could not be used together and the container exited on startup. [#27754](https://github.com/open-webui/open-webui/pull/27754), [#27752](https://github.com/open-webui/open-webui/issues/27752)
- 📗 **Opening knowledge attached to a model or folder.** Clicking a knowledge item attached to a model or a folder opens it again, so it can be read and its retrieval mode changed between focused retrieval and the whole document, where since 0.11.0 neither was possible outside a chat. [#27686](https://github.com/open-webui/open-webui/pull/27686), [#27684](https://github.com/open-webui/open-webui/issues/27684), [#27801](https://github.com/open-webui/open-webui/issues/27801), [#28825](https://github.com/open-webui/open-webui/issues/28825)
- 🎚 **Retrieval mode shown from another item.** The retrieval mode shown when opening a knowledge item is now that item's own, where it could show the setting of whichever item was opened before it. [#27686](https://github.com/open-webui/open-webui/pull/27686), [#27684](https://github.com/open-webui/open-webui/issues/27684), [#27801](https://github.com/open-webui/open-webui/issues/27801), [#28825](https://github.com/open-webui/open-webui/issues/28825)
- 🔩 **A new chat shown as nearly full.** The indicator of how full a conversation is now counts tokens the same way the shortening does, where the two read different figures from providers that report both and a fresh chat could appear close to its limit. [#27620](https://github.com/open-webui/open-webui/pull/27620), [#27608](https://github.com/open-webui/open-webui/issues/27608), [Commit](https://github.com/open-webui/open-webui/commit/978d2572140e4fe31ebbc33274f699c71b8cdc29)
- 🫱 **Abandoned changes to a group's sharing setting.** Closing the edit dialog for a user group without saving now discards a change to who can share to that group, where the change stayed on screen and was written to the database the next time anything else about the group was saved. [#28076](https://github.com/open-webui/open-webui/pull/28076), [#28075](https://github.com/open-webui/open-webui/issues/28075)
- 📛 **Editing the wrong group.** The dialog for editing a user group now stays with the group it was opened for, where a reordering of the list beneath it could leave it saving to a different group. [#28076](https://github.com/open-webui/open-webui/pull/28076), [#28075](https://github.com/open-webui/open-webui/issues/28075)
- 🪢 **Signing people out from the identity provider.** A sign-out sent by an identity provider to end someone's session now works, where the check that the message was genuine could fail against providers whose signing keys need the same authentication as everything else, leaving the person signed in. [Commit](https://github.com/open-webui/open-webui/commit/aeda6ff13a25d3b3ba1b303609f35382db22142c)
- 🪣 **Files left behind when a knowledge base is emptied.** Emptying a knowledge base now removes the files it held, along with their stored copies and their search data, where all three were left behind with nothing in the interface to clear them. [Commit](https://github.com/open-webui/open-webui/commit/363ad352fec9553469852d111bc0506b896504a6), [#27988](https://github.com/open-webui/open-webui/issues/27988)
- ♻ **Blank errors when a tool server's saved sign-in cannot be read.** A tool server whose stored sign-in details cannot be decrypted, which happens when "WEBUI_SECRET_KEY" changes since the key protecting them follows it, now names the server and says to reconnect it, where every startup logged two errors with no message at all. [Commit](https://github.com/open-webui/open-webui/commit/91917b23952af8ca4457f8bb3db70c75ab838fbf), [#28666](https://github.com/open-webui/open-webui/pull/28666), [#28665](https://github.com/open-webui/open-webui/issues/28665)
- 🖍 **Clearing the supported media types.** Emptying the supported media types in the documents settings now stays empty, where the previous value came back on the next visit, so images kept being sent to the extraction engine instead of straight to a model that can read them. [#28750](https://github.com/open-webui/open-webui/pull/28750), [#28747](https://github.com/open-webui/open-webui/issues/28747), [Commit](https://github.com/open-webui/open-webui/commit/ecad20b77f9bd700bd818a2f447176f5994320c3)
- 🫳 **Dropping a chat where it already was.** Dragging a chat in the sidebar and releasing it where it started no longer reloads the whole sidebar, which took around a dozen requests for a move that changed nothing. [#28664](https://github.com/open-webui/open-webui/pull/28664), [#28663](https://github.com/open-webui/open-webui/issues/28663)
- 🖲 **Filtering while on a later page.** Changing a filter in the workspace, such as showing only what you created, now returns to the first page, where the list could come back empty with the page controls gone and no way back. [#28734](https://github.com/open-webui/open-webui/issues/28734), [Commit](https://github.com/open-webui/open-webui/commit/a914868e3c83b08be1399dbae1e4da25a0c00c87)
- 🪺 **Workspace counts left behind.** The number beside each workspace tab now follows its list, where creating, copying, importing, or deleting something left the old number in place until you moved to another tab or reloaded, and the tools count ignored its search entirely. [#28983](https://github.com/open-webui/open-webui/pull/28983), [#28981](https://github.com/open-webui/open-webui/issues/28981)
- 📞 **Links that start a voice call.** Opening a link that starts a call now starts one, where it opened the controls pane and stopped there, leaving the only way to begin a call from outside the application broken. [#28721](https://github.com/open-webui/open-webui/pull/28721), [#28677](https://github.com/open-webui/open-webui/issues/28677), [Commit](https://github.com/open-webui/open-webui/commit/8c1f3d382470edb81fa1efdf6cf4fb6bb188c460)
- 🆔 **Signing in where the provider uses a numeric account id.** Signing in through GitHub, or any other provider that identifies people by a number, works again on PostgreSQL, where every attempt failed outright since 0.10.2, and on SQLite the existing account was not matched; accounts stored the old way are corrected on the next sign-in. [Commit](https://github.com/open-webui/open-webui/commit/a6834f089bc2980fced99a617670e785c988cc9f), [#28954](https://github.com/open-webui/open-webui/pull/28954), [#27760](https://github.com/open-webui/open-webui/issues/27760)
- ☎ **Chats opening halfway up.** A conversation whose recent messages are short now opens at the latest message rather than somewhere in the middle, and older messages loaded while scrolling up no longer shift what you were reading. [#28657](https://github.com/open-webui/open-webui/pull/28657), [#28656](https://github.com/open-webui/open-webui/issues/28656)
- 🛤 **The terminal picker vanishing mid-reply.** The terminal picker now stays in place while a reply is being written, greyed out until it finishes, where it disappeared from the message box entirely and took the name of the selected terminal with it. [Commit](https://github.com/open-webui/open-webui/commit/7a533d0d5b85981f8c5668636f931f8b4a97d604)
- 🎧 **Who a channel says its members are.** The member list of a channel now includes its owner and everyone in a group that was granted access, where the owner never appeared and granting both a person and a group they were not in listed nobody at all while the count beside it said two. [#28289](https://github.com/open-webui/open-webui/pull/28289), [#28288](https://github.com/open-webui/open-webui/issues/28288), [Commit](https://github.com/open-webui/open-webui/commit/e3e82b14714661ba5bd24e2c47e42cef5594b132)
- ⛳ **Embedding servers behind a password.** An embedding server protected by a username and password rather than a key now works, where an empty key still sent an authorisation header, which such a server rejected and which made document uploads fail. [#28684](https://github.com/open-webui/open-webui/pull/28684), [#28683](https://github.com/open-webui/open-webui/issues/28683), [Commit](https://github.com/open-webui/open-webui/commit/97466deea105d6fdde60bf4d9401703fd32e03e9)
- ✒ **Note edits lost without warning.** Typing in a note now reaches the database, where a save already waiting could be cancelled without a replacement, leaving the editor showing text that was never stored and nothing to say so; this affected notes written by a model or through the API, and any instance without Redis after a restart. [#28669](https://github.com/open-webui/open-webui/pull/28669), [#28667](https://github.com/open-webui/open-webui/issues/28667)
- 🗄 **Browsing outside the home directory in a terminal.** A terminal server set to allow browsing the whole filesystem can be browsed above the home directory again, where the file panel pinned itself there and opening a file elsewhere quietly did nothing. [#29006](https://github.com/open-webui/open-webui/pull/29006), [#29000](https://github.com/open-webui/open-webui/issues/29000)
- 🏞 **Clearing a model's picture.** A model's picture can be reset to the default logo again, where since 0.11.0 a custom one could only ever be replaced. [#29007](https://github.com/open-webui/open-webui/pull/29007), [#27685](https://github.com/open-webui/open-webui/issues/27685), [Commit](https://github.com/open-webui/open-webui/commit/06e7aac219b3027a9d578624c473b041ad26a692)
- 👤 **Fallback profile pictures.** A profile picture that fails to load, such as one belonging to a deleted account, now falls back to the default avatar instead of showing clipped placeholder text beside the message. [#28270](https://github.com/open-webui/open-webui/pull/28270), [#28269](https://github.com/open-webui/open-webui/issues/28269)
- ⏱️ **Unanswered prompts in tools.** On deployments that set "WEBSOCKET_EVENT_CALLER_TIMEOUT", a question a tool asks you that goes unanswered now reports a timeout rather than an empty reply, and waiting too long no longer risks disconnecting a tab that is still open. [#28311](https://github.com/open-webui/open-webui/pull/28311)
- 👪 **Group member counts updating.** Adding or removing someone from a group in the admin panel now updates that group's member count straight away, where it stayed at the old number until the page was reloaded. [Commit](https://github.com/open-webui/open-webui/commit/18bf0ade7b35fb25ff7920ad3713624d93f04401)
- 🕳️ **Blank messages left in a conversation.** Streamed events that never fill in an item no longer leave an empty assistant message saved in the conversation and sent back to the model on every later turn. [Commit](https://github.com/open-webui/open-webui/commit/2a0274a0a039dbe0a1ad4d24003b085aae7b896b)
- 📓 **Notes opening blank.** A note whose shared editing session has not been started yet now opens with its stored content even when several people open it at once, where previously anyone but a lone first viewer got an empty document. [Commit](https://github.com/open-webui/open-webui/commit/5078d987f83943671f1d23ebe66bcc8d2af902c3)
- 🤫 **Replies stopping silently after a tool ran.** When a provider rejects the follow-up request made after a tool finishes, the reason is now shown in the chat instead of the reply simply ending with nothing said. [Commit](https://github.com/open-webui/open-webui/commit/a610d77137fabf60a2e0daa961a8ef8f7320293a), [#28633](https://github.com/open-webui/open-webui/issues/28633)
- 🎟️ **Connections whose tags were saved as plain text.** A connection with tags stored as plain text no longer breaks its editor panel or silently blanks the tags on every model coming from it. [Commit](https://github.com/open-webui/open-webui/commit/8be4c5fa6a849a9ff4a74f6a2870a115a836e68a), [#28749](https://github.com/open-webui/open-webui/issues/28749)
- 🔌 **Stream filters on direct API calls.** A filter reading a streamed event as an object now works on requests made straight to the chat completions endpoint, matching every other path, where it used to raise and end the reply partway. [Commit](https://github.com/open-webui/open-webui/commit/684111715f742f56a3c7efac9eb7a1e68e54a3ec)
- 🔇 **Filter failures that said nothing.** When a filter's outlet or stream hook raises, the failure is now reported with the filter's name and a traceback at the default log level, where it was swallowed and left plugin authors with nothing to go on. [Commit](https://github.com/open-webui/open-webui/commit/35fbde0a3fb303b7076b03465b1a2b0ba832df53)
- 🪜 **Falling back when a base model is gone.** Chatting with a workspace model whose base model has been removed now falls back to the default model for everyone, where the fallback previously applied only for administrators and anyone else was refused whenever that base model had no workspace entry of its own, such as one supplied by a pipe. [Commit](https://github.com/open-webui/open-webui/commit/20fe43d9da621957c48fc92104bc8f1cc0d691b7)
- 🧠 **Conversations breaking after a model switch.** Switching to a different model once a reasoning model has answered no longer leaves the conversation unusable, where the earlier model's stored reasoning was replayed to a provider that then rejected every later request. [Commit](https://github.com/open-webui/open-webui/commit/c4b3e6840f34a5d267bc12c0ec1a82156a324d45), [#28240](https://github.com/open-webui/open-webui/issues/28240)
- 👻 **Messages disappearing from a conversation.** Two things saving the same conversation at once no longer discard each other's changes, where a message could vanish from the screen while still counting toward the model's context and a file attached during a reply was lost on the next save. [Commit](https://github.com/open-webui/open-webui/commit/1c13fedb16e74c5888c52ccd928a5eb5bbb1068d), [#28742](https://github.com/open-webui/open-webui/issues/28742)
- 🧬 **Workspace models pointing at themselves.** A workspace model whose base model was set to its own identifier is now saved without that reference, where the entry was thrown away while models were being combined so none of its settings ever took effect. [Commit](https://github.com/open-webui/open-webui/commit/eadce55e343df69520e6bc86dbbdfdcdfdcafb30), [#28952](https://github.com/open-webui/open-webui/issues/28952), [#28923](https://github.com/open-webui/open-webui/issues/28923)
- 📭 **A short model list sticking on every worker.** When one worker briefly reports fewer models than it should, that shorter list is no longer left in place for the whole deployment, where chatting with one of the missing models failed on every server until something restarted. [Commit](https://github.com/open-webui/open-webui/commit/6330350a406d8c1fd603f725e59a88324c8e2256), [#28777](https://github.com/open-webui/open-webui/issues/28777)
- 📋 **Copying from the terminal file browser.** Copying a file path or a file's contents now works again, where it silently did nothing on deployments the browser does not treat as a secure origin. [Commit](https://github.com/open-webui/open-webui/commit/aa3d56961049618b28717d87acff1d9360cad47a), [Commit](https://github.com/open-webui/open-webui/commit/2f97c9fce36b3794d14c925086eca78ff92adf4a), [#29015](https://github.com/open-webui/open-webui/issues/29015)
- 🗺️ **Browsing the web through Microsoft Web IQ.** Fetching a page with the web loader set to Microsoft Web IQ now works, where every attempt failed before a single request was made and had done so ever since that loader was added. [Commit](https://github.com/open-webui/open-webui/commit/6dcc2d52692c3fa1993ff1b3ae84c07d35fd9f6b), [Commit](https://github.com/open-webui/open-webui/commit/140d2cf4b59e71d2e9f4986d4d8649fccb47c83d), [#28688](https://github.com/open-webui/open-webui/issues/28688)
- 🔛 **Enable or disable all automations only reaching the ones on screen.** Turning every automation on or off now covers every automation matching your current search and filter, where it only ever touched the ones loaded on the page you were looking at and left the rest running as they were. [Commit](https://github.com/open-webui/open-webui/commit/f4a0d3c9734d1662a3f78891f21934f2b82aed1e)
### Changed
- ⚠️ **Database Migrations**: This release includes database schema changes; we strongly recommend backing up your database and all associated data before upgrading in production environments. If you are running a multi-worker, multi-server, or load-balanced deployment, all instances must be updated simultaneously, rolling updates are not supported and will cause application failures due to schema incompatibility.
- 🏋️ **What "THREAD_POOL_SIZE" now sizes.** The setting now governs both of the pools that background work runs in, where it previously governed only one and the other, carrying most of the blocking work including knowledge searches, sign-ins and file storage, was fixed at a small ceiling no setting could raise, so an instance that set it high will now use more threads than before, up to twice the configured value across the two pools, and one that relied on the old ceiling to hold thread use down should review it. [Commit](https://github.com/open-webui/open-webui/commit/4ec6ee14418edd04eaba9e34bd5868453f61df40), [#28168](https://github.com/open-webui/open-webui/issues/28168)
- 💾 **Saving replies as they stream.** The "ENABLE_REALTIME_CHAT_SAVE" setting no longer has any effect, because a reply in progress is now held outside the database and written once when it finishes. [Commit](https://github.com/open-webui/open-webui/commit/a1579a01ff43cacb357269707d36267ad35e01d6)
- 🎭 **Playwright web loader egress.** When a page is fetched with the Playwright loader, the page's own requests for its images, scripts, and stylesheets are now made by the Open WebUI backend instead of by the browser, so administrators using a remote browser through "PLAYWRIGHT_WS_URL" should expect that traffic to leave from the backend's address rather than the browser host, those using a private certificate authority should expect it to be trusted through "AIOHTTP_CLIENT_SSL_CERT_FILE" rather than the browser's own store, and those who set a proxy on the loader should know it no longer applies to these requests, which follow the environment's proxy settings instead. [#28634](https://github.com/open-webui/open-webui/pull/28634)
- 🐌 **Slower Playwright page loads without async.** A page fetched with the Playwright loader on the synchronous path now fetches its images, scripts, and stylesheets one at a time rather than together, which in the change's own measurements took a page with thirty assets from 2.0 to 3.0 seconds, and one with eight slow assets from 1.1 to 4.5 seconds; the asynchronous path is unaffected. [#28634](https://github.com/open-webui/open-webui/pull/28634)
- 🐳 **Test-only packages removed from the image.** The container no longer ships pytest, pytest-docker, the Docker SDK, or netcat, none of which anything in Open WebUI used, so the image is smaller; anyone whose own tools or functions relied on those being present will need to install them themselves. [#28726](https://github.com/open-webui/open-webui/pull/28726)
- 🛡 **Forms in embedded pages now work by default.** A page shown inside a chat, such as an artifact or an HTML preview, may now submit forms unless you turn that off, where it was blocked unless you turned it on. [Commit](https://github.com/open-webui/open-webui/commit/3c66d639e31ba8a7477d337b37d8d671ea2430cf)
- 🅰 **Connection prefixes now show in model names.** A connection's prefix appears in the name shown in the model picker, not only in its identifier, where whether it did depended on the shape of the provider's reply and so worked on some connections and not others; models on prefixed connections will now read differently than before. [#28950](https://github.com/open-webui/open-webui/pull/28950), [#28929](https://github.com/open-webui/open-webui/issues/28929)
- 📊 **Default usage statistics range.** Usage statistics now cover the past two years by default for everyone, instead of starting from the date the account was created. [Commit](https://github.com/open-webui/open-webui/commit/8dbbc206c5a0706789722c42827479e5db10bb2b)
## [0.11.0] - 2026-07-27 ## [0.11.0] - 2026-07-27
### Added ### Added

View file

@ -20,7 +20,7 @@ Examples of behavior that contribute to a positive and professional community in
- **Respecting others.** Be considerate, listen actively, and engage with empathy toward others' viewpoints and experiences. - **Respecting others.** Be considerate, listen actively, and engage with empathy toward others' viewpoints and experiences.
- **Constructive feedback.** Provide actionable, thoughtful, and respectful feedback that helps improve the project and encourages collaboration. Avoid unproductive negativity or hypercriticism. - **Constructive feedback.** Provide actionable, thoughtful, and respectful feedback that helps improve the project and encourages collaboration. Avoid unproductive negativity or hypercriticism.
- **Recognizing volunteer contributions.** Appreciate that contributors dedicate their free time and resources selflessly. Approach them with gratitude and patience. - **Recognizing volunteer contributions.** Appreciate that **contributors dedicate their free time and resources selflessly**. Approach them with gratitude and patience.
- **Focusing on shared goals.** Collaborate in ways that prioritize the health, success, and sustainability of the community over individual agendas. - **Focusing on shared goals.** Collaborate in ways that prioritize the health, success, and sustainability of the community over individual agendas.
Examples of unacceptable behavior include: Examples of unacceptable behavior include:
@ -32,11 +32,23 @@ Examples of unacceptable behavior include:
- **Entitlement, demand, or aggression toward contributors.** Volunteers are under no obligation to provide immediate or personalized support. Rude or dismissive behavior will not be tolerated. - **Entitlement, demand, or aggression toward contributors.** Volunteers are under no obligation to provide immediate or personalized support. Rude or dismissive behavior will not be tolerated.
- **Unproductive or destructive behavior.** This includes venting frustration as hostility ("tantrums"), hypercriticism, attention-seeking negativity, or anything that distracts from the project's goals. - **Unproductive or destructive behavior.** This includes venting frustration as hostility ("tantrums"), hypercriticism, attention-seeking negativity, or anything that distracts from the project's goals.
- **Spamming and promotional exploitation.** Sharing irrelevant product promotions or self-promotion in the community is not allowed unless it directly contributes value to the discussion. - **Spamming and promotional exploitation.** Sharing irrelevant product promotions or self-promotion in the community is not allowed unless it directly contributes value to the discussion.
- Posting low-effort, hard to read, essay-length AI generated comments or other forms of low-quality, hard to parse content that puts the burden of understanding on the reader.
### How We Develop the Project
Development is led by the maintainers, and code pull requests are reserved for work we explicitly request or exceptional contributions we choose to consider at our discretion. We use actionable reports and concrete use cases to understand problems, then evaluate, revise, and implement the appropriate approach internally. We assess each change against the project's architecture, existing behavior, quality standards, and future direction before settling on an implementation. Resolving a reported problem requires that broader context, and a working external patch usually requires substantial rewriting to meet the project's standards. Reviewing the patch, explaining the required changes, and coordinating successive revisions usually takes more effort than developing the solution internally. Fragmented commit histories, branches that have not been rebased, unresolved conflicts, and lengthy or unverified AI-generated comments add cleanup and discussion that delay the underlying work. Maintainers remain responsible for testing, documenting, supporting, and maintaining every accepted change, so we choose the approach based on the whole product and its ongoing maintenance. Clear reports, reproduction details, and relevant context give us what we need to make those decisions and develop the solution. A polished implementation, clean commit history, or completed checklist does not establish an exception to this process, and opening an issue or discussion is not an invitation to submit a PR. Wait for an explicit maintainer request before investing in a PR; unsolicited submissions are generally closed without review, and requested PRs remain subject to maintainer judgment.
### Feedback and Community Engagement ### Feedback and Community Engagement
- **Constructive feedback is encouraged, but hostile or entitled behavior will result in immediate action.** If you disagree with elements of the project, we encourage you to offer meaningful improvements or fork the project if necessary. Healthy discussions and technical disagreements are welcome only when handled with professionalism. Participation should help maintainers understand a concrete problem while respecting the project's priorities and available capacity. Please follow the [issue templates](.github/ISSUE_TEMPLATE) and [pull request policy](.github/pull_request_template.md) before submitting anything.
- **Respect contributors' time and efforts.** No one is entitled to personalized or on-demand assistance. This is a community built on collaboration and shared effort; demanding or demeaning behavior undermines that trust and will not be allowed.
- **Make reports actionable.** Search existing issues and discussions, check the latest version and whether the problem is already addressed on `dev`, and use the appropriate template. Bug reports should describe a reproducible problem, the affected workflow, expected and actual behavior, and relevant evidence. Feature requests should explain the user-facing need; broader product, UX, architecture, or maintenance questions belong in Discussions. Report security concerns privately through the [security reporting process](https://github.com/open-webui/open-webui/security).
- **Share the problem before investing in code.** Start with an actionable issue or discussion and leave implementation planning to the maintainers. An issue or discussion alone is not an invitation to submit a PR. Please wait for an explicit request before opening one; any exception is at the maintainers' discretion. Implementation notes, local diffs, or patches may be shared as reference in the relevant issue or discussion.
- **Respect maintainers' discretion.** Submitting an issue, proposal, or pull request does not create an obligation to respond, review, implement, or merge it. Maintainers set the project's direction and defer or close submissions based on scope, quality, maintenance cost, or available capacity. Unsolicited pull requests are generally closed without review.
- **Keep discussion focused and concise.** Provide new information when it helps evaluate the problem. Repeated bumps, duplicate submissions, unsolicited direct messages seeking attention, or pressure for timelines place an unnecessary burden on contributors.
- **Respect decisions and boundaries.** Technical disagreement is welcome when expressed professionally. Reopening a declined request or continuing to press for a different outcome without new, relevant information is not constructive. You are free to explore a different direction in your own fork.
Participants are expected to respect maintainers' decisions and the contribution process. Harassment, hostility, or repeated disregard for these boundaries will result in enforcement under this Code of Conduct.
### Zero Tolerance: No Warnings, Immediate Action ### Zero Tolerance: No Warnings, Immediate Action

View file

@ -26,6 +26,9 @@ ARG GID=0
######## WebUI frontend ######## ######## WebUI frontend ########
FROM --platform=$BUILDPLATFORM node:22-alpine3.20 AS build FROM --platform=$BUILDPLATFORM node:22-alpine3.20 AS build
ARG BUILD_HASH ARG BUILD_HASH
ARG USE_SLIM
ARG UID
ARG GID
# Set Node.js options (heap limit Allocation failed - JavaScript heap out of memory) # Set Node.js options (heap limit Allocation failed - JavaScript heap out of memory)
# ENV NODE_OPTIONS="--max-old-space-size=4096" # ENV NODE_OPTIONS="--max-old-space-size=4096"
@ -40,7 +43,14 @@ RUN npm ci --force
COPY . . COPY . .
ENV APP_BUILD_HASH=${BUILD_HASH} ENV APP_BUILD_HASH=${BUILD_HASH}
RUN npm run build RUN npm run build && \
if [ "$USE_SLIM" = "true" ]; then find build -type f -name '*.map' -delete; fi
# Prepare backend ownership before the final copy so static assets occupy one layer.
# Group 0 write access lets arbitrary OpenShift UIDs update these assets at startup.
RUN chown -R $UID:$GID /app/backend && \
chgrp -R 0 /app/backend/open_webui/static && \
chmod -R g=u /app/backend/open_webui/static
######## WebUI backend ######## ######## WebUI backend ########
FROM python:3.11-slim-bookworm AS base FROM python:3.11-slim-bookworm AS base
@ -123,24 +133,33 @@ RUN echo -n 00000000-0000-0000-0000-000000000000 > $HOME/.cache/chroma/telemetry
# Make sure the user has access to the app and root directory # Make sure the user has access to the app and root directory
RUN chown -R $UID:$GID /app $HOME RUN chown -R $UID:$GID /app $HOME
# Install common system dependencies # Slim cannot bundle a local model server or GPU runtime.
RUN if [ "$USE_SLIM" = "true" ] && { [ "$USE_CUDA" = "true" ] || [ "$USE_OLLAMA" = "true" ]; }; then \
echo "USE_SLIM cannot be combined with USE_CUDA or USE_OLLAMA" >&2; exit 1; fi
# Keep the slim runtime free of local document/audio processing tools.
# Git-based tool requirements require the standard image.
RUN apt-get update && \ RUN apt-get update && \
apt-get install -y --no-install-recommends \ apt-get install -y --no-install-recommends \
git build-essential pandoc gcc netcat-openbsd curl jq ca-certificates \ curl jq ca-certificates \
libmariadb-dev \ && if [ "$USE_SLIM" != "true" ]; then \
python3-dev \ apt-get install -y --no-install-recommends \
ffmpeg libsm6 libxext6 zstd \ git build-essential pandoc gcc libmariadb-dev ffmpeg libsm6 libxext6; \
&& rm -rf /var/lib/apt/lists/* fi && if [ "$USE_OLLAMA" = "true" ]; then \
apt-get install -y --no-install-recommends zstd; \
fi && rm -rf /var/lib/apt/lists/*
# install python dependencies # install python dependencies
COPY --chown=$UID:$GID ./backend/requirements.txt ./requirements.txt COPY --chown=$UID:$GID ./backend/requirements*.txt ./
# Set UV_LINK_MODE to copy to prevent 0-byte file corruption in QEMU arm64 cross-builds # Set UV_LINK_MODE to copy to prevent 0-byte file corruption in QEMU arm64 cross-builds
ENV UV_LINK_MODE=copy ENV UV_LINK_MODE=copy
RUN set -e; \ RUN --mount=from=ghcr.io/astral-sh/uv:0.12.10,source=/uv,target=/bin/uv \
pip3 install --no-cache-dir uv; \ set -e; \
if [ "$USE_CUDA" = "true" ]; then \ if [ "$USE_SLIM" = "true" ]; then \
uv pip install --system -r requirements-slim.txt --no-cache-dir; \
elif [ "$USE_CUDA" = "true" ]; then \
# If you use CUDA the whisper and embedding model will be downloaded on first use # If you use CUDA the whisper and embedding model will be downloaded on first use
# fix: pin torch<=2.9.1 - torch 2.10.0 aarch64 wheels cause SIGILL on ARM devices (RPi 4 Cortex-A72) #21349 # fix: pin torch<=2.9.1 - torch 2.10.0 aarch64 wheels cause SIGILL on ARM devices (RPi 4 Cortex-A72) #21349
pip3 install 'torch<=2.9.1' torchvision torchaudio --index-url https://download.pytorch.org/whl/$USE_CUDA_DOCKER_VER --no-cache-dir; \ pip3 install 'torch<=2.9.1' torchvision torchaudio --index-url https://download.pytorch.org/whl/$USE_CUDA_DOCKER_VER --no-cache-dir; \
@ -149,7 +168,6 @@ RUN set -e; \
python -c "import os; from sentence_transformers import SentenceTransformer; SentenceTransformer(os.environ.get('AUXILIARY_EMBEDDING_MODEL', 'TaylorAI/bge-micro-v2'), device='cpu')"; \ python -c "import os; from sentence_transformers import SentenceTransformer; SentenceTransformer(os.environ.get('AUXILIARY_EMBEDDING_MODEL', 'TaylorAI/bge-micro-v2'), device='cpu')"; \
python -c "import os; from faster_whisper import WhisperModel; WhisperModel(os.environ['WHISPER_MODEL'], device='cpu', compute_type='int8', download_root=os.environ['WHISPER_MODEL_DIR'])"; \ python -c "import os; from faster_whisper import WhisperModel; WhisperModel(os.environ['WHISPER_MODEL'], device='cpu', compute_type='int8', download_root=os.environ['WHISPER_MODEL_DIR'])"; \
python -c "import os; import tiktoken; tiktoken.get_encoding(os.environ['TIKTOKEN_ENCODING_NAME'])"; \ python -c "import os; import tiktoken; tiktoken.get_encoding(os.environ['TIKTOKEN_ENCODING_NAME'])"; \
python -c "import nltk; nltk.download('punkt_tab')"; \
else \ else \
pip3 install 'torch<=2.9.1' torchvision torchaudio --index-url https://download.pytorch.org/whl/cpu --no-cache-dir; \ pip3 install 'torch<=2.9.1' torchvision torchaudio --index-url https://download.pytorch.org/whl/cpu --no-cache-dir; \
uv pip install --system -r requirements.txt --no-cache-dir; \ uv pip install --system -r requirements.txt --no-cache-dir; \
@ -158,12 +176,17 @@ RUN set -e; \
python -c "import os; from sentence_transformers import SentenceTransformer; SentenceTransformer(os.environ.get('AUXILIARY_EMBEDDING_MODEL', 'TaylorAI/bge-micro-v2'), device='cpu')"; \ python -c "import os; from sentence_transformers import SentenceTransformer; SentenceTransformer(os.environ.get('AUXILIARY_EMBEDDING_MODEL', 'TaylorAI/bge-micro-v2'), device='cpu')"; \
python -c "import os; from faster_whisper import WhisperModel; WhisperModel(os.environ['WHISPER_MODEL'], device='cpu', compute_type='int8', download_root=os.environ['WHISPER_MODEL_DIR'])"; \ python -c "import os; from faster_whisper import WhisperModel; WhisperModel(os.environ['WHISPER_MODEL'], device='cpu', compute_type='int8', download_root=os.environ['WHISPER_MODEL_DIR'])"; \
python -c "import os; import tiktoken; tiktoken.get_encoding(os.environ['TIKTOKEN_ENCODING_NAME'])"; \ python -c "import os; import tiktoken; tiktoken.get_encoding(os.environ['TIKTOKEN_ENCODING_NAME'])"; \
python -c "import nltk; nltk.download('punkt_tab')"; \
fi; \ fi; \
fi; \ fi; \
mkdir -p /app/backend/data; chown -R $UID:$GID /app/backend/data/; \ mkdir -p /app/backend/data; chown -R $UID:$GID /app/backend/data/; \
if [ -d /app/backend/data/cache ]; then chmod -R a+rX /app/backend/data/cache; fi; \
rm -rf /var/lib/apt/lists/*; rm -rf /var/lib/apt/lists/*;
# Optional: PPTX parsing through unstructured may need spaCy's English model.
# Keep this out of the default image to avoid the extra image bloat; deployments
# with read-only site-packages can uncomment it and bake the model in.
# RUN python -m spacy download en_core_web_sm
# Install Ollama if requested # Install Ollama if requested
RUN if [ "$USE_OLLAMA" = "true" ]; then \ RUN if [ "$USE_OLLAMA" = "true" ]; then \
date +%s > /tmp/ollama_build_hash && \ date +%s > /tmp/ollama_build_hash && \
@ -181,19 +204,8 @@ COPY --chown=$UID:$GID --from=build /app/build /app/build
COPY --chown=$UID:$GID --from=build /app/CHANGELOG.md /app/CHANGELOG.md COPY --chown=$UID:$GID --from=build /app/CHANGELOG.md /app/CHANGELOG.md
COPY --chown=$UID:$GID --from=build /app/package.json /app/package.json COPY --chown=$UID:$GID --from=build /app/package.json /app/package.json
# copy backend files # copy backend files with the ownership and static permissions prepared above
COPY --chown=$UID:$GID ./backend . COPY --from=build /app/backend .
# The backend rewrites its bundled static assets (favicons, splash, manifest,
# loader.js, ...) under open_webui/static at startup. Make that directory
# writable by an arbitrary UID -- which under OpenShift's restricted SCC is
# always a member of GID 0 -- so those writes don't fail with EACCES and crash
# the boot log with "[Errno 13] Permission denied". `chmod -R g=u` mirrors the
# owner bits onto the group (the Red Hat arbitrary-UID idiom). This is applied
# unconditionally because it targets a directory the app writes on every start;
# the broader, opt-in USE_PERMISSION_HARDENING below covers the rest of /app.
RUN chgrp -R 0 /app/backend/open_webui/static && \
chmod -R g=u /app/backend/open_webui/static
EXPOSE 8080 EXPOSE 8080

View file

@ -10,9 +10,7 @@
[![Discord](https://img.shields.io/badge/Discord-Open_WebUI-blue?logo=discord&logoColor=white)](https://discord.gg/5rJgQTnV4s) [![Discord](https://img.shields.io/badge/Discord-Open_WebUI-blue?logo=discord&logoColor=white)](https://discord.gg/5rJgQTnV4s)
[![](https://img.shields.io/static/v1?label=Sponsor&message=%E2%9D%A4&logo=GitHub&color=%23fe8e86)](https://github.com/sponsors/open-webui) [![](https://img.shields.io/static/v1?label=Sponsor&message=%E2%9D%A4&logo=GitHub&color=%23fe8e86)](https://github.com/sponsors/open-webui)
![Open WebUI Banner](./banner.png) Open WebUI is **a home for AI**, a self-hosted AI platform that's **[extensible](https://docs.openwebui.com/features/extensibility/plugin/)**, **[feature-rich](https://docs.openwebui.com/features/)**, user-friendly, and built to run **[entirely offline](https://openwebui.com/sovereign-ai)**. With support for **Ollama** and **OpenAI-compatible APIs**, it gives you a powerful, provider-agnostic interface for both local and cloud-based models.
**Open WebUI is an [extensible](https://docs.openwebui.com/features/extensibility/plugin), feature-rich, and user-friendly self-hosted AI platform designed to operate entirely offline.** It supports various LLM runners like **Ollama** and **OpenAI-compatible APIs**, with **built-in inference engine** for RAG, making it a **powerful AI deployment solution**.
Passionate about open-source AI? [Join our team →](https://careers.openwebui.com/) Passionate about open-source AI? [Join our team →](https://careers.openwebui.com/)
@ -20,8 +18,6 @@ Passionate about open-source AI? [Join our team →](https://careers.openwebui.c
> [!TIP] > [!TIP]
> **Looking for an [Enterprise Plan](https://docs.openwebui.com/enterprise)?** – **[Speak with Our Sales Team Today!](https://docs.openwebui.com/enterprise)** > **Looking for an [Enterprise Plan](https://docs.openwebui.com/enterprise)?** – **[Speak with Our Sales Team Today!](https://docs.openwebui.com/enterprise)**
>
> Get **enhanced capabilities**, including **custom theming and branding**, **Service Level Agreement (SLA) support**, **Long-Term Support (LTS) versions**, and **more!**
For more information, be sure to check out our [Open WebUI Documentation](https://docs.openwebui.com/). For more information, be sure to check out our [Open WebUI Documentation](https://docs.openwebui.com/).
@ -37,6 +33,8 @@ For more information, be sure to check out our [Open WebUI Documentation](https:
- 🤖 **Models & Agents**: Wrap any base model with custom instructions, tools, and knowledge to build specialized agents. Supports dynamic variables, per-user/group access control, and community preset imports via [Open WebUI Community](https://openwebui.com/). - 🤖 **Models & Agents**: Wrap any base model with custom instructions, tools, and knowledge to build specialized agents. Supports dynamic variables, per-user/group access control, and community preset imports via [Open WebUI Community](https://openwebui.com/).
- ⚡ **Agentic Execution with [Open Terminal](https://github.com/open-webui/open-terminal)**: Give your agents a terminal and filesystem to carry out multi-step tasks. Let them analyze data, run scripts, fix errors, and produce files directly in chat. Scale to teams with **[Terminals (Enterprise)](https://github.com/open-webui/terminals)** for per-user isolated environments, resource limits, and automatic lifecycle management.
- 📝 **Notes**: A dedicated workspace for content outside conversations. Draft with a rich editor, use AI to rewrite selected text, and attach notes to any chat for full-context injection. - 📝 **Notes**: A dedicated workspace for content outside conversations. Draft with a rich editor, use AI to rewrite selected text, and attach notes to any chat for full-context injection.
- 📢 **Channels**: Real-time shared spaces where your team and AI models collaborate in one timeline. Tag models to draft or critique, with threads, reactions, pins, and access control. - 📢 **Channels**: Real-time shared spaces where your team and AI models collaborate in one timeline. Tag models to draft or critique, with threads, reactions, pins, and access control.

View file

@ -1,3 +1,3 @@
export CORS_ALLOW_ORIGIN="http://localhost:5173;http://localhost:8080" export CORS_ALLOW_ORIGIN="http://localhost:5173;http://localhost:8080"
PORT="${PORT:-8080}" PORT="${PORT:-8080}"
uvicorn open_webui.main:app --port $PORT --host 0.0.0.0 --forwarded-allow-ips "${FORWARDED_ALLOW_IPS:-*}" --reload uvicorn open_webui.main:app --port $PORT --host 0.0.0.0 --forwarded-allow-ips "${FORWARDED_ALLOW_IPS:-*}" --ws-per-message-deflate "${UVICORN_WS_PER_MESSAGE_DEFLATE:-true}" --reload

View file

@ -18,6 +18,9 @@ def version_callback(value: bool) -> None:
if value: if value:
from open_webui.env import VERSION from open_webui.env import VERSION
# LICENSE covers this Open WebUI CLI identifier.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
typer.echo(f'Open WebUI version: {VERSION}') typer.echo(f'Open WebUI version: {VERSION}')
raise typer.Exit() raise typer.Exit()
@ -71,7 +74,7 @@ def serve(
os.environ['LD_LIBRARY_PATH'] = ':'.join(LD_LIBRARY_PATH) os.environ['LD_LIBRARY_PATH'] = ':'.join(LD_LIBRARY_PATH)
import open_webui.main # noqa: F401 import open_webui.main # noqa: F401
from open_webui.env import UVICORN_WORKERS # Import the workers setting from open_webui.env import UVICORN_WORKERS, UVICORN_WS_PER_MESSAGE_DEFLATE
# On Windows, uvicorn's default loop factory hardcodes ProactorEventLoop, # On Windows, uvicorn's default loop factory hardcodes ProactorEventLoop,
# which is incompatible with psycopg v3 async. Setting loop='none' lets # which is incompatible with psycopg v3 async. Setting loop='none' lets
@ -84,6 +87,7 @@ def serve(
port=port, port=port,
forwarded_allow_ips='*', forwarded_allow_ips='*',
workers=UVICORN_WORKERS, workers=UVICORN_WORKERS,
ws_per_message_deflate=UVICORN_WS_PER_MESSAGE_DEFLATE,
loop=loop, loop=loop,
) )
@ -94,12 +98,15 @@ def dev(
port: int = 8080, port: int = 8080,
reload: bool = True, reload: bool = True,
): ):
from open_webui.env import UVICORN_WS_PER_MESSAGE_DEFLATE
uvicorn.run( uvicorn.run(
'open_webui.main:app', 'open_webui.main:app',
host=host, host=host,
port=port, port=port,
reload=reload, reload=reload,
forwarded_allow_ips='*', forwarded_allow_ips='*',
ws_per_message_deflate=UVICORN_WS_PER_MESSAGE_DEFLATE,
) )

View file

@ -1,7 +1,6 @@
from __future__ import annotations from __future__ import annotations
import base64 import base64
import json
import logging import logging
import os import os
import shutil import shutil
@ -18,8 +17,10 @@ from authlib.integrations.starlette_client import OAuth
from pydantic import BaseModel from pydantic import BaseModel
from open_webui.env import ( from open_webui.env import (
USE_SLIM,
DATA_DIR, DATA_DIR,
DATABASE_URL, DATABASE_URL,
ENABLE_ADMIN_CHAT_ACCESS,
ENABLE_DB_MIGRATIONS, ENABLE_DB_MIGRATIONS,
ENV, ENV,
FRONTEND_BUILD_DIR, FRONTEND_BUILD_DIR,
@ -35,11 +36,12 @@ from open_webui.env import (
log, log,
) )
from open_webui.models.config import Config from open_webui.models.config import Config
from open_webui.utils.json_codec import JSONCodec
async def seed_registered_defaults(): async def seed_registered_defaults():
await Config.rename_prefix('rag.web', 'web') await Config.rename_prefix('rag.web', 'web')
await Config.repair_flattened_dict_configs() await Config.repair_config_rows()
await Config.seed_defaults(DEFAULT_CONFIG) await Config.seed_defaults(DEFAULT_CONFIG)
@ -73,6 +75,7 @@ def run_migrations():
command.upgrade(alembic_cfg, 'head') command.upgrade(alembic_cfg, 'head')
except Exception as e: except Exception as e:
log.exception(f'Error running migrations: {e}') log.exception(f'Error running migrations: {e}')
raise
if ENABLE_DB_MIGRATIONS: if ENABLE_DB_MIGRATIONS:
@ -84,7 +87,7 @@ async def import_legacy_config_json():
if not os.path.exists(f'{DATA_DIR}/config.json'): if not os.path.exists(f'{DATA_DIR}/config.json'):
return return
with open(f'{DATA_DIR}/config.json', 'r') as _f: with open(f'{DATA_DIR}/config.json', 'r') as _f:
await Config.upsert(json.load(_f)) await Config.upsert(JSONCodec.loads(_f.read()))
os.rename(f'{DATA_DIR}/config.json', f'{DATA_DIR}/old_config.json') os.rename(f'{DATA_DIR}/config.json', f'{DATA_DIR}/old_config.json')
@ -114,6 +117,9 @@ for file_path in (FRONTEND_BUILD_DIR / 'static').glob('**/*'):
except Exception as e: except Exception as e:
logging.error(f'An error occurred: {e}') logging.error(f'An error occurred: {e}')
# LICENSE covers copied Open WebUI logo/favicon assets.
# Do not alter, remove, obscure, or replace them except as LICENSE permits:
# https://docs.openwebui.com/license.
frontend_favicon = FRONTEND_BUILD_DIR / 'static' / 'favicon.png' frontend_favicon = FRONTEND_BUILD_DIR / 'static' / 'favicon.png'
if frontend_favicon.exists(): if frontend_favicon.exists():
@ -181,6 +187,9 @@ CACHE_DIR.mkdir(parents=True, exist_ok=True)
# CUSTOM_NAME (Legacy) # CUSTOM_NAME (Legacy)
#################################### ####################################
# LICENSE covers this legacy Open WebUI branding path.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
CUSTOM_NAME = os.getenv('CUSTOM_NAME', '') CUSTOM_NAME = os.getenv('CUSTOM_NAME', '')
if CUSTOM_NAME: if CUSTOM_NAME:
@ -219,6 +228,7 @@ if CUSTOM_NAME:
#################################### ####################################
ENABLE_DIRECT_CONNECTIONS = os.getenv('ENABLE_DIRECT_CONNECTIONS', 'False').lower() == 'true' ENABLE_DIRECT_CONNECTIONS = os.getenv('ENABLE_DIRECT_CONNECTIONS', 'False').lower() == 'true'
ENABLE_DIRECT_INTEGRATIONS = os.getenv('ENABLE_DIRECT_INTEGRATIONS', 'False').lower() == 'true'
#################################### ####################################
# OLLAMA_BASE_URL # OLLAMA_BASE_URL
@ -270,9 +280,9 @@ def _resolve_ollama_base_url(url: str) -> str:
if not default.result() and fallback.result(): if not default.result() and fallback.result():
url = url.replace(':11434', ':12434') url = url.replace(':11434', ':12434')
log.info(f'Ollama port 11434 unreachable on {host}, falling back to 12434') log.info('Ollama port 11434 unreachable on %s, falling back to 12434', host)
elif not default.result(): elif not default.result():
log.info(f'Ollama ports 11434 and 12434 both unreachable on {host}') log.info('Ollama ports 11434 and 12434 both unreachable on %s', host)
return url return url
@ -293,12 +303,12 @@ OLLAMA_API_CONFIGS = {}
_ollama_api_configs = os.getenv('OLLAMA_API_CONFIGS', '') _ollama_api_configs = os.getenv('OLLAMA_API_CONFIGS', '')
if _ollama_api_configs: if _ollama_api_configs:
try: try:
parsed = json.loads(_ollama_api_configs) parsed = JSONCodec.loads(_ollama_api_configs)
if isinstance(parsed, dict): if isinstance(parsed, dict):
OLLAMA_API_CONFIGS = parsed OLLAMA_API_CONFIGS = parsed
else: else:
log.warning('OLLAMA_API_CONFIGS must be a JSON object, ignoring') log.warning('OLLAMA_API_CONFIGS must be a JSON object, ignoring')
except (json.JSONDecodeError, TypeError): except (JSONCodec.JSONDecodeError, TypeError):
log.warning('OLLAMA_API_CONFIGS is not valid JSON, ignoring') log.warning('OLLAMA_API_CONFIGS is not valid JSON, ignoring')
#################################### ####################################
@ -340,12 +350,12 @@ OPENAI_API_CONFIGS = {}
_openai_api_configs = os.getenv('OPENAI_API_CONFIGS', '') _openai_api_configs = os.getenv('OPENAI_API_CONFIGS', '')
if _openai_api_configs: if _openai_api_configs:
try: try:
parsed = json.loads(_openai_api_configs) parsed = JSONCodec.loads(_openai_api_configs)
if isinstance(parsed, dict): if isinstance(parsed, dict):
OPENAI_API_CONFIGS = parsed OPENAI_API_CONFIGS = parsed
else: else:
log.warning('OPENAI_API_CONFIGS must be a JSON object, ignoring') log.warning('OPENAI_API_CONFIGS must be a JSON object, ignoring')
except (json.JSONDecodeError, TypeError): except (JSONCodec.JSONDecodeError, TypeError):
log.warning('OPENAI_API_CONFIGS is not valid JSON, ignoring') log.warning('OPENAI_API_CONFIGS is not valid JSON, ignoring')
# Get the actual OpenAI API key based on the base URL # Get the actual OpenAI API key based on the base URL
@ -369,7 +379,7 @@ ENABLE_BASE_MODELS_CACHE = os.getenv('ENABLE_BASE_MODELS_CACHE', 'False').lower(
#################################### ####################################
try: try:
tool_server_connections = json.loads(os.getenv('TOOL_SERVER_CONNECTIONS', '[]')) tool_server_connections = JSONCodec.loads(os.getenv('TOOL_SERVER_CONNECTIONS', '[]'))
except Exception as e: except Exception as e:
log.exception(f'Error loading TOOL_SERVER_CONNECTIONS: {e}') log.exception(f'Error loading TOOL_SERVER_CONNECTIONS: {e}')
tool_server_connections = [] tool_server_connections = []
@ -383,12 +393,12 @@ OAUTH_CLIENT_TIMEOUT = os.getenv('OAUTH_CLIENT_TIMEOUT', '')
# TERMINAL_SERVER # TERMINAL_SERVER
#################################### ####################################
terminal_server_connections = json.loads(os.getenv('TERMINAL_SERVER_CONNECTIONS', '[]')) terminal_server_connections = JSONCodec.loads(os.getenv('TERMINAL_SERVER_CONNECTIONS', '[]'))
TERMINAL_SERVER_CONNECTIONS = terminal_server_connections TERMINAL_SERVER_CONNECTIONS = terminal_server_connections
try: try:
TERMINAL_PROXY_HEADERS = json.loads(os.getenv('TERMINAL_PROXY_HEADERS', '{}')) TERMINAL_PROXY_HEADERS = JSONCodec.loads(os.getenv('TERMINAL_PROXY_HEADERS', '{}'))
except Exception: except Exception:
TERMINAL_PROXY_HEADERS = {} TERMINAL_PROXY_HEADERS = {}
@ -490,12 +500,12 @@ CODE_INTERPRETER_PYODIDE_PROMPT = """
# Vector Database # Vector Database
#################################### ####################################
VECTOR_DB = os.getenv('VECTOR_DB', 'chroma') VECTOR_DB = os.getenv('VECTOR_DB', 'pgvector' if USE_SLIM else 'chroma')
# Chroma # Chroma
CHROMA_DATA_PATH = f'{DATA_DIR}/vector_db' CHROMA_DATA_PATH = f'{DATA_DIR}/vector_db'
if VECTOR_DB == 'chroma': if VECTOR_DB == 'chroma' and not USE_SLIM:
import chromadb import chromadb
CHROMA_TENANT = os.getenv('CHROMA_TENANT', chromadb.DEFAULT_TENANT) CHROMA_TENANT = os.getenv('CHROMA_TENANT', chromadb.DEFAULT_TENANT)
@ -636,7 +646,7 @@ SSL_ASSERT_FINGERPRINT = os.getenv('SSL_ASSERT_FINGERPRINT', None)
ELASTICSEARCH_INDEX_PREFIX = os.getenv('ELASTICSEARCH_INDEX_PREFIX', 'open_webui_collections') ELASTICSEARCH_INDEX_PREFIX = os.getenv('ELASTICSEARCH_INDEX_PREFIX', 'open_webui_collections')
# Pgvector # Pgvector
PGVECTOR_DB_URL = os.getenv('PGVECTOR_DB_URL', DATABASE_URL) PGVECTOR_DB_URL = os.getenv('PGVECTOR_DB_URL', DATABASE_URL)
if VECTOR_DB == 'pgvector' and not PGVECTOR_DB_URL.startswith('postgres'): if not USE_SLIM and VECTOR_DB == 'pgvector' and not PGVECTOR_DB_URL.startswith('postgres'):
raise ValueError( raise ValueError(
'Pgvector requires setting PGVECTOR_DB_URL or using Postgres with vector extension as the primary database.' 'Pgvector requires setting PGVECTOR_DB_URL or using Postgres with vector extension as the primary database.'
) )
@ -731,6 +741,10 @@ else:
except Exception: except Exception:
PGVECTOR_IVFFLAT_LISTS = 100 PGVECTOR_IVFFLAT_LISTS = 100
PGVECTOR_ITERATIVE_SCAN = os.getenv('PGVECTOR_ITERATIVE_SCAN', 'relaxed_order').strip().lower()
if PGVECTOR_ITERATIVE_SCAN not in ('off', 'relaxed_order', 'strict_order'):
PGVECTOR_ITERATIVE_SCAN = 'relaxed_order'
# openGauss # openGauss
OPENGAUSS_DB_URL = os.getenv('OPENGAUSS_DB_URL', DATABASE_URL) OPENGAUSS_DB_URL = os.getenv('OPENGAUSS_DB_URL', DATABASE_URL)
@ -797,7 +811,7 @@ ORACLE_DB_POOL_MAX = int(os.getenv('ORACLE_DB_POOL_MAX', 10))
ORACLE_DB_POOL_INCREMENT = int(os.getenv('ORACLE_DB_POOL_INCREMENT', 1)) ORACLE_DB_POOL_INCREMENT = int(os.getenv('ORACLE_DB_POOL_INCREMENT', 1))
if VECTOR_DB == 'oracle23ai': if not USE_SLIM and VECTOR_DB == 'oracle23ai':
if not ORACLE_DB_USER or not ORACLE_DB_PASSWORD or not ORACLE_DB_DSN: if not ORACLE_DB_USER or not ORACLE_DB_PASSWORD or not ORACLE_DB_DSN:
raise ValueError('Oracle23ai requires setting ORACLE_DB_USER, ORACLE_DB_PASSWORD, and ORACLE_DB_DSN.') raise ValueError('Oracle23ai requires setting ORACLE_DB_USER, ORACLE_DB_PASSWORD, and ORACLE_DB_DSN.')
if ORACLE_DB_USE_WALLET and (not ORACLE_WALLET_DIR or not ORACLE_WALLET_PASSWORD): if ORACLE_DB_USE_WALLET and (not ORACLE_WALLET_DIR or not ORACLE_WALLET_PASSWORD):
@ -805,7 +819,7 @@ if VECTOR_DB == 'oracle23ai':
'Oracle23ai requires setting ORACLE_WALLET_DIR and ORACLE_WALLET_PASSWORD when using wallet authentication.' 'Oracle23ai requires setting ORACLE_WALLET_DIR and ORACLE_WALLET_PASSWORD when using wallet authentication.'
) )
log.info(f'VECTOR_DB: {VECTOR_DB}') log.info('VECTOR_DB: %s', VECTOR_DB)
# S3 Vector # S3 Vector
S3_VECTOR_BUCKET_NAME = os.getenv('S3_VECTOR_BUCKET_NAME', None) S3_VECTOR_BUCKET_NAME = os.getenv('S3_VECTOR_BUCKET_NAME', None)
@ -894,8 +908,8 @@ MINERU_API_KEY = os.getenv('MINERU_API_KEY', '')
mineru_params = os.getenv('MINERU_PARAMS', '') mineru_params = os.getenv('MINERU_PARAMS', '')
try: try:
mineru_params = json.loads(mineru_params) mineru_params = JSONCodec.loads(mineru_params)
except json.JSONDecodeError: except JSONCodec.JSONDecodeError:
mineru_params = {} mineru_params = {}
MINERU_PARAMS = mineru_params MINERU_PARAMS = mineru_params
@ -908,8 +922,8 @@ EXTERNAL_DOCUMENT_LOADER_API_KEY = os.getenv('EXTERNAL_DOCUMENT_LOADER_API_KEY',
external_document_loader_headers = os.getenv('EXTERNAL_DOCUMENT_LOADER_HEADERS', '') external_document_loader_headers = os.getenv('EXTERNAL_DOCUMENT_LOADER_HEADERS', '')
try: try:
external_document_loader_headers = json.loads(external_document_loader_headers) external_document_loader_headers = JSONCodec.loads(external_document_loader_headers)
except json.JSONDecodeError: except JSONCodec.JSONDecodeError:
external_document_loader_headers = {} external_document_loader_headers = {}
if not isinstance(external_document_loader_headers, dict): if not isinstance(external_document_loader_headers, dict):
external_document_loader_headers = {} external_document_loader_headers = {}
@ -918,14 +932,16 @@ EXTERNAL_DOCUMENT_LOADER_HEADERS = external_document_loader_headers
TIKA_SERVER_URL = os.getenv('TIKA_SERVER_URL', 'http://tika:9998') TIKA_SERVER_URL = os.getenv('TIKA_SERVER_URL', 'http://tika:9998')
TIKA_SERVER_VERSION = os.getenv('TIKA_SERVER_VERSION', '3')
DOCLING_SERVER_URL = os.getenv('DOCLING_SERVER_URL', 'http://docling:5001') DOCLING_SERVER_URL = os.getenv('DOCLING_SERVER_URL', 'http://docling:5001')
DOCLING_API_KEY = os.getenv('DOCLING_API_KEY', '') DOCLING_API_KEY = os.getenv('DOCLING_API_KEY', '')
docling_params = os.getenv('DOCLING_PARAMS', '') docling_params = os.getenv('DOCLING_PARAMS', '')
try: try:
docling_params = json.loads(docling_params) docling_params = JSONCodec.loads(docling_params)
except json.JSONDecodeError: except JSONCodec.JSONDecodeError:
docling_params = {} docling_params = {}
DOCLING_PARAMS = docling_params DOCLING_PARAMS = docling_params
@ -966,6 +982,8 @@ RAG_FILE_MAX_COUNT = int(os.getenv('RAG_FILE_MAX_COUNT')) if os.getenv('RAG_FILE
RAG_FILE_MAX_SIZE = int(os.getenv('RAG_FILE_MAX_SIZE')) if os.getenv('RAG_FILE_MAX_SIZE') else None RAG_FILE_MAX_SIZE = int(os.getenv('RAG_FILE_MAX_SIZE')) if os.getenv('RAG_FILE_MAX_SIZE') else None
ENABLE_KNOWLEDGE_FILE_RETENTION = os.getenv('ENABLE_KNOWLEDGE_FILE_RETENTION', 'False').lower() == 'true'
RAG_FILE_CONTENT_SEARCH_MAX_CHARS = int(os.getenv('RAG_FILE_CONTENT_SEARCH_MAX_CHARS', str(64 * 1024 * 1024))) RAG_FILE_CONTENT_SEARCH_MAX_CHARS = int(os.getenv('RAG_FILE_CONTENT_SEARCH_MAX_CHARS', str(64 * 1024 * 1024)))
FILE_IMAGE_COMPRESSION_WIDTH = ( FILE_IMAGE_COMPRESSION_WIDTH = (
@ -988,7 +1006,7 @@ PDF_EXTRACT_IMAGES = os.getenv('PDF_EXTRACT_IMAGES', 'False').lower() == 'true'
PDF_LOADER_MODE = os.getenv('PDF_LOADER_MODE', 'page') PDF_LOADER_MODE = os.getenv('PDF_LOADER_MODE', 'page')
RAG_EMBEDDING_MODEL = os.getenv('RAG_EMBEDDING_MODEL', 'sentence-transformers/all-MiniLM-L6-v2') RAG_EMBEDDING_MODEL = os.getenv('RAG_EMBEDDING_MODEL', 'sentence-transformers/all-MiniLM-L6-v2')
log.info(f'Embedding model set: {RAG_EMBEDDING_MODEL}') log.info('Embedding model set: %s', RAG_EMBEDDING_MODEL)
RAG_TOKENIZER_MODEL = os.getenv('RAG_TOKENIZER_MODEL', '') RAG_TOKENIZER_MODEL = os.getenv('RAG_TOKENIZER_MODEL', '')
@ -1016,7 +1034,7 @@ RAG_RERANKING_ENGINE = os.getenv('RAG_RERANKING_ENGINE', '')
RAG_RERANKING_MODEL = os.getenv('RAG_RERANKING_MODEL', '') RAG_RERANKING_MODEL = os.getenv('RAG_RERANKING_MODEL', '')
if RAG_RERANKING_MODEL != '': if RAG_RERANKING_MODEL != '':
log.info(f'Reranking model set: {RAG_RERANKING_MODEL}') log.info('Reranking model set: %s', RAG_RERANKING_MODEL)
RAG_RERANKING_MODEL_AUTO_UPDATE = ( RAG_RERANKING_MODEL_AUTO_UPDATE = (
@ -1100,12 +1118,26 @@ ENABLE_LOCAL_WEB_FETCH = (
ENABLE_RAG_LOCAL_WEB_FETCH = ENABLE_LOCAL_WEB_FETCH ENABLE_RAG_LOCAL_WEB_FETCH = ENABLE_LOCAL_WEB_FETCH
# Operators extend this through WEB_FETCH_FILTER_LIST.
DEFAULT_WEB_FETCH_FILTER_LIST = [ DEFAULT_WEB_FETCH_FILTER_LIST = [
'!169.254.169.254', '!169.254.169.254',
'!fd00:ec2::254', '!fd00:ec2::254',
'!metadata.google.internal', '!metadata.google.internal',
'!metadata.azure.com', '!metadata.azure.com',
'!100.100.100.200', '!100.100.100.200',
'!168.63.129.16', # Azure platform channel, reachable from every Azure VM
'!192.88.99.0/24', # 6to4 relay anycast, deprecated by RFC 7526
'!224.0.0.0/4', # IPv4 multicast
'!::ffff:0:0:0/96', # IPv4-translated (SIIT, RFC 2765), never routed
'!64:ff9b:1::/48', # NAT64 local-use prefix, RFC 8215, not a public destination
'!100:0:0:1::/64', # dummy prefix, RFC 9780
'!2001:1::1', # PCP anycast, RFC 7723, answered by the local network's own edge device
'!2001:1::2', # TURN anycast, RFC 8155, likewise
'!2001:20::/28', # ORCHIDv2, RFC 7343, never routed
'!2001:30::/28', # DRIP, RFC 9374, never routed
'!5f00::/16', # SRv6 SIDs, RFC 9602, internal to one segment routing domain
'!fec0::/10', # IPv6 site-local, deprecated by RFC 3879
'!ff00::/8', # IPv6 multicast
] ]
web_fetch_filter_list = os.getenv('WEB_FETCH_FILTER_LIST', '') web_fetch_filter_list = os.getenv('WEB_FETCH_FILTER_LIST', '')
@ -1148,7 +1180,7 @@ WEB_SEARCH_RESULT_COUNT = int(os.getenv('WEB_SEARCH_RESULT_COUNT', '3'))
try: try:
web_search_domain_filter_list = json.loads(os.getenv('WEB_SEARCH_DOMAIN_FILTER_LIST', '[]')) web_search_domain_filter_list = JSONCodec.loads(os.getenv('WEB_SEARCH_DOMAIN_FILTER_LIST', '[]'))
except Exception as e: except Exception as e:
web_search_domain_filter_list = [ web_search_domain_filter_list = [
# "wikipedia.com", # "wikipedia.com",
@ -1244,6 +1276,9 @@ AZURE_AI_SEARCH_ENDPOINT = os.getenv('AZURE_AI_SEARCH_ENDPOINT', '')
AZURE_AI_SEARCH_INDEX_NAME = os.getenv('AZURE_AI_SEARCH_INDEX_NAME', '') AZURE_AI_SEARCH_INDEX_NAME = os.getenv('AZURE_AI_SEARCH_INDEX_NAME', '')
EXA_API_KEY = os.getenv('EXA_API_KEY', '') EXA_API_KEY = os.getenv('EXA_API_KEY', '')
EXA_MAX_CONTENT_LENGTH = int(os.environ['EXA_MAX_CONTENT_LENGTH']) if os.getenv('EXA_MAX_CONTENT_LENGTH') else None
if EXA_MAX_CONTENT_LENGTH is not None and EXA_MAX_CONTENT_LENGTH <= 0:
raise ValueError('EXA_MAX_CONTENT_LENGTH must be a positive integer or unset')
PERPLEXITY_API_KEY = os.getenv('PERPLEXITY_API_KEY', '') PERPLEXITY_API_KEY = os.getenv('PERPLEXITY_API_KEY', '')
@ -1267,6 +1302,12 @@ TAVILY_API_KEY = os.getenv('TAVILY_API_KEY', '')
TAVILY_EXTRACT_DEPTH = os.getenv('TAVILY_EXTRACT_DEPTH', 'basic') TAVILY_EXTRACT_DEPTH = os.getenv('TAVILY_EXTRACT_DEPTH', 'basic')
STAAN_API_KEY = os.getenv('STAAN_API_KEY', '')
STAAN_MARKET = os.getenv('STAAN_MARKET', 'en-us')
STAAN_MAX_SNIPPETS = int(os.getenv('STAAN_MAX_SNIPPETS', '0'))
PLAYWRIGHT_WS_URL = os.getenv('PLAYWRIGHT_WS_URL', '') PLAYWRIGHT_WS_URL = os.getenv('PLAYWRIGHT_WS_URL', '')
PLAYWRIGHT_TIMEOUT = int(os.getenv('PLAYWRIGHT_TIMEOUT', '10000')) PLAYWRIGHT_TIMEOUT = int(os.getenv('PLAYWRIGHT_TIMEOUT', '10000'))
@ -1297,8 +1338,8 @@ LINKUP_API_KEY = os.getenv('LINKUP_API_KEY', '')
linkup_search_params = os.getenv('LINKUP_SEARCH_PARAMS', '') linkup_search_params = os.getenv('LINKUP_SEARCH_PARAMS', '')
try: try:
linkup_search_params = json.loads(linkup_search_params) linkup_search_params = JSONCodec.loads(linkup_search_params)
except json.JSONDecodeError: except JSONCodec.JSONDecodeError:
linkup_search_params = {} linkup_search_params = {}
LINKUP_SEARCH_PARAMS = linkup_search_params LINKUP_SEARCH_PARAMS = linkup_search_params
@ -1330,8 +1371,8 @@ AUTOMATIC1111_API_AUTH = os.getenv('AUTOMATIC1111_API_AUTH', '')
automatic1111_params = os.getenv('AUTOMATIC1111_PARAMS', '') automatic1111_params = os.getenv('AUTOMATIC1111_PARAMS', '')
try: try:
automatic1111_params = json.loads(automatic1111_params) automatic1111_params = JSONCodec.loads(automatic1111_params)
except json.JSONDecodeError: except JSONCodec.JSONDecodeError:
automatic1111_params = {} automatic1111_params = {}
AUTOMATIC1111_PARAMS = automatic1111_params AUTOMATIC1111_PARAMS = automatic1111_params
@ -1455,8 +1496,8 @@ COMFYUI_WORKFLOW = os.getenv('COMFYUI_WORKFLOW', COMFYUI_DEFAULT_WORKFLOW)
comfyui_workflow_nodes = os.getenv('COMFYUI_WORKFLOW_NODES', '') comfyui_workflow_nodes = os.getenv('COMFYUI_WORKFLOW_NODES', '')
try: try:
comfyui_workflow_nodes = json.loads(comfyui_workflow_nodes) comfyui_workflow_nodes = JSONCodec.loads(comfyui_workflow_nodes)
except json.JSONDecodeError: except JSONCodec.JSONDecodeError:
comfyui_workflow_nodes = [] comfyui_workflow_nodes = []
COMFYUI_WORKFLOW_NODES = comfyui_workflow_nodes COMFYUI_WORKFLOW_NODES = comfyui_workflow_nodes
@ -1468,8 +1509,8 @@ IMAGES_OPENAI_API_KEY = os.getenv('IMAGES_OPENAI_API_KEY', OPENAI_API_KEY)
images_openai_params = os.getenv('IMAGES_OPENAI_PARAMS', '') images_openai_params = os.getenv('IMAGES_OPENAI_PARAMS', '')
try: try:
images_openai_params = json.loads(images_openai_params) images_openai_params = JSONCodec.loads(images_openai_params)
except json.JSONDecodeError: except JSONCodec.JSONDecodeError:
images_openai_params = {} images_openai_params = {}
@ -1507,8 +1548,8 @@ IMAGES_EDIT_COMFYUI_WORKFLOW = os.getenv('IMAGES_EDIT_COMFYUI_WORKFLOW', '')
images_edit_comfyui_workflow_nodes = os.getenv('IMAGES_EDIT_COMFYUI_WORKFLOW_NODES', '') images_edit_comfyui_workflow_nodes = os.getenv('IMAGES_EDIT_COMFYUI_WORKFLOW_NODES', '')
try: try:
images_edit_comfyui_workflow_nodes = json.loads(images_edit_comfyui_workflow_nodes) images_edit_comfyui_workflow_nodes = JSONCodec.loads(images_edit_comfyui_workflow_nodes)
except json.JSONDecodeError: except JSONCodec.JSONDecodeError:
images_edit_comfyui_workflow_nodes = [] images_edit_comfyui_workflow_nodes = []
IMAGES_EDIT_COMFYUI_WORKFLOW_NODES = images_edit_comfyui_workflow_nodes IMAGES_EDIT_COMFYUI_WORKFLOW_NODES = images_edit_comfyui_workflow_nodes
@ -1582,8 +1623,8 @@ AUDIO_TTS_OPENAI_API_KEY = os.getenv('AUDIO_TTS_OPENAI_API_KEY', OPENAI_API_KEY)
audio_tts_openai_params = os.getenv('AUDIO_TTS_OPENAI_PARAMS', '') audio_tts_openai_params = os.getenv('AUDIO_TTS_OPENAI_PARAMS', '')
try: try:
audio_tts_openai_params = json.loads(audio_tts_openai_params) audio_tts_openai_params = JSONCodec.loads(audio_tts_openai_params)
except json.JSONDecodeError: except JSONCodec.JSONDecodeError:
audio_tts_openai_params = {} audio_tts_openai_params = {}
AUDIO_TTS_OPENAI_PARAMS = audio_tts_openai_params AUDIO_TTS_OPENAI_PARAMS = audio_tts_openai_params
@ -1634,46 +1675,17 @@ DEFAULT_MODELS = os.getenv('DEFAULT_MODELS', None)
DEFAULT_PINNED_MODELS = os.getenv('DEFAULT_PINNED_MODELS', None) DEFAULT_PINNED_MODELS = os.getenv('DEFAULT_PINNED_MODELS', None)
# None uses the frontend's localized defaults; an empty list disables suggestions.
try: try:
default_prompt_suggestions = json.loads(os.getenv('DEFAULT_PROMPT_SUGGESTIONS', '[]')) DEFAULT_PROMPT_SUGGESTIONS = JSONCodec.loads(os.getenv('DEFAULT_PROMPT_SUGGESTIONS', 'null'))
except Exception as e: except Exception as e:
log.exception(f'Error loading DEFAULT_PROMPT_SUGGESTIONS: {e}') log.exception(f'Error loading DEFAULT_PROMPT_SUGGESTIONS: {e}')
default_prompt_suggestions = [] DEFAULT_PROMPT_SUGGESTIONS = None
if default_prompt_suggestions == []:
default_prompt_suggestions = [
{
'title': ['Help me study', 'vocabulary for a college entrance exam'],
'content': "Help me study vocabulary: write a sentence for me to fill in the blank, and I'll try to pick the correct option.",
},
{
'title': ['Give me ideas', "for what to do with my kids' art"],
'content': "What are 5 creative things I could do with my kids' art? I don't want to throw them away, but it's also so much clutter.",
},
{
'title': ['Tell me a fun fact', 'about the Roman Empire'],
'content': 'Tell me a random fun fact about the Roman Empire',
},
{
'title': ['Show me a code snippet', "of a website's sticky header"],
'content': "Show me a code snippet of a website's sticky header in CSS and JavaScript.",
},
{
'title': [
'Explain options trading',
"if I'm familiar with buying and selling stocks",
],
'content': "Explain options trading in simple terms if I'm familiar with buying and selling stocks.",
},
{
'title': ['Overcome procrastination', 'give me tips'],
'content': 'Could you start by asking me about instances when I procrastinate the most and then give me some suggestions to overcome it?',
},
]
DEFAULT_PROMPT_SUGGESTIONS = default_prompt_suggestions DEFAULT_PROMPT_SUGGESTIONS_I18N = {}
try: try:
model_order_list = json.loads(os.getenv('MODEL_ORDER_LIST', '[]')) model_order_list = JSONCodec.loads(os.getenv('MODEL_ORDER_LIST', '[]'))
except Exception as e: except Exception as e:
log.exception(f'Error loading MODEL_ORDER_LIST: {e}') log.exception(f'Error loading MODEL_ORDER_LIST: {e}')
model_order_list = [] model_order_list = []
@ -1681,7 +1693,7 @@ except Exception as e:
MODEL_ORDER_LIST = model_order_list MODEL_ORDER_LIST = model_order_list
try: try:
default_model_metadata = json.loads(os.getenv('DEFAULT_MODEL_METADATA', '{}')) default_model_metadata = JSONCodec.loads(os.getenv('DEFAULT_MODEL_METADATA', '{}'))
except Exception as e: except Exception as e:
log.exception(f'Error loading DEFAULT_MODEL_METADATA: {e}') log.exception(f'Error loading DEFAULT_MODEL_METADATA: {e}')
default_model_metadata = {} default_model_metadata = {}
@ -1689,13 +1701,22 @@ except Exception as e:
DEFAULT_MODEL_METADATA = default_model_metadata DEFAULT_MODEL_METADATA = default_model_metadata
try: try:
default_model_params = json.loads(os.getenv('DEFAULT_MODEL_PARAMS', '{}')) default_model_params = JSONCodec.loads(os.getenv('DEFAULT_MODEL_PARAMS', '{}'))
except Exception as e: except Exception as e:
log.exception(f'Error loading DEFAULT_MODEL_PARAMS: {e}') log.exception(f'Error loading DEFAULT_MODEL_PARAMS: {e}')
default_model_params = {} default_model_params = {}
DEFAULT_MODEL_PARAMS = default_model_params DEFAULT_MODEL_PARAMS = default_model_params
try:
default_interface_settings = JSONCodec.loads(os.getenv('DEFAULT_INTERFACE_SETTINGS', '{}'))
except Exception as e:
log.exception(f'Error loading DEFAULT_INTERFACE_SETTINGS: {e}')
default_interface_settings = {}
DEFAULT_INTERFACE_SETTINGS = default_interface_settings if isinstance(default_interface_settings, dict) else {}
DEFAULT_USER_ROLE = os.getenv('DEFAULT_USER_ROLE', 'pending') DEFAULT_USER_ROLE = os.getenv('DEFAULT_USER_ROLE', 'pending')
DEFAULT_GROUP_ID = os.getenv('DEFAULT_GROUP_ID', '') DEFAULT_GROUP_ID = os.getenv('DEFAULT_GROUP_ID', '')
@ -2032,7 +2053,7 @@ ENABLE_USER_STATUS = os.getenv('ENABLE_USER_STATUS', 'True').lower() == 'true'
ENABLE_EVALUATION_ARENA_MODELS = os.getenv('ENABLE_EVALUATION_ARENA_MODELS', 'True').lower() == 'true' ENABLE_EVALUATION_ARENA_MODELS = os.getenv('ENABLE_EVALUATION_ARENA_MODELS', 'True').lower() == 'true'
try: try:
evaluation_arena_models = json.loads(os.getenv('EVALUATION_ARENA_MODELS', '[]')) evaluation_arena_models = JSONCodec.loads(os.getenv('EVALUATION_ARENA_MODELS', '[]'))
if not isinstance(evaluation_arena_models, list) or not all( if not isinstance(evaluation_arena_models, list) or not all(
isinstance(model, dict) for model in evaluation_arena_models isinstance(model, dict) for model in evaluation_arena_models
): ):
@ -2047,6 +2068,9 @@ DEFAULT_ARENA_MODEL = {
'id': 'arena-model', 'id': 'arena-model',
'name': 'Arena Model', 'name': 'Arena Model',
'meta': { 'meta': {
# LICENSE covers this Open WebUI fallback logo.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
'profile_image_url': '/favicon.png', 'profile_image_url': '/favicon.png',
'description': 'Submit your questions to anonymous AI chatbots and vote on the best response.', 'description': 'Submit your questions to anonymous AI chatbots and vote on the best response.',
'model_ids': None, 'model_ids': None,
@ -2067,8 +2091,6 @@ BYPASS_ADMIN_ACCESS_CONTROL = (
== 'true' == 'true'
) )
ENABLE_ADMIN_CHAT_ACCESS = os.getenv('ENABLE_ADMIN_CHAT_ACCESS', 'True').lower() == 'true'
ENABLE_ADMIN_ANALYTICS = os.getenv('ENABLE_ADMIN_ANALYTICS', 'True').lower() == 'true' ENABLE_ADMIN_ANALYTICS = os.getenv('ENABLE_ADMIN_ANALYTICS', 'True').lower() == 'true'
ENABLE_COMMUNITY_SHARING = os.getenv('ENABLE_COMMUNITY_SHARING', 'True').lower() == 'true' ENABLE_COMMUNITY_SHARING = os.getenv('ENABLE_COMMUNITY_SHARING', 'True').lower() == 'true'
@ -2079,6 +2101,7 @@ ENABLE_USER_WEBHOOKS = os.getenv('ENABLE_USER_WEBHOOKS', 'False').lower() == 'tr
# FastAPI / AnyIO settings # FastAPI / AnyIO settings
THREAD_POOL_SIZE = os.getenv('THREAD_POOL_SIZE', None) THREAD_POOL_SIZE = os.getenv('THREAD_POOL_SIZE', None)
THREAD_POOL_THREAD_NAME_PREFIX = os.getenv('THREAD_POOL_THREAD_NAME_PREFIX', '')
if THREAD_POOL_SIZE is not None and isinstance(THREAD_POOL_SIZE, str): if THREAD_POOL_SIZE is not None and isinstance(THREAD_POOL_SIZE, str):
try: try:
@ -2125,6 +2148,7 @@ else:
class BannerModel(BaseModel): class BannerModel(BaseModel):
i18n: dict[str, dict[str, str]] | None = None
id: str id: str
type: str type: str
title: str | None = None title: str | None = None
@ -2134,7 +2158,7 @@ class BannerModel(BaseModel):
try: try:
banners = json.loads(os.getenv('WEBUI_BANNERS', '[]')) banners = JSONCodec.loads(os.getenv('WEBUI_BANNERS', '[]'))
banners = [BannerModel(**banner) for banner in banners] banners = [BannerModel(**banner) for banner in banners]
except Exception as e: except Exception as e:
log.exception(f'Error loading WEBUI_BANNERS: {e}') log.exception(f'Error loading WEBUI_BANNERS: {e}')
@ -2157,10 +2181,20 @@ TASK_MODEL = os.getenv('TASK_MODEL', '')
TASK_MODEL_EXTERNAL = os.getenv('TASK_MODEL_EXTERNAL', '') TASK_MODEL_EXTERNAL = os.getenv('TASK_MODEL_EXTERNAL', '')
try:
task_model_params = JSONCodec.loads(os.getenv('TASK_MODEL_PARAMS', '{}'))
except Exception as e:
log.exception(f'Error loading TASK_MODEL_PARAMS: {e}')
task_model_params = {}
TASK_MODEL_PARAMS = task_model_params
CONTEXT_COMPACTION_MODEL = os.getenv('CONTEXT_COMPACTION_MODEL', '') CONTEXT_COMPACTION_MODEL = os.getenv('CONTEXT_COMPACTION_MODEL', '')
ENABLE_CONTEXT_COMPACTION = os.getenv('ENABLE_CONTEXT_COMPACTION', 'False').lower() == 'true' ENABLE_CONTEXT_COMPACTION = os.getenv('ENABLE_CONTEXT_COMPACTION', 'False').lower() == 'true'
ENABLE_TOOL_PERMISSIONS = os.getenv('ENABLE_TOOL_PERMISSIONS', 'False').lower() == 'true'
CONTEXT_COMPACTION_TOKEN_THRESHOLD = int(os.getenv('CONTEXT_COMPACTION_TOKEN_THRESHOLD', '80000')) CONTEXT_COMPACTION_TOKEN_THRESHOLD = int(os.getenv('CONTEXT_COMPACTION_TOKEN_THRESHOLD', '80000'))
_CONTEXT_COMPACTION_TOKEN_CAP = os.getenv('CONTEXT_COMPACTION_TOKEN_CAP') _CONTEXT_COMPACTION_TOKEN_CAP = os.getenv('CONTEXT_COMPACTION_TOKEN_CAP')
@ -2470,12 +2504,12 @@ GOOGLE_OAUTH_AUTHORIZE_PARAMS = {}
_google_oauth_authorize_params = os.getenv('GOOGLE_OAUTH_AUTHORIZE_PARAMS', '') _google_oauth_authorize_params = os.getenv('GOOGLE_OAUTH_AUTHORIZE_PARAMS', '')
if _google_oauth_authorize_params: if _google_oauth_authorize_params:
try: try:
_parsed = json.loads(_google_oauth_authorize_params) _parsed = JSONCodec.loads(_google_oauth_authorize_params)
if isinstance(_parsed, dict): if isinstance(_parsed, dict):
GOOGLE_OAUTH_AUTHORIZE_PARAMS = _parsed GOOGLE_OAUTH_AUTHORIZE_PARAMS = _parsed
else: else:
log.warning('GOOGLE_OAUTH_AUTHORIZE_PARAMS must be a JSON object, ignoring') log.warning('GOOGLE_OAUTH_AUTHORIZE_PARAMS must be a JSON object, ignoring')
except (json.JSONDecodeError, TypeError): except (JSONCodec.JSONDecodeError, TypeError):
log.warning('GOOGLE_OAUTH_AUTHORIZE_PARAMS is not valid JSON, ignoring') log.warning('GOOGLE_OAUTH_AUTHORIZE_PARAMS is not valid JSON, ignoring')
MICROSOFT_CLIENT_ID = os.getenv('MICROSOFT_CLIENT_ID', '') MICROSOFT_CLIENT_ID = os.getenv('MICROSOFT_CLIENT_ID', '')
@ -2590,12 +2624,12 @@ OAUTH_AUTHORIZE_PARAMS = {}
_oauth_authorize_params = os.getenv('OAUTH_AUTHORIZE_PARAMS', '') _oauth_authorize_params = os.getenv('OAUTH_AUTHORIZE_PARAMS', '')
if _oauth_authorize_params: if _oauth_authorize_params:
try: try:
_parsed = json.loads(_oauth_authorize_params) _parsed = JSONCodec.loads(_oauth_authorize_params)
if isinstance(_parsed, dict): if isinstance(_parsed, dict):
OAUTH_AUTHORIZE_PARAMS = _parsed OAUTH_AUTHORIZE_PARAMS = _parsed
else: else:
log.warning('OAUTH_AUTHORIZE_PARAMS must be a JSON object, ignoring') log.warning('OAUTH_AUTHORIZE_PARAMS must be a JSON object, ignoring')
except (json.JSONDecodeError, TypeError): except (JSONCodec.JSONDecodeError, TypeError):
log.warning('OAUTH_AUTHORIZE_PARAMS is not valid JSON, ignoring') log.warning('OAUTH_AUTHORIZE_PARAMS is not valid JSON, ignoring')
@ -2787,6 +2821,7 @@ LDAP_ATTRIBUTE_FOR_GROUPS = os.getenv('LDAP_ATTRIBUTE_FOR_GROUPS', 'memberOf')
DEFAULT_CONFIG = { DEFAULT_CONFIG = {
'direct.enable': ENABLE_DIRECT_CONNECTIONS, 'direct.enable': ENABLE_DIRECT_CONNECTIONS,
'direct.integrations.enable': ENABLE_DIRECT_INTEGRATIONS,
'ollama.enable': ENABLE_OLLAMA_API, 'ollama.enable': ENABLE_OLLAMA_API,
'ollama.base_urls': OLLAMA_BASE_URLS, 'ollama.base_urls': OLLAMA_BASE_URLS,
'ollama.api_configs': OLLAMA_API_CONFIGS, 'ollama.api_configs': OLLAMA_API_CONFIGS,
@ -2848,6 +2883,7 @@ DEFAULT_CONFIG = {
'rag.external_document_loader_api_key': EXTERNAL_DOCUMENT_LOADER_API_KEY, 'rag.external_document_loader_api_key': EXTERNAL_DOCUMENT_LOADER_API_KEY,
'rag.external_document_loader_headers': EXTERNAL_DOCUMENT_LOADER_HEADERS, 'rag.external_document_loader_headers': EXTERNAL_DOCUMENT_LOADER_HEADERS,
'rag.tika_server_url': TIKA_SERVER_URL, 'rag.tika_server_url': TIKA_SERVER_URL,
'rag.tika_server_version': TIKA_SERVER_VERSION,
'rag.docling_server_url': DOCLING_SERVER_URL, 'rag.docling_server_url': DOCLING_SERVER_URL,
'rag.docling_api_key': DOCLING_API_KEY, 'rag.docling_api_key': DOCLING_API_KEY,
'rag.docling_params': DOCLING_PARAMS, 'rag.docling_params': DOCLING_PARAMS,
@ -2950,6 +2986,7 @@ DEFAULT_CONFIG = {
'web.search.azure_ai_search_endpoint': AZURE_AI_SEARCH_ENDPOINT, 'web.search.azure_ai_search_endpoint': AZURE_AI_SEARCH_ENDPOINT,
'web.search.azure_ai_search_index_name': AZURE_AI_SEARCH_INDEX_NAME, 'web.search.azure_ai_search_index_name': AZURE_AI_SEARCH_INDEX_NAME,
'web.search.exa_api_key': EXA_API_KEY, 'web.search.exa_api_key': EXA_API_KEY,
'web.search.exa_max_content_length': EXA_MAX_CONTENT_LENGTH,
'web.search.perplexity_api_key': PERPLEXITY_API_KEY, 'web.search.perplexity_api_key': PERPLEXITY_API_KEY,
'web.search.perplexity_model': PERPLEXITY_MODEL, 'web.search.perplexity_model': PERPLEXITY_MODEL,
'web.search.perplexity_search_context_usage': PERPLEXITY_SEARCH_CONTEXT_USAGE, 'web.search.perplexity_search_context_usage': PERPLEXITY_SEARCH_CONTEXT_USAGE,
@ -2961,6 +2998,9 @@ DEFAULT_CONFIG = {
'web.search.sougou_api_sk': SOUGOU_API_SK, 'web.search.sougou_api_sk': SOUGOU_API_SK,
'web.search.tavily_api_key': TAVILY_API_KEY, 'web.search.tavily_api_key': TAVILY_API_KEY,
'web.search.tavily_extract_depth': TAVILY_EXTRACT_DEPTH, 'web.search.tavily_extract_depth': TAVILY_EXTRACT_DEPTH,
'web.search.staan_api_key': STAAN_API_KEY,
'web.search.staan_market': STAAN_MARKET,
'web.search.staan_max_snippets': STAAN_MAX_SNIPPETS,
'web.loader.playwright_ws_url': PLAYWRIGHT_WS_URL, 'web.loader.playwright_ws_url': PLAYWRIGHT_WS_URL,
'web.loader.playwright_timeout': PLAYWRIGHT_TIMEOUT, 'web.loader.playwright_timeout': PLAYWRIGHT_TIMEOUT,
'web.loader.firecrawl_api_key': FIRECRAWL_API_KEY, 'web.loader.firecrawl_api_key': FIRECRAWL_API_KEY,
@ -3046,7 +3086,10 @@ DEFAULT_CONFIG = {
'ui.default_locale': DEFAULT_LOCALE, 'ui.default_locale': DEFAULT_LOCALE,
'ui.default_models': DEFAULT_MODELS, 'ui.default_models': DEFAULT_MODELS,
'ui.default_pinned_models': DEFAULT_PINNED_MODELS, 'ui.default_pinned_models': DEFAULT_PINNED_MODELS,
'ui.default_interface_settings': DEFAULT_INTERFACE_SETTINGS,
'ui.i18n': {},
'ui.prompt_suggestions': DEFAULT_PROMPT_SUGGESTIONS, 'ui.prompt_suggestions': DEFAULT_PROMPT_SUGGESTIONS,
'ui.prompt_suggestions_i18n': DEFAULT_PROMPT_SUGGESTIONS_I18N,
'ui.model_order_list': MODEL_ORDER_LIST, 'ui.model_order_list': MODEL_ORDER_LIST,
'models.default_metadata': DEFAULT_MODEL_METADATA, 'models.default_metadata': DEFAULT_MODEL_METADATA,
'models.default_params': DEFAULT_MODEL_PARAMS, 'models.default_params': DEFAULT_MODEL_PARAMS,
@ -3085,12 +3128,14 @@ DEFAULT_CONFIG = {
'auth.admin.email': ADMIN_EMAIL, 'auth.admin.email': ADMIN_EMAIL,
'task.model.default': TASK_MODEL, 'task.model.default': TASK_MODEL,
'task.model.external': TASK_MODEL_EXTERNAL, 'task.model.external': TASK_MODEL_EXTERNAL,
'task.model.params': TASK_MODEL_PARAMS,
'chat.context_compaction.model': CONTEXT_COMPACTION_MODEL, 'chat.context_compaction.model': CONTEXT_COMPACTION_MODEL,
'chat.context_compaction.enable': ENABLE_CONTEXT_COMPACTION, 'chat.context_compaction.enable': ENABLE_CONTEXT_COMPACTION,
'chat.context_compaction.token_threshold': CONTEXT_COMPACTION_TOKEN_THRESHOLD, 'chat.context_compaction.token_threshold': CONTEXT_COMPACTION_TOKEN_THRESHOLD,
'chat.context_compaction.token_cap': CONTEXT_COMPACTION_TOKEN_CAP, 'chat.context_compaction.token_cap': CONTEXT_COMPACTION_TOKEN_CAP,
'chat.context_compaction.retention_percentage': CONTEXT_COMPACTION_RETENTION_PERCENTAGE, 'chat.context_compaction.retention_percentage': CONTEXT_COMPACTION_RETENTION_PERCENTAGE,
'chat.context_compaction.prompt_template': CONTEXT_COMPACTION_PROMPT_TEMPLATE, 'chat.context_compaction.prompt_template': CONTEXT_COMPACTION_PROMPT_TEMPLATE,
'chat.tool_permissions.enable': ENABLE_TOOL_PERMISSIONS,
'task.title.prompt_template': TITLE_GENERATION_PROMPT_TEMPLATE, 'task.title.prompt_template': TITLE_GENERATION_PROMPT_TEMPLATE,
'task.tags.prompt_template': TAGS_GENERATION_PROMPT_TEMPLATE, 'task.tags.prompt_template': TAGS_GENERATION_PROMPT_TEMPLATE,
'task.image.prompt_template': IMAGE_PROMPT_GENERATION_PROMPT_TEMPLATE, 'task.image.prompt_template': IMAGE_PROMPT_GENERATION_PROMPT_TEMPLATE,

View file

@ -3,7 +3,6 @@ from __future__ import annotations
import errno import errno
from enum import Enum from enum import Enum
_ERRNO_MESSAGES = { _ERRNO_MESSAGES = {
errno.ENAMETOOLONG: 'File name is too long.', errno.ENAMETOOLONG: 'File name is too long.',
errno.ENOSPC: 'The server is out of storage space.', errno.ENOSPC: 'The server is out of storage space.',
@ -99,7 +98,7 @@ class ERROR_MESSAGES(str, Enum):
INVALID_URL = 'The URL you provided is invalid. Please double-check and try again.' INVALID_URL = 'The URL you provided is invalid. Please double-check and try again.'
WEB_SEARCH_ERROR = lambda err='': err if err else 'Something went wrong while searching the web.' WEB_SEARCH_ERROR = 'Something went wrong while searching the web.'
OLLAMA_API_DISABLED = 'The Ollama API is disabled. Please enable it to use this feature.' OLLAMA_API_DISABLED = 'The Ollama API is disabled. Please enable it to use this feature.'
@ -118,9 +117,16 @@ class ERROR_MESSAGES(str, Enum):
AUTOMATION_TOO_FREQUENT = lambda interval='': f'Schedule too frequent. Minimum interval is {interval} seconds.' AUTOMATION_TOO_FREQUENT = lambda interval='': f'Schedule too frequent. Minimum interval is {interval} seconds.'
AUTOMATION_INVALID_RRULE = lambda err='': f'Invalid RRULE: {err}' AUTOMATION_INVALID_RRULE = lambda err='': f'Invalid RRULE: {err}'
AUTOMATION_NO_FUTURE_RUNS = 'RRULE has no future occurrences' AUTOMATION_NO_FUTURE_RUNS = 'RRULE has no future occurrences'
AUTOMATION_COUNT_REQUIRES_DTSTART = (
'RRULE with COUNT requires an explicit DTSTART line to anchor the occurrence window'
)
CALENDAR_RRULE_TOO_FREQUENT = 'Recurring events cannot repeat more often than daily'
FEATURE_DISABLED = lambda name='': f'{name} is disabled' FEATURE_DISABLED = lambda name='': f'{name} is disabled'
INPUT_TOO_LONG = lambda size='': f'Input prompt exceeds maximum length of {size}' INPUT_TOO_LONG = lambda size='': f'Input prompt exceeds maximum length of {size}'
# LICENSE covers this Open WebUI error identifier.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
SERVER_CONNECTION_ERROR = 'Open WebUI: Server Connection Error' SERVER_CONNECTION_ERROR = 'Open WebUI: Server Connection Error'
REQUIRED_FIELD_EMPTY = lambda name='': f'Required field {name} is empty' REQUIRED_FIELD_EMPTY = lambda name='': f'Required field {name} is empty'
OAUTH_NOT_CONFIGURED = lambda name='': f"Provider '{name}' is not configured" OAUTH_NOT_CONFIGURED = lambda name='': f"Provider '{name}' is not configured"

View file

@ -7,7 +7,9 @@ import pkgutil
import re import re
import shutil import shutil
import sys import sys
import threading
import traceback import traceback
from contextlib import nullcontext
from pathlib import Path from pathlib import Path
from typing import Any, Optional from typing import Any, Optional
from uuid import uuid4 from uuid import uuid4
@ -40,12 +42,13 @@ except ImportError:
print('dotenv not installed, skipping...') print('dotenv not installed, skipping...')
DOCKER = os.getenv('DOCKER', 'False').lower() == 'true' DOCKER = os.getenv('DOCKER', 'False').lower() == 'true'
USE_SLIM = os.getenv('USE_SLIM_DOCKER', 'False').lower() == 'true'
USE_CUDA = os.getenv('USE_CUDA_DOCKER', 'false') USE_CUDA = os.getenv('USE_CUDA_DOCKER', 'false')
DEVICE_TYPE = 'cpu' DEVICE_TYPE = 'cpu'
_cuda_error: Optional[str] = None _cuda_error: Optional[str] = None
if USE_CUDA.lower() == 'true': if not USE_SLIM and USE_CUDA.lower() == 'true':
try: try:
import torch # noqa: E402 import torch # noqa: E402
@ -57,7 +60,7 @@ if USE_CUDA.lower() == 'true':
os.environ['USE_CUDA_DOCKER'] = 'false' os.environ['USE_CUDA_DOCKER'] = 'false'
USE_CUDA = 'false' USE_CUDA = 'false'
if sys.platform == 'darwin' and DEVICE_TYPE == 'cpu': if not USE_SLIM and sys.platform == 'darwin' and DEVICE_TYPE == 'cpu':
try: try:
import torch # noqa: E402 import torch # noqa: E402
@ -66,6 +69,9 @@ if sys.platform == 'darwin' and DEVICE_TYPE == 'cpu':
except Exception: except Exception:
pass pass
# Torch MPS inference is not thread-safe and a concurrent call kills the whole process.
MPS_INFERENCE_LOCK = threading.Lock() if DEVICE_TYPE == 'mps' else nullcontext()
#################################### ####################################
# LOGGING # LOGGING
#################################### ####################################
@ -227,7 +233,7 @@ if FROM_INIT_PY:
# Check if the data directory exists in the package directory # Check if the data directory exists in the package directory
if DATA_DIR.exists() and DATA_DIR != NEW_DATA_DIR: if DATA_DIR.exists() and DATA_DIR != NEW_DATA_DIR:
log.info(f'Moving {DATA_DIR} to {NEW_DATA_DIR}') log.info('Moving %s to %s', DATA_DIR, NEW_DATA_DIR)
for item in DATA_DIR.iterdir(): for item in DATA_DIR.iterdir():
dest = NEW_DATA_DIR / item.name dest = NEW_DATA_DIR / item.name
if item.is_dir(): if item.is_dir():
@ -245,8 +251,6 @@ if FROM_INIT_PY:
STATIC_DIR = Path(os.getenv('STATIC_DIR', OPEN_WEBUI_DIR / 'static')) STATIC_DIR = Path(os.getenv('STATIC_DIR', OPEN_WEBUI_DIR / 'static'))
FONTS_DIR = Path(os.getenv('FONTS_DIR', OPEN_WEBUI_DIR / 'static' / 'fonts'))
FRONTEND_BUILD_DIR = Path(os.getenv('FRONTEND_BUILD_DIR', BASE_DIR / 'build')).resolve() FRONTEND_BUILD_DIR = Path(os.getenv('FRONTEND_BUILD_DIR', BASE_DIR / 'build')).resolve()
if FROM_INIT_PY: if FROM_INIT_PY:
@ -353,20 +357,23 @@ DATABASE_SQLITE_PRAGMA_MMAP_SIZE = os.getenv('DATABASE_SQLITE_PRAGMA_MMAP_SIZE',
# truncated. 67108864 ≈ 64 MB. Set to -1 for no limit (SQLite default). # truncated. 67108864 ≈ 64 MB. Set to -1 for no limit (SQLite default).
DATABASE_SQLITE_PRAGMA_JOURNAL_SIZE_LIMIT = os.getenv('DATABASE_SQLITE_PRAGMA_JOURNAL_SIZE_LIMIT', '67108864') DATABASE_SQLITE_PRAGMA_JOURNAL_SIZE_LIMIT = os.getenv('DATABASE_SQLITE_PRAGMA_JOURNAL_SIZE_LIMIT', '67108864')
DATABASE_USER_ACTIVE_STATUS_UPDATE_INTERVAL = os.getenv('DATABASE_USER_ACTIVE_STATUS_UPDATE_INTERVAL', None) # Seconds between presence writes per user per worker; keep under the 180s active-user window. 0 disables.
if DATABASE_USER_ACTIVE_STATUS_UPDATE_INTERVAL is not None: try:
try: DATABASE_USER_ACTIVE_STATUS_UPDATE_INTERVAL = float(os.getenv('DATABASE_USER_ACTIVE_STATUS_UPDATE_INTERVAL', '60'))
DATABASE_USER_ACTIVE_STATUS_UPDATE_INTERVAL = float(DATABASE_USER_ACTIVE_STATUS_UPDATE_INTERVAL) except ValueError:
except Exception: DATABASE_USER_ACTIVE_STATUS_UPDATE_INTERVAL = 60.0
DATABASE_USER_ACTIVE_STATUS_UPDATE_INTERVAL = 0.0
DATABASE_ENABLE_SESSION_SHARING = os.getenv('DATABASE_ENABLE_SESSION_SHARING', 'False').lower() == 'true' DATABASE_ENABLE_SESSION_SHARING = os.getenv('DATABASE_ENABLE_SESSION_SHARING', 'False').lower() == 'true'
ENABLE_PUBLIC_ACTIVE_USERS_COUNT = os.getenv('ENABLE_PUBLIC_ACTIVE_USERS_COUNT', 'True').lower() == 'true' ENABLE_PUBLIC_ACTIVE_USERS_COUNT = os.getenv('ENABLE_PUBLIC_ACTIVE_USERS_COUNT', 'True').lower() == 'true'
RESET_CONFIG_ON_START = os.getenv('RESET_CONFIG_ON_START', 'False').lower() == 'true' RESET_CONFIG_ON_START = os.getenv('RESET_CONFIG_ON_START', 'False').lower() == 'true'
ENABLE_REALTIME_CHAT_SAVE = os.getenv('ENABLE_REALTIME_CHAT_SAVE', 'False').lower() == 'true' ENABLE_REALTIME_CHAT_SAVE = os.getenv('ENABLE_REALTIME_CHAT_SAVE', 'False').lower() == 'true'
ENABLE_QUERIES_CACHE = os.getenv('ENABLE_QUERIES_CACHE', 'False').lower() == 'true' ENABLE_QUERIES_CACHE = os.getenv('ENABLE_QUERIES_CACHE', 'False').lower() == 'true'
ENABLE_ADMIN_CHAT_ACCESS = os.getenv('ENABLE_ADMIN_CHAT_ACCESS', 'True').lower() == 'true'
RAG_SYSTEM_CONTEXT = os.getenv('RAG_SYSTEM_CONTEXT', 'False').lower() == 'true' RAG_SYSTEM_CONTEXT = os.getenv('RAG_SYSTEM_CONTEXT', 'False').lower() == 'true'
# Empty by default: chunk metadata also holds internal bookkeeping (file hashes, collection names, scores).
RAG_SOURCE_METADATA_KEYS = [key.strip() for key in os.getenv('RAG_SOURCE_METADATA_KEYS', '').split(',') if key.strip()]
#################################### ####################################
# REDIS # REDIS
#################################### ####################################
@ -376,6 +383,19 @@ REDIS_CLUSTER = os.getenv('REDIS_CLUSTER', 'False').lower() == 'true'
REDIS_KEY_PREFIX = os.getenv('REDIS_KEY_PREFIX', 'open-webui') REDIS_KEY_PREFIX = os.getenv('REDIS_KEY_PREFIX', 'open-webui')
try:
REDIS_RESPONSE_STREAM_TTL = int(os.getenv('REDIS_RESPONSE_STREAM_TTL', '3600'))
except ValueError:
REDIS_RESPONSE_STREAM_TTL = 3600
# Seconds a task survives without a heartbeat. 0 disables expiry.
try:
REDIS_TASK_TTL = int(os.getenv('REDIS_TASK_TTL', '300'))
if REDIS_TASK_TTL != 0 and REDIS_TASK_TTL < 60:
REDIS_TASK_TTL = 300
except ValueError:
REDIS_TASK_TTL = 300
REDIS_SENTINEL_HOSTS = os.getenv('REDIS_SENTINEL_HOSTS', '') REDIS_SENTINEL_HOSTS = os.getenv('REDIS_SENTINEL_HOSTS', '')
REDIS_SENTINEL_PORT = os.getenv('REDIS_SENTINEL_PORT', '26379') REDIS_SENTINEL_PORT = os.getenv('REDIS_SENTINEL_PORT', '26379')
@ -443,6 +463,9 @@ try:
except (ValueError, TypeError): except (ValueError, TypeError):
UVICORN_WORKERS = 1 UVICORN_WORKERS = 1
# tiny delta-stream frames make per-frame websocket compression CPU-bound, allow opting out (true/false)
UVICORN_WS_PER_MESSAGE_DEFLATE = os.getenv('UVICORN_WS_PER_MESSAGE_DEFLATE', 'True').lower() == 'true'
#################################### ####################################
# WEBSOCKET SUPPORT # WEBSOCKET SUPPORT
#################################### ####################################
@ -499,6 +522,15 @@ try:
except ValueError: except ValueError:
WEBSOCKET_SERVER_PING_INTERVAL = 25 WEBSOCKET_SERVER_PING_INTERVAL = 25
WEBSOCKET_HEARTBEAT_INTERVAL = os.getenv('WEBSOCKET_HEARTBEAT_INTERVAL', '')
if WEBSOCKET_HEARTBEAT_INTERVAL == '':
WEBSOCKET_HEARTBEAT_INTERVAL = None
else:
try:
WEBSOCKET_HEARTBEAT_INTERVAL = min(max(int(WEBSOCKET_HEARTBEAT_INTERVAL), 5), 90)
except ValueError:
WEBSOCKET_HEARTBEAT_INTERVAL = 30
WEBSOCKET_EVENT_CALLER_TIMEOUT = os.getenv('WEBSOCKET_EVENT_CALLER_TIMEOUT', '') WEBSOCKET_EVENT_CALLER_TIMEOUT = os.getenv('WEBSOCKET_EVENT_CALLER_TIMEOUT', '')
if WEBSOCKET_EVENT_CALLER_TIMEOUT == '': if WEBSOCKET_EVENT_CALLER_TIMEOUT == '':
@ -571,6 +603,8 @@ def _parse_ssl_env(value: str) -> 'bool | _ssl.SSLContext':
REQUESTS_VERIFY = os.getenv('REQUESTS_VERIFY', 'True').lower() == 'true' REQUESTS_VERIFY = os.getenv('REQUESTS_VERIFY', 'True').lower() == 'true'
TAVILY_API_BASE_URL = os.getenv('TAVILY_API_BASE_URL', 'https://api.tavily.com').rstrip('/')
_aiohttp_timeout_raw = os.getenv('AIOHTTP_CLIENT_TIMEOUT', '') _aiohttp_timeout_raw = os.getenv('AIOHTTP_CLIENT_TIMEOUT', '')
try: try:
AIOHTTP_CLIENT_TIMEOUT = int(_aiohttp_timeout_raw) if _aiohttp_timeout_raw else None AIOHTTP_CLIENT_TIMEOUT = int(_aiohttp_timeout_raw) if _aiohttp_timeout_raw else None
@ -602,6 +636,18 @@ SEARXNG_CLIENT_KEY_FILE = os.getenv('SEARXNG_CLIENT_KEY_FILE', '').strip()
# When False (default), outbound HTTP requests do not follow 3xx redirects. # When False (default), outbound HTTP requests do not follow 3xx redirects.
AIOHTTP_CLIENT_ALLOW_REDIRECTS = os.getenv('AIOHTTP_CLIENT_ALLOW_REDIRECTS', 'False').lower() == 'true' AIOHTTP_CLIENT_ALLOW_REDIRECTS = os.getenv('AIOHTTP_CLIENT_ALLOW_REDIRECTS', 'False').lower() == 'true'
# Opt-in c-ares DNS resolution (aiodns). Off by default: c-ares breaks name
# resolution in some environments (#28013, #28215). Must run before any
# TCPConnector is constructed.
AIOHTTP_CLIENT_ASYNC_DNS_RESOLVER = os.getenv('AIOHTTP_CLIENT_ASYNC_DNS_RESOLVER', 'False').lower() == 'true'
if not AIOHTTP_CLIENT_ASYNC_DNS_RESOLVER:
import aiohttp
aiohttp.DefaultResolver = aiohttp.resolver.ThreadedResolver # for plugin code
aiohttp.resolver.DefaultResolver = aiohttp.resolver.ThreadedResolver
aiohttp.connector.DefaultResolver = aiohttp.resolver.ThreadedResolver
# Optional User-Agent override for outbound web-loader fetches. When set, # Optional User-Agent override for outbound web-loader fetches. When set,
# SafeWebBaseLoader sends this value instead of the default python-requests UA # SafeWebBaseLoader sends this value instead of the default python-requests UA
# which is aggressively blocked by Cloudflare, Wikipedia, and similar services. # which is aggressively blocked by Cloudflare, Wikipedia, and similar services.
@ -791,6 +837,16 @@ BYPASS_RETRIEVAL_ACCESS_CONTROL = os.getenv('BYPASS_RETRIEVAL_ACCESS_CONTROL', '
# for non-admin users. When False (default), unknown collection names are # for non-admin users. When False (default), unknown collection names are
# denied — closing the legacy unscoped namespace. # denied — closing the legacy unscoped namespace.
ENABLE_RETRIEVAL_UNSCOPED_COLLECTIONS = os.getenv('ENABLE_RETRIEVAL_UNSCOPED_COLLECTIONS', 'False').lower() == 'true' ENABLE_RETRIEVAL_UNSCOPED_COLLECTIONS = os.getenv('ENABLE_RETRIEVAL_UNSCOPED_COLLECTIONS', 'False').lower() == 'true'
# Falls back to the upload size limit, because a document cannot legitimately carry more metadata
# than the file itself is allowed to be. Left unbounded, a small archive that expands enormously
# during extraction can exhaust memory. RAG_FILE_MAX_SIZE is in MB.
RAG_METADATA_MAX_VALUE_CHARS = (
int(os.getenv('RAG_METADATA_MAX_VALUE_CHARS'))
if os.getenv('RAG_METADATA_MAX_VALUE_CHARS')
else ((int(os.getenv('RAG_FILE_MAX_SIZE', '0')) or 0) * 1024 * 1024 or None)
)
MINERU_MAX_MARKDOWN_BYTES = ( MINERU_MAX_MARKDOWN_BYTES = (
int(os.getenv('MINERU_MAX_MARKDOWN_BYTES')) if os.getenv('MINERU_MAX_MARKDOWN_BYTES') else None int(os.getenv('MINERU_MAX_MARKDOWN_BYTES')) if os.getenv('MINERU_MAX_MARKDOWN_BYTES') else None
) )
@ -798,7 +854,7 @@ MINERU_MAX_MARKDOWN_BYTES = (
# When enabled, skips pydub-based preprocessing (format conversion, compression, # When enabled, skips pydub-based preprocessing (format conversion, compression,
# and chunked splitting) before sending files to processing engines. Useful when # and chunked splitting) before sending files to processing engines. Useful when
# the upstream provider handles these steps or when ffmpeg is unavailable. # the upstream provider handles these steps or when ffmpeg is unavailable.
BYPASS_PYDUB_PREPROCESSING = os.getenv('BYPASS_PYDUB_PREPROCESSING', 'False').lower() == 'true' BYPASS_PYDUB_PREPROCESSING = USE_SLIM or os.getenv('BYPASS_PYDUB_PREPROCESSING', 'False').lower() == 'true'
# When disabled (default), the OpenAI catch-all proxy endpoint (/{path:path}) # When disabled (default), the OpenAI catch-all proxy endpoint (/{path:path})
# is blocked. Enable only if you need direct passthrough to upstream OpenAI- # is blocked. Enable only if you need direct passthrough to upstream OpenAI-
@ -888,10 +944,18 @@ if LICENSE_PUBLIC_KEY:
# WEBUI Identity # WEBUI Identity
#################################### ####################################
# LICENSE covers this Open WebUI branding surface, including name, logo,
# visual, textual, symbolic identifiers, metadata, and surrounding UI.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
WEBUI_NAME = os.getenv('WEBUI_NAME', 'Open WebUI') WEBUI_NAME = os.getenv('WEBUI_NAME', 'Open WebUI')
if WEBUI_NAME != 'Open WebUI': if WEBUI_NAME != 'Open WebUI':
WEBUI_NAME += ' (Open WebUI)' WEBUI_NAME += ' (Open WebUI)'
# LICENSE covers this Open WebUI branding surface, including this favicon
# and any visual, textual, or symbolic identifiers it preserves.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
WEBUI_FAVICON_URL = 'https://openwebui.com/favicon.png' WEBUI_FAVICON_URL = 'https://openwebui.com/favicon.png'
WEBUI_BUILD_HASH = os.getenv('WEBUI_BUILD_HASH', 'dev-build') WEBUI_BUILD_HASH = os.getenv('WEBUI_BUILD_HASH', 'dev-build')
TRUSTED_SIGNATURE_KEY = os.getenv('TRUSTED_SIGNATURE_KEY', '') TRUSTED_SIGNATURE_KEY = os.getenv('TRUSTED_SIGNATURE_KEY', '')
@ -930,6 +994,7 @@ FORWARD_USER_INFO_HEADER_USER_NAME = os.getenv('FORWARD_USER_INFO_HEADER_USER_NA
FORWARD_USER_INFO_HEADER_USER_ID = os.getenv('FORWARD_USER_INFO_HEADER_USER_ID', 'X-OpenWebUI-User-Id') FORWARD_USER_INFO_HEADER_USER_ID = os.getenv('FORWARD_USER_INFO_HEADER_USER_ID', 'X-OpenWebUI-User-Id')
FORWARD_USER_INFO_HEADER_USER_EMAIL = os.getenv('FORWARD_USER_INFO_HEADER_USER_EMAIL', 'X-OpenWebUI-User-Email') FORWARD_USER_INFO_HEADER_USER_EMAIL = os.getenv('FORWARD_USER_INFO_HEADER_USER_EMAIL', 'X-OpenWebUI-User-Email')
FORWARD_USER_INFO_HEADER_USER_ROLE = os.getenv('FORWARD_USER_INFO_HEADER_USER_ROLE', 'X-OpenWebUI-User-Role') FORWARD_USER_INFO_HEADER_USER_ROLE = os.getenv('FORWARD_USER_INFO_HEADER_USER_ROLE', 'X-OpenWebUI-User-Role')
FORWARD_USER_INFO_HEADER_AUTH_TYPE = os.getenv('FORWARD_USER_INFO_HEADER_AUTH_TYPE', 'X-OpenWebUI-Auth-Type')
FORWARD_SESSION_INFO_HEADER_MESSAGE_ID = os.getenv('FORWARD_SESSION_INFO_HEADER_MESSAGE_ID', 'X-OpenWebUI-Message-Id') FORWARD_SESSION_INFO_HEADER_MESSAGE_ID = os.getenv('FORWARD_SESSION_INFO_HEADER_MESSAGE_ID', 'X-OpenWebUI-Message-Id')
FORWARD_SESSION_INFO_HEADER_CHAT_ID = os.getenv('FORWARD_SESSION_INFO_HEADER_CHAT_ID', 'X-OpenWebUI-Chat-Id') FORWARD_SESSION_INFO_HEADER_CHAT_ID = os.getenv('FORWARD_SESSION_INFO_HEADER_CHAT_ID', 'X-OpenWebUI-Chat-Id')
@ -948,6 +1013,10 @@ except ValueError:
# Progressive Web App # Progressive Web App
#################################### ####################################
# LICENSE covers this install-time Open WebUI branding surface, including
# names, logos, manifests, metadata, and surrounding UI.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
EXTERNAL_PWA_MANIFEST_URL = os.getenv('EXTERNAL_PWA_MANIFEST_URL', None) EXTERNAL_PWA_MANIFEST_URL = os.getenv('EXTERNAL_PWA_MANIFEST_URL', None)
#################################### ####################################
@ -984,6 +1053,13 @@ ENABLE_CHAT_RESPONSE_BASE64_IMAGE_URL_CONVERSION = (
) )
ENABLE_API_OUTLET_FILTERS = os.getenv('ENABLE_API_OUTLET_FILTERS', 'True').lower() == 'true' ENABLE_API_OUTLET_FILTERS = os.getenv('ENABLE_API_OUTLET_FILTERS', 'True').lower() == 'true'
# Opt in to CPython's in-place string append optimization for streamed responses.
# Off by default for a staged rollout. Only a host already out of memory can lose
# text here; the default path (a full copy per chunk) raises there too.
ENABLE_CHAT_RESPONSE_STREAM_INPLACE_APPEND = (
os.getenv('ENABLE_CHAT_RESPONSE_STREAM_INPLACE_APPEND', 'False').lower() == 'true'
)
# When enabled, uses a hardcoded extension-to-MIME dictionary as a last-resort # When enabled, uses a hardcoded extension-to-MIME dictionary as a last-resort
# fallback when both mimetypes.guess_type() and file.meta.content_type fail to # fallback when both mimetypes.guess_type() and file.meta.content_type fail to
# determine the content type. This can help on minimal container images (e.g. # determine the content type. This can help on minimal container images (e.g.

View file

@ -411,6 +411,11 @@ class EventDefinitions(BaseModel):
description='Retrieval content was processed.', description='Retrieval content was processed.',
message='Retrieval Content processed', message='Retrieval Content processed',
) )
RETRIEVAL_CONTENT_PROCESS_FAILED: EventDefinition = EventDefinition(
name='retrieval.content.process_failed',
description='Retrieval content processing failed.',
message='Retrieval Content process failed',
)
RETRIEVAL_COLLECTION_DELETED: EventDefinition = EventDefinition( RETRIEVAL_COLLECTION_DELETED: EventDefinition = EventDefinition(
name='retrieval.collection.deleted', name='retrieval.collection.deleted',
description='A retrieval collection was deleted.', description='A retrieval collection was deleted.',
@ -666,6 +671,7 @@ NOTIFICATION_EVENTS = (
EVENTS.CHAT_FAILED.name, EVENTS.CHAT_FAILED.name,
EVENTS.CHANNEL_MESSAGE.name, EVENTS.CHANNEL_MESSAGE.name,
EVENTS.CALENDAR_ALERT.name, EVENTS.CALENDAR_ALERT.name,
EVENTS.RETRIEVAL_CONTENT_PROCESS_FAILED.name,
) )
@ -1026,6 +1032,9 @@ def build_event(
async def dispatch_webhook_event(app: Any, event: Event) -> None: async def dispatch_webhook_event(app: Any, event: Event) -> None:
# LICENSE covers this Open WebUI webhook identifier.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
name = getattr(getattr(app, 'state', None), 'WEBUI_NAME', 'Open WebUI') name = getattr(getattr(app, 'state', None), 'WEBUI_NAME', 'Open WebUI')
subject = event.subject or {} subject = event.subject or {}
subject_id = subject.get('id') subject_id = subject.get('id')
@ -1077,6 +1086,20 @@ class NotificationEventSink:
schedule_notification_dispatch(app, event) schedule_notification_dispatch(app, event)
class SocketSessionEventSink:
async def handle_event(self, app: Any, event: Event, request: Any | None = None) -> None:
if event.event not in {EVENTS.USER_DELETED.name, EVENTS.USER_ROLE_UPDATED.name}:
return
subject = event.subject or {}
if subject.get('type') != 'user' or not subject.get('id'):
return
from open_webui.socket.main import disconnect_user_sessions
await disconnect_user_sessions(str(subject['id']))
async def dispatch_event_functions( async def dispatch_event_functions(
app: Any, event: Event, request: Any | None = None, extra_function_ids: list[str] | None = None app: Any, event: Event, request: Any | None = None, extra_function_ids: list[str] | None = None
) -> None: ) -> None:
@ -1145,7 +1168,7 @@ class EventFunctionSink:
schedule_event_function_dispatch(app, event, request) schedule_event_function_dispatch(app, event, request)
EVENT_SINKS = [EventFunctionSink(), WebhookEventSink(), NotificationEventSink()] EVENT_SINKS = [SocketSessionEventSink(), EventFunctionSink(), WebhookEventSink(), NotificationEventSink()]
async def publish_event( async def publish_event(

View file

@ -1,6 +1,5 @@
import asyncio import asyncio
import inspect import inspect
import json
import logging import logging
import sys import sys
from typing import AsyncGenerator, Generator, Iterator from typing import AsyncGenerator, Generator, Iterator
@ -29,6 +28,7 @@ from open_webui.socket.main import (
get_event_emitter, get_event_emitter,
) )
from open_webui.utils.access_control import check_model_access from open_webui.utils.access_control import check_model_access
from open_webui.utils.json_codec import JSONCodec
from open_webui.utils.misc import ( from open_webui.utils.misc import (
add_or_update_system_message, add_or_update_system_message,
get_last_user_message, get_last_user_message,
@ -100,7 +100,7 @@ async def get_function_models(request):
log.exception(e) log.exception(e)
sub_pipes = [] sub_pipes = []
log.debug(f"get_function_models: function '{pipe.id}' is a manifold of {sub_pipes}") log.debug("get_function_models: function '%s' is a manifold of %s", pipe.id, sub_pipes)
for p in sub_pipes: for p in sub_pipes:
sub_pipe_id = f'{pipe.id}.{p["id"]}' sub_pipe_id = f'{pipe.id}.{p["id"]}'
@ -126,7 +126,10 @@ async def get_function_models(request):
pipe_flag = {'type': 'pipe'} pipe_flag = {'type': 'pipe'}
log.debug( log.debug(
f"get_function_models: function '{pipe.id}' is a single pipe {{ 'id': {pipe.id}, 'name': {pipe.name} }}" "get_function_models: function '%s' is a single pipe { 'id': %s, 'name': %s }",
pipe.id,
pipe.id,
pipe.name,
) )
pipe_models.append( pipe_models.append(
@ -170,7 +173,7 @@ async def generate_function_chat_completion(request, form_data, user, models: di
line = line.model_dump_json() line = line.model_dump_json()
line = f'data: {line}' line = f'data: {line}'
if isinstance(line, dict): if isinstance(line, dict):
line = f'data: {json.dumps(line)}' line = f'data: {JSONCodec.dumps(line)}'
try: try:
line = line.decode('utf-8') line = line.decode('utf-8')
@ -181,7 +184,7 @@ async def generate_function_chat_completion(request, form_data, user, models: di
return f'{line}\n\n' return f'{line}\n\n'
else: else:
line = openai_chat_chunk_message_template(form_data['model'], line) line = openai_chat_chunk_message_template(form_data['model'], line)
return f'data: {json.dumps(line)}\n\n' return f'data: {JSONCodec.dumps(line)}\n\n'
def get_pipe_id(form_data: dict) -> str: def get_pipe_id(form_data: dict) -> str:
pipe_id = form_data['model'] pipe_id = form_data['model']
@ -209,6 +212,9 @@ async def generate_function_chat_completion(request, form_data, user, models: di
return params return params
# Set server-side by utils/chat.py, never by client input. Mirrors the routers.
bypass_system_prompt = getattr(request.state, 'bypass_system_prompt', False)
# Copy so the base-model substitution below doesn't leak into the caller's # Copy so the base-model substitution below doesn't leak into the caller's
# payload, which the tool-call continuation re-submits. Mirrors the routers. # payload, which the tool-call continuation re-submits. Mirrors the routers.
form_data = {**form_data} form_data = {**form_data}
@ -288,7 +294,8 @@ async def generate_function_chat_completion(request, form_data, user, models: di
if params: if params:
system = params.pop('system', None) system = params.pop('system', None)
form_data = apply_model_params_to_body_openai(params, form_data) form_data = apply_model_params_to_body_openai(params, form_data)
form_data = await apply_system_prompt_to_body(system, form_data, metadata, user) if not bypass_system_prompt:
form_data = await apply_system_prompt_to_body(system, form_data, metadata, user)
pipe_id = get_pipe_id(form_data) pipe_id = get_pipe_id(form_data)
function_module = await get_function_module_by_id(request, pipe_id) function_module = await get_function_module_by_id(request, pipe_id)
@ -308,17 +315,17 @@ async def generate_function_chat_completion(request, form_data, user, models: di
yield data yield data
return return
if isinstance(res, dict): if isinstance(res, dict):
yield f'data: {json.dumps(res)}\n\n' yield f'data: {JSONCodec.dumps(res)}\n\n'
return return
except Exception as e: except Exception as e:
log.error(f'Error: {e}') log.error(f'Error: {e}')
yield f'data: {json.dumps({"error": {"detail": str(e)}})}\n\n' yield f'data: {JSONCodec.dumps({"error": {"detail": str(e)}})}\n\n'
return return
if isinstance(res, str): if isinstance(res, str):
message = openai_chat_chunk_message_template(form_data['model'], res) message = openai_chat_chunk_message_template(form_data['model'], res)
yield f'data: {json.dumps(message)}\n\n' yield f'data: {JSONCodec.dumps(message)}\n\n'
if isinstance(res, Iterator): if isinstance(res, Iterator):
for line in res: for line in res:
@ -330,7 +337,7 @@ async def generate_function_chat_completion(request, form_data, user, models: di
finish_message = openai_chat_chunk_message_template(form_data['model'], '') finish_message = openai_chat_chunk_message_template(form_data['model'], '')
finish_message['choices'][0]['finish_reason'] = 'stop' finish_message['choices'][0]['finish_reason'] = 'stop'
yield f'data: {json.dumps(finish_message)}\n\n' yield f'data: {JSONCodec.dumps(finish_message)}\n\n'
yield 'data: [DONE]' yield 'data: [DONE]'
return StreamingResponse(stream_content(), media_type='text/event-stream') return StreamingResponse(stream_content(), media_type='text/event-stream')

View file

@ -1,8 +1,8 @@
from __future__ import annotations from __future__ import annotations
import json
import logging import logging
import os import os
import re
import sys import sys
from contextlib import asynccontextmanager, contextmanager from contextlib import asynccontextmanager, contextmanager
from datetime import datetime, timedelta, timezone from datetime import datetime, timedelta, timezone
@ -27,7 +27,9 @@ from open_webui.env import (
DATABASE_URL, DATABASE_URL,
ENABLE_DB_MIGRATIONS, ENABLE_DB_MIGRATIONS,
OPEN_WEBUI_DIR, OPEN_WEBUI_DIR,
USE_SLIM,
) )
from open_webui.utils.json_codec import JSONCodec
from sqlalchemy import Dialect, MetaData, create_engine, event, types from sqlalchemy import Dialect, MetaData, create_engine, event, types
from sqlalchemy.engine.url import make_url from sqlalchemy.engine.url import make_url
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker, create_async_engine from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker, create_async_engine
@ -124,23 +126,34 @@ class JSONField(types.TypeDecorator): # TEXT-backed JSON storage
"""Store arbitrary Python objects as JSON-encoded TEXT. """Store arbitrary Python objects as JSON-encoded TEXT.
Used instead of native JSON columns for portability across SQLite and Used instead of native JSON columns for portability across SQLite and
PostgreSQL. Values are serialized with ``json.dumps`` on write and PostgreSQL. Values are serialized with ``JSONCodec.dumps`` on write and
deserialized with ``json.loads`` on read. deserialized with ``JSONCodec.loads`` on read.
""" """
impl = types.UnicodeText impl = types.UnicodeText
cache_ok = True cache_ok = True
def process_bind_param(self, value: _T | None, dialect: Dialect) -> Any: def process_bind_param(self, value: _T | None, dialect: Dialect) -> Any:
return json.dumps(value) if value is not None else None return JSONCodec.dumps(value) if value is not None else None
def process_result_value(self, value: _T | None, dialect: Dialect) -> Any: def process_result_value(self, value: _T | None, dialect: Dialect) -> Any:
return json.loads(value) if value is not None else None return JSONCodec.loads(value) if value is not None else None
def copy(self, **kwargs: Any) -> Self: def copy(self, **kwargs: Any) -> Self:
return JSONField(length=self.impl.length) return JSONField(length=self.impl.length)
if USE_SLIM:
if make_url(DATABASE_URL).get_backend_name() not in ('sqlite', 'postgresql', 'postgres'):
raise ValueError(
'Slim requires SQLite or PostgreSQL for DATABASE_URL. Use the standard image for other databases.'
)
if DATABASE_ENABLE_IAM_TOKEN_AUTH:
raise ValueError(
'AWS RDS IAM authentication requires the standard image. Slim supports PostgreSQL database credentials.'
)
# Normalize SSL params from the URL once; the sync engine needs them # Normalize SSL params from the URL once; the sync engine needs them
# reattached in canonical libpq form for psycopg2. # reattached in canonical libpq form for psycopg2.
_url_without_ssl, _ssl_dict = extract_ssl_params_from_url(DATABASE_URL) _url_without_ssl, _ssl_dict = extract_ssl_params_from_url(DATABASE_URL)
@ -202,6 +215,20 @@ def enable_iam_token_auth(connectable) -> None:
return return
engine = getattr(connectable, 'sync_engine', connectable) engine = getattr(connectable, 'sync_engine', connectable)
url = engine.url
auth = _rds_iam_token_auth
# The token is bound to one host/port/user pair; leave other databases on their own credentials.
if (url.host, url.port or 5432, url.username) != (auth.host, auth.port, auth.username):
log.warning(
'AWS RDS IAM token auth not applied to %s: the token is issued for %s@%s:%s, '
'so this connection uses the password from its own URL',
url.render_as_string(hide_password=True),
auth.username,
auth.host,
auth.port,
)
return
if not event.contains(engine, 'do_connect', _set_iam_token_password): if not event.contains(engine, 'do_connect', _set_iam_token_password):
event.listen(engine, 'do_connect', _set_iam_token_password) event.listen(engine, 'do_connect', _set_iam_token_password)
@ -232,6 +259,27 @@ def _make_async_url(url: str) -> str:
return url return url
def _json_codec_kwargs(kwargs: dict) -> dict:
"""Default an engine to JSONCodec for native ``JSON`` columns.
Unlike ``JSONField``, those serialize through the engine, which otherwise uses
stdlib ``json``. With ``ENABLE_ORJSON`` off JSONCodec is stdlib ``json`` anyway.
"""
kwargs.setdefault('json_serializer', JSONCodec.dumps)
kwargs.setdefault('json_deserializer', JSONCodec.loads)
return kwargs
def _create_engine(*args, **kwargs):
"""``create_engine`` with the app JSON codec wired in."""
return create_engine(*args, **_json_codec_kwargs(kwargs))
def _create_async_engine(*args, **kwargs):
"""``create_async_engine`` with the app JSON codec wired in."""
return create_async_engine(*args, **_json_codec_kwargs(kwargs))
# ============================================================ # ============================================================
# SYNC ENGINE (used only for: startup migrations, config loading, # SYNC ENGINE (used only for: startup migrations, config loading,
# Alembic, peewee migration, health checks) # Alembic, peewee migration, health checks)
@ -260,7 +308,7 @@ if SQLALCHEMY_DATABASE_URL.startswith('sqlite+sqlcipher://'):
# in the native sqlcipher3 C library. Use NullPool by default for safety, # in the native sqlcipher3 C library. Use NullPool by default for safety,
# or QueuePool if DATABASE_POOL_SIZE is explicitly configured. # or QueuePool if DATABASE_POOL_SIZE is explicitly configured.
if isinstance(DATABASE_POOL_SIZE, int) and DATABASE_POOL_SIZE > 0: if isinstance(DATABASE_POOL_SIZE, int) and DATABASE_POOL_SIZE > 0:
engine = create_engine( engine = _create_engine(
'sqlite://', 'sqlite://',
creator=create_sqlcipher_connection, creator=create_sqlcipher_connection,
pool_size=DATABASE_POOL_SIZE, pool_size=DATABASE_POOL_SIZE,
@ -272,7 +320,7 @@ if SQLALCHEMY_DATABASE_URL.startswith('sqlite+sqlcipher://'):
echo=False, echo=False,
) )
else: else:
engine = create_engine( engine = _create_engine(
'sqlite://', 'sqlite://',
creator=create_sqlcipher_connection, creator=create_sqlcipher_connection,
poolclass=NullPool, poolclass=NullPool,
@ -282,10 +330,49 @@ if SQLALCHEMY_DATABASE_URL.startswith('sqlite+sqlcipher://'):
log.info('Connected to encrypted SQLite database using SQLCipher') log.info('Connected to encrypted SQLite database using SQLCipher')
elif 'sqlite' in SQLALCHEMY_DATABASE_URL: elif 'sqlite' in SQLALCHEMY_DATABASE_URL:
engine = create_engine(SQLALCHEMY_DATABASE_URL, connect_args={'check_same_thread': False}) engine = _create_engine(SQLALCHEMY_DATABASE_URL, connect_args={'check_same_thread': False})
def _apply_sqlite_pragmas(dbapi_connection): def _apply_sqlite_pragmas(dbapi_connection):
"""Apply all configured SQLite PRAGMAs to a raw DBAPI connection.""" """Apply all configured SQLite PRAGMAs to a raw DBAPI connection."""
# SQLite LIKE folds ASCII only; SQLAlchemy SQLite ILIKE compiles to lower(x) LIKE lower(?).
compiled_patterns = {}
def like(pattern, value, escape=None):
if pattern is None or value is None:
return None
pattern = str(pattern).lower()
escape = str(escape).lower() if escape is not None else None
key = (pattern, escape)
compiled = compiled_patterns.get(key)
if compiled is False:
return False
if compiled is None:
regex = []
escaped = False
for char in pattern:
if escape and not escaped and char == escape:
escaped = True
continue
regex.append(
'.*' if not escaped and char == '%' else '.' if not escaped and char == '_' else re.escape(char)
)
escaped = False
if escaped:
compiled = False
if len(compiled_patterns) >= 512:
compiled_patterns.clear()
compiled_patterns[key] = compiled
return False
compiled = re.compile(''.join(regex), re.DOTALL)
if len(compiled_patterns) >= 512:
compiled_patterns.clear()
compiled_patterns[key] = compiled
return compiled.fullmatch(str(value).lower()) is not None
dbapi_connection.create_function('like', 2, like, deterministic=True)
dbapi_connection.create_function('like', 3, like, deterministic=True)
cursor = dbapi_connection.cursor() cursor = dbapi_connection.cursor()
if DATABASE_ENABLE_SQLITE_WAL: if DATABASE_ENABLE_SQLITE_WAL:
cursor.execute('PRAGMA journal_mode=WAL') cursor.execute('PRAGMA journal_mode=WAL')
@ -314,7 +401,7 @@ elif 'sqlite' in SQLALCHEMY_DATABASE_URL:
else: else:
if isinstance(DATABASE_POOL_SIZE, int): if isinstance(DATABASE_POOL_SIZE, int):
if DATABASE_POOL_SIZE > 0: if DATABASE_POOL_SIZE > 0:
engine = create_engine( engine = _create_engine(
SQLALCHEMY_DATABASE_URL, SQLALCHEMY_DATABASE_URL,
pool_size=DATABASE_POOL_SIZE, pool_size=DATABASE_POOL_SIZE,
max_overflow=DATABASE_POOL_MAX_OVERFLOW, max_overflow=DATABASE_POOL_MAX_OVERFLOW,
@ -324,9 +411,9 @@ else:
poolclass=QueuePool, poolclass=QueuePool,
) )
else: else:
engine = create_engine(SQLALCHEMY_DATABASE_URL, pool_pre_ping=True, poolclass=NullPool) engine = _create_engine(SQLALCHEMY_DATABASE_URL, pool_pre_ping=True, poolclass=NullPool)
else: else:
engine = create_engine(SQLALCHEMY_DATABASE_URL, pool_pre_ping=True) engine = _create_engine(SQLALCHEMY_DATABASE_URL, pool_pre_ping=True)
enable_iam_token_auth(engine) enable_iam_token_auth(engine)
@ -373,7 +460,7 @@ if 'sqlite' in ASYNC_SQLALCHEMY_DATABASE_URL:
# No pool_pre_ping: a local SQLite file cannot drop connections, and the # No pool_pre_ping: a local SQLite file cannot drop connections, and the
# ping costs a worker-thread hop plus a SELECT 1 on every checkout. # ping costs a worker-thread hop plus a SELECT 1 on every checkout.
_sqlite_pool_size = DATABASE_POOL_SIZE if isinstance(DATABASE_POOL_SIZE, int) and DATABASE_POOL_SIZE > 0 else 512 _sqlite_pool_size = DATABASE_POOL_SIZE if isinstance(DATABASE_POOL_SIZE, int) and DATABASE_POOL_SIZE > 0 else 512
async_engine = create_async_engine( async_engine = _create_async_engine(
ASYNC_SQLALCHEMY_DATABASE_URL, ASYNC_SQLALCHEMY_DATABASE_URL,
connect_args={'check_same_thread': False}, connect_args={'check_same_thread': False},
pool_size=_sqlite_pool_size, pool_size=_sqlite_pool_size,
@ -387,7 +474,7 @@ if 'sqlite' in ASYNC_SQLALCHEMY_DATABASE_URL:
else: else:
if isinstance(DATABASE_POOL_SIZE, int): if isinstance(DATABASE_POOL_SIZE, int):
if DATABASE_POOL_SIZE > 0: if DATABASE_POOL_SIZE > 0:
async_engine = create_async_engine( async_engine = _create_async_engine(
ASYNC_SQLALCHEMY_DATABASE_URL, ASYNC_SQLALCHEMY_DATABASE_URL,
pool_size=DATABASE_POOL_SIZE, pool_size=DATABASE_POOL_SIZE,
max_overflow=DATABASE_POOL_MAX_OVERFLOW, max_overflow=DATABASE_POOL_MAX_OVERFLOW,
@ -396,13 +483,13 @@ else:
pool_pre_ping=True, pool_pre_ping=True,
) )
else: else:
async_engine = create_async_engine( async_engine = _create_async_engine(
ASYNC_SQLALCHEMY_DATABASE_URL, ASYNC_SQLALCHEMY_DATABASE_URL,
pool_pre_ping=True, pool_pre_ping=True,
poolclass=NullPool, poolclass=NullPool,
) )
else: else:
async_engine = create_async_engine( async_engine = _create_async_engine(
ASYNC_SQLALCHEMY_DATABASE_URL, ASYNC_SQLALCHEMY_DATABASE_URL,
pool_pre_ping=True, pool_pre_ping=True,
) )

View file

@ -1,17 +1,19 @@
from __future__ import annotations from __future__ import annotations
import asyncio import asyncio
import json import copy
import logging import logging
import mimetypes import mimetypes
import os import os
import sys import sys
import time import time
from concurrent.futures import ThreadPoolExecutor
from contextlib import asynccontextmanager from contextlib import asynccontextmanager
from uuid import uuid4 from uuid import uuid4
import aiohttp import aiohttp
import anyio.to_thread import anyio.to_thread
from cryptography.fernet import InvalidToken
from fastapi import ( from fastapi import (
Depends, Depends,
FastAPI, FastAPI,
@ -64,6 +66,7 @@ from open_webui.config import (
ONEDRIVE_SHAREPOINT_URL, ONEDRIVE_SHAREPOINT_URL,
STATIC_DIR, STATIC_DIR,
THREAD_POOL_SIZE, THREAD_POOL_SIZE,
THREAD_POOL_THREAD_NAME_PREFIX,
WEBUI_AUTH, WEBUI_AUTH,
WEBUI_NAME, WEBUI_NAME,
async_reset_config, async_reset_config,
@ -71,7 +74,9 @@ from open_webui.config import (
seed_registered_defaults, seed_registered_defaults,
) )
from open_webui.constants import ERROR_MESSAGES, TASKS from open_webui.constants import ERROR_MESSAGES, TASKS
from open_webui.utils.recurrence import RecurrenceEvaluationTimeout
from open_webui.env import ( from open_webui.env import (
USE_SLIM,
AIOHTTP_CLIENT_SESSION_SSL, AIOHTTP_CLIENT_SESSION_SSL,
AUDIT_EXCLUDED_PATHS, AUDIT_EXCLUDED_PATHS,
AUDIT_INCLUDED_PATHS, AUDIT_INCLUDED_PATHS,
@ -83,19 +88,19 @@ from open_webui.env import (
ENABLE_COMPRESSION_MIDDLEWARE, ENABLE_COMPRESSION_MIDDLEWARE,
ENABLE_CUSTOM_MODEL_FALLBACK, ENABLE_CUSTOM_MODEL_FALLBACK,
ENABLE_EASTER_EGGS, ENABLE_EASTER_EGGS,
ENABLE_PLUGINS,
EXTERNAL_PWA_MANIFEST_URL,
# OAuth Back-Channel Logout # OAuth Back-Channel Logout
ENABLE_OAUTH_BACKCHANNEL_LOGOUT, ENABLE_OAUTH_BACKCHANNEL_LOGOUT,
ENABLE_OTEL, ENABLE_OTEL,
ENABLE_PLUGINS,
ENABLE_PUBLIC_ACTIVE_USERS_COUNT, ENABLE_PUBLIC_ACTIVE_USERS_COUNT,
ENABLE_PYODIDE_FILE_PERSISTENCE,
# SCIM # SCIM
ENABLE_SCIM, ENABLE_SCIM,
ENABLE_SIGNUP_PASSWORD_CONFIRMATION, ENABLE_SIGNUP_PASSWORD_CONFIRMATION,
ENABLE_STAR_SESSIONS_MIDDLEWARE, ENABLE_STAR_SESSIONS_MIDDLEWARE,
ENABLE_PYODIDE_FILE_PERSISTENCE,
ENABLE_VERSION_UPDATE_CHECK, ENABLE_VERSION_UPDATE_CHECK,
ENABLE_WEBSOCKET_SUPPORT, ENABLE_WEBSOCKET_SUPPORT,
EXTERNAL_PWA_MANIFEST_URL,
GLOBAL_LOG_LEVEL, GLOBAL_LOG_LEVEL,
INSTANCE_ID, INSTANCE_ID,
LICENSE_KEY, LICENSE_KEY,
@ -103,11 +108,14 @@ from open_webui.env import (
MAX_BODY_LOG_SIZE, MAX_BODY_LOG_SIZE,
# Redis # Redis
REDIS_KEY_PREFIX, REDIS_KEY_PREFIX,
REDIS_TASK_TTL,
REDIS_URL, REDIS_URL,
RESET_CONFIG_ON_START, RESET_CONFIG_ON_START,
SAFE_MODE, SAFE_MODE,
SCIM_TOKEN, SCIM_TOKEN,
VERSION, VERSION,
WEBSOCKET_HEARTBEAT_INTERVAL,
WEBSOCKET_MANAGER,
# Admin Account Runtime Creation # Admin Account Runtime Creation
WEBUI_ADMIN_EMAIL, WEBUI_ADMIN_EMAIL,
WEBUI_ADMIN_NAME, WEBUI_ADMIN_NAME,
@ -121,12 +129,14 @@ from open_webui.env import (
from open_webui.events import ( from open_webui.events import (
EVENTS, EVENTS,
delete_event_webhook, delete_event_webhook,
get_event_catalog as get_event_catalog_items,
get_event_webhooks, get_event_webhooks,
migrate_legacy_webhook_config, migrate_legacy_webhook_config,
publish_event, publish_event,
upsert_event_webhook, upsert_event_webhook,
) )
from open_webui.events import (
get_event_catalog as get_event_catalog_items,
)
from open_webui.internal.db import engine, get_async_session from open_webui.internal.db import engine, get_async_session
from open_webui.models.access_grants import AccessGrants from open_webui.models.access_grants import AccessGrants
from open_webui.models.channels import Channels from open_webui.models.channels import Channels
@ -134,7 +144,7 @@ from open_webui.models.chats import ChatForm, Chats
from open_webui.models.config import Config from open_webui.models.config import Config
from open_webui.models.functions import Functions from open_webui.models.functions import Functions
from open_webui.models.messages import Messages from open_webui.models.messages import Messages
from open_webui.models.models import Models from open_webui.models.models import Models, normalize_model_tags
from open_webui.models.users import Users from open_webui.models.users import Users
from open_webui.routers import ( from open_webui.routers import (
analytics, analytics,
@ -154,8 +164,8 @@ from open_webui.routers import (
knowledge, knowledge,
memories, memories,
models, models,
notifications,
notes, notes,
notifications,
ollama, ollama,
openai, openai,
pipelines, pipelines,
@ -182,6 +192,7 @@ from open_webui.socket.main import (
get_user_id_from_session_pool, get_user_id_from_session_pool,
periodic_session_pool_cleanup, periodic_session_pool_cleanup,
periodic_usage_pool_cleanup, periodic_usage_pool_cleanup,
redis_event_listener,
) )
from open_webui.socket.main import ( from open_webui.socket.main import (
app as socket_app, app as socket_app,
@ -193,18 +204,15 @@ from open_webui.tasks import (
list_task_ids_by_item_id, list_task_ids_by_item_id,
list_tasks, list_tasks,
redis_task_command_listener, redis_task_command_listener,
redis_task_heartbeat,
stop_item_tasks, stop_item_tasks,
stop_task, stop_task,
) # Import from tasks.py ) # Import from tasks.py
from open_webui.utils import logger from open_webui.utils import logger
from open_webui.utils.access_control import has_permission from open_webui.utils.access_control import has_permission
from open_webui.utils.access_control.folders import has_folder_write_access
from open_webui.utils.actions import chat_action as chat_action_handler from open_webui.utils.actions import chat_action as chat_action_handler
from open_webui.utils.asgi_middleware import ( from open_webui.utils.asgi_middleware import AppHTTPMiddleware
AuthTokenMiddleware,
CommitSessionMiddleware,
RedirectMiddleware,
WebsocketUpgradeGuardMiddleware,
)
from open_webui.utils.audit import AuditLevel, AuditLoggingMiddleware from open_webui.utils.audit import AuditLevel, AuditLoggingMiddleware
from open_webui.utils.auth import ( from open_webui.utils.auth import (
create_admin_user, create_admin_user,
@ -229,14 +237,17 @@ from open_webui.utils.chat_variables import (
normalize_chat_variables, normalize_chat_variables,
) )
from open_webui.utils.embeddings import generate_embeddings from open_webui.utils.embeddings import generate_embeddings
from open_webui.utils.json_codec import JSONCodec
from open_webui.utils.json_response import apply_orjson_http_json from open_webui.utils.json_response import apply_orjson_http_json
from open_webui.utils.logger import start_logger from open_webui.utils.logger import start_logger
from open_webui.utils.middleware import ( from open_webui.utils.middleware import (
background_tasks_handler, background_tasks_handler,
build_chat_response_context, build_chat_response_context,
drain_approved_tool_calls,
process_chat_payload, process_chat_payload,
process_chat_response, process_chat_response,
) )
from open_webui.utils.misc import get_response_error_detail, merge_model_params
from open_webui.utils.model_ids import strip_provider_model_prefix from open_webui.utils.model_ids import strip_provider_model_prefix
from open_webui.utils.models import ( from open_webui.utils.models import (
check_model_access, check_model_access,
@ -258,8 +269,12 @@ from open_webui.utils.oauth import (
) )
from open_webui.utils.plugin import install_tool_and_function_dependencies from open_webui.utils.plugin import install_tool_and_function_dependencies
from open_webui.utils.redis import get_redis_client from open_webui.utils.redis import get_redis_client
from open_webui.utils.security_headers import SecurityHeadersMiddleware from open_webui.utils.session_pool import cleanup_response, get_client_timeout, get_session, stream_wrapper
from open_webui.utils.session_pool import cleanup_response, get_session, stream_wrapper from open_webui.utils.tool_approval import (
ResolveToolCallForm,
build_tool_approval_resume_payload,
resolve_tool_call_output,
)
from open_webui.utils.tools import set_terminal_servers, set_tool_servers from open_webui.utils.tools import set_terminal_servers, set_tool_servers
if SAFE_MODE: if SAFE_MODE:
@ -320,6 +335,9 @@ https://github.com/open-webui/open-webui
print(banner) print(banner)
except UnicodeEncodeError: except UnicodeEncodeError:
# Stdout can't encode the box-drawing banner (Windows cp1252, redirected/headless stdout); fall back to ASCII. # Stdout can't encode the box-drawing banner (Windows cp1252, redirected/headless stdout); fall back to ASCII.
# LICENSE covers this Open WebUI CLI identifier.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
print(f'Open WebUI v{VERSION} - building the best AI user interface.\nhttps://github.com/open-webui/open-webui') print(f'Open WebUI v{VERSION} - building the best AI user interface.\nhttps://github.com/open-webui/open-webui')
@ -329,6 +347,16 @@ async def lifespan(app: FastAPI):
# This allows sync functions to schedule work on the main loop without blocking health checks # This allows sync functions to schedule work on the main loop without blocking health checks
app.state.main_loop = asyncio.get_running_loop() app.state.main_loop = asyncio.get_running_loop()
if THREAD_POOL_SIZE and THREAD_POOL_SIZE > 0:
# asyncio offloads bypass AnyIO's limiter, so configure both before the first offload.
anyio.to_thread.current_default_thread_limiter().total_tokens = THREAD_POOL_SIZE
app.state.main_loop.set_default_executor(
ThreadPoolExecutor(
max_workers=THREAD_POOL_SIZE,
thread_name_prefix=THREAD_POOL_THREAD_NAME_PREFIX,
)
)
app.state.instance_id = INSTANCE_ID app.state.instance_id = INSTANCE_ID
start_logger() start_logger()
@ -363,17 +391,18 @@ async def lifespan(app: FastAPI):
if app.state.redis is not None: if app.state.redis is not None:
app.state.redis_task_command_listener = asyncio.create_task(redis_task_command_listener(app)) app.state.redis_task_command_listener = asyncio.create_task(redis_task_command_listener(app))
if REDIS_TASK_TTL > 0:
app.state.redis_task_heartbeat = asyncio.create_task(redis_task_heartbeat(app))
if THREAD_POOL_SIZE and THREAD_POOL_SIZE > 0: if WEBSOCKET_MANAGER == 'redis':
limiter = anyio.to_thread.current_default_thread_limiter() app.state.redis_event_listener = asyncio.create_task(redis_event_listener())
limiter.total_tokens = THREAD_POOL_SIZE
asyncio.create_task(periodic_usage_pool_cleanup()) app.state.periodic_usage_pool_cleanup = asyncio.create_task(periodic_usage_pool_cleanup())
asyncio.create_task(periodic_session_pool_cleanup()) app.state.periodic_session_pool_cleanup = asyncio.create_task(periodic_session_pool_cleanup())
from open_webui.utils.automations import scheduler_worker_loop from open_webui.utils.automations import scheduler_worker_loop
asyncio.create_task(scheduler_worker_loop(app)) app.state.scheduler_worker_loop = asyncio.create_task(scheduler_worker_loop(app))
if await Config.get('models.base_models_cache'): if await Config.get('models.base_models_cache'):
try: try:
@ -420,13 +449,13 @@ async def lifespan(app: FastAPI):
log.info('Initializing tool servers...') log.info('Initializing tool servers...')
try: try:
await set_tool_servers(mock_request) await set_tool_servers(mock_request)
log.info(f'Initialized {len(app.state.TOOL_SERVERS)} tool server(s)') log.info('Initialized %s tool server(s)', len(app.state.TOOL_SERVERS))
except Exception as e: except Exception as e:
log.warning(f'Failed to initialize tool servers at startup: {e}') log.warning(f'Failed to initialize tool servers at startup: {e}')
try: try:
await set_terminal_servers(mock_request) await set_terminal_servers(mock_request)
log.info(f'Initialized {len(app.state.TERMINAL_SERVERS)} terminal server(s)') log.info('Initialized %s terminal server(s)', len(app.state.TERMINAL_SERVERS))
except Exception as e: except Exception as e:
log.warning(f'Failed to initialize terminal servers at startup: {e}') log.warning(f'Failed to initialize terminal servers at startup: {e}')
@ -454,6 +483,16 @@ async def lifespan(app: FastAPI):
if hasattr(app.state, 'redis_task_command_listener'): if hasattr(app.state, 'redis_task_command_listener'):
app.state.redis_task_command_listener.cancel() app.state.redis_task_command_listener.cancel()
if hasattr(app.state, 'redis_task_heartbeat'):
app.state.redis_task_heartbeat.cancel()
if hasattr(app.state, 'redis_event_listener'):
app.state.redis_event_listener.cancel()
app.state.periodic_usage_pool_cleanup.cancel()
app.state.periodic_session_pool_cleanup.cancel()
app.state.scheduler_worker_loop.cancel()
await publish_event(app, EVENTS.SYSTEM_SHUTDOWN_COMPLETED, source='system') await publish_event(app, EVENTS.SYSTEM_SHUTDOWN_COMPLETED, source='system')
@ -461,6 +500,9 @@ async def lifespan(app: FastAPI):
# response_model routes keep FastAPI's Pydantic fast path either way. # response_model routes keep FastAPI's Pydantic fast path either way.
apply_orjson_http_json() apply_orjson_http_json()
# LICENSE covers this Open WebUI API metadata identifier.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
app = FastAPI( app = FastAPI(
title='Open WebUI', title='Open WebUI',
docs_url='/docs' if ENV == 'dev' else None, docs_url='/docs' if ENV == 'dev' else None,
@ -469,6 +511,12 @@ app = FastAPI(
lifespan=lifespan, lifespan=lifespan,
) )
@app.exception_handler(RecurrenceEvaluationTimeout)
async def recurrence_timeout_handler(request: Request, exc: RecurrenceEvaluationTimeout):
return JSONResponse(status_code=400, content={'detail': str(exc)})
# Used by readiness checks to gate traffic until startup work is done. # Used by readiness checks to gate traffic until startup work is done.
app.state.startup_complete = False app.state.startup_complete = False
@ -483,6 +531,10 @@ app.state.oauth_client_manager = oauth_client_manager
app.state.instance_id = None app.state.instance_id = None
app.state.redis = None app.state.redis = None
# LICENSE covers this Open WebUI branding surface, including name, logo,
# visual, textual, symbolic identifiers, metadata, and surrounding UI.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
app.state.WEBUI_NAME = WEBUI_NAME app.state.WEBUI_NAME = WEBUI_NAME
app.state.LICENSE_METADATA = None app.state.LICENSE_METADATA = None
app.state.USER_COUNT = None app.state.USER_COUNT = None
@ -592,8 +644,18 @@ async def initialize_runtime_config(app: FastAPI):
f'mcp:{server_id}', f'mcp:{server_id}',
OAuthClientInformationFull(**oauth_client_info), OAuthClientInformationFull(**oauth_client_info),
) )
except InvalidToken:
log.error(
'Error adding OAuth client for MCP tool server %s: InvalidToken. '
'Stored OAuth client data is invalid; reconnect this tool server.',
server_id,
)
except Exception as e: except Exception as e:
log.error(f'Error adding OAuth client for MCP tool server {server_id}: {e}') log.error(
'Error adding OAuth client for MCP tool server %s: %s',
server_id,
f'{type(e).__name__}: {e}' if str(e) else type(e).__name__,
)
arena_models = await Config.get('evaluation.arena.models', []) or [] arena_models = await Config.get('evaluation.arena.models', []) or []
if any('access_control' in m.get('meta', {}) for m in arena_models): if any('access_control' in m.get('meta', {}) for m in arena_models):
@ -763,11 +825,7 @@ if ENABLE_COMPRESSION_MIDDLEWARE:
# `terminate_force_close` tracebacks under aiosqlite and as random # `terminate_force_close` tracebacks under aiosqlite and as random
# CancelledError storms across the request path. See # CancelledError storms across the request path. See
# `open_webui.utils.asgi_middleware` for the rationale. # `open_webui.utils.asgi_middleware` for the rationale.
app.add_middleware(RedirectMiddleware) app.add_middleware(AppHTTPMiddleware)
app.add_middleware(SecurityHeadersMiddleware)
app.add_middleware(CommitSessionMiddleware)
app.add_middleware(AuthTokenMiddleware, fastapi_app=app)
app.add_middleware(WebsocketUpgradeGuardMiddleware)
app.add_middleware( app.add_middleware(
@ -855,19 +913,17 @@ async def get_models(request: Request, refresh: bool = False, user=Depends(get_v
models = await get_filtered_models(models, user) models = await get_filtered_models(models, user)
for model in models: for model in models:
info = model.get('info') if isinstance(model.get('info'), dict) else {}
meta = info.get('meta') if isinstance(info.get('meta'), dict) else {}
# Remove profile image URL to reduce payload size # Remove profile image URL to reduce payload size
if model.get('info', {}).get('meta', {}).get('profile_image_url'): meta.pop('profile_image_url', None)
model['info']['meta'].pop('profile_image_url', None)
try: if 'tags' in meta:
model_tags = [tag.get('name') for tag in model.get('info', {}).get('meta', {}).get('tags', [])] meta['tags'] = normalize_model_tags(meta['tags'])
tags = [tag.get('name') for tag in model.get('tags', [])]
tags = list(set(model_tags + tags)) tags = normalize_model_tags(meta.get('tags')) + normalize_model_tags(model.get('tags'))
model['tags'] = [{'name': tag} for tag in tags] model['tags'] = list({tag['name']: tag for tag in tags}.values())
except Exception as e:
log.debug(f'Error processing model tags: {e}')
model['tags'] = []
model_order_list = await Config.get('ui.model_order_list') model_order_list = await Config.get('ui.model_order_list')
if model_order_list: if model_order_list:
@ -882,7 +938,7 @@ async def get_models(request: Request, refresh: bool = False, user=Depends(get_v
if log.isEnabledFor(logging.DEBUG): if log.isEnabledFor(logging.DEBUG):
log.debug( log.debug(
f'/api/models returned filtered models accessible to the user: {json.dumps([model.get("id") for model in models])}' f'/api/models returned filtered models accessible to the user: {JSONCodec.dumps([model.get("id") for model in models])}'
) )
return {'data': models} return {'data': models}
@ -935,7 +991,7 @@ async def unload_model(request: Request, form_data: ModelUnloadForm, user=Depend
prefix_id = api_config.get('prefix_id', None) prefix_id = api_config.get('prefix_id', None)
actual_model = strip_provider_model_prefix(model_id, prefix_id) actual_model = strip_provider_model_prefix(model_id, prefix_id)
payload = json.dumps({'model': actual_model, 'keep_alive': 0, 'prompt': ''}) payload = JSONCodec.dumps({'model': actual_model, 'keep_alive': 0, 'prompt': ''})
try: try:
timeout = aiohttp.ClientTimeout(total=30) timeout = aiohttp.ClientTimeout(total=30)
@ -1066,53 +1122,67 @@ async def chat_completion(
metadata = {} metadata = {}
try: try:
model_info = None model_info = None
fallback_model = None
missing_base_model = False
if not model_item.get('direct', False): if not model_item.get('direct', False):
if model_id not in request.app.state.MODELS: if model_id not in request.app.state.MODELS:
raise Exception('Model not found') raise Exception('Model not found')
model = request.app.state.MODELS[model_id] model = request.app.state.MODELS[model_id]
model_info = await Models.get_model_by_id(model_id) model_info = await Models.get_model_by_id(model_id)
missing_base_model = bool(
model_info and model_info.base_model_id and model_info.base_model_id not in request.app.state.MODELS
)
if missing_base_model and ENABLE_CUSTOM_MODEL_FALLBACK:
fallback_model_id = next(
(
model_id.strip()
for model_id in ((await Config.get('ui.default_models')) or '').split(',')
if model_id.strip()
),
None,
)
if fallback_model_id:
fallback_model = request.app.state.MODELS.get(fallback_model_id)
# Check if user has access to the model # Check if user has access to the model
if not BYPASS_MODEL_ACCESS_CONTROL and (user.role != 'admin' or not BYPASS_ADMIN_ACCESS_CONTROL): if not BYPASS_MODEL_ACCESS_CONTROL and (user.role != 'admin' or not BYPASS_ADMIN_ACCESS_CONTROL):
try: try:
await check_model_access(user, model, model_info=model_info) access_model_info = (
model_info.model_copy(update={'base_model_id': None})
if fallback_model is not None
else model_info
)
await check_model_access(user, model, model_info=access_model_info)
if fallback_model is not None:
await check_model_access(user, fallback_model)
except Exception as e: except Exception as e:
raise e raise e
else: else:
model = model_item model = model_item
await _set_direct_model(request, model, user) await _set_direct_model(request, model, user)
# Read before the fallback below can rebind model to a different one.
model_capabilities = ((model.get('info') or {}).get('meta') or {}).get('capabilities') or {}
# Model params: global defaults as base, per-model overrides win # Model params: global defaults as base, per-model overrides win
default_model_params = await Config.get('models.default_params', {}) or {} default_model_params = copy.deepcopy(await Config.get('models.default_params', {}) or {})
model_info_params = { model_info_params = merge_model_params(
**default_model_params, default_model_params,
**(model_info.params.model_dump() if model_info and model_info.params else {}), model_info.params.model_dump() if model_info and model_info.params else {},
} )
request_params = {key: value for key, value in (form_data.get('params') or {}).items() if value is not None} request_params = {key: value for key, value in (form_data.get('params') or {}).items() if value is not None}
if model_info_params or request_params: if model_info_params or request_params:
form_data['params'] = { form_data['params'] = merge_model_params(model_info_params, request_params)
**model_info_params,
**request_params,
}
# Check base model existence for custom models # Check base model existence for custom models
if model_info and model_info.base_model_id: if missing_base_model:
base_model_id = model_info.base_model_id if fallback_model is None:
if base_model_id not in request.app.state.MODELS: raise Exception('Model not found')
if ENABLE_CUSTOM_MODEL_FALLBACK: # Update model and form_data so routing uses the fallback model's type
default_models = ((await Config.get('ui.default_models')) or '').split(',') model = fallback_model
form_data['model'] = fallback_model['id']
fallback_model_id = default_models[0].strip() if default_models[0] else None
if fallback_model_id and fallback_model_id in request.app.state.MODELS:
# Update model and form_data so routing uses the fallback model's type
model = request.app.state.MODELS[fallback_model_id]
form_data['model'] = fallback_model_id
else:
raise Exception('Model not found')
else:
raise Exception('Model not found')
# Chat Params # Chat Params
stream_delta_chunk_size = form_data.get('params', {}).get('stream_delta_chunk_size') stream_delta_chunk_size = form_data.get('params', {}).get('stream_delta_chunk_size')
@ -1123,6 +1193,10 @@ async def chat_completion(
if model_info_params.get('stream_response') is not None: if model_info_params.get('stream_response') is not None:
form_data['stream'] = model_info_params.get('stream_response') form_data['stream'] = model_info_params.get('stream_response')
# Providers only report token counts when asked, so ask on every caller's behalf.
if form_data.get('stream') and model_capabilities.get('usage'):
form_data['stream_options'] = {**(form_data.get('stream_options') or {}), 'include_usage': True}
if model_info_params.get('stream_delta_chunk_size'): if model_info_params.get('stream_delta_chunk_size'):
stream_delta_chunk_size = model_info_params.get('stream_delta_chunk_size') stream_delta_chunk_size = model_info_params.get('stream_delta_chunk_size')
@ -1155,7 +1229,7 @@ async def chat_completion(
message_ids = [{'model_id': model_id, 'message_id': form_data.pop('id', None)}] message_ids = [{'model_id': model_id, 'message_id': form_data.pop('id', None)}]
user_message = form_data.pop('user_message', None) or form_data.pop('parent_message', None) user_message = form_data.pop('user_message', None) or form_data.pop('parent_message', None)
chat_id = form_data.get('chat_id') or '' chat_id = form_data.pop('chat_id', None) or ''
chat_variables = form_data.pop('chat_variables', None) chat_variables = form_data.pop('chat_variables', None)
if chat_variables is None: if chat_variables is None:
existing_chat = await Chats.get_chat_by_id(chat_id) if is_saved_chat_id(chat_id) else None existing_chat = await Chats.get_chat_by_id(chat_id) if is_saved_chat_id(chat_id) else None
@ -1177,15 +1251,28 @@ async def chat_completion(
): ):
tool_servers = None tool_servers = None
automation_id = form_data.pop('automation_id', None)
tool_approval_mode = (
'full'
if automation_id or chat_id.startswith('channel:')
else (
form_data.get('params', {}).get('tool_approval_mode')
if await Config.get('chat.tool_permissions.enable', False)
else 'full'
)
or 'full'
)
metadata = { metadata = {
'user_id': user.id, 'user_id': user.id,
'user_agent': request.headers.get('user-agent', '') or '', 'user_agent': request.headers.get('user-agent', '') or '',
'internal': getattr(request.state, 'internal', False) is True, 'internal': getattr(request.state, 'internal', False) is True,
'chat_id': form_data.pop('chat_id', None) or '', 'chat_id': chat_id,
'user_message': user_message, 'user_message': user_message,
'user_message_id': user_message.get('id') if user_message else None, 'user_message_id': user_message.get('id') if user_message else None,
'assistant_message_id': form_data.pop('assistant_message_id', None), 'assistant_message_id': form_data.pop('assistant_message_id', None),
'session_id': form_data.pop('session_id', None), 'session_id': form_data.pop('session_id', None),
'automation_id': automation_id,
'folder_id': form_data.pop('folder_id', None), 'folder_id': form_data.pop('folder_id', None),
'filter_ids': form_data.pop('filter_ids', []), 'filter_ids': form_data.pop('filter_ids', []),
'tool_ids': form_data.get('tool_ids', None), 'tool_ids': form_data.get('tool_ids', None),
@ -1205,6 +1292,7 @@ async def chat_completion(
or model_info_params.get('function_calling') or model_info_params.get('function_calling')
or 'native' or 'native'
), ),
'tool_approval_mode': tool_approval_mode,
}, },
} }
@ -1218,8 +1306,8 @@ async def chat_completion(
if metadata.get('chat_id') and user: if metadata.get('chat_id') and user:
chat_id = metadata['chat_id'] chat_id = metadata['chat_id']
# Gate channel: branch — caller needs write access on the channel # Gate channel: branch — caller needs write access on the channel, and the
# and the supplied message_id must belong to that channel. # supplied message_id must belong to that channel and be the caller's own.
if chat_id.startswith('channel:'): if chat_id.startswith('channel:'):
channel_id = chat_id.removeprefix('channel:') channel_id = chat_id.removeprefix('channel:')
channel = await Channels.get_channel_by_id(channel_id) channel = await Channels.get_channel_by_id(channel_id)
@ -1251,7 +1339,11 @@ async def chat_completion(
if not target_message_id: if not target_message_id:
continue continue
target_message = await Messages.get_message_by_id(target_message_id) target_message = await Messages.get_message_by_id(target_message_id)
if target_message and target_message.channel_id != channel.id: if target_message and (
target_message.channel_id != channel.id
# Write access is not authorship — block cross-member edits.
or (user.role != 'admin' and target_message.user_id != user.id)
):
raise HTTPException( raise HTTPException(
status_code=status.HTTP_403_FORBIDDEN, status_code=status.HTTP_403_FORBIDDEN,
detail=ERROR_MESSAGES.DEFAULT(), detail=ERROR_MESSAGES.DEFAULT(),
@ -1259,6 +1351,14 @@ async def chat_completion(
if is_saved_chat_id(chat_id): if is_saved_chat_id(chat_id):
if is_new_chat: if is_new_chat:
# The chat created below is persisted with this folder_id.
folder_id = metadata['folder_id']
if folder_id is not None and not await has_folder_write_access(user.id, folder_id):
raise HTTPException(
status_code=status.HTTP_404_NOT_FOUND,
detail=ERROR_MESSAGES.NOT_FOUND,
)
# Build the full history upfront with ALL assistant placeholders # Build the full history upfront with ALL assistant placeholders
user_message = metadata.get('user_message') or {} user_message = metadata.get('user_message') or {}
user_message_id = user_message.get('id') if user_message else None user_message_id = user_message.get('id') if user_message else None
@ -1365,7 +1465,7 @@ async def chat_completion(
user.id, user.id,
) )
except Exception as e: except Exception as e:
log.debug(f'Error inserting chat files: {e}') log.debug('Error inserting chat files: %s', e)
pass pass
if initial_title_generation is not None and all_assistant_ids: if initial_title_generation is not None and all_assistant_ids:
@ -1386,8 +1486,8 @@ async def chat_completion(
async def run_initial_title_generation(): async def run_initial_title_generation():
try: try:
await background_tasks_handler(title_ctx) await background_tasks_handler(title_ctx)
except Exception as e: except Exception:
log.debug(f'Error generating initial chat title: {e}') log.exception('Error generating initial chat title')
asyncio.create_task(run_initial_title_generation()) asyncio.create_task(run_initial_title_generation())
else: else:
@ -1407,15 +1507,13 @@ async def chat_completion(
# The old frontend saveChatHandler did this on every message; # The old frontend saveChatHandler did this on every message;
# now the backend owns persistence. # now the backend owns persistence.
chat_files = metadata.get('files') chat_files = metadata.get('files')
if chat_files is not None or selected_chat_models: chat_fields = {}
existing_chat = await Chats.get_chat_by_id(chat_id) if chat_files is not None:
if existing_chat: chat_fields['files'] = chat_files
updated = {**existing_chat.chat} if selected_chat_models:
if chat_files is not None: chat_fields['models'] = selected_chat_models
updated['files'] = chat_files if chat_fields:
if selected_chat_models: await Chats.update_chat_by_id(chat_id, chat_fields, touch=False)
updated['models'] = selected_chat_models
await Chats.update_chat_by_id(chat_id, updated, touch=False)
await Chats.update_chat_variables_by_id(chat_id, chat_variables) await Chats.update_chat_variables_by_id(chat_id, chat_variables)
@ -1475,7 +1573,7 @@ async def chat_completion(
user.id, user.id,
) )
except Exception as e: except Exception as e:
log.debug(f'Error inserting chat files: {e}') log.debug('Error inserting chat files: %s', e)
pass pass
# Save ALL assistant placeholders # Save ALL assistant placeholders
@ -1500,6 +1598,8 @@ async def chat_completion(
for entry in message_ids: for entry in message_ids:
target_model_id = entry['model_id'] target_model_id = entry['model_id']
assistant_message_id = entry['message_id'] assistant_message_id = entry['message_id']
if assistant_message_id and assistant_message_id == metadata.get('assistant_message_id'):
continue
if assistant_message_id: if assistant_message_id:
assistant_message = { assistant_message = {
'id': assistant_message_id, 'id': assistant_message_id,
@ -1546,26 +1646,28 @@ async def chat_completion(
async def process_chat(request, form_data, user, metadata, model, tasks=None): async def process_chat(request, form_data, user, metadata, model, tasks=None):
try: try:
ctx = None
if metadata.get('assistant_message_id'):
ctx = await build_chat_response_context(request, form_data, user, model, metadata, tasks, [])
form_data, metadata, events = await process_chat_payload(request, form_data, user, metadata, model) form_data, metadata, events = await process_chat_payload(request, form_data, user, metadata, model)
if await drain_approved_tool_calls(request, form_data, user, model, metadata):
return {'status': True, 'chat_id': metadata.get('chat_id'), 'paused': True}
response = await chat_completion_handler(request, form_data, user) response = await chat_completion_handler(request, form_data, user)
# When the upstream provider returns an error (e.g. HTTP 400 # When the upstream provider returns an error (e.g. HTTP 400
# content-filter, quota exceeded), generate_chat_completion # content-filter, quota exceeded), generate_chat_completion
# returns a JSONResponse instead of raising. Detect this and # returns a JSONResponse instead of raising. Detect this and
# raise so the except-block below emits chat:message:error + # raise so the except-block below emits a terminal
# chat:tasks:cancel, unblocking the frontend. # chat:message:error, unblocking the frontend.
if isinstance(response, JSONResponse) and response.status_code >= 400: if isinstance(response, JSONResponse) and response.status_code >= 400:
try: raise Exception(get_response_error_detail(response))
error_body = json.loads(response.body.decode('utf-8', 'replace'))
detail = error_body.get('error', error_body) if isinstance(error_body, dict) else error_body
if isinstance(detail, dict):
detail = detail.get('message', detail.get('detail', str(detail)))
except Exception:
detail = f'Provider returned HTTP {response.status_code}'
raise Exception(detail)
ctx = await build_chat_response_context(request, form_data, user, model, metadata, tasks, events) if ctx is None:
ctx = await build_chat_response_context(request, form_data, user, model, metadata, tasks, events)
else:
ctx.update(form_data=form_data, metadata=metadata, events=events)
return await process_chat_response(response, ctx) return await process_chat_response(response, ctx)
except asyncio.CancelledError: except asyncio.CancelledError:
@ -1594,6 +1696,7 @@ async def chat_completion(
{ {
'parentId': metadata.get('user_message_id', None), 'parentId': metadata.get('user_message_id', None),
'error': {'content': error_detail}, 'error': {'content': error_detail},
'done': True,
}, },
) )
@ -1602,12 +1705,9 @@ async def chat_completion(
await event_emitter( await event_emitter(
{ {
'type': 'chat:message:error', 'type': 'chat:message:error',
'data': {'error': {'content': error_detail}}, 'data': {'error': {'content': error_detail}, 'done': True},
} }
) )
await event_emitter(
{'type': 'chat:tasks:cancel'},
)
except Exception: except Exception:
pass pass
@ -1641,9 +1741,9 @@ async def chat_completion(
try: try:
await client.disconnect() await client.disconnect()
except BaseException as e: except BaseException as e:
log.debug(f'Error disconnecting MCP client: {e}') log.debug('Error disconnecting MCP client: %s', e)
except BaseException as e: except BaseException as e:
log.debug(f'Error cleaning up MCP clients: {e}') log.debug('Error cleaning up MCP clients: %s', e)
# Deregister this task, then emit chat:active=false if no others remain # Deregister this task, then emit chat:active=false if no others remain
try: try:
@ -1789,6 +1889,27 @@ async def chat_completion(
generate_chat_completions = chat_completion generate_chat_completions = chat_completion
generate_chat_completion = chat_completion generate_chat_completion = chat_completion
@app.post('/api/v1/chats/{id}/messages/{message_id}/resolve')
async def resolve_chat_message_tool_call(
request: Request,
id: str,
message_id: str,
form_data: ResolveToolCallForm,
user=Depends(get_verified_user),
db: AsyncSession = Depends(get_async_session),
):
resolution = await resolve_tool_call_output(id, message_id, form_data, user, db=db)
payload = await build_tool_approval_resume_payload(id, message_id, chat=resolution['chat'])
result = await chat_completion(request, payload, user)
return {
'status': True,
'chat_id': id,
'message_id': message_id,
**(result if isinstance(result, dict) else {}),
}
# Expose as app.state so internal callers (e.g. automations) can # Expose as app.state so internal callers (e.g. automations) can
# use the full pipeline without importing from main.py (avoids circular deps). # use the full pipeline without importing from main.py (avoids circular deps).
app.state.CHAT_COMPLETION_HANDLER = chat_completion app.state.CHAT_COMPLETION_HANDLER = chat_completion
@ -1820,7 +1941,7 @@ async def count_message_tokens(
async def passthrough_anthropic_messages(request: Request, form_data: dict, user) -> Response | dict: async def passthrough_anthropic_messages(request: Request, form_data: dict, user) -> Response | dict:
requested_model, payload, url, key, headers, cookies = await openai.get_anthropic_token_count_target( requested_model, payload, url, key, headers, cookies = await openai.get_anthropic_request_target(
request, form_data, user request, form_data, user
) )
request_url = f'{url.rstrip("/")}/messages' request_url = f'{url.rstrip("/")}/messages'
@ -1832,11 +1953,11 @@ async def passthrough_anthropic_messages(request: Request, form_data: dict, user
response = await session.request( response = await session.request(
method='POST', method='POST',
url=request_url, url=request_url,
data=json.dumps(payload), data=JSONCodec.dumps(payload),
headers=headers, headers=headers,
cookies=cookies, cookies=cookies,
ssl=AIOHTTP_CLIENT_SESSION_SSL, ssl=AIOHTTP_CLIENT_SESSION_SSL,
timeout=aiohttp.ClientTimeout(total=openai.AIOHTTP_CLIENT_TIMEOUT), timeout=get_client_timeout(stream=bool(payload.get('stream'))),
) )
if 'text/event-stream' in response.headers.get('Content-Type', ''): if 'text/event-stream' in response.headers.get('Content-Type', ''):
@ -1925,13 +2046,6 @@ async def generate_messages(
# Convert Anthropic payload to OpenAI format # Convert Anthropic payload to OpenAI format
openai_payload = convert_anthropic_to_openai_payload(form_data, passthrough_params) openai_payload = convert_anthropic_to_openai_payload(form_data, passthrough_params)
model_meta = model_info.meta.model_dump() if model_info and model_info.meta else {}
if (model_meta.get('capabilities') or {}).get('usage') is True:
if openai_payload.get('stream'):
stream_options = openai_payload.get('stream_options')
if not isinstance(stream_options, dict):
stream_options = {}
openai_payload['stream_options'] = {**stream_options, 'include_usage': True}
# Route through the existing chat_completion handler # Route through the existing chat_completion handler
response = await chat_completion(request, openai_payload, user) response = await chat_completion(request, openai_payload, user)
@ -2039,13 +2153,14 @@ async def list_tasks_by_chat_id_endpoint(request: Request, chat_id: str, user=De
task_ids = await list_task_ids_by_item_id(request.app.state.redis, chat_id) task_ids = await list_task_ids_by_item_id(request.app.state.redis, chat_id)
log.debug(f'Task IDs for chat {chat_id}: {task_ids}') log.debug('Task IDs for chat %s: %s', chat_id, task_ids)
return {'task_ids': task_ids} return {'task_ids': task_ids}
@app.post('/api/tasks/chat/{chat_id:path}/stop') @app.post('/api/tasks/chat/{chat_id:path}/stop')
async def stop_tasks_by_chat_id_endpoint(request: Request, chat_id: str, user=Depends(get_verified_user)): async def stop_tasks_by_chat_id_endpoint(request: Request, chat_id: str, user=Depends(get_verified_user)):
socket_id = get_temporary_chat_session_id(chat_id) socket_id = get_temporary_chat_session_id(chat_id)
chat = None
if socket_id: if socket_id:
owner_id = get_user_id_from_session_pool(socket_id) owner_id = get_user_id_from_session_pool(socket_id)
if owner_id != user.id and user.role != 'admin': if owner_id != user.id and user.role != 'admin':
@ -2055,6 +2170,47 @@ async def stop_tasks_by_chat_id_endpoint(request: Request, chat_id: str, user=De
if chat is None or (chat.user_id != user.id and user.role != 'admin'): if chat is None or (chat.user_id != user.id and user.role != 'admin'):
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail=ERROR_MESSAGES.NOT_FOUND) raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail=ERROR_MESSAGES.NOT_FOUND)
result = await stop_item_tasks(request.app.state.redis, chat_id) result = await stop_item_tasks(request.app.state.redis, chat_id)
if not socket_id and str(result.get('message', '')).startswith('No tasks found'):
messages_map = await Chats.get_messages_map_by_chat_id(chat_id) or {}
for message_id, message in messages_map.items():
if message.get('role') != 'assistant' or message.get('done') is not False:
continue
output = message.get('output')
if isinstance(output, list):
for item in output:
if item.get('type') == 'function_call' and item.get('status') in {
'pending',
'queued',
'requires_approval',
}:
item['status'] = 'rejected'
item.pop('approved', None)
await Chats.upsert_message_to_chat_by_id_and_message_id(
chat_id,
message_id,
{'done': True, **({'output': output} if isinstance(output, list) else {})},
touch=False,
)
result = {
'status': True,
'message': 'Finalized pending approval message.',
}
event_emitter = await get_event_emitter(
{
'user_id': chat.user_id,
'chat_id': chat_id,
'message_id': message_id,
},
update_db=False,
)
if event_emitter:
await event_emitter({'type': 'chat:completion', 'data': {'done': True, 'output': output}})
await event_emitter({'type': 'chat:tasks:cancel'})
return result return result
@ -2106,6 +2262,7 @@ async def get_app_config(request: Request):
'auth.enable_api_keys', 'auth.enable_api_keys',
'ui.enable_password_change_form', 'ui.enable_password_change_form',
'direct.enable', 'direct.enable',
'direct.integrations.enable',
'folders.enable', 'folders.enable',
'folders.max_file_count', 'folders.max_file_count',
'channels.enable', 'channels.enable',
@ -2113,6 +2270,7 @@ async def get_app_config(request: Request):
'automations.enable', 'automations.enable',
'notes.enable', 'notes.enable',
'chat.context_compaction.enable', 'chat.context_compaction.enable',
'chat.tool_permissions.enable',
'web.search.enable', 'web.search.enable',
'web.search.confirmation.enable', 'web.search.confirmation.enable',
'web.search.confirmation.content', 'web.search.confirmation.content',
@ -2129,7 +2287,10 @@ async def get_app_config(request: Request):
'memories.enable', 'memories.enable',
'ui.default_models', 'ui.default_models',
'ui.default_pinned_models', 'ui.default_pinned_models',
'ui.default_interface_settings',
'ui.i18n',
'ui.prompt_suggestions', 'ui.prompt_suggestions',
'ui.prompt_suggestions_i18n',
'code_execution.engine', 'code_execution.engine',
'code_interpreter.engine', 'code_interpreter.engine',
'audio.tts.engine', 'audio.tts.engine',
@ -2152,6 +2313,7 @@ async def get_app_config(request: Request):
'name': app.state.WEBUI_NAME, 'name': app.state.WEBUI_NAME,
'version': VERSION, 'version': VERSION,
'default_locale': str(DEFAULT_LOCALE), 'default_locale': str(DEFAULT_LOCALE),
'i18n': config.get('ui.i18n') or {},
'oauth': { 'oauth': {
# Hide providers (and thus the login buttons / auto-redirect) when OAuth # Hide providers (and thus the login buttons / auto-redirect) when OAuth
# is disabled, without clearing the admin's provider configuration. # is disabled, without clearing the admin's provider configuration.
@ -2163,6 +2325,7 @@ async def get_app_config(request: Request):
'auto_redirect': config.get('oauth.auto_redirect'), 'auto_redirect': config.get('oauth.auto_redirect'),
}, },
'features': { 'features': {
'slim': USE_SLIM,
# --- Public: required by login/signup page pre-auth --- # --- Public: required by login/signup page pre-auth ---
'auth': WEBUI_AUTH, 'auth': WEBUI_AUTH,
'auth_trusted_header': bool(WEBUI_AUTH_TRUSTED_EMAIL_HEADER), 'auth_trusted_header': bool(WEBUI_AUTH_TRUSTED_EMAIL_HEADER),
@ -2171,6 +2334,11 @@ async def get_app_config(request: Request):
'enable_signup': config.get('ui.enable_signup'), 'enable_signup': config.get('ui.enable_signup'),
'enable_login_form': config.get('ui.enable_login_form'), 'enable_login_form': config.get('ui.enable_login_form'),
'enable_websocket': ENABLE_WEBSOCKET_SUPPORT, 'enable_websocket': ENABLE_WEBSOCKET_SUPPORT,
**(
{'websocket_heartbeat_interval': WEBSOCKET_HEARTBEAT_INTERVAL}
if WEBSOCKET_HEARTBEAT_INTERVAL is not None
else {}
),
# --- Authenticated: only consumed by logged-in frontend --- # --- Authenticated: only consumed by logged-in frontend ---
**( **(
{ {
@ -2181,6 +2349,7 @@ async def get_app_config(request: Request):
'enable_public_active_users_count': ENABLE_PUBLIC_ACTIVE_USERS_COUNT, 'enable_public_active_users_count': ENABLE_PUBLIC_ACTIVE_USERS_COUNT,
'enable_easter_eggs': ENABLE_EASTER_EGGS, 'enable_easter_eggs': ENABLE_EASTER_EGGS,
'enable_direct_connections': config.get('direct.enable'), 'enable_direct_connections': config.get('direct.enable'),
'enable_direct_integrations': config.get('direct.integrations.enable', False),
'enable_plugins': ENABLE_PLUGINS, 'enable_plugins': ENABLE_PLUGINS,
'enable_folders': config.get('folders.enable'), 'enable_folders': config.get('folders.enable'),
'folder_max_file_count': config.get('folders.max_file_count'), 'folder_max_file_count': config.get('folders.max_file_count'),
@ -2189,6 +2358,7 @@ async def get_app_config(request: Request):
'enable_automations': config.get('automations.enable'), 'enable_automations': config.get('automations.enable'),
'enable_notes': config.get('notes.enable'), 'enable_notes': config.get('notes.enable'),
'enable_context_compaction': config.get('chat.context_compaction.enable'), 'enable_context_compaction': config.get('chat.context_compaction.enable'),
'enable_tool_permissions': config.get('chat.tool_permissions.enable'),
'enable_web_search': config.get('web.search.enable'), 'enable_web_search': config.get('web.search.enable'),
'enable_web_search_confirmation': config.get('web.search.confirmation.enable'), 'enable_web_search_confirmation': config.get('web.search.confirmation.enable'),
'web_search_confirmation_content': config.get('web.search.confirmation.content'), 'web_search_confirmation_content': config.get('web.search.confirmation.content'),
@ -2224,6 +2394,7 @@ async def get_app_config(request: Request):
'default_models': config.get('ui.default_models'), 'default_models': config.get('ui.default_models'),
'default_pinned_models': config.get('ui.default_pinned_models'), 'default_pinned_models': config.get('ui.default_pinned_models'),
'default_prompt_suggestions': config.get('ui.prompt_suggestions'), 'default_prompt_suggestions': config.get('ui.prompt_suggestions'),
'default_prompt_suggestions_i18n': config.get('ui.prompt_suggestions_i18n'),
**({'user_count': user_count} if user_count is not None else {}), **({'user_count': user_count} if user_count is not None else {}),
'code': { 'code': {
'engine': config.get('code_execution.engine'), 'engine': config.get('code_execution.engine'),
@ -2259,6 +2430,7 @@ async def get_app_config(request: Request):
'sharepoint_tenant_id': ONEDRIVE_SHAREPOINT_TENANT_ID, 'sharepoint_tenant_id': ONEDRIVE_SHAREPOINT_TENANT_ID,
}, },
'ui': { 'ui': {
'default_interface_settings': config.get('ui.default_interface_settings'),
'pending_user_overlay_title': config.get('ui.pending_user_overlay_title'), 'pending_user_overlay_title': config.get('ui.pending_user_overlay_title'),
'pending_user_overlay_content': config.get('ui.pending_user_overlay_content'), 'pending_user_overlay_content': config.get('ui.pending_user_overlay_content'),
'response_watermark': config.get('ui.watermark'), 'response_watermark': config.get('ui.watermark'),
@ -2437,8 +2609,8 @@ async def get_app_latest_release_version(user=Depends(get_verified_user)):
return {'current': VERSION, 'latest': latest_version[1:]} return {'current': VERSION, 'latest': latest_version[1:]}
except Exception as e: except Exception as e:
log.debug(e) log.warning(f'Version update check failed: {e}')
return {'current': VERSION, 'latest': VERSION} return {'current': VERSION, 'latest': None}
@app.get('/api/changelog') @app.get('/api/changelog')
@ -2559,8 +2731,19 @@ async def register_client(request, client_id: str) -> bool:
oauth_server_key, oauth_server_key,
oauth_scope=oauth_scope, oauth_scope=oauth_scope,
) )
except InvalidToken:
log.error(
'OAuth client re-registration failed for %s: InvalidToken. '
'Stored OAuth client data is invalid; reconnect this tool server.',
client_id,
)
return False
except Exception as e: except Exception as e:
log.error(f'OAuth client re-registration failed for {client_id}: {e}') log.error(
'OAuth client re-registration failed for %s: %s',
client_id,
f'{type(e).__name__}: {e}' if str(e) else type(e).__name__,
)
return False return False
try: try:
@ -2582,7 +2765,7 @@ async def register_client(request, client_id: str) -> bool:
**apply_connection_oauth_options(connection, oauth_client_info.model_dump(mode='json')) **apply_connection_oauth_options(connection, oauth_client_info.model_dump(mode='json'))
) )
oauth_client_manager.add_client(client_id, oauth_client_info) oauth_client_manager.add_client(client_id, oauth_client_info)
log.info(f'Re-registered OAuth client {client_id} for tool server') log.info('Re-registered OAuth client %s for tool server', client_id)
return True return True
@ -2626,7 +2809,7 @@ async def oauth_client_authorize(
detail='OAuth client registration is still invalid after re-registration', detail='OAuth client registration is still invalid after re-registration',
) )
return await oauth_client_manager.handle_authorize(request, client_id=client_id) return await oauth_client_manager.handle_authorize(request, client_id=client_id, user_id=user.id)
@app.get('/oauth/clients/{client_id}/callback') @app.get('/oauth/clients/{client_id}/callback')
@ -2634,12 +2817,10 @@ async def oauth_client_callback(
client_id: str, client_id: str,
request: Request, request: Request,
response: Response, response: Response,
user=Depends(get_verified_user),
): ):
return await oauth_client_manager.handle_callback( return await oauth_client_manager.handle_callback(
request, request,
client_id=client_id, client_id=client_id,
user_id=user.id if user else None,
response=response, response=response,
) )
@ -2688,6 +2869,10 @@ async def oauth_backchannel_logout(
async def get_manifest_json(): async def get_manifest_json():
external_pwa_manifest_url = getattr(app.state, 'EXTERNAL_PWA_MANIFEST_URL', None) external_pwa_manifest_url = getattr(app.state, 'EXTERNAL_PWA_MANIFEST_URL', None)
if external_pwa_manifest_url: if external_pwa_manifest_url:
# LICENSE covers this install-time Open WebUI branding surface, including
# names, logos, manifests, metadata, and surrounding UI.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
session = await get_session() session = await get_session()
async with session.get( async with session.get(
external_pwa_manifest_url, external_pwa_manifest_url,
@ -2696,6 +2881,10 @@ async def get_manifest_json():
r.raise_for_status() r.raise_for_status()
return await r.json() return await r.json()
else: else:
# LICENSE covers this generated Open WebUI install branding surface,
# including names, logos, manifests, metadata, and surrounding UI.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
return { return {
'name': app.state.WEBUI_NAME, 'name': app.state.WEBUI_NAME,
'short_name': app.state.WEBUI_NAME, 'short_name': app.state.WEBUI_NAME,
@ -2704,6 +2893,9 @@ async def get_manifest_json():
'display': 'standalone', 'display': 'standalone',
'background_color': '#343541', 'background_color': '#343541',
'icons': [ 'icons': [
# LICENSE covers this Open WebUI install icon.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
{ {
'src': '/static/logo.png', 'src': '/static/logo.png',
'type': 'image/png', 'type': 'image/png',
@ -2728,6 +2920,9 @@ async def get_manifest_json():
@app.get('/opensearch.xml') @app.get('/opensearch.xml')
async def get_opensearch_xml(): async def get_opensearch_xml():
webui_url = await Config.get('webui.url') webui_url = await Config.get('webui.url')
# LICENSE covers this Open WebUI search identifier.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
xml_content = rf""" xml_content = rf"""
<OpenSearchDescription xmlns="http://a9.com/-/spec/opensearch/1.1/" xmlns:moz="http://www.mozilla.org/2006/browser/search/"> <OpenSearchDescription xmlns="http://a9.com/-/spec/opensearch/1.1/" xmlns:moz="http://www.mozilla.org/2006/browser/search/">
<ShortName>{app.state.WEBUI_NAME}</ShortName> <ShortName>{app.state.WEBUI_NAME}</ShortName>
@ -2745,7 +2940,7 @@ def _sync_db_ping() -> None:
"""Verify the database is reachable with a simple SELECT 1. """Verify the database is reachable with a simple SELECT 1.
Uses a raw connection from the engine pool instead of the thread-local Uses a raw connection from the engine pool instead of the thread-local
ScopedSession. This is necessary because CommitSessionMiddleware ScopedSession. This is necessary because AppHTTPMiddleware
deliberately skips healthcheck paths (/health, /ready, /health/db), deliberately skips healthcheck paths (/health, /ready, /health/db),
so any ScopedSession opened on a healthcheck worker thread is never so any ScopedSession opened on a healthcheck worker thread is never
rolled back or removed. If the session ever enters an invalid state rolled back or removed. If the session ever enters an invalid state
@ -2819,6 +3014,11 @@ async def check_db_health():
# --- static assets & files --- # --- static assets & files ---
# Windows registry entries can override these with text/plain, which breaks module and wasm loading
mimetypes.add_type('text/javascript', '.js')
mimetypes.add_type('text/javascript', '.mjs')
mimetypes.add_type('application/wasm', '.wasm')
# Serve build-time static assets (CSS, JS, images, favicon, etc.) # Serve build-time static assets (CSS, JS, images, favicon, etc.)
app.mount('/static', StaticFiles(directory=STATIC_DIR), name='static') app.mount('/static', StaticFiles(directory=STATIC_DIR), name='static')
@ -2864,7 +3064,6 @@ def swagger_ui_html(*args, **kwargs):
applications.get_swagger_ui_html = swagger_ui_html applications.get_swagger_ui_html = swagger_ui_html
if os.path.exists(FRONTEND_BUILD_DIR): if os.path.exists(FRONTEND_BUILD_DIR):
mimetypes.add_type('text/javascript', '.js')
pyodide_dir = FRONTEND_BUILD_DIR / 'pyodide' pyodide_dir = FRONTEND_BUILD_DIR / 'pyodide'
if os.path.exists(pyodide_dir): if os.path.exists(pyodide_dir):
app.mount('/pyodide', CORSStaticFiles(directory=pyodide_dir), name='pyodide') app.mount('/pyodide', CORSStaticFiles(directory=pyodide_dir), name='pyodide')

View file

@ -0,0 +1,28 @@
"""Add group_member user_id index
Revision ID: 1ce6ade7d93b
Revises: f0bd01a18a3d
Create Date: 2026-07-31 03:00:00.000000
"""
import sqlalchemy as sa
from alembic import op
revision = '1ce6ade7d93b'
down_revision = 'f0bd01a18a3d'
branch_labels = None
depends_on = None
def upgrade():
conn = op.get_bind()
inspector = sa.inspect(conn)
existing_indexes = {idx['name'] for idx in inspector.get_indexes('group_member')}
if 'ix_group_member_user_id_group_id' not in existing_indexes:
op.create_index('ix_group_member_user_id_group_id', 'group_member', ['user_id', 'group_id'])
def downgrade():
op.drop_index('ix_group_member_user_id_group_id', table_name='group_member')

View file

@ -0,0 +1,64 @@
"""repair double encoded user oauth
Revision ID: 6d09d1bf1f23
Revises: 1ce6ade7d93b
Create Date: 2026-08-10 23:20:20.374826
"""
import json
from typing import Sequence, Union
from alembic import op
import sqlalchemy as sa
import open_webui.internal.db
# revision identifiers, used by Alembic.
revision: str = '6d09d1bf1f23'
down_revision: Union[str, None] = '1ce6ade7d93b'
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
_user = sa.table(
'user',
sa.column('id', sa.Text),
sa.column('oauth', sa.JSON),
)
def _decode_json_object(value: str) -> dict | None:
try:
decoded = json.loads(value)
except Exception:
return None
return decoded if isinstance(decoded, dict) else None
def upgrade() -> None:
conn = op.get_bind()
inspector = sa.inspect(conn)
if 'user' not in inspector.get_table_names():
return
user_columns = {c['name'] for c in inspector.get_columns('user')}
if 'oauth' not in user_columns:
return
rows = conn.execute(sa.select(_user.c.id, _user.c.oauth).where(_user.c.oauth.is_not(None))).fetchall()
for uid, oauth in rows:
if not isinstance(oauth, str):
continue
decoded = _decode_json_object(oauth)
if decoded is None:
continue
conn.execute(sa.update(_user).where(_user.c.id == uid).values(oauth=decoded))
def downgrade() -> None:
pass

View file

@ -185,9 +185,7 @@ def upgrade() -> None:
for uid, oauth_sub in rows: for uid, oauth_sub in rows:
if oauth_sub: if oauth_sub:
provider, sub = oauth_sub.split('@', 1) if '@' in oauth_sub else ('oidc', oauth_sub) provider, sub = oauth_sub.split('@', 1) if '@' in oauth_sub else ('oidc', oauth_sub)
conn.execute( conn.execute(sa.update(_user).where(_user.c.id == uid).values(oauth={provider: {'sub': sub}}))
sa.update(_user).where(_user.c.id == uid).values(oauth=json.dumps({provider: {'sub': sub}}))
)
# ── Migrate api_key column → api_key table (only if old column still exists) # ── Migrate api_key column → api_key table (only if old column still exists)
if 'api_key' in user_columns: if 'api_key' in user_columns:
@ -226,7 +224,7 @@ def downgrade() -> None:
for uid, oauth in rows: for uid, oauth in rows:
try: try:
data = json.loads(oauth) data = oauth if isinstance(oauth, dict) else json.loads(oauth)
provider = list(data.keys())[0] provider = list(data.keys())[0]
sub = data[provider].get('sub') sub = data[provider].get('sub')
oauth_sub = f'{provider}@{sub}' oauth_sub = f'{provider}@{sub}'

View file

@ -0,0 +1,70 @@
"""add chat timer_at and chat list, unread and timer indexes
Revision ID: d4c1a8e37b62
Revises: 6d09d1bf1f23
Create Date: 2026-08-23 18:05:12.441907
"""
from collections.abc import Sequence
import sqlalchemy as sa
from alembic import op
# revision identifiers, used by Alembic.
revision: str = 'd4c1a8e37b62'
down_revision: str | None = '6d09d1bf1f23'
branch_labels: str | Sequence[str] | None = None
depends_on: str | Sequence[str] | None = None
def upgrade() -> None:
op.add_column('chat', sa.Column('timer_at', sa.BigInteger(), nullable=True))
op.create_index(
'timer_at_idx',
'chat',
['timer_at'],
sqlite_where=sa.text('timer_at IS NOT NULL'),
postgresql_where=sa.text('timer_at IS NOT NULL'),
)
# Timers created before this migration carry their due time in meta only, and would never fire.
chat = sa.table(
'chat', sa.column('id', sa.String), sa.column('meta', sa.JSON), sa.column('timer_at', sa.BigInteger)
)
conn = op.get_bind()
pending = conn.execute(
sa.select(chat.c.id, chat.c.meta)
.where(chat.c.meta['type'].as_string() == 'timer')
.where(chat.c.meta['status'].as_string() == 'pending')
).all()
for chat_id, meta in pending:
try: # imported chats can carry any meta, and a non-numeric due time must not abort the migration
due_at = int(meta.get('timer_at'))
except (TypeError, ValueError):
continue
conn.execute(chat.update().where(chat.c.id == chat_id).values(timer_at=due_at))
op.create_index('user_id_updated_at_id_idx', 'chat', ['user_id', sa.text('updated_at DESC'), 'id'])
op.create_index(
'user_id_timer_at_idx',
'chat',
['user_id', 'timer_at'],
sqlite_where=sa.text('timer_at IS NOT NULL'),
postgresql_where=sa.text('timer_at IS NOT NULL'),
)
op.create_index(
'user_id_folder_unread_idx',
'chat',
['user_id', 'folder_id', 'archived', 'updated_at', 'last_read_at', 'id'],
)
op.create_index('chat_message_chat_role_done_idx', 'chat_message', ['chat_id', 'role', 'done'])
def downgrade() -> None:
op.drop_index('chat_message_chat_role_done_idx', table_name='chat_message')
op.drop_index('user_id_folder_unread_idx', table_name='chat')
op.drop_index('user_id_timer_at_idx', table_name='chat')
op.drop_index('user_id_updated_at_id_idx', table_name='chat')
op.drop_index('timer_at_idx', table_name='chat')
op.drop_column('chat', 'timer_at')

View file

@ -839,7 +839,8 @@ class AccessGrantsTable:
): ):
""" """
Filter for items where user has read BUT NOT write access. Filter for items where user has read BUT NOT write access.
Public items are NOT considered read_only. A public (user:*) read grant counts as read access, so publicly shared
read-only items are listed rather than being reachable only by direct link.
Note: This method builds SQLAlchemy expressions and does NOT perform I/O itself, Note: This method builds SQLAlchemy expressions and does NOT perform I/O itself,
so it remains synchronous. The caller is responsible for executing the query so it remains synchronous. The caller is responsible for executing the query
@ -850,7 +851,6 @@ class AccessGrantsTable:
from sqlalchemy import exists as sa_exists from sqlalchemy import exists as sa_exists
# Has read grant (not public)
read_grant_exists = ( read_grant_exists = (
select(AccessGrant.id) select(AccessGrant.id)
.where( .where(
@ -858,6 +858,10 @@ class AccessGrantsTable:
AccessGrant.resource_id == DocumentModel.id, AccessGrant.resource_id == DocumentModel.id,
AccessGrant.permission == 'read', AccessGrant.permission == 'read',
or_( or_(
and_(
AccessGrant.principal_type == 'user',
AccessGrant.principal_id == '*',
),
*( *(
[ [
and_( and_(
@ -884,7 +888,6 @@ class AccessGrantsTable:
.exists() .exists()
) )
# Does NOT have write grant
write_grant_exists = ( write_grant_exists = (
select(AccessGrant.id) select(AccessGrant.id)
.where( .where(
@ -892,6 +895,10 @@ class AccessGrantsTable:
AccessGrant.resource_id == DocumentModel.id, AccessGrant.resource_id == DocumentModel.id,
AccessGrant.permission == 'write', AccessGrant.permission == 'write',
or_( or_(
and_(
AccessGrant.principal_type == 'user',
AccessGrant.principal_id == '*',
),
*( *(
[ [
and_( and_(
@ -918,21 +925,7 @@ class AccessGrantsTable:
.exists() .exists()
) )
# Is NOT public conditions = [read_grant_exists, ~write_grant_exists]
public_grant_exists = (
select(AccessGrant.id)
.where(
AccessGrant.resource_type == resource_type,
AccessGrant.resource_id == DocumentModel.id,
AccessGrant.permission == 'read',
AccessGrant.principal_type == 'user',
AccessGrant.principal_id == '*',
)
.correlate(DocumentModel)
.exists()
)
conditions = [read_grant_exists, ~write_grant_exists, ~public_grant_exists]
# Not owner # Not owner
if user_id: if user_id:

View file

@ -9,7 +9,7 @@ from typing import Optional
import bcrypt import bcrypt
from open_webui.internal.db import Base, JSONField, get_async_db_context from open_webui.internal.db import Base, JSONField, get_async_db_context
from open_webui.models.users import User, UserModel, UserProfileImageResponse, Users from open_webui.models.users import User, UserModel, UserProfileImageResponse, Users
from open_webui.utils.validate import validate_profile_image_url from open_webui.utils.validate import validate_image_url
from pydantic import BaseModel, field_validator from pydantic import BaseModel, field_validator
from sqlalchemy import Boolean, Column, String, Text, delete, select, update from sqlalchemy import Boolean, Column, String, Text, delete, select, update
from sqlalchemy.exc import IntegrityError from sqlalchemy.exc import IntegrityError
@ -87,7 +87,7 @@ class SignupForm(BaseModel):
@classmethod @classmethod
def check_profile_image_url(cls, v: str | None) -> str | None: def check_profile_image_url(cls, v: str | None) -> str | None:
if v is not None: if v is not None:
return validate_profile_image_url(v) return validate_image_url(v)
return v return v

View file

@ -1,9 +1,10 @@
import logging import logging
import time import time
from typing import Optional from typing import Literal, Optional
from uuid import uuid4 from uuid import uuid4
from open_webui.internal.db import Base, get_async_db_context from open_webui.internal.db import Base, get_async_db_context
from open_webui.utils.misc import json_text_variants
from pydantic import BaseModel, ConfigDict from pydantic import BaseModel, ConfigDict
from sqlalchemy import JSON, BigInteger, Boolean, Column, Index, String, Text, cast, delete, func, or_, select, update from sqlalchemy import JSON, BigInteger, Boolean, Column, Index, String, Text, cast, delete, func, or_, select, update
from sqlalchemy.ext.asyncio import AsyncSession from sqlalchemy.ext.asyncio import AsyncSession
@ -64,11 +65,17 @@ class AutomationTerminalConfig(BaseModel):
cwd: Optional[str] = None cwd: Optional[str] = None
class AutomationTarget(BaseModel):
type: Literal['chat', 'channel'] = 'chat'
channel_id: Optional[str] = None
class AutomationData(BaseModel): class AutomationData(BaseModel):
prompt: str prompt: str
model_id: str model_id: str
rrule: str rrule: str
terminal: Optional[AutomationTerminalConfig] = None terminal: Optional[AutomationTerminalConfig] = None
target: Optional[AutomationTarget] = None
class AutomationModel(BaseModel): class AutomationModel(BaseModel):
@ -179,16 +186,16 @@ class AutomationTable:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
stmt = select(Automation).filter_by(user_id=user_id) stmt = select(Automation).filter_by(user_id=user_id)
if folder_id is not None: if folder_id:
stmt = stmt.filter(Automation.folder_id == (folder_id or None)) stmt = stmt.filter(Automation.folder_id == folder_id)
if query: if query:
search = f'%{query}%' # Search the name column and the prompt inside the JSON data.
# Search in name and prompt inside JSON data data_text = cast(Automation.data, String)
stmt = stmt.filter( stmt = stmt.filter(
or_( or_(
Automation.name.ilike(search), Automation.name.ilike(f'%{query}%'),
cast(Automation.data, String).ilike(search), *(data_text.ilike(f'%{variant}%') for variant in json_text_variants(query)),
) )
) )
@ -305,6 +312,7 @@ class AutomationTable:
rows = result.scalars().all() rows = result.scalars().all()
from open_webui.utils.automations import next_run_ns from open_webui.utils.automations import next_run_ns
from open_webui.utils.recurrence import RecurrenceEvaluationTimeout
# Batch-fetch user timezones so rescheduling respects each # Batch-fetch user timezones so rescheduling respects each
# user's local timezone instead of falling back to server time. # user's local timezone instead of falling back to server time.
@ -316,13 +324,20 @@ class AutomationTable:
tz_result = await db.execute(select(User.id, User.timezone).where(User.id.in_(user_ids))) tz_result = await db.execute(select(User.id, User.timezone).where(User.id.in_(user_ids)))
timezone_by_user_id = {uid: tz for uid, tz in tz_result.all()} timezone_by_user_id = {uid: tz for uid, tz in tz_result.all()}
claimed = []
for row in rows: for row in rows:
try:
next_run_at = await next_run_ns(row.data.get('rrule', ''), tz=timezone_by_user_id.get(row.user_id))
except RecurrenceEvaluationTimeout:
log.warning('Skipping automation %s: recurrence evaluation timed out', row.id)
continue
row.last_run_at = now_ns row.last_run_at = now_ns
row.next_run_at = next_run_ns(row.data.get('rrule', ''), tz=timezone_by_user_id.get(row.user_id)) row.next_run_at = next_run_at
claimed.append(row)
await db.commit() await db.commit()
return [AutomationModel.model_validate(r) for r in rows] return [AutomationModel.model_validate(r) for r in claimed]
#################### ####################

View file

@ -4,6 +4,7 @@ from typing import Optional
from uuid import uuid4 from uuid import uuid4
from open_webui.internal.db import Base, get_async_db_context from open_webui.internal.db import Base, get_async_db_context
from open_webui.constants import ERROR_MESSAGES
from open_webui.models.access_grants import AccessGrantModel, AccessGrants from open_webui.models.access_grants import AccessGrantModel, AccessGrants
from open_webui.models.groups import Groups from open_webui.models.groups import Groups
from open_webui.models.users import User, UserModel, UserResponse from open_webui.models.users import User, UserModel, UserResponse
@ -26,6 +27,7 @@ from sqlalchemy import (
from sqlalchemy.ext.asyncio import AsyncSession from sqlalchemy.ext.asyncio import AsyncSession
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
MIN_CALENDAR_RRULE_INTERVAL_SECONDS = 24 * 60 * 60
#################### ####################
@ -177,6 +179,20 @@ class CalendarUpdateForm(BaseModel):
access_grants: Optional[list[dict]] = None access_grants: Optional[list[dict]] = None
async def validate_calendar_rrule(value: Optional[str]) -> None:
if value:
from open_webui.utils.recurrence import rrule_interval_seconds
try:
interval = await rrule_interval_seconds(value)
except ValueError:
raise
except Exception as e:
raise ValueError(ERROR_MESSAGES.AUTOMATION_INVALID_RRULE(e)) from e
if interval is not None and interval < MIN_CALENDAR_RRULE_INTERVAL_SECONDS:
raise ValueError(ERROR_MESSAGES.CALENDAR_RRULE_TOO_FREQUENT)
class CalendarEventForm(BaseModel): class CalendarEventForm(BaseModel):
calendar_id: str calendar_id: str
title: str title: str
@ -431,6 +447,7 @@ class CalendarEventTable:
async def insert_new_event( async def insert_new_event(
self, user_id: str, form_data: CalendarEventForm, db: Optional[AsyncSession] = None self, user_id: str, form_data: CalendarEventForm, db: Optional[AsyncSession] = None
) -> Optional[CalendarEventModel]: ) -> Optional[CalendarEventModel]:
await validate_calendar_rrule(form_data.rrule)
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
now = int(time.time_ns()) now = int(time.time_ns())
event = CalendarEvent( event = CalendarEvent(
@ -533,7 +550,8 @@ class CalendarEventTable:
& (CalendarEvent.start_at < end) & (CalendarEvent.start_at < end)
& or_( & or_(
CalendarEvent.end_at.is_(None) & (CalendarEvent.start_at >= start), CalendarEvent.end_at.is_(None) & (CalendarEvent.start_at >= start),
CalendarEvent.end_at.isnot(None) & (CalendarEvent.end_at > start), CalendarEvent.end_at.isnot(None)
& ((CalendarEvent.end_at > start) | (CalendarEvent.start_at >= start)),
) )
), ),
# Recurring: fetch all (expansion in Python) # Recurring: fetch all (expansion in Python)
@ -660,6 +678,7 @@ class CalendarEventTable:
async def update_event_by_id( async def update_event_by_id(
self, id: str, form_data: CalendarEventUpdateForm, db: Optional[AsyncSession] = None self, id: str, form_data: CalendarEventUpdateForm, db: Optional[AsyncSession] = None
) -> Optional[CalendarEventModel]: ) -> Optional[CalendarEventModel]:
await validate_calendar_rrule(form_data.rrule)
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
result = await db.execute(select(CalendarEvent).filter(CalendarEvent.id == id)) result = await db.execute(select(CalendarEvent).filter(CalendarEvent.id == id))
event = result.scalars().first() event = result.scalars().first()
@ -734,10 +753,10 @@ class CalendarEventTable:
events = [] events = []
for event, tz in rows: for event, tz in rows:
model = CalendarEventModel.model_validate(event) model = CalendarEventModel.model_validate(event)
# Determine per-event alert window # meta is user-writable and this poll is shared by every user.
alert_minutes = None alert_minutes = (model.meta or {}).get('alert_minutes')
if model.meta and 'alert_minutes' in model.meta: if not isinstance(alert_minutes, (int, float)):
alert_minutes = model.meta['alert_minutes'] alert_minutes = None
if alert_minutes is not None: if alert_minutes is not None:
if alert_minutes < 0: if alert_minutes < 0:

View file

@ -1,4 +1,3 @@
import json
import secrets import secrets
import time import time
import uuid import uuid
@ -10,7 +9,8 @@ from open_webui.models.access_grants import (
AccessGrants, AccessGrants,
) )
from open_webui.models.groups import Groups from open_webui.models.groups import Groups
from open_webui.utils.validate import validate_profile_image_url from open_webui.models.users import User
from open_webui.utils.validate import validate_image_url
from pydantic import BaseModel, ConfigDict, Field, field_validator from pydantic import BaseModel, ConfigDict, Field, field_validator
from sqlalchemy import ( from sqlalchemy import (
JSON, JSON,
@ -253,7 +253,7 @@ class ChannelWebhookForm(BaseModel):
def check_profile_image_url(cls, v: Optional[str]) -> Optional[str]: def check_profile_image_url(cls, v: Optional[str]) -> Optional[str]:
if v is None: if v is None:
return v return v
return validate_profile_image_url(v) return validate_image_url(v)
class ChannelTable: class ChannelTable:
@ -438,17 +438,17 @@ class ChannelTable:
match_count = func.sum( match_count = func.sum(
case( case(
(ChannelMember.user_id.in_(unique_user_ids), 1), (User.id.in_(unique_user_ids), 1),
else_=0, else_=0,
) )
) )
subquery = ( subquery = (
select(ChannelMember.channel_id) select(ChannelMember.channel_id)
.join(User, User.id == ChannelMember.user_id)
.group_by(ChannelMember.channel_id) .group_by(ChannelMember.channel_id)
# 1. Channel must have exactly len(user_ids) members # Match the exact set of accounts that still exist.
.having(func.count(ChannelMember.user_id) == len(unique_user_ids)) .having(func.count(User.id) == len(unique_user_ids))
# 2. All those members must be in unique_user_ids
.having(match_count == len(unique_user_ids)) .having(match_count == len(unique_user_ids))
.subquery() .subquery()
) )

View file

@ -1,4 +1,3 @@
import json
import time import time
import uuid import uuid
from collections import Counter from collections import Counter
@ -169,6 +168,7 @@ class ChatMessage(Base):
Index('chat_message_chat_parent_idx', 'chat_id', 'parent_id'), Index('chat_message_chat_parent_idx', 'chat_id', 'parent_id'),
Index('chat_message_model_created_idx', 'model_id', 'created_at'), Index('chat_message_model_created_idx', 'model_id', 'created_at'),
Index('chat_message_user_created_idx', 'user_id', 'created_at'), Index('chat_message_user_created_idx', 'user_id', 'created_at'),
Index('chat_message_chat_role_done_idx', 'chat_id', 'role', 'done'), # unfinished-assistant probe
) )
@ -207,6 +207,66 @@ class ChatMessageModel(BaseModel):
class ChatMessageTable: class ChatMessageTable:
@staticmethod
def _apply_message_data(message: ChatMessage, data: dict, now: int) -> None:
"""Overwrite only the fields the payload carries."""
if 'role' in data:
message.role = data['role']
if 'parent_id' in data or 'parentId' in data:
message.parent_id = data.get('parent_id') or data.get('parentId')
if 'content' in data:
message.content = data.get('content')
if 'output' in data:
message.output = data.get('output')
if 'model_id' in data or 'model' in data:
message.model_id = data.get('model_id') or data.get('model')
if 'files' in data:
message.files = data.get('files')
if 'sources' in data:
message.sources = data.get('sources')
if 'embeds' in data:
message.embeds = data.get('embeds')
if 'meta' in data:
message.meta = data.get('meta')
if 'done' in data:
message.done = data['done']
if 'status_history' in data or 'statusHistory' in data:
message.status_history = data.get('status_history') or data.get('statusHistory')
if 'error' in data:
message.error = data.get('error')
if 'context_summary' in data or 'contextSummary' in data:
message.context_summary = data.get('context_summary') or data.get('contextSummary')
usage = get_usage(data)
if usage:
existing_usage = normalize_usage(message.usage)
message.usage = existing_usage if usage == existing_usage else merge_usage(existing_usage, usage)
message.updated_at = now
@staticmethod
def _build_message(composite_id: str, chat_id: str, user_id: str, data: dict, now: int) -> ChatMessage:
return ChatMessage(
id=composite_id,
chat_id=chat_id,
user_id=user_id,
role=data.get('role', 'user'),
parent_id=data.get('parent_id') or data.get('parentId'),
content=data.get('content'),
output=data.get('output'),
model_id=data.get('model_id') or data.get('model'),
files=data.get('files'),
sources=data.get('sources'),
embeds=data.get('embeds'),
meta=data.get('meta'),
done=data.get('done', True),
status_history=data.get('status_history') or data.get('statusHistory'),
error=data.get('error'),
usage=get_usage(data),
context_summary=data.get('context_summary') or data.get('contextSummary'),
created_at=data.get('timestamp', now),
updated_at=now,
)
async def upsert_message( async def upsert_message(
self, self,
message_id: str, message_id: str,
@ -218,76 +278,46 @@ class ChatMessageTable:
"""Insert or update a chat message.""" """Insert or update a chat message."""
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
now = int(time.time()) now = int(time.time())
timestamp = data.get('timestamp', now)
# Use composite ID: {chat_id}-{message_id} # Use composite ID: {chat_id}-{message_id}
composite_id = f'{chat_id}-{message_id}' composite_id = f'{chat_id}-{message_id}'
existing = await db.get(ChatMessage, composite_id) message = await db.get(ChatMessage, composite_id)
if existing: if message:
# Update existing self._apply_message_data(message, data, now)
if 'role' in data:
existing.role = data['role']
if 'parent_id' in data or 'parentId' in data:
existing.parent_id = data.get('parent_id') or data.get('parentId')
if 'content' in data:
existing.content = data.get('content')
if 'output' in data:
existing.output = data.get('output')
if 'model_id' in data or 'model' in data:
existing.model_id = data.get('model_id') or data.get('model')
if 'files' in data:
existing.files = data.get('files')
if 'sources' in data:
existing.sources = data.get('sources')
if 'embeds' in data:
existing.embeds = data.get('embeds')
if 'meta' in data:
existing.meta = data.get('meta')
if 'done' in data:
existing.done = data.get('done', True)
if 'status_history' in data or 'statusHistory' in data:
existing.status_history = data.get('status_history') or data.get('statusHistory')
if 'error' in data:
existing.error = data.get('error')
if 'context_summary' in data or 'contextSummary' in data:
existing.context_summary = data.get('context_summary') or data.get('contextSummary')
# Extract and normalize usage
usage = get_usage(data)
if usage:
existing_usage = normalize_usage(existing.usage or {}) if existing.usage else {}
existing.usage = existing_usage if usage == existing_usage else merge_usage(existing_usage, usage)
existing.updated_at = now
await db.commit()
return ChatMessageModel.model_validate(existing)
else: else:
# Insert new message = self._build_message(composite_id, chat_id, user_id, data, now)
# Extract and normalize usage
usage = get_usage(data)
message = ChatMessage(
id=composite_id,
chat_id=chat_id,
user_id=user_id,
role=data.get('role', 'user'),
parent_id=data.get('parent_id') or data.get('parentId'),
content=data.get('content'),
output=data.get('output'),
model_id=data.get('model_id') or data.get('model'),
files=data.get('files'),
sources=data.get('sources'),
embeds=data.get('embeds'),
meta=data.get('meta'),
done=data.get('done', True),
status_history=data.get('status_history') or data.get('statusHistory'),
error=data.get('error'),
usage=usage,
context_summary=data.get('context_summary') or data.get('contextSummary'),
created_at=timestamp,
updated_at=now,
)
db.add(message) db.add(message)
await db.commit()
return ChatMessageModel.model_validate(message) await db.commit()
return ChatMessageModel.model_validate(message)
async def upsert_messages(
self,
chat_id: str,
user_id: str,
messages: dict[str, dict],
db: AsyncSession | None = None,
) -> None:
"""Insert or update the given messages of one chat."""
if not messages:
return
async with get_async_db_context(db) as db:
now = int(time.time())
result = await db.execute(
select(ChatMessage).filter(ChatMessage.id.in_([f'{chat_id}-{message_id}' for message_id in messages]))
)
existing_by_id = {row.id: row for row in result.scalars().all()}
for message_id, data in messages.items():
composite_id = f'{chat_id}-{message_id}'
message = existing_by_id.get(composite_id)
if message:
self._apply_message_data(message, data, now)
else:
db.add(self._build_message(composite_id, chat_id, user_id, data, now))
await db.commit()
async def get_message_by_id(self, id: str, db: Optional[AsyncSession] = None) -> Optional[ChatMessageModel]: async def get_message_by_id(self, id: str, db: Optional[AsyncSession] = None) -> Optional[ChatMessageModel]:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:

View file

@ -2,13 +2,16 @@
from __future__ import annotations from __future__ import annotations
import json
import logging import logging
import re
import time import time
import uuid import uuid
from typing import Any, Literal
# local imports # local imports
from open_webui.env import ENABLE_ADMIN_CHAT_ACCESS
from open_webui.internal.db import Base, JSONField, get_async_db_context from open_webui.internal.db import Base, JSONField, get_async_db_context
from open_webui.models.access_grants import AccessGrants
from open_webui.models.automations import AutomationRun from open_webui.models.automations import AutomationRun
from open_webui.models.chat_messages import ChatMessage, ChatMessages from open_webui.models.chat_messages import ChatMessage, ChatMessages
from open_webui.models.folders import Folders from open_webui.models.folders import Folders
@ -41,6 +44,62 @@ from sqlalchemy.sql.expression import bindparam
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
ACTIVE_CHAT_GAP_SECONDS = 30 * 60 ACTIVE_CHAT_GAP_SECONDS = 30 * 60
CHAT_SEARCH_FILTER_PREFIXES = ('tag:', 'folder:', 'pinned:', 'archived:', 'shared:')
def chat_search_content_query(text: str) -> str:
words = sanitize_text_for_db(text).lower().strip().split()
return ' '.join(word for word in words if not word.startswith(CHAT_SEARCH_FILTER_PREFIXES)).strip()
def chat_search_terms(text: str) -> list[str]:
return list(dict.fromkeys(re.findall(r'[a-z0-9]+', text.lower())))
def chat_search_message_content_match_sql(dialect_name: str, key: str) -> str:
if dialect_name == 'sqlite':
return f"""
(
EXISTS (
SELECT 1
FROM json_each(Chat.chat, '$.history.messages') AS history_message
WHERE LOWER(history_message.value->>'content') LIKE '%' || :{key} || '%'
)
OR EXISTS (
SELECT 1
FROM json_each(Chat.chat, '$.messages') AS legacy_message
WHERE LOWER(legacy_message.value->>'content') LIKE '%' || :{key} || '%'
)
)
"""
if dialect_name == 'postgresql':
return f"""
(
EXISTS (
SELECT 1
FROM chat_message AS message
WHERE message.chat_id = Chat.id
AND message.user_id = Chat.user_id
AND json_typeof(message.content) = 'string'
AND LOWER(message.content #>> '{{}}') LIKE '%' || :{key} || '%'
)
OR EXISTS (
SELECT 1
FROM json_each(Chat.chat#>'{{history,messages}}') AS history_message
WHERE json_typeof(history_message.value->'content') = 'string'
AND LOWER(history_message.value->>'content') LIKE '%' || :{key} || '%'
)
OR EXISTS (
SELECT 1
FROM json_array_elements(Chat.chat->'messages') AS legacy_message
WHERE json_typeof(legacy_message->'content') = 'string'
AND LOWER(legacy_message->>'content') LIKE '%' || :{key} || '%'
)
)
"""
raise NotImplementedError(f'Unsupported dialect: {dialect_name}')
def chat_list_order(sort_by: str = 'updated_at', sort_dir: str = 'desc', user_id: str | None = None): def chat_list_order(sort_by: str = 'updated_at', sort_dir: str = 'desc', user_id: str | None = None):
@ -91,6 +150,7 @@ class Chat(Base): # database table mapping for chat entity
current_message_id = Column(Text, nullable=True) current_message_id = Column(Text, nullable=True)
last_read_at = Column(BigInteger, nullable=True) last_read_at = Column(BigInteger, nullable=True)
timer_at = Column(BigInteger, nullable=True) # ns due time, set only while a timer chat waits to be claimed
__table_args__ = ( __table_args__ = (
# Performance indexes for common queries # Performance indexes for common queries
@ -99,6 +159,23 @@ class Chat(Base): # database table mapping for chat entity
Index('user_id_archived_idx', 'user_id', 'archived'), Index('user_id_archived_idx', 'user_id', 'archived'),
Index('updated_at_user_id_idx', 'updated_at', 'user_id'), Index('updated_at_user_id_idx', 'updated_at', 'user_id'),
Index('folder_id_user_id_idx', 'folder_id', 'user_id'), Index('folder_id_user_id_idx', 'folder_id', 'user_id'),
Index('user_id_updated_at_id_idx', 'user_id', updated_at.desc(), 'id'),
Index(
'timer_at_idx',
'timer_at',
sqlite_where=text('timer_at IS NOT NULL'),
postgresql_where=text('timer_at IS NOT NULL'),
),
# timer_at key column turns the IS NOT NULL into a seek, so this beats the plain user_id indexes
Index(
'user_id_timer_at_idx',
'user_id',
'timer_at',
sqlite_where=text('timer_at IS NOT NULL'),
postgresql_where=text('timer_at IS NOT NULL'),
),
# covering index: lets SQLite serve count_unread_by_folder_ids without reading chat rows
Index('user_id_folder_unread_idx', 'user_id', 'folder_id', 'archived', 'updated_at', 'last_read_at', 'id'),
) )
@ -129,6 +206,7 @@ class ChatModel(BaseModel):
current_message_id: str | None = None current_message_id: str | None = None
last_read_at: int | None = None last_read_at: int | None = None
timer_at: int | None = None
@field_validator('variables', mode='before') @field_validator('variables', mode='before')
@classmethod @classmethod
@ -180,6 +258,7 @@ class ChatForm(BaseModel):
class ChatImportForm(ChatForm): class ChatImportForm(ChatForm):
meta: dict | None = {} meta: dict | None = {}
pinned: bool | None = False pinned: bool | None = False
archived: bool | None = False
current_message_id: str | None = None current_message_id: str | None = None
created_at: int | None = None created_at: int | None = None
updated_at: int | None = None updated_at: int | None = None
@ -189,11 +268,6 @@ class ChatsImportForm(BaseModel):
chats: list[ChatImportForm] chats: list[ChatImportForm]
class ChatTitleMessagesForm(BaseModel):
title: str
messages: list[dict]
class ChatTitleForm(BaseModel): class ChatTitleForm(BaseModel):
title: str title: str
@ -231,6 +305,7 @@ class ChatTitleIdResponse(BaseModel):
last_read_at: int | None = None last_read_at: int | None = None
snippet: str | None = None snippet: str | None = None
active: bool = False active: bool = False
archived: bool = False
class SharedChatResponse(BaseModel): class SharedChatResponse(BaseModel):
@ -284,7 +359,7 @@ class MessageStats(BaseModel):
token_count: int | None = None token_count: int | None = None
timestamp: int | None = None timestamp: int | None = None
rating: int | None = None # Derived from message.annotation.rating rating: int | None = None # Derived from message.annotation.rating
tags: list[str | None] = None # Derived from message.annotation.tags tags: list[str] | None = None # Derived from message.annotation.tags
class ChatHistoryStats(BaseModel): class ChatHistoryStats(BaseModel):
@ -361,6 +436,37 @@ class ChatTable:
return changed return changed
@staticmethod
def _last_descendant_id(messages: dict, message_id: str) -> str:
seen_ids = set()
while message_id in messages and message_id not in seen_ids:
seen_ids.add(message_id)
message = messages[message_id]
child_ids = message.get('childrenIds') if isinstance(message, dict) else []
child_ids = child_ids if isinstance(child_ids, list) else []
next_id = next((child_id for child_id in reversed(child_ids) if child_id in messages), None)
if not next_id:
break
message_id = next_id
return message_id
@staticmethod
def _add_child_id_to_parent(messages: dict, parent_id: str | None, child_id: str) -> bool:
parent = messages.get(parent_id) if parent_id else None
if not isinstance(parent, dict):
return False
child_ids = parent.get('childrenIds')
if not isinstance(child_ids, list):
child_ids = []
parent['childrenIds'] = child_ids
if child_id in child_ids:
return False
child_ids.append(child_id)
return True
def _repair_chat_current_id(self, chat: dict) -> bool: def _repair_chat_current_id(self, chat: dict) -> bool:
history = chat.get('history') history = chat.get('history')
if not isinstance(history, dict): if not isinstance(history, dict):
@ -370,6 +476,12 @@ class ChatTable:
if not isinstance(messages, dict): if not isinstance(messages, dict):
return False return False
changed = False
for message_id, message in messages.items():
if not isinstance(message, dict):
continue
changed = self._add_child_id_to_parent(messages, message.get('parentId'), message_id) or changed
current_id = history.get('currentId') current_id = history.get('currentId')
current_message = messages.get(current_id) current_message = messages.get(current_id)
output = [] output = []
@ -393,7 +505,13 @@ class ChatTable:
and current_message.get('role') and current_message.get('role')
and not current_is_bad_leaf and not current_is_bad_leaf
): ):
return False if current_message.get('contextSummary') or current_message.get('context_summary'):
last_descendant_id = self._last_descendant_id(messages, current_id)
if last_descendant_id != current_id:
history['currentId'] = last_descendant_id
return True
return changed
latest_leaf_id = None latest_leaf_id = None
latest_timestamp = -1 latest_timestamp = -1
@ -408,7 +526,7 @@ class ChatTable:
latest_timestamp = timestamp latest_timestamp = timestamp
if not latest_leaf_id or latest_leaf_id == current_id: if not latest_leaf_id or latest_leaf_id == current_id:
return False return changed
history['currentId'] = latest_leaf_id history['currentId'] = latest_leaf_id
return True return True
@ -421,6 +539,7 @@ class ChatTable:
db: AsyncSession | None = None, db: AsyncSession | None = None,
*, *,
internal_meta: dict | None = None, internal_meta: dict | None = None,
timer_at: int | None = None,
) -> ChatModel | None: ) -> ChatModel | None:
async with get_async_db_context(db) as session: async with get_async_db_context(db) as session:
chat = ChatModel( chat = ChatModel(
@ -433,6 +552,7 @@ class ChatTable:
'chat': self._clean_null_bytes(form_data.chat), 'chat': self._clean_null_bytes(form_data.chat),
'folder_id': form_data.folder_id, 'folder_id': form_data.folder_id,
'meta': internal_meta or {}, 'meta': internal_meta or {},
'timer_at': timer_at,
'variables': form_data.variables or {}, 'variables': form_data.variables or {},
'current_message_id': self.get_current_message_id(form_data.chat), 'current_message_id': self.get_current_message_id(form_data.chat),
'created_at': int(time.time()), 'created_at': int(time.time()),
@ -523,6 +643,7 @@ class ChatTable:
'meta': form_data.meta, 'meta': form_data.meta,
'variables': form_data.variables or {}, 'variables': form_data.variables or {},
'pinned': form_data.pinned, 'pinned': form_data.pinned,
'archived': form_data.archived,
'folder_id': form_data.folder_id, 'folder_id': form_data.folder_id,
'current_message_id': form_data.current_message_id or self.get_current_message_id(form_data.chat), 'current_message_id': form_data.current_message_id or self.get_current_message_id(form_data.chat),
'created_at': (form_data.created_at if form_data.created_at else int(time.time())), 'created_at': (form_data.created_at if form_data.created_at else int(time.time())),
@ -596,17 +717,29 @@ class ChatTable:
*, *,
touch: bool = True, touch: bool = True,
) -> ChatModel | None: ) -> ChatModel | None:
"""Persist updated chat content, sanitizing null bytes.""" """Patch top-level chat keys; history is merged so stale writers don't drop messages."""
try: # load the chat record for in-place mutation try:
async with get_async_db_context(db) as session: async with get_async_db_context(db) as session:
chat_item = await session.get(Chat, id) chat_item = await session.get(
Chat,
id,
populate_existing=True,
with_for_update=session.bind.dialect.name == 'postgresql',
)
if chat_item is None: if chat_item is None:
return None return None
chat_item.chat = self._clean_null_bytes(chat) stored = chat_item.chat or {}
chat_item.title = self._clean_null_bytes(chat['title']) if 'title' in chat else 'New Chat' updated = {**stored, **chat}
if 'history' in chat:
# The caller built its history from an earlier read; merge so messages saved since then survive.
updated['history'] = self.merge_history(stored.get('history'), chat['history'])
updated = self._clean_null_bytes(updated)
chat_item.chat = updated
chat_item.title = updated.get('title', 'New Chat')
if any(key in chat for key in ('history', 'messages', 'currentId', 'branchPointMessageId')): if any(key in chat for key in ('history', 'messages', 'currentId', 'branchPointMessageId')):
chat_item.current_message_id = self.get_current_message_id(chat) chat_item.current_message_id = self.get_current_message_id(updated)
if touch: if touch:
chat_item.updated_at = int(time.time()) chat_item.updated_at = int(time.time())
@ -713,7 +846,12 @@ class ChatTable:
async def update_chat_title_by_id(self, id: str, title: str) -> ChatModel | None: async def update_chat_title_by_id(self, id: str, title: str) -> ChatModel | None:
try: try:
async with get_async_db_context() as session: async with get_async_db_context() as session:
chat_item = await session.get(Chat, id) chat_item = await session.get(
Chat,
id,
populate_existing=True,
with_for_update=session.bind.dialect.name == 'postgresql',
)
if chat_item is None: if chat_item is None:
return None return None
clean_title = self._clean_null_bytes(title) clean_title = self._clean_null_bytes(title)
@ -774,11 +912,12 @@ class ChatTable:
def merge_history(existing_history: dict | None, incoming_history: dict | None) -> dict: def merge_history(existing_history: dict | None, incoming_history: dict | None) -> dict:
existing = (existing_history or {}).get('messages') or {} existing = (existing_history or {}).get('messages') or {}
incoming = (incoming_history or {}).get('messages') or {} incoming = (incoming_history or {}).get('messages') or {}
merged = {**existing, **incoming} merged = {
merged = {message_id: message for message_id, message in merged.items() if isinstance(message, dict)} message_id: {**message, 'childrenIds': []}
for message_id, message in {**existing, **incoming}.items()
if isinstance(message, dict)
}
for message in merged.values():
message['childrenIds'] = []
for message_id, message in merged.items(): for message_id, message in merged.items():
parent_id = message.get('parentId') parent_id = message.get('parentId')
if parent_id in merged: if parent_id in merged:
@ -826,8 +965,10 @@ class ChatTable:
if current_id is None if current_id is None
else messages.get(current_id, {}).get('childrenIds', []) else messages.get(current_id, {}).get('childrenIds', [])
) )
while child_ids: visited_ids = set()
while child_ids and child_ids[-1] not in visited_ids:
current_id = child_ids[-1] current_id = child_ids[-1]
visited_ids.add(current_id)
child_ids = messages.get(current_id, {}).get('childrenIds', []) child_ids = messages.get(current_id, {}).get('childrenIds', [])
history['currentId'] = current_id if current_id in messages else None history['currentId'] = current_id if current_id in messages else None
return deleted_ids return deleted_ids
@ -874,26 +1015,24 @@ class ChatTable:
'role': role, 'role': role,
'timestamp': message.get('timestamp') or int(time.time()), 'timestamp': message.get('timestamp') or int(time.time()),
} }
history['currentId'] = message_id
history['currentId'] = message_id ChatTable._add_child_id_to_parent(messages, messages[message_id].get('parentId'), message_id)
return messages[message_id] return messages[message_id]
async def backfill_messages_by_chat_id(self, chat_id: str, user_id: str, messages: dict[str, dict]) -> None: async def backfill_messages_by_chat_id(self, chat_id: str, user_id: str, messages: dict[str, dict]) -> None:
"""Write messages to the ``chat_message`` table so future lookups """Write messages to the ``chat_message`` table so future lookups
use the fast path. Errors are logged but never raised. use the fast path. Errors are logged but never raised.
""" """
for message_id, message in messages.items(): writable = {
if not isinstance(message, dict) or not message.get('role'): message_id: message
continue for message_id, message in messages.items()
try: if isinstance(message, dict) and message.get('role')
await ChatMessages.upsert_message( }
message_id=message_id, try:
chat_id=chat_id, await ChatMessages.upsert_messages(chat_id, user_id, writable)
user_id=user_id, except Exception as e:
data=message, log.warning('Backfill failed for chat %s: %s', chat_id, e)
)
except Exception as e:
log.warning('Backfill failed for message %s in chat %s: %s', message_id, chat_id, e)
async def reconcile_messages_by_chat_id(self, chat_id: str, user_id: str, messages: dict[str, dict]) -> None: async def reconcile_messages_by_chat_id(self, chat_id: str, user_id: str, messages: dict[str, dict]) -> None:
"""Sync ``chat_message`` rows with the committed JSON blob. """Sync ``chat_message`` rows with the committed JSON blob.
@ -960,11 +1099,44 @@ class ChatTable:
return history_messages return history_messages
async def get_message_by_id_and_message_id(self, id: str, message_id: str) -> dict | None: async def get_message_by_id_and_message_id(self, id: str, message_id: str) -> dict | None:
chat = await self.get_chat_by_id(id) messages_map = await ChatMessages.get_messages_map_by_chat_id(id)
if messages_map and message_id in messages_map:
return messages_map[message_id]
# Messages the frontend saved straight into the chat blob have no chat_message row yet.
async with get_async_db_context() as session:
result = await session.execute(select(Chat.chat[('history', 'messages')]).filter_by(id=id))
row = result.one_or_none()
if row is None:
return None
messages = row[0] or {}
return self._clean_null_bytes(messages.get(message_id, {}))
async def get_message_metadata(
self,
chat_id: str,
message_id: str,
metadata_key: Literal['files', 'sources', 'embeds'],
) -> Any | None:
"""Read one message metadata field without rebuilding the whole history."""
async with get_async_db_context() as db:
# Read the column directly; some stored rows cannot be validated as full ChatMessageModel objects.
result = await db.execute(
select(getattr(ChatMessage, metadata_key)).where(ChatMessage.id == f'{chat_id}-{message_id}')
)
metadata_row = result.first()
if metadata_row is not None:
return metadata_row[0]
chat = await self.get_chat_by_id(chat_id)
if chat is None: if chat is None:
return None return None
return chat.chat.get('history', {}).get('messages', {}).get(message_id, {}) message = chat.chat.get('history', {}).get('messages', {}).get(message_id, {})
return message.get(metadata_key)
async def upsert_message_to_chat_by_id_and_message_id( async def upsert_message_to_chat_by_id_and_message_id(
self, id: str, message_id: str, message: dict, *, touch: bool = True self, id: str, message_id: str, message: dict, *, touch: bool = True
@ -974,27 +1146,29 @@ class ChatTable:
if output_text: if output_text:
message['content'] = output_text message['content'] = output_text
# Sanitize message content for null characters before upserting message = self._clean_null_bytes(message)
if isinstance(message.get('content'), str): message_id = self._clean_null_bytes(message_id)
message['content'] = sanitize_text_for_db(message['content'])
try: try:
async with get_async_db_context() as session: async with get_async_db_context() as session:
chat_item = await session.get(Chat, id) chat_item = await session.get(
Chat,
id,
populate_existing=True,
with_for_update=session.bind.dialect.name == 'postgresql',
)
if chat_item is None: if chat_item is None:
return None return None
self._sanitize_chat_row(chat_item)
chat = chat_item.chat or {} chat = chat_item.chat or {}
self._repair_chat_current_id(chat) self._repair_chat_current_id(chat)
history = chat.get('history', {}) history = chat.get('history', {})
saved_message = self.upsert_message_to_history(history, message_id, message) saved_message = self.upsert_message_to_history(history, message_id, message)
chat['history'] = history chat['history'] = history
clean_chat = self._clean_null_bytes(chat) chat_item.chat = chat # chat is a fresh dict when the column was empty
chat_item.chat = clean_chat chat_item.title = self._clean_null_bytes(chat.get('title', 'New Chat'))
chat_item.title = self._clean_null_bytes(clean_chat['title']) if 'title' in clean_chat else 'New Chat' chat_item.current_message_id = self.get_current_message_id(chat)
chat_item.current_message_id = self.get_current_message_id(clean_chat)
flag_modified(chat_item, 'chat') flag_modified(chat_item, 'chat')
if touch: if touch:
@ -1022,33 +1196,33 @@ class ChatTable:
async def delete_message_from_chat_by_id_and_message_id(self, id: str, message_id: str) -> ChatModel | None: async def delete_message_from_chat_by_id_and_message_id(self, id: str, message_id: str) -> ChatModel | None:
try: try:
async with get_async_db_context() as session: async with get_async_db_context() as session:
chat_item = await session.get(Chat, id) chat_item = await session.get(
Chat,
id,
populate_existing=True,
with_for_update=session.bind.dialect.name == 'postgresql',
)
if chat_item is None: if chat_item is None:
return None return None
self._sanitize_chat_row(chat_item)
chat = chat_item.chat or {} chat = chat_item.chat or {}
self._repair_chat_current_id(chat) self._repair_chat_current_id(chat)
history = chat.get('history', {}) history = chat.get('history', {})
deleted_ids = self.delete_message_from_history(history, message_id) deleted_ids = self.delete_message_from_history(history, message_id)
if not deleted_ids: if not deleted_ids:
clean_chat = self._clean_null_bytes(chat) chat_item.chat = chat
chat_item.chat = clean_chat chat_item.title = self._clean_null_bytes(chat.get('title', 'New Chat'))
chat_item.title = ( chat_item.current_message_id = self.get_current_message_id(chat)
self._clean_null_bytes(clean_chat['title']) if 'title' in clean_chat else 'New Chat'
)
chat_item.current_message_id = self.get_current_message_id(clean_chat)
flag_modified(chat_item, 'chat') flag_modified(chat_item, 'chat')
await session.commit() await session.commit()
return ChatModel.model_validate(chat_item) return ChatModel.model_validate(chat_item)
messages = history.get('messages') or {} messages = history.get('messages') or {}
chat['history'] = history chat['history'] = history
clean_chat = self._clean_null_bytes(chat) chat_item.chat = chat
chat_item.chat = clean_chat chat_item.title = self._clean_null_bytes(chat.get('title', 'New Chat'))
chat_item.title = self._clean_null_bytes(clean_chat['title']) if 'title' in clean_chat else 'New Chat' chat_item.current_message_id = self.get_current_message_id(chat)
chat_item.current_message_id = self.get_current_message_id(clean_chat)
flag_modified(chat_item, 'chat') flag_modified(chat_item, 'chat')
chat_item.updated_at = int(time.time()) chat_item.updated_at = int(time.time())
await session.commit() await session.commit()
@ -1066,12 +1240,17 @@ class ChatTable:
self, id: str, message_id: str, status: dict self, id: str, message_id: str, status: dict
) -> ChatModel | None: ) -> ChatModel | None:
try: try:
status = self._clean_null_bytes(status)
async with get_async_db_context() as session: async with get_async_db_context() as session:
chat_item = await session.get(Chat, id) chat_item = await session.get(
Chat,
id,
populate_existing=True,
with_for_update=session.bind.dialect.name == 'postgresql',
)
if chat_item is None: if chat_item is None:
return None return None
self._sanitize_chat_row(chat_item)
chat = chat_item.chat or {} chat = chat_item.chat or {}
self._repair_chat_current_id(chat) self._repair_chat_current_id(chat)
history = chat.get('history', {}) history = chat.get('history', {})
@ -1082,10 +1261,9 @@ class ChatTable:
history['messages'][message_id]['statusHistory'] = status_history history['messages'][message_id]['statusHistory'] = status_history
chat['history'] = history chat['history'] = history
clean_chat = self._clean_null_bytes(chat) chat_item.chat = chat
chat_item.chat = clean_chat chat_item.title = self._clean_null_bytes(chat.get('title', 'New Chat'))
chat_item.title = self._clean_null_bytes(clean_chat['title']) if 'title' in clean_chat else 'New Chat' chat_item.current_message_id = self.get_current_message_id(chat)
chat_item.current_message_id = self.get_current_message_id(clean_chat)
flag_modified(chat_item, 'chat') flag_modified(chat_item, 'chat')
await session.commit() await session.commit()
@ -1093,13 +1271,20 @@ class ChatTable:
except Exception: except Exception:
return None return None
async def add_message_files_by_id_and_message_id(self, id: str, message_id: str, files: list[dict]) -> list[dict]: async def add_message_files_by_id_and_message_id(
self, id: str, message_id: str, files: list[dict]
) -> list[dict] | None:
async with get_async_db_context() as session: async with get_async_db_context() as session:
chat = await self.get_chat_by_id(id, db=session) chat_item = await session.get(
if chat is None: Chat,
id,
populate_existing=True,
with_for_update=session.bind.dialect.name == 'postgresql',
)
if chat_item is None:
return None return None
chat = chat.chat chat = chat_item.chat or {}
history = chat.get('history', {}) history = chat.get('history', {})
message_files = [] message_files = []
@ -1109,8 +1294,14 @@ class ChatTable:
message_files = message_files + files message_files = message_files + files
history['messages'][message_id]['files'] = message_files history['messages'][message_id]['files'] = message_files
# Written here rather than through update_chat_by_id: with session sharing off that opens a second
# connection, which then blocks on the lock this one holds.
chat['history'] = history chat['history'] = history
await self.update_chat_by_id(id, chat, db=session) chat_item.chat = self._clean_null_bytes(chat)
# History was mutated in place, so the new blob compares equal to the loaded one.
flag_modified(chat_item, 'chat')
chat_item.updated_at = int(time.time())
await session.commit()
return message_files return message_files
async def insert_shared_chat_by_chat_id(self, chat_id: str, db: AsyncSession | None = None) -> ChatModel | None: async def insert_shared_chat_by_chat_id(self, chat_id: str, db: AsyncSession | None = None) -> ChatModel | None:
@ -1508,6 +1699,7 @@ class ChatTable:
repaired_history = self._repair_chat_current_id(chat_item.chat or {}) repaired_history = self._repair_chat_current_id(chat_item.chat or {})
if repaired_history: if repaired_history:
chat_item.current_message_id = self.get_current_message_id(chat_item.chat)
flag_modified(chat_item, 'chat') flag_modified(chat_item, 'chat')
if self._sanitize_chat_row(chat_item) or repaired_history: if self._sanitize_chat_row(chat_item) or repaired_history:
await session.commit() await session.commit()
@ -1549,6 +1741,7 @@ class ChatTable:
repaired_history = self._repair_chat_current_id(chat.chat or {}) repaired_history = self._repair_chat_current_id(chat.chat or {})
if repaired_history: if repaired_history:
chat.current_message_id = self.get_current_message_id(chat.chat)
flag_modified(chat, 'chat') flag_modified(chat, 'chat')
if self._sanitize_chat_row(chat) or repaired_history: if self._sanitize_chat_row(chat) or repaired_history:
await session.commit() await session.commit()
@ -1557,6 +1750,41 @@ class ChatTable:
except Exception: except Exception:
return None return None
async def get_chat_by_id_for_user(
self,
id: str,
user,
db: AsyncSession | None = None,
) -> ChatModel | None:
chat = await self.get_chat_by_id_and_user_id(id, user.id, db=db)
if chat:
return chat
chat = await self.get_chat_by_id(id, db=db)
if not chat:
return None
if user.role == 'admin' and (ENABLE_ADMIN_CHAT_ACCESS or is_internal_chat(chat.meta)):
return chat
if await AccessGrants.has_access(
user_id=user.id,
resource_type='shared_chat',
resource_id=id,
permission='read',
db=db,
):
return chat
if chat.folder_id:
from open_webui.utils.access_control.folders import has_folder_access
folder = await Folders.get_folder_by_id(chat.folder_id, db=db)
if folder and await has_folder_access(user.id, folder, 'read', db):
return chat
return None
async def is_chat_owner(self, id: str, user_id: str, db: AsyncSession | None = None) -> bool: async def is_chat_owner(self, id: str, user_id: str, db: AsyncSession | None = None) -> bool:
""" """
Lightweight ownership check — uses EXISTS subquery instead of loading Lightweight ownership check — uses EXISTS subquery instead of loading
@ -1732,7 +1960,7 @@ class ChatTable:
return [ChatModel.model_validate(chat) for chat in result.scalars().all()] return [ChatModel.model_validate(chat) for chat in result.scalars().all()]
# search user conversations # search user conversations
async def get_chats_by_user_id_and_search_text( async def get_chats_by_user_id_and_search_text( # noqa: C901
self, self,
user_id: str, user_id: str,
search_text: str, search_text: str,
@ -1751,7 +1979,7 @@ class ChatTable:
user_id, include_archived, filter={}, skip=skip, limit=limit, db=db user_id, include_archived, filter={}, skip=skip, limit=limit, db=db
) )
search_text_words = search_text.split(' ') search_text_words = search_text.split()
# search_text might contain 'tag:tag_name' format so we need to extract the tag_name # search_text might contain 'tag:tag_name' format so we need to extract the tag_name
tag_ids = [ tag_ids = [
@ -1759,10 +1987,8 @@ class ChatTable:
] ]
# Extract folder names # Extract folder names
folders = await Folders.search_folders_by_names( folder_names = [word.replace('folder:', '') for word in search_text_words if word.startswith('folder:')]
user_id, folders = await Folders.search_folders_by_names(user_id, folder_names)
[word.replace('folder:', '') for word in search_text_words if word.startswith('folder:')],
)
folder_ids = [folder.id for folder in folders] folder_ids = [folder.id for folder in folders]
is_pinned = None is_pinned = None
@ -1783,19 +2009,10 @@ class ChatTable:
elif 'shared:false' in search_text_words: elif 'shared:false' in search_text_words:
is_shared = False is_shared = False
search_text_words = [ search_text_words = [word for word in search_text_words if not word.startswith(CHAT_SEARCH_FILTER_PREFIXES)]
word
for word in search_text_words
if (
not word.startswith('tag:')
and not word.startswith('folder:')
and not word.startswith('pinned:')
and not word.startswith('archived:')
and not word.startswith('shared:')
)
]
search_text = ' '.join(search_text_words) phrase_query = ' '.join(search_text_words).strip()
search_terms = chat_search_terms(phrase_query)
async with get_async_db_context(db) as session: async with get_async_db_context(db) as session:
stmt = select(Chat).filter(Chat.user_id == user_id) stmt = select(Chat).filter(Chat.user_id == user_id)
@ -1815,30 +2032,46 @@ class ChatTable:
else: else:
stmt = stmt.filter(Chat.share_id.is_(None)) stmt = stmt.filter(Chat.share_id.is_(None))
if folder_ids: if folder_names:
stmt = stmt.filter(Chat.folder_id.in_(folder_ids)) stmt = stmt.filter(Chat.folder_id.in_(folder_ids))
stmt = stmt.order_by(Chat.updated_at.desc(), Chat.id)
# Check if the database dialect is either 'sqlite' or 'postgresql' # Check if the database dialect is either 'sqlite' or 'postgresql'
bind = await session.connection() bind = await session.connection()
dialect_name = bind.dialect.name dialect_name = bind.dialect.name
if dialect_name == 'sqlite':
# SQLite case: using JSON1 extension for JSON searching search_params = {}
sqlite_content_sql = ( exact_match_clause = None
'EXISTS (' if phrase_query:
' SELECT 1 ' exact_match_clause = or_(
" FROM json_each(Chat.chat, '$.messages') AS message " Chat.title.ilike(bindparam('phrase_title_key')),
" WHERE LOWER(message.value->>'content') LIKE '%' || :content_key || '%'" text(chat_search_message_content_match_sql(dialect_name, 'phrase_content_key')),
')'
) )
sqlite_content_clause = text(sqlite_content_sql) search_params.update(
stmt = stmt.filter( {
or_(Chat.title.ilike(bindparam('title_key')), sqlite_content_clause).params( 'phrase_title_key': f'%{phrase_query}%',
title_key=f'%{search_text}%', content_key=search_text 'phrase_content_key': phrase_query,
) }
) )
term_clauses = []
for term_idx, term in enumerate(search_terms):
title_key = f'term_title_key_{term_idx}'
content_key = f'term_content_key_{term_idx}'
term_clauses.append(
or_(
Chat.title.ilike(bindparam(title_key)),
text(chat_search_message_content_match_sql(dialect_name, content_key)),
)
)
search_params[title_key] = f'%{term}%'
search_params[content_key] = term
if term_clauses:
stmt = stmt.filter(or_(exact_match_clause, and_(*term_clauses)))
else:
stmt = stmt.filter(exact_match_clause)
if dialect_name == 'sqlite':
# Check if there are any tags to filter # Check if there are any tags to filter
if 'none' in tag_ids: if 'none' in tag_ids:
stmt = stmt.filter( stmt = stmt.filter(
@ -1872,38 +2105,6 @@ class ChatTable:
# Safety filter: title must not contain actual null bytes # Safety filter: title must not contain actual null bytes
stmt = stmt.filter(text("Chat.title::text NOT LIKE '%\\x00%'")) stmt = stmt.filter(text("Chat.title::text NOT LIKE '%\\x00%'"))
postgres_content_sql = """
EXISTS (
SELECT 1
FROM chat_message AS message
WHERE message.chat_id = Chat.id
AND message.user_id = Chat.user_id
AND json_typeof(message.content) = 'string'
AND LOWER(message.content #>> '{}') LIKE '%' || :content_key || '%'
)
OR EXISTS (
SELECT 1
FROM json_each(Chat.chat#>'{history,messages}') AS history_message
WHERE json_typeof(history_message.value->'content') = 'string'
AND LOWER(history_message.value->>'content') LIKE '%' || :content_key || '%'
)
OR EXISTS (
SELECT 1
FROM json_array_elements(Chat.chat->'messages') AS legacy_message
WHERE json_typeof(legacy_message->'content') = 'string'
AND LOWER(legacy_message->>'content') LIKE '%' || :content_key || '%'
)
"""
postgres_content_clause = text(postgres_content_sql)
stmt = stmt.filter(
or_(
Chat.title.ilike(bindparam('title_key')),
postgres_content_clause,
)
).params(title_key=f'%{search_text}%', content_key=search_text.lower())
if 'none' in tag_ids: if 'none' in tag_ids:
stmt = stmt.filter( stmt = stmt.filter(
text(""" text("""
@ -1931,12 +2132,20 @@ class ChatTable:
else: else:
raise NotImplementedError(f'Unsupported dialect: {dialect_name}') raise NotImplementedError(f'Unsupported dialect: {dialect_name}')
if exact_match_clause is not None:
stmt = stmt.order_by(case((exact_match_clause, 0), else_=1), Chat.updated_at.desc(), Chat.id)
else:
stmt = stmt.order_by(Chat.updated_at.desc(), Chat.id)
if search_params:
stmt = stmt.params(**search_params)
# Perform pagination at the SQL level # Perform pagination at the SQL level
stmt = stmt.offset(skip).limit(limit) stmt = stmt.offset(skip).limit(limit)
result = await session.execute(stmt) result = await session.execute(stmt)
all_chats = result.scalars().all() all_chats = result.scalars().all()
log.info(f'The number of chats: {len(all_chats)}') log.info('The number of chats: %s', len(all_chats))
# Validate and return chats # Validate and return chats
return [ChatModel.model_validate(chat) for chat in all_chats] return [ChatModel.model_validate(chat) for chat in all_chats]
@ -2101,7 +2310,7 @@ class ChatTable:
bind = await session.connection() bind = await session.connection()
dialect_name = bind.dialect.name dialect_name = bind.dialect.name
log.info(f'DB dialect name: {dialect_name}') log.info('DB dialect name: %s', dialect_name)
if dialect_name == 'sqlite': if dialect_name == 'sqlite':
stmt = stmt.filter( stmt = stmt.filter(
text(f"EXISTS (SELECT 1 FROM json_each(Chat.meta, '$.tags') WHERE json_each.value = :tag_id)") text(f"EXISTS (SELECT 1 FROM json_each(Chat.meta, '$.tags') WHERE json_each.value = :tag_id)")
@ -2228,7 +2437,7 @@ class ChatTable:
result = await session.execute(stmt.where(Chat.meta['internal'].as_boolean().is_not(True))) result = await session.execute(stmt.where(Chat.meta['internal'].as_boolean().is_not(True)))
count = result.scalar() count = result.scalar()
log.info(f"Count of chats for folder '{folder_id}': {count}") log.info("Count of chats for folder '%s': %s", folder_id, count)
return count return count
async def count_chats_by_folder_ids_and_user_id( async def count_chats_by_folder_ids_and_user_id(
@ -2242,7 +2451,7 @@ class ChatTable:
result = await session.execute(stmt.where(Chat.meta['internal'].as_boolean().is_not(True))) result = await session.execute(stmt.where(Chat.meta['internal'].as_boolean().is_not(True)))
count = result.scalar() count = result.scalar()
log.info(f"Count of chats for folders '{folder_ids}': {count}") log.info("Count of chats for folders '%s': %s", folder_ids, count)
return count return count
async def delete_tag_by_id_and_user_id_and_tag_name( async def delete_tag_by_id_and_user_id_and_tag_name(
@ -2326,18 +2535,15 @@ class ChatTable:
except Exception: except Exception:
return False return False
async def move_chats_by_user_id_and_folder_id( async def move_chats_by_folder_id(
self, self,
user_id: str,
folder_id: str, folder_id: str,
new_folder_id: str | None, new_folder_id: str | None,
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> bool: ) -> bool:
try: try:
async with get_async_db_context(db) as session: async with get_async_db_context(db) as session:
await session.execute( await session.execute(update(Chat).filter_by(folder_id=folder_id).values(folder_id=new_folder_id))
update(Chat).filter_by(user_id=user_id, folder_id=folder_id).values(folder_id=new_folder_id)
)
await session.commit() await session.commit()
return True return True
@ -2369,7 +2575,7 @@ class ChatTable:
file_ids: list[str], file_ids: list[str],
user_id: str, user_id: str,
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> list[ChatFileModel | None]: ) -> list[ChatFileModel] | None:
if not file_ids: if not file_ids:
return None return None

View file

@ -32,6 +32,8 @@ DICT_CONFIG_KEY_ALIASES = {
'audio.tts.openai.params': ('AUDIO_TTS_OPENAI_PARAMS',), 'audio.tts.openai.params': ('AUDIO_TTS_OPENAI_PARAMS',),
'models.default_metadata': ('DEFAULT_MODEL_METADATA',), 'models.default_metadata': ('DEFAULT_MODEL_METADATA',),
'models.default_params': ('DEFAULT_MODEL_PARAMS',), 'models.default_params': ('DEFAULT_MODEL_PARAMS',),
'task.model.params': ('TASK_MODEL_PARAMS',),
'ui.default_interface_settings': ('DEFAULT_INTERFACE_SETTINGS',),
'user.permissions': ('USER_PERMISSIONS',), 'user.permissions': ('USER_PERMISSIONS',),
} }
DICT_CONFIG_KEYS = tuple(DICT_CONFIG_KEY_ALIASES) DICT_CONFIG_KEYS = tuple(DICT_CONFIG_KEY_ALIASES)
@ -298,14 +300,16 @@ class Config(Base):
) )
@staticmethod @staticmethod
async def repair_flattened_dict_configs() -> None: async def repair_config_rows() -> None:
"""Reassemble dict config values flattened by the per-key migration.""" """Repair known legacy config row shapes."""
if not Config.PERSISTENT_ENABLED: if not Config.PERSISTENT_ENABLED:
return return
async with get_async_db() as db: async with get_async_db() as db:
repaired_keys: list[str] = [] repaired_keys: list[str] = []
orphan_keys: list[str] = [] orphan_keys: list[str] = []
default_model_keys: list[str] = []
now = int(time.time())
for config_key, aliases in DICT_CONFIG_KEY_ALIASES.items(): for config_key, aliases in DICT_CONFIG_KEY_ALIASES.items():
prefixes = (config_key, *aliases) prefixes = (config_key, *aliases)
@ -353,14 +357,26 @@ class Config(Base):
if existing: if existing:
existing.value = repaired existing.value = repaired
existing.updated_at = int(time.time()) existing.updated_at = now
else: else:
db.add(Config(key=config_key, value=repaired, updated_at=int(time.time()))) db.add(Config(key=config_key, value=repaired, updated_at=now))
repaired_keys.append(config_key) repaired_keys.append(config_key)
if orphan_keys: if orphan_keys:
await db.execute(delete(Config).where(Config.key.in_(orphan_keys))) await db.execute(delete(Config).where(Config.key.in_(orphan_keys)))
if repaired_keys or orphan_keys: for key in ('ui.default_models', 'ui.default_pinned_models'):
row = await db.get(Config, key)
if not row or not isinstance(row.value, list):
continue
row.value = ','.join(model_id for model_id in (str(item).strip() for item in row.value) if model_id)
row.updated_at = now
default_model_keys.append(key)
if repaired_keys or orphan_keys or default_model_keys:
await db.commit() await db.commit()
log.info('Repaired flattened dict config rows for %s', ', '.join(repaired_keys)) if repaired_keys or orphan_keys:
log.info('Repaired flattened dict config rows for %s', ', '.join(repaired_keys))
if default_model_keys:
log.info('Repaired default model config rows for %s', ', '.join(default_model_keys))

View file

@ -157,6 +157,11 @@ class FolderTable:
except Exception: except Exception:
return None return None
async def get_folders_by_ids(self, ids: list[str], db: AsyncSession | None = None) -> list[FolderModel]:
async with get_async_db_context(db) as db:
result = await db.execute(select(Folder).filter(Folder.id.in_(ids)).order_by(Folder.updated_at.desc()))
return [FolderModel.model_validate(folder) for folder in result.scalars().all()]
async def get_shared_folder_ids_for_user( async def get_shared_folder_ids_for_user(
self, user_id: str, user_group_ids: set[str], db: Optional[AsyncSession] = None self, user_id: str, user_group_ids: set[str], db: Optional[AsyncSession] = None
) -> dict[str, str]: ) -> dict[str, str]:
@ -197,10 +202,14 @@ class FolderTable:
try: try:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
folders = [] folders = []
seen_ids = {id}
async def get_children(folder): async def get_children(folder):
children = await self.get_folders_by_parent_id_and_user_id(folder.id, user_id, db=db) children = await self.get_folders_by_parent_id_and_user_id(folder.id, user_id, db=db)
for child in children: for child in children:
if child.id in seen_ids:
continue
seen_ids.add(child.id)
await get_children(child) await get_children(child)
folders.append(child) folders.append(child)
@ -230,7 +239,9 @@ class FolderTable:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
# Check if folder exists # Check if folder exists
result = await db.execute( result = await db.execute(
select(Folder).filter_by(parent_id=parent_id, user_id=user_id).filter(Folder.name.ilike(name)) select(Folder)
.filter_by(parent_id=parent_id, user_id=user_id)
.filter(func.lower(Folder.name) == func.lower(name))
) )
folder = result.scalars().first() folder = result.scalars().first()
@ -246,7 +257,9 @@ class FolderTable:
self, parent_id: Optional[str], user_id: str, db: Optional[AsyncSession] = None self, parent_id: Optional[str], user_id: str, db: Optional[AsyncSession] = None
) -> list[FolderModel]: ) -> list[FolderModel]:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
result = await db.execute(select(Folder).filter_by(parent_id=parent_id, user_id=user_id)) result = await db.execute(
select(Folder).filter_by(parent_id=parent_id, user_id=user_id).order_by(Folder.updated_at.desc())
)
return [FolderModel.model_validate(folder) for folder in result.scalars().all()] return [FolderModel.model_validate(folder) for folder in result.scalars().all()]
async def get_folder_ids_by_id_and_user_id_in_subtree( async def get_folder_ids_by_id_and_user_id_in_subtree(
@ -258,15 +271,17 @@ class FolderTable:
if not folder: if not folder:
return [] return []
folder_ids = [folder.id] folder_ids = {folder.id}
folders = [FolderModel.model_validate(folder)] folders = [FolderModel.model_validate(folder)]
while folders: while folders:
current_folder = folders.pop() current_folder = folders.pop()
children = await self.get_folders_by_parent_id_and_user_id(current_folder.id, user_id, db=db) children = await self.get_folders_by_parent_id_and_user_id(current_folder.id, user_id, db=db)
folder_ids.extend(child.id for child in children) for child in children:
folders.extend(children) if child.id not in folder_ids:
folder_ids.add(child.id)
folders.append(child)
return folder_ids return list(folder_ids)
async def update_folder_parent_id_by_id_and_user_id( async def update_folder_parent_id_by_id_and_user_id(
self, self,
@ -376,11 +391,15 @@ class FolderTable:
return folder_ids return folder_ids
folder_ids.append(folder.id) folder_ids.append(folder.id)
seen_ids = {folder.id}
# Delete all children folders # Delete all children folders
async def delete_children(folder): async def delete_children(folder):
folder_children = await self.get_folders_by_parent_id_and_user_id(folder.id, user_id, db=db) folder_children = await self.get_folders_by_parent_id_and_user_id(folder.id, user_id, db=db)
for folder_child in folder_children: for folder_child in folder_children:
if folder_child.id in seen_ids:
continue
seen_ids.add(folder_child.id)
await delete_children(folder_child) await delete_children(folder_child)
folder_ids.append(folder_child.id) folder_ids.append(folder_child.id)

View file

@ -1,4 +1,3 @@
import json
import logging import logging
import time import time
import uuid import uuid
@ -13,6 +12,7 @@ from sqlalchemy import (
BigInteger, BigInteger,
Column, Column,
ForeignKey, ForeignKey,
Index,
String, String,
Text, Text,
and_, and_,
@ -72,6 +72,8 @@ class GroupModel(BaseModel):
class GroupMember(Base): class GroupMember(Base):
__tablename__ = 'group_member' __tablename__ = 'group_member'
# The table's (group_id, user_id) unique constraint cannot serve user_id lookups.
__table_args__ = (Index('ix_group_member_user_id_group_id', 'user_id', 'group_id'),)
id = Column(Text, unique=True, primary_key=True) id = Column(Text, unique=True, primary_key=True)
group_id = Column( group_id = Column(

View file

@ -1,4 +1,3 @@
import json
import logging import logging
import time import time
import uuid import uuid
@ -468,14 +467,22 @@ class KnowledgeTable:
print('search_knowledge_files error:', e) print('search_knowledge_files error:', e)
return KnowledgeFileListResponse(items=[], total=0) return KnowledgeFileListResponse(items=[], total=0)
async def check_access_by_user_id(self, id, user_id, permission='write', db: Optional[AsyncSession] = None) -> bool: async def check_access_by_user_id(
self,
id,
user_id,
permission='write',
db: Optional[AsyncSession] = None,
user_group_ids: set[str] | None = None,
) -> bool:
knowledge = await self.get_knowledge_by_id(id, db=db) knowledge = await self.get_knowledge_by_id(id, db=db)
if not knowledge: if not knowledge:
return False return False
if knowledge.user_id == user_id: if knowledge.user_id == user_id:
return True return True
user_groups = await Groups.get_groups_by_member_id(user_id, db=db) if user_group_ids is None:
user_group_ids = {group.id for group in user_groups} user_groups = await Groups.get_groups_by_member_id(user_id, db=db)
user_group_ids = {group.id for group in user_groups}
return await AccessGrants.has_access( return await AccessGrants.has_access(
user_id=user_id, user_id=user_id,
resource_type='knowledge', resource_type='knowledge',
@ -485,24 +492,6 @@ class KnowledgeTable:
db=db, db=db,
) )
async def get_knowledge_bases_by_user_id(
self, user_id: str, permission: str = 'write', db: Optional[AsyncSession] = None
) -> list[KnowledgeUserModel]:
knowledge_bases = await self.get_knowledge_bases(db=db)
user_groups = await Groups.get_groups_by_member_id(user_id, db=db)
user_group_ids = {group.id for group in user_groups}
# One grants query for all non-owned knowledge bases instead of one each
accessible_ids = await AccessGrants.get_accessible_resource_ids(
user_id=user_id,
resource_type='knowledge',
resource_ids=[kb.id for kb in knowledge_bases if kb.user_id != user_id],
permission=permission,
user_group_ids=user_group_ids,
db=db,
)
return [kb for kb in knowledge_bases if kb.user_id == user_id or kb.id in accessible_ids]
async def get_knowledge_by_id(self, id: str, db: Optional[AsyncSession] = None) -> Optional[KnowledgeModel]: async def get_knowledge_by_id(self, id: str, db: Optional[AsyncSession] = None) -> Optional[KnowledgeModel]:
try: try:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
@ -512,29 +501,6 @@ class KnowledgeTable:
except Exception: except Exception:
return None return None
async def get_knowledge_by_id_and_user_id(
self, id: str, user_id: str, db: Optional[AsyncSession] = None
) -> Optional[KnowledgeModel]:
knowledge = await self.get_knowledge_by_id(id, db=db)
if not knowledge:
return None
if knowledge.user_id == user_id:
return knowledge
user_groups = await Groups.get_groups_by_member_id(user_id, db=db)
user_group_ids = {group.id for group in user_groups}
if await AccessGrants.has_access(
user_id=user_id,
resource_type='knowledge',
resource_id=knowledge.id,
permission='write',
user_group_ids=user_group_ids,
db=db,
):
return knowledge
return None
async def get_knowledges_by_file_id(self, file_id: str, db: Optional[AsyncSession] = None) -> list[KnowledgeModel]: async def get_knowledges_by_file_id(self, file_id: str, db: Optional[AsyncSession] = None) -> list[KnowledgeModel]:
try: try:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
@ -659,6 +625,7 @@ class KnowledgeTable:
db=db, db=db,
), ),
breadcrumbs=await self.get_directory_breadcrumbs( breadcrumbs=await self.get_directory_breadcrumbs(
knowledge_id,
filter.get('directory_id') if filter else None, filter.get('directory_id') if filter else None,
db=db, db=db,
), ),
@ -809,25 +776,6 @@ class KnowledgeTable:
log.exception(e) log.exception(e)
return None return None
async def update_knowledge_data_by_id(
self, id: str, data: dict, db: Optional[AsyncSession] = None
) -> Optional[KnowledgeModel]:
try:
async with get_async_db_context(db) as db:
await db.execute(
update(Knowledge)
.filter_by(id=id)
.values(
data=data,
updated_at=int(time.time()),
)
)
await db.commit()
return await self.get_knowledge_by_id(id=id, db=db)
except Exception as e:
log.exception(e)
return None
async def update_knowledge_meta_by_id( async def update_knowledge_meta_by_id(
self, id: str, meta: dict, db: Optional[AsyncSession] = None self, id: str, meta: dict, db: Optional[AsyncSession] = None
) -> Optional[KnowledgeModel]: ) -> Optional[KnowledgeModel]:
@ -961,6 +909,7 @@ class KnowledgeTable:
async def get_directory_breadcrumbs( async def get_directory_breadcrumbs(
self, self,
knowledge_id: str,
directory_id: Optional[str], directory_id: Optional[str],
db: Optional[AsyncSession] = None, db: Optional[AsyncSession] = None,
) -> list[KnowledgeDirectoryModel]: ) -> list[KnowledgeDirectoryModel]:
@ -975,7 +924,10 @@ class KnowledgeTable:
while current_id and current_id not in seen: while current_id and current_id not in seen:
seen.add(current_id) seen.add(current_id)
result = await db.execute(select(KnowledgeDirectory).filter_by(id=current_id)) # Scoped by knowledge base so a caller-supplied id cannot walk another one's tree.
result = await db.execute(
select(KnowledgeDirectory).filter_by(id=current_id, knowledge_id=knowledge_id)
)
directory = result.scalars().first() directory = result.scalars().first()
if not directory: if not directory:
break break
@ -1122,6 +1074,26 @@ class KnowledgeTable:
for child_id in child_ids: for child_id in child_ids:
await self._delete_files_in_subtree(child_id, db=db) await self._delete_files_in_subtree(child_id, db=db)
async def get_files_by_id_and_directory_id(
self,
knowledge_id: str,
directory_id: str,
db: Optional[AsyncSession] = None,
) -> list[FileModel]:
"""Get all files in a directory and its subdirectories."""
async with get_async_db_context(db) as db:
directory_ids = [directory_id]
for parent_id in directory_ids:
result = await db.execute(select(KnowledgeDirectory.id).filter_by(parent_id=parent_id))
directory_ids.extend(result.scalars().all())
result = await db.execute(
select(File)
.join(KnowledgeFile, File.id == KnowledgeFile.file_id)
.filter(KnowledgeFile.knowledge_id == knowledge_id)
.filter(KnowledgeFile.directory_id.in_(directory_ids))
)
return [FileModel.model_validate(file) for file in result.scalars().all()]
async def move_file_to_directory( async def move_file_to_directory(
self, self,
knowledge_id: str, knowledge_id: str,

View file

@ -1,4 +1,3 @@
import json
import time import time
import uuid import uuid
from typing import Optional from typing import Optional

View file

@ -1,26 +1,33 @@
from __future__ import annotations from __future__ import annotations
import json
import logging import logging
import time import time
from copy import deepcopy from copy import deepcopy
from typing import Any, Optional from typing import Any
from open_webui.internal.db import Base, JSONField, get_async_db_context from open_webui.internal.db import Base, JSONField, get_async_db_context
from open_webui.models.access_grants import AccessGrantModel, AccessGrants from open_webui.models.access_grants import AccessGrantModel, AccessGrants
from open_webui.models.groups import Groups from open_webui.models.groups import Groups
from open_webui.models.users import User, UserModel, UserResponse, Users from open_webui.models.users import User, UserModel, UserResponse, Users
from open_webui.utils.validate import validate_profile_image_url from open_webui.utils.misc import json_text_variants
from pydantic import BaseModel, ConfigDict, Field, field_validator, model_validator from open_webui.utils.validate import validate_image_url
from pydantic import BaseModel, ConfigDict, Field, ValidationInfo, field_validator, model_validator
from sqlalchemy import BigInteger, Boolean, Column, String, Text, cast, delete, func, or_, select, update from sqlalchemy import BigInteger, Boolean, Column, String, Text, cast, delete, func, or_, select, update
from sqlalchemy.dialects.postgresql import JSONB
from sqlalchemy.ext.asyncio import AsyncSession from sqlalchemy.ext.asyncio import AsyncSession
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
# Track invalid profile_image_url values we've already warned about so we
# don't flood the logs on every DB read (the validator fires per-row). def normalize_model_tags(tags: Any) -> list[dict[str, str]]:
_warned_profile_urls: set[str] = set() if not isinstance(tags, list):
return []
normalized = []
for tag in tags:
name = tag.get('name') if isinstance(tag, dict) else tag
if isinstance(name, str) and name.strip():
normalized.append({'name': name.strip()})
return normalized
def strip_extracted_content_from_model_knowledge(knowledge: Any) -> Any: def strip_extracted_content_from_model_knowledge(knowledge: Any) -> Any:
@ -68,26 +75,24 @@ class ModelMeta(BaseModel):
"""Metadata for a workspace model entry (profile, description, tags, capabilities).""" """Metadata for a workspace model entry (profile, description, tags, capabilities)."""
profile_image_url: str | None = None profile_image_url: str | None = None
background_image_url: str | None = None
description: str | None = Field(default=None, description='User-facing description of the model.') description: str | None = Field(default=None, description='User-facing description of the model.')
i18n: dict[str, Any] | None = None
capabilities: dict | None = None capabilities: dict | None = None
knowledge: list[Any] | None = None knowledge: list[Any] | None = None
model_config = ConfigDict(extra='allow') model_config = ConfigDict(extra='allow')
@field_validator('profile_image_url', mode='before') @field_validator('profile_image_url', 'background_image_url', mode='before')
@classmethod @classmethod
def check_profile_image_url(cls, v: str | None) -> str | None: def check_image_url(cls, v: str | None, info: ValidationInfo) -> str | None:
if v is None: if v is None:
return v return v
try: try:
return validate_profile_image_url(v) return validate_image_url(v, file_only=info.field_name == 'background_image_url')
except ValueError: except ValueError:
if v not in _warned_profile_urls: if info.field_name == 'background_image_url':
_warned_profile_urls.add(v) raise
log.warning(
'Clearing invalid profile_image_url stored in DB (likely a legacy SVG data-URI): %.80s…',
v,
)
return None return None
@field_validator('knowledge', mode='before') @field_validator('knowledge', mode='before')
@ -99,15 +104,7 @@ class ModelMeta(BaseModel):
@classmethod @classmethod
def normalize_tags(cls, data): def normalize_tags(cls, data):
if isinstance(data, dict) and 'tags' in data: if isinstance(data, dict) and 'tags' in data:
raw_tags = data['tags'] data['tags'] = normalize_model_tags(data['tags'])
if isinstance(raw_tags, list):
normalized = []
for tag in raw_tags:
if isinstance(tag, str):
normalized.append({'name': tag})
elif isinstance(tag, dict) and 'name' in tag:
normalized.append(tag)
data['tags'] = normalized
return data return data
@ -172,12 +169,12 @@ class ModelAccessListResponse(BaseModel):
class ModelForm(BaseModel): class ModelForm(BaseModel):
model_config = ConfigDict(extra='ignore') model_config = ConfigDict(extra='ignore')
id: str id: str = Field(pattern=r'^\S+$')
base_model_id: str | None = None base_model_id: str | None = None
name: str name: str
meta: ModelMeta meta: ModelMeta
params: ModelParams params: ModelParams
access_grants: list[dict | None] = None access_grants: list[dict] | None = None
is_active: bool = True is_active: bool = True
@ -188,7 +185,7 @@ class ModelsTable:
async def _to_model_model( async def _to_model_model(
self, self,
model: Model, model: Model,
access_grants: list[AccessGrantModel | None] = None, access_grants: list[AccessGrantModel] | None = None,
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> ModelModel: ) -> ModelModel:
if isinstance(model.meta, dict): if isinstance(model.meta, dict):
@ -244,9 +241,24 @@ class ModelsTable:
log.error('Skipping model %r during get_all_models due to error: %s', model.id, exc) log.error('Skipping model %r during get_all_models due to error: %s', model.id, exc)
return models return models
async def get_models(self, db: AsyncSession | None = None) -> list[ModelUserResponse]: async def get_models(
self, writable_by_user_id: str | None = None, db: AsyncSession | None = None, ids: list[str] | None = None
) -> list[ModelUserResponse]:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
result = await db.execute(select(Model).filter(Model.base_model_id != None)) stmt = select(Model).filter(Model.base_model_id != None)
if ids is not None:
stmt = stmt.filter(Model.id.in_(ids))
if writable_by_user_id:
user_group_ids = {
group.id for group in await Groups.get_groups_by_member_id(writable_by_user_id, db=db)
}
stmt = self._has_permission(
db, stmt, {'user_id': writable_by_user_id, 'group_ids': user_group_ids}, permission='write'
)
result = await db.execute(stmt)
all_models = result.scalars().all() all_models = result.scalars().all()
user_ids = list(set(model.user_id for model in all_models)) user_ids = list(set(model.user_id for model in all_models))
@ -275,6 +287,27 @@ class ModelsTable:
) )
return models return models
async def get_model_owner_ids_by_file_id(
self, file_id: str, db: AsyncSession | None = None, include_background: bool = False
) -> dict[str, str]:
"""Return model IDs mapped to owner IDs for models referencing the file."""
async with get_async_db_context(db) as db:
# File ids are server-generated uuids, so the text match can only over-match.
result = await db.execute(
select(Model.id, Model.user_id, Model.meta).filter(
Model.base_model_id.is_not(None), cast(Model.meta, String).like(f'%{file_id}%')
)
)
return {
model_id: user_id
for model_id, user_id, meta in result.all()
if any(
isinstance(item, dict) and item.get('type') == 'file' and item.get('id') == file_id
for item in meta.get('knowledge') or []
)
or (include_background and meta.get('background_image_url') == f'/api/v1/files/{file_id}/content')
}
@staticmethod @staticmethod
def _meta_has_tag(meta: dict | None, tag: str) -> bool: def _meta_has_tag(meta: dict | None, tag: str) -> bool:
if not meta: if not meta:
@ -301,28 +334,6 @@ class ModelsTable:
for model in all_models for model in all_models
] ]
async def get_models_by_user_id(
self,
user_id: str,
permission: str = 'write',
db: AsyncSession | None = None,
user_group_ids: set[str] | None = None,
) -> list[ModelUserResponse]:
models = await self.get_models(db=db)
if user_group_ids is None:
user_group_ids = {group.id for group in await Groups.get_groups_by_member_id(user_id, db=db)}
# One grants query for all non-owned models instead of one per model
accessible_ids = await AccessGrants.get_accessible_resource_ids(
user_id=user_id,
resource_type='model',
resource_ids=[model.id for model in models if model.user_id != user_id],
permission=permission,
user_group_ids=user_group_ids,
db=db,
)
return [model for model in models if model.user_id == user_id or model.id in accessible_ids]
def _has_permission(self, db, query, filter: dict, permission: str = 'read'): def _has_permission(self, db, query, filter: dict, permission: str = 'read'):
return AccessGrants.has_permission_filter( return AccessGrants.has_permission_filter(
db=db, db=db,
@ -374,20 +385,14 @@ class ModelsTable:
tag = filter.get('tag') tag = filter.get('tag')
if tag: if tag:
# SQLite stores JSON text via json.dumps(ensure_ascii=True), if db.bind.dialect.name == 'sqlite' and not tag.isascii():
# so non-ASCII chars are \uXXXX-escaped. PostgreSQL native JSONB # SQLite's LOWER() is ASCII-only, so match non-ASCII tags exact-case.
# stores literal Unicode. Use the right pattern for each. meta_text = cast(Model.meta, String)
if db.bind.dialect.name == 'sqlite': variants = json_text_variants(tag)
if tag.isascii():
meta_text = func.lower(cast(Model.meta, String))
pattern = f'%{json.dumps(tag.lower())}%'
else:
meta_text = cast(Model.meta, String)
pattern = f'%{json.dumps(tag)}%'
else: else:
meta_text = func.lower(cast(Model.meta, String)) meta_text = func.lower(cast(Model.meta, String))
pattern = f'%{json.dumps(tag.lower(), ensure_ascii=False)}%' variants = json_text_variants(tag.lower())
stmt = stmt.filter(meta_text.like(pattern)) stmt = stmt.filter(or_(*(meta_text.like(f'%"{variant}"%') for variant in variants)))
order_by = filter.get('order_by') order_by = filter.get('order_by')
direction = filter.get('direction') direction = filter.get('direction')
@ -443,11 +448,13 @@ class ModelsTable:
return ModelListResponse(items=models, total=total) return ModelListResponse(items=models, total=total)
async def get_model_meta_by_id(self, id: str, db: AsyncSession | None = None) -> tuple[dict, int | None]: async def get_model_meta_by_id(
"""Return (meta, updated_at) for a model, skipping access grant resolution.""" self, id: str, db: AsyncSession | None = None
) -> tuple[dict, str, int | None] | None:
"""Return (meta, user_id, updated_at) for a model, skipping access grant resolution."""
try: try:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
result = await db.execute(select(Model.meta, Model.updated_at).filter_by(id=id)) result = await db.execute(select(Model.meta, Model.user_id, Model.updated_at).filter_by(id=id))
return result.first() return result.first()
except Exception: except Exception:
return None return None
@ -605,25 +612,16 @@ class ModelsTable:
# Update or insert models # Update or insert models
for model in models: for model in models:
model_data = {
**model.model_dump(exclude={'access_grants'}),
'user_id': user_id,
'updated_at': int(time.time()),
}
if model.id in existing_ids: if model.id in existing_ids:
await db.execute( await db.execute(update(Model).filter_by(id=model.id).values(**model_data))
update(Model)
.filter_by(id=model.id)
.values(
**model.model_dump(exclude={'access_grants'}),
user_id=user_id,
updated_at=int(time.time()),
)
)
else: else:
new_model = Model( db.add(Model(**model_data))
**{
**model.model_dump(exclude={'access_grants'}),
'user_id': user_id,
'updated_at': int(time.time()),
}
)
db.add(new_model)
await AccessGrants.set_access_grants('model', model.id, model.access_grants, db=db) await AccessGrants.set_access_grants('model', model.id, model.access_grants, db=db)
# Remove models that are no longer present # Remove models that are no longer present

View file

@ -1,4 +1,3 @@
import json
import time import time
import uuid import uuid
from functools import lru_cache from functools import lru_cache
@ -8,7 +7,8 @@ from open_webui.internal.db import Base, get_async_db_context
from open_webui.models.access_grants import AccessGrantModel, AccessGrants from open_webui.models.access_grants import AccessGrantModel, AccessGrants
from open_webui.models.groups import Groups from open_webui.models.groups import Groups
from open_webui.models.users import User, UserModel, UserResponse, Users from open_webui.models.users import User, UserModel, UserResponse, Users
from pydantic import BaseModel, ConfigDict, Field from open_webui.utils.json_codec import JSONCodec
from pydantic import BaseModel, ConfigDict, Field, field_validator
from sqlalchemy import JSON, BigInteger, Boolean, Column, ForeignKey, Text, delete, func, or_, select, update from sqlalchemy import JSON, BigInteger, Boolean, Column, ForeignKey, Text, delete, func, or_, select, update
from sqlalchemy.ext.asyncio import AsyncSession from sqlalchemy.ext.asyncio import AsyncSession
@ -31,6 +31,32 @@ class Note(Base):
updated_at = Column(BigInteger) updated_at = Column(BigInteger)
def sanitize_note_data(data: Optional[dict]) -> Optional[dict]:
"""Sanitize malformed note.data so content.md is always markdown text."""
if data is None:
return None
if not isinstance(data, dict):
return {'content': {'md': str(data)}}
content = data.get('content')
if not isinstance(content, dict) or 'md' not in content or isinstance(content.get('md'), str):
return data
md = content.get('md') if content.get('md') is not None else ''
if isinstance(md, (dict, list)):
md = f'```json\n{JSONCodec.dumps(md, indent=2, ensure_ascii=False)}\n```'
else:
md = str(md)
return {
**data,
'content': {
**content,
'md': md,
},
}
class NoteModel(BaseModel): class NoteModel(BaseModel):
model_config = ConfigDict(from_attributes=True) model_config = ConfigDict(from_attributes=True)
@ -47,6 +73,11 @@ class NoteModel(BaseModel):
created_at: int # timestamp in epoch created_at: int # timestamp in epoch
updated_at: int # timestamp in epoch updated_at: int # timestamp in epoch
@field_validator('data', mode='before')
@classmethod
def sanitize_data(cls, data):
return sanitize_note_data(data)
class PinnedNote(Base): class PinnedNote(Base):
__tablename__ = 'pinned_note' __tablename__ = 'pinned_note'
@ -68,6 +99,11 @@ class NoteForm(BaseModel):
meta: Optional[dict] = None meta: Optional[dict] = None
access_grants: Optional[list[dict]] = None access_grants: Optional[list[dict]] = None
@field_validator('data', mode='before')
@classmethod
def sanitize_data(cls, data):
return sanitize_note_data(data)
class NoteUpdateForm(BaseModel): class NoteUpdateForm(BaseModel):
title: Optional[str] = None title: Optional[str] = None
@ -75,6 +111,11 @@ class NoteUpdateForm(BaseModel):
meta: Optional[dict] = None meta: Optional[dict] = None
access_grants: Optional[list[dict]] = None access_grants: Optional[list[dict]] = None
@field_validator('data', mode='before')
@classmethod
def sanitize_data(cls, data):
return sanitize_note_data(data)
class NoteUserResponse(NoteModel): class NoteUserResponse(NoteModel):
user: Optional[UserResponse] = None user: Optional[UserResponse] = None
@ -305,6 +346,7 @@ class NoteTable:
return None return None
form_data = form_data.model_dump(exclude_unset=True) form_data = form_data.model_dump(exclude_unset=True)
note.data = sanitize_note_data(note.data) or {}
if 'title' in form_data: if 'title' in form_data:
note.title = form_data['title'] note.title = form_data['title']

View file

@ -1,6 +1,5 @@
import base64 import base64
import hashlib import hashlib
import json
import logging import logging
import time import time
import uuid import uuid
@ -9,6 +8,7 @@ from typing import List, Optional
from cryptography.fernet import Fernet from cryptography.fernet import Fernet
from open_webui.env import OAUTH_SESSION_TOKEN_ENCRYPTION_KEY from open_webui.env import OAUTH_SESSION_TOKEN_ENCRYPTION_KEY
from open_webui.internal.db import Base, get_async_db_context from open_webui.internal.db import Base, get_async_db_context
from open_webui.utils.json_codec import JSONCodec
from pydantic import BaseModel, ConfigDict from pydantic import BaseModel, ConfigDict
from sqlalchemy import BigInteger, Column, Index, String, Text, delete, select, update from sqlalchemy import BigInteger, Column, Index, String, Text, delete, select, update
from sqlalchemy.ext.asyncio import AsyncSession from sqlalchemy.ext.asyncio import AsyncSession
@ -85,7 +85,7 @@ class OAuthSessionTable:
def _encrypt_token(self, token) -> str: def _encrypt_token(self, token) -> str:
"""Encrypt OAuth tokens for storage""" """Encrypt OAuth tokens for storage"""
try: try:
token_json = json.dumps(token) token_json = JSONCodec.dumps(token)
encrypted = self.fernet.encrypt(token_json.encode()).decode() encrypted = self.fernet.encrypt(token_json.encode()).decode()
return encrypted return encrypted
except Exception as e: except Exception as e:
@ -96,7 +96,7 @@ class OAuthSessionTable:
"""Decrypt OAuth tokens from storage""" """Decrypt OAuth tokens from storage"""
try: try:
decrypted = self.fernet.decrypt(token.encode()).decode() decrypted = self.fernet.decrypt(token.encode()).decode()
return json.loads(decrypted) return JSONCodec.loads(decrypted)
except Exception as e: except Exception as e:
log.error(f'Error decrypting tokens: {type(e).__name__}: {e}') log.error(f'Error decrypting tokens: {type(e).__name__}: {e}')
raise raise

View file

@ -1,7 +1,6 @@
"""Prompt history model for version tracking.""" """Prompt history model for version tracking."""
import difflib import difflib
import json
import time import time
import uuid import uuid
from typing import Optional from typing import Optional

View file

@ -2,7 +2,6 @@
from __future__ import annotations from __future__ import annotations
import json
import logging import logging
import time import time
import uuid import uuid
@ -15,6 +14,7 @@ from open_webui.models.access_grants import AccessGrantModel, AccessGrants
from open_webui.models.groups import Groups from open_webui.models.groups import Groups
from open_webui.models.prompt_history import PromptHistories from open_webui.models.prompt_history import PromptHistories
from open_webui.models.users import User, UserModel, UserResponse, Users from open_webui.models.users import User, UserModel, UserResponse, Users
from open_webui.utils.misc import json_text_variants
from pydantic import BaseModel, ConfigDict, Field from pydantic import BaseModel, ConfigDict, Field
from sqlalchemy import JSON, BigInteger, Boolean, Column, String, Text, cast, delete, func, or_, select, text, update from sqlalchemy import JSON, BigInteger, Boolean, Column, String, Text, cast, delete, func, or_, select, text, update
from sqlalchemy.ext.asyncio import AsyncSession from sqlalchemy.ext.asyncio import AsyncSession
@ -47,7 +47,7 @@ class PromptModel(BaseModel):
content: str content: str
data: dict | None = None data: dict | None = None
meta: dict | None = None meta: dict | None = None
tags: list[str | None] = None tags: list[str] | None = None
is_active: bool | None = True is_active: bool | None = True
version_id: str | None = None version_id: str | None = None
created_at: int | None = None created_at: int | None = None
@ -86,8 +86,8 @@ class PromptForm(BaseModel):
content: str content: str
data: dict | None = None data: dict | None = None
meta: dict | None = None meta: dict | None = None
tags: list[str | None] = None tags: list[str] | None = None
access_grants: list[dict | None] = None access_grants: list[dict] | None = None
version_id: str | None = None # Active version version_id: str | None = None # Active version
commit_message: str | None = None # For history tracking commit_message: str | None = None # For history tracking
is_production: bool | None = True # Whether to set new version as production is_production: bool | None = True # Whether to set new version as production
@ -100,7 +100,7 @@ class PromptsTable:
async def _to_prompt_model( async def _to_prompt_model(
self, self,
prompt: Prompt, prompt: Prompt,
access_grants: list[AccessGrantModel | None] = None, access_grants: list[AccessGrantModel] | None = None,
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> PromptModel: ) -> PromptModel:
prompt_model = PromptModel.model_validate(prompt) prompt_model = PromptModel.model_validate(prompt)
@ -334,17 +334,19 @@ class PromptsTable:
tag_lower = tag.lower() tag_lower = tag.lower()
if dialect_name == 'sqlite': if dialect_name == 'sqlite':
tag_lower = tag.replace('\\', '\\\\').replace('%', '\\%').replace('_', '\\_')
tag_clause = text( tag_clause = text(
'EXISTS (SELECT 1 FROM json_each(prompt.tags) t WHERE LOWER(t.value) = :tag_val)' "EXISTS (SELECT 1 FROM json_each(prompt.tags) t WHERE t.value LIKE :tag_val ESCAPE '\\')"
) )
elif dialect_name == 'postgresql': elif dialect_name == 'postgresql':
tag_clause = text( tag_clause = text(
'EXISTS (SELECT 1 FROM json_array_elements_text(prompt.tags) t WHERE LOWER(t) = :tag_val)' 'EXISTS (SELECT 1 FROM json_array_elements_text(prompt.tags) t WHERE LOWER(t) = :tag_val)'
) )
else: else:
# Fallback: LIKE on serialised JSON text (ASCII-safe only) # Fallback for dialects with no JSON array function: LIKE on the text.
tag_clause = func.lower(cast(Prompt.tags, String)).like( tags_text = func.lower(cast(Prompt.tags, String))
f'%{json.dumps(tag_lower, ensure_ascii=False)}%' tag_clause = or_(
*(tags_text.like(f'%"{variant}"%') for variant in json_text_variants(tag_lower))
) )
tag_lower = None tag_lower = None
@ -504,14 +506,16 @@ class PromptsTable:
) )
# Update prompt fields # Update prompt fields
prompt.name = form_data.name
prompt.command = form_data.command prompt.command = form_data.command
prompt.content = form_data.content
prompt.data = form_data.data or prompt.data
prompt.meta = form_data.meta or prompt.meta
if form_data.tags is not None: if form_data.is_production:
prompt.tags = form_data.tags prompt.name = form_data.name
prompt.content = form_data.content
prompt.data = form_data.data or prompt.data
prompt.meta = form_data.meta or prompt.meta
if form_data.tags is not None:
prompt.tags = form_data.tags
if form_data.access_grants is not None: if form_data.access_grants is not None:
await AccessGrants.set_access_grants('prompt', prompt.id, form_data.access_grants, db=session) await AccessGrants.set_access_grants('prompt', prompt.id, form_data.access_grants, db=session)
@ -529,7 +533,7 @@ class PromptsTable:
'command': prompt.command, 'command': prompt.command,
'data': form_data.data or {}, 'data': form_data.data or {},
'meta': form_data.meta or {}, 'meta': form_data.meta or {},
'tags': prompt.tags or [], 'tags': form_data.tags if form_data.tags is not None else (prompt.tags or []),
'access_grants': [grant.model_dump() for grant in current_access_grants], 'access_grants': [grant.model_dump() for grant in current_access_grants],
} }
@ -556,7 +560,7 @@ class PromptsTable:
prompt_id: str, prompt_id: str,
name: str, name: str,
command: str, command: str,
tags: list[str | None] = None, tags: list[str] | None = None,
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> PromptModel | None: ) -> PromptModel | None:
"""Update only name, command, and tags (no history created).""" """Update only name, command, and tags (no history created)."""

View file

@ -33,6 +33,7 @@ class Skill(Base):
class SkillMeta(BaseModel): class SkillMeta(BaseModel):
i18n: dict[str, dict[str, str]] | None = None
tags: Optional[list[str]] = [] tags: Optional[list[str]] = []
@ -163,9 +164,30 @@ class SkillsTable:
except Exception: except Exception:
return None return None
async def get_skills(self, db: Optional[AsyncSession] = None) -> list[SkillUserModel]: async def get_skills(
self,
user_id: str | None = None,
ids: list[str] | None = None,
db: AsyncSession | None = None,
) -> list[SkillUserModel]:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
result = await db.execute(select(Skill).order_by(Skill.updated_at.desc())) stmt = select(Skill).order_by(Skill.updated_at.desc())
if ids is not None:
stmt = stmt.filter(Skill.id.in_(ids))
if user_id is not None:
user_group_ids = {group.id for group in await Groups.get_groups_by_member_id(user_id, db=db)}
stmt = AccessGrants.has_permission_filter(
db=db,
query=stmt,
DocumentModel=Skill,
filter={'user_id': user_id, 'group_ids': user_group_ids},
resource_type='skill',
permission='read',
)
result = await db.execute(stmt)
all_skills = result.scalars().all() all_skills = result.scalars().all()
user_ids = list(set(skill.user_id for skill in all_skills)) user_ids = list(set(skill.user_id for skill in all_skills))
@ -194,28 +216,6 @@ class SkillsTable:
) )
return skills return skills
async def get_skills_by_user_id(
self, user_id: str, permission: str = 'write', db: Optional[AsyncSession] = None
) -> list[SkillUserModel]:
skills = await self.get_skills(db=db)
user_groups = await Groups.get_groups_by_member_id(user_id, db=db)
user_group_ids = {group.id for group in user_groups}
result = []
for skill in skills:
if skill.user_id == user_id:
result.append(skill)
elif await AccessGrants.has_access(
user_id=user_id,
resource_type='skill',
resource_id=skill.id,
permission=permission,
user_group_ids=user_group_ids,
db=db,
):
result.append(skill)
return result
async def search_skills( async def search_skills(
self, self,
user_id: str, user_id: str,

View file

@ -97,7 +97,7 @@ class TagTable:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
id = name.replace(' ', '_').lower() id = name.replace(' ', '_').lower()
result = await db.execute(delete(Tag).filter_by(id=id, user_id=user_id)) result = await db.execute(delete(Tag).filter_by(id=id, user_id=user_id))
log.debug(f'res: {result.rowcount}') log.debug('res: %s', result.rowcount)
await db.commit() await db.commit()
return True return True
except Exception as e: except Exception as e:

View file

@ -34,6 +34,7 @@ class Tool(Base): # database table definition
class ToolMeta(BaseModel): class ToolMeta(BaseModel):
i18n: dict[str, dict[str, str]] | None = None
description: str | None = None description: str | None = None
manifest: dict | None = {} manifest: dict | None = {}
has_user_valves: bool = False has_user_valves: bool = False
@ -89,7 +90,7 @@ class ToolForm(BaseModel):
name: str name: str
content: str content: str
meta: ToolMeta meta: ToolMeta
access_grants: list[dict | None] = None access_grants: list[dict] | None = None
class ToolValves(BaseModel): class ToolValves(BaseModel):
@ -103,7 +104,7 @@ class ToolsTable:
async def _to_tool_model( async def _to_tool_model(
self, self,
tool: Tool, tool: Tool,
access_grants: list[AccessGrantModel | None] = None, access_grants: list[AccessGrantModel] | None = None,
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> ToolModel: ) -> ToolModel:
tool_model = ToolModel.model_validate(tool) tool_model = ToolModel.model_validate(tool)
@ -169,20 +170,37 @@ class ToolsTable:
for tool in tools for tool in tools
} }
async def get_tools(self, defer_content: bool = False, db: AsyncSession | None = None) -> list[ToolUserModel]: async def get_tools(
self,
defer_content: bool = False,
db: AsyncSession | None = None,
user_id: str | None = None,
user_group_ids: set[str] | None = None,
permission: str = 'read',
) -> list[ToolUserModel]:
async with get_async_db_context(db) as db: async with get_async_db_context(db) as db:
if defer_content: # Skip Tool.content (plugin source, potentially large) via a
# Skip Tool.content (plugin source, potentially large) via a # column select; Row attributes satisfy from_attributes.
# column select; Row attributes satisfy from_attributes. stmt = (
result = await db.execute( select(Tool.id, Tool.user_id, Tool.name, Tool.specs, Tool.meta, Tool.updated_at, Tool.created_at)
select( if defer_content
Tool.id, Tool.user_id, Tool.name, Tool.specs, Tool.meta, Tool.updated_at, Tool.created_at else select(Tool)
).order_by(Tool.updated_at.desc()) ).order_by(Tool.updated_at.desc())
if user_id is not None:
if user_group_ids is None:
user_group_ids = {group.id for group in await Groups.get_groups_by_member_id(user_id, db=db)}
stmt = AccessGrants.has_permission_filter(
db=db,
query=stmt,
DocumentModel=Tool,
filter={'user_id': user_id, 'group_ids': user_group_ids},
resource_type='tool',
permission=permission,
) )
all_tools = result.all()
else: result = await db.execute(stmt)
result = await db.execute(select(Tool).order_by(Tool.updated_at.desc())) all_tools = result.all() if defer_content else result.scalars().all()
all_tools = result.scalars().all()
user_ids = list(set(tool.user_id for tool in all_tools)) user_ids = list(set(tool.user_id for tool in all_tools))
tool_ids = [tool.id for tool in all_tools] tool_ids = [tool.id for tool in all_tools]
@ -217,20 +235,15 @@ class ToolsTable:
defer_content: bool = False, defer_content: bool = False,
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> list[ToolUserModel]: ) -> list[ToolUserModel]:
tools = await self.get_tools(defer_content=defer_content, db=db)
user_groups = await Groups.get_groups_by_member_id(user_id, db=db) user_groups = await Groups.get_groups_by_member_id(user_id, db=db)
user_group_ids = {group.id for group in user_groups} user_group_ids = {group.id for group in user_groups}
return await self.get_tools(
# One grants query for all non-owned tools instead of one per tool defer_content=defer_content,
accessible_ids = await AccessGrants.get_accessible_resource_ids(
user_id=user_id,
resource_type='tool',
resource_ids=[tool.id for tool in tools if tool.user_id != user_id],
permission=permission,
user_group_ids=user_group_ids,
db=db, db=db,
user_id=user_id,
user_group_ids=user_group_ids,
permission=permission,
) )
return [tool for tool in tools if tool.user_id == user_id or tool.id in accessible_ids]
async def get_tool_valves_by_id(self, id: str, db: AsyncSession | None = None) -> dict | None: async def get_tool_valves_by_id(self, id: str, db: AsyncSession | None = None) -> dict | None:
try: try:

View file

@ -4,12 +4,18 @@ from __future__ import annotations
import datetime import datetime
import time import time
from typing import Optional from typing import Literal, Optional
from open_webui.env import DATABASE_USER_ACTIVE_STATUS_UPDATE_INTERVAL from open_webui.env import DATABASE_USER_ACTIVE_STATUS_UPDATE_INTERVAL
from open_webui.internal.db import Base, JSONField, get_async_db_context from open_webui.internal.db import Base, JSONField, get_async_db_context
from open_webui.utils.misc import throttle from open_webui.utils.misc import throttle
from open_webui.utils.validate import validate_profile_image_url from open_webui.utils.validate import validate_image_url
from pydantic import BaseModel, ConfigDict, Field, field_validator, model_validator from pydantic import (
BaseModel,
ConfigDict,
Field,
field_validator,
model_validator,
)
from sqlalchemy import ( from sqlalchemy import (
JSON, JSON,
BigInteger, BigInteger,
@ -27,7 +33,6 @@ from sqlalchemy import (
select, select,
update, update,
) )
from sqlalchemy.dialects.postgresql import JSONB
from sqlalchemy.ext.asyncio import AsyncSession from sqlalchemy.ext.asyncio import AsyncSession
#################### ####################
@ -37,6 +42,97 @@ from sqlalchemy.ext.asyncio import AsyncSession
#################### ####################
class InterfaceTitleSettings(BaseModel):
model_config = ConfigDict(extra='forbid')
auto: bool | None = None
class InterfaceImageCompressionSize(BaseModel):
model_config = ConfigDict(extra='forbid')
width: int | float | Literal[''] | None = None
height: int | float | Literal[''] | None = None
class InterfaceFloatingActionButton(BaseModel):
model_config = ConfigDict(extra='forbid')
id: str
label: str
input: bool
prompt: str
class InterfaceSettings(BaseModel):
"""Fields owned by the Interface settings panel; not the entire user UI dict."""
model_config = ConfigDict(extra='forbid')
autoTags: bool | None = None
autoFollowUps: bool | None = None
highContrastMode: bool | None = None
detectArtifacts: bool | None = None
responseAutoCopy: bool | None = None
showUsername: bool | None = None
showUpdateToast: bool | None = None
showChangelog: bool | None = None
showEmojiInCall: bool | None = None
voiceInterruption: bool | None = None
displayMultiModelResponsesInTabs: bool | None = None
chatFadeStreamingText: bool | None = None
richTextInput: bool | None = None
showFormattingToolbar: bool | None = None
insertPromptAsRichText: bool | None = None
promptAutocomplete: bool | None = None
insertSuggestionPrompt: bool | None = None
keepFollowUpPrompts: bool | None = None
insertFollowUpPrompt: bool | None = None
regenerateMenu: bool | None = None
enableMessageQueue: bool | None = None
largeTextAsFile: bool | None = None
copyFormatted: bool | None = None
collapseCodeBlocks: bool | None = None
renderMarkdownInUserMessages: bool | None = None
renderMarkdownInAssistantMessages: bool | None = None
expandDetails: bool | None = None
chatHoverPreview: bool | None = None
renderMarkdownInPreviews: bool | None = None
chatBubble: bool | None = None
widescreenMode: bool | None = None
splitLargeChunks: bool | None = None
scrollOnBranchChange: bool | None = None
scrollOnResponseGeneration: bool | None = None
showFilesOnTerminalSelect: bool | None = None
temporaryChatByDefault: bool | None = None
userLocation: bool | None = None
showChatTitleInTab: bool | None = None
iframeSandboxAllowScripts: bool | None = None
iframeSandboxAllowSameOrigin: bool | None = None
iframeSandboxAllowForms: bool | None = None
iframeSandboxAllowDownloads: bool | None = None
terminalPreviewAllowSameOrigin: bool | None = None
stylizedPdfExport: bool | None = None
hapticFeedback: bool | None = None
ctrlEnterToSend: bool | None = None
showFloatingActionButtons: bool | None = None
imageCompression: bool | None = None
imageCompressionInChannels: bool | None = None
landingPageMode: Literal['', 'chat'] | None = None
chatDirection: Literal['LTR', 'RTL', 'auto'] | None = None
terminalFileDisplay: Literal['sidebar', 'inline'] | None = None
defaultUploadContext: Literal['full', 'focused'] | None = None
webSearch: Literal['always'] | None = None
models: list[str] | None = None
backgroundImageUrl: str | None = None
fontFamily: str | None = None
textScale: float | None = None
title: InterfaceTitleSettings | None = None
imageCompressionSize: InterfaceImageCompressionSize | None = None
floatingActionButtons: list[InterfaceFloatingActionButton] | None = None
class UserSettings(BaseModel): class UserSettings(BaseModel):
ui: dict | None = {} ui: dict | None = {}
model_config = ConfigDict(extra='allow') model_config = ConfigDict(extra='allow')
@ -181,7 +277,7 @@ class UpdateProfileForm(BaseModel):
@field_validator('profile_image_url') @field_validator('profile_image_url')
@classmethod @classmethod
def check_profile_image_url(cls, v: str) -> str: def check_profile_image_url(cls, v: str) -> str:
return validate_profile_image_url(v) return validate_image_url(v)
class UserGroupIdsModel(UserModel): class UserGroupIdsModel(UserModel):
@ -271,7 +367,7 @@ class UserUpdateForm(BaseModel):
def check_profile_image_url(cls, v: str | None) -> str | None: def check_profile_image_url(cls, v: str | None) -> str | None:
if v is None: if v is None:
return v return v
return validate_profile_image_url(v) return validate_image_url(v)
class UsersTable: class UsersTable:
@ -287,7 +383,7 @@ class UsersTable:
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> UserModel | None: ) -> UserModel | None:
try: try:
profile_image_url = validate_profile_image_url(profile_image_url) profile_image_url = validate_image_url(profile_image_url)
except ValueError: except ValueError:
profile_image_url = '/user.png' profile_image_url = '/user.png'
@ -360,16 +456,17 @@ class UsersTable:
sub: str, sub: str,
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> UserModel | None: ) -> UserModel | None:
"""Look up a user by OAuth provider + subject claim (dialect-aware JSON filter).""" """Look up a user by OAuth provider + subject claim."""
sub = str(sub)
async with get_async_db_context(db) as session: async with get_async_db_context(db) as session:
dialect = session.bind.dialect.name # Subscript, never contains(): on a JSON column contains() degrades to a substring LIKE.
query = select(User) sub_expr = User.oauth[provider]['sub'].as_string()
if dialect == 'sqlite': query = select(User).where(sub_expr == sub)
oauth_match = User.oauth.contains({provider: {'sub': sub}}) # SQLite preserves JSON numeric type here; Postgres ->> already compares numeric JSON as text.
query = query.where(oauth_match) if session.get_bind().dialect.name == 'sqlite' and sub.isdecimal():
elif dialect == 'postgresql': sub_int = int(sub)
oauth_match = User.oauth[provider].cast(JSONB)['sub'].astext == sub if str(sub_int) == sub and sub_int <= 2**63 - 1:
query = query.where(oauth_match) query = select(User).where(or_(sub_expr == sub, sub_expr == sub_int))
row = (await session.execute(query)).scalars().first() row = (await session.execute(query)).scalars().first()
return UserModel.model_validate(row) if row else None return UserModel.model_validate(row) if row else None
@ -379,27 +476,76 @@ class UsersTable:
external_id: str, external_id: str,
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> UserModel | None: ) -> UserModel | None:
"""Look up a user by SCIM provider + external ID (dialect-aware JSON filter).""" """Look up a user by SCIM provider + external ID."""
async with get_async_db_context(db) as session: async with get_async_db_context(db) as session:
dialect = session.bind.dialect.name # Subscript, never contains(): on a JSON column contains() degrades to a substring LIKE.
query = select(User) query = select(User).where(User.scim[provider]['external_id'].as_string() == external_id)
if dialect == 'sqlite':
scim_match = User.scim.contains({provider: {'external_id': external_id}})
query = query.where(scim_match)
elif dialect == 'postgresql':
scim_match = User.scim[provider].cast(JSONB)['external_id'].astext == external_id
query = query.where(scim_match)
row = (await session.execute(query)).scalars().first() row = (await session.execute(query)).scalars().first()
return UserModel.model_validate(row) if row else None return UserModel.model_validate(row) if row else None
async def get_users( async def get_scim_users(
self, self,
filter: dict | None = None, filter: dict | None = None,
sort: dict | None = None,
skip: int | None = None, skip: int | None = None,
limit: int | None = None, limit: int | None = None,
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> dict: ) -> dict:
"""Paginated user listing with optional filters for role, group, and channel.""" async with get_async_db_context(db) as session:
stmt = select(User).where(or_(User.oauth.cast(String) != 'null', User.scim.cast(String) != 'null'))
if filter:
user_id = filter.get('id')
if user_id:
stmt = stmt.where(User.id == user_id)
email = filter.get('email')
if email:
stmt = stmt.where(func.lower(User.email) == email.lower())
order_by = sort.get('order_by') if sort else None
direction = sort.get('direction') if sort else None
if order_by == 'created_at':
stmt = stmt.order_by(User.created_at.asc() if direction == 'asc' else User.created_at.desc())
count_result = await session.execute(select(func.count()).select_from(stmt.subquery()))
total = count_result.scalar()
if skip is not None:
stmt = stmt.offset(skip)
if limit is not None:
stmt = stmt.limit(limit)
result = await session.execute(stmt)
users = result.scalars().all()
return {
'users': [UserModel.model_validate(user) for user in users],
'total': total,
}
async def get_scim_user_by_id(
self,
id: str,
db: AsyncSession | None = None,
) -> UserModel | None:
async with get_async_db_context(db) as session:
stmt = select(User).where(
User.id == id,
or_(User.oauth.cast(String) != 'null', User.scim.cast(String) != 'null'),
)
user = (await session.execute(stmt)).scalars().first()
return UserModel.model_validate(user) if user else None
async def get_users(
self,
filter: dict | None = None,
sort: dict | None = None,
skip: int | None = None,
limit: int | None = None,
db: AsyncSession | None = None,
) -> dict:
"""Paginated user listing with optional filters and sort."""
async with get_async_db_context(db) as session: async with get_async_db_context(db) as session:
# Deferred imports to avoid circular dependencies # Deferred imports to avoid circular dependencies
from open_webui.models.channels import ChannelMember from open_webui.models.channels import ChannelMember
@ -460,64 +606,63 @@ class UsersTable:
if exclude_roles: if exclude_roles:
stmt = stmt.filter(~User.role.in_(exclude_roles)) stmt = stmt.filter(~User.role.in_(exclude_roles))
order_by = filter.get('order_by') order_by = sort.get('order_by') if sort else None
direction = filter.get('direction') direction = sort.get('direction') if sort else None
if order_by and order_by.startswith('group_id:'): if order_by and order_by.startswith('group_id:'):
group_id = order_by.split(':', 1)[1] group_id = order_by.split(':', 1)[1]
# Subquery that checks if the user belongs to the group # Subquery that checks if the user belongs to the group
membership_exists = exists( membership_exists = exists(
select(GroupMember.id).where( select(GroupMember.id).where(
GroupMember.user_id == User.id, GroupMember.user_id == User.id,
GroupMember.group_id == group_id, GroupMember.group_id == group_id,
)
) )
)
# CASE: user in group → 1, user not in group → 0 # CASE: user in group → 1, user not in group → 0
group_sort = case((membership_exists, 1), else_=0) group_sort = case((membership_exists, 1), else_=0)
if direction == 'asc': if direction == 'asc':
stmt = stmt.order_by(group_sort.asc(), User.name.asc()) stmt = stmt.order_by(group_sort.asc(), User.name.asc())
else: else:
stmt = stmt.order_by(group_sort.desc(), User.name.asc()) stmt = stmt.order_by(group_sort.desc(), User.name.asc())
elif order_by == 'name': elif order_by == 'name':
if direction == 'asc': if direction == 'asc':
stmt = stmt.order_by(User.name.asc()) stmt = stmt.order_by(User.name.asc())
else: else:
stmt = stmt.order_by(User.name.desc()) stmt = stmt.order_by(User.name.desc())
elif order_by == 'email': elif order_by == 'email':
if direction == 'asc': if direction == 'asc':
stmt = stmt.order_by(User.email.asc()) stmt = stmt.order_by(User.email.asc())
else: else:
stmt = stmt.order_by(User.email.desc()) stmt = stmt.order_by(User.email.desc())
elif order_by == 'created_at': elif order_by == 'created_at':
if direction == 'asc': if direction == 'asc':
stmt = stmt.order_by(User.created_at.asc()) stmt = stmt.order_by(User.created_at.asc())
else: else:
stmt = stmt.order_by(User.created_at.desc()) stmt = stmt.order_by(User.created_at.desc())
elif order_by == 'last_active_at': elif order_by == 'last_active_at':
if direction == 'asc': if direction == 'asc':
stmt = stmt.order_by(User.last_active_at.asc()) stmt = stmt.order_by(User.last_active_at.asc())
else: else:
stmt = stmt.order_by(User.last_active_at.desc()) stmt = stmt.order_by(User.last_active_at.desc())
elif order_by == 'updated_at': elif order_by == 'updated_at':
if direction == 'asc': if direction == 'asc':
stmt = stmt.order_by(User.updated_at.asc()) stmt = stmt.order_by(User.updated_at.asc())
else: else:
stmt = stmt.order_by(User.updated_at.desc()) stmt = stmt.order_by(User.updated_at.desc())
elif order_by == 'role': elif order_by == 'role':
if direction == 'asc': if direction == 'asc':
stmt = stmt.order_by(User.role.asc()) stmt = stmt.order_by(User.role.asc())
else: else:
stmt = stmt.order_by(User.role.desc()) stmt = stmt.order_by(User.role.desc())
elif not filter:
else:
stmt = stmt.order_by(User.created_at.desc()) stmt = stmt.order_by(User.created_at.desc())
# Count BEFORE pagination # Count BEFORE pagination
@ -609,7 +754,7 @@ class UsersTable:
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> UserModel | None: ) -> UserModel | None:
try: try:
profile_image_url = validate_profile_image_url(profile_image_url) profile_image_url = validate_image_url(profile_image_url)
except ValueError: except ValueError:
profile_image_url = '/user.png' profile_image_url = '/user.png'
@ -636,7 +781,10 @@ class UsersTable:
if not user: if not user:
return None return None
oauth = dict(user.oauth or {}) oauth = dict(user.oauth or {})
oauth[provider] = {'sub': sub} provider_oauth = oauth.get(provider)
provider_oauth = dict(provider_oauth) if isinstance(provider_oauth, dict) else {}
provider_oauth['sub'] = str(sub)
oauth[provider] = provider_oauth
user.oauth = oauth user.oauth = oauth
await session.commit() await session.commit()
return UserModel.model_validate(user) return UserModel.model_validate(user)
@ -645,7 +793,7 @@ class UsersTable:
self, self,
id: str, id: str,
provider: str, provider: str,
external_id: str, external_id: str | None,
db: AsyncSession | None = None, db: AsyncSession | None = None,
) -> UserModel | None: ) -> UserModel | None:
"""Update or insert a SCIM provider/external_id pair into the user's scim JSON field.""" """Update or insert a SCIM provider/external_id pair into the user's scim JSON field."""
@ -655,7 +803,9 @@ class UsersTable:
return None return None
scim = dict(user.scim or {}) scim = dict(user.scim or {})
scim[provider] = {'external_id': external_id} scim[provider] = {'external_id': external_id}
user.scim = scim if scim != user.scim:
user.scim = scim
user.updated_at = int(time.time())
await session.commit() await session.commit()
return UserModel.model_validate(user) return UserModel.model_validate(user)
@ -678,7 +828,18 @@ class UsersTable:
if not user: if not user:
return None return None
user_settings = dict(user.settings or {}) user_settings = dict(user.settings or {})
updated = dict(updated)
ui_settings = updated.pop('ui', None)
user_settings.update(updated) user_settings.update(updated)
if ui_settings is not None:
# UI updates are field-level patches: omission keeps a value; null resets it.
current_ui_settings = dict(user_settings.get('ui') or {})
for key, value in ui_settings.items():
if value is None:
current_ui_settings.pop(key, None)
else:
current_ui_settings[key] = value
user_settings['ui'] = current_ui_settings
user.settings = user_settings user.settings = user_settings
await session.commit() await session.commit()
return UserModel.model_validate(user) return UserModel.model_validate(user)

View file

@ -7,6 +7,7 @@ from typing import List, Optional
import requests import requests
from fastapi import HTTPException, status from fastapi import HTTPException, status
from langchain_core.documents import Document from langchain_core.documents import Document
from open_webui.utils.json_codec import JSONCodec
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@ -64,25 +65,6 @@ class DatalabMarkerLoader:
} }
return mime_map.get(ext, 'application/octet-stream') return mime_map.get(ext, 'application/octet-stream')
def check_marker_request_status(self, request_id: str) -> dict:
url = f'{self.api_base_url}/{request_id}'
headers = {'X-Api-Key': self.api_key}
try:
response = requests.get(url, headers=headers)
response.raise_for_status()
result = response.json()
log.info(f'Marker API status check for request {request_id}: {result}')
return result
except requests.HTTPError as e:
log.error(f'Error checking Marker request status: {e}')
raise HTTPException(
status.HTTP_502_BAD_GATEWAY,
detail=f'Failed to check Marker request: {e}',
)
except ValueError as e:
log.error(f'Invalid JSON checking Marker request: {e}')
raise HTTPException(status.HTTP_502_BAD_GATEWAY, detail=f'Invalid JSON: {e}')
def load(self) -> List[Document]: def load(self) -> List[Document]:
filename = os.path.basename(self.file_path) filename = os.path.basename(self.file_path)
mime_type = self._get_mime_type(filename) mime_type = self._get_mime_type(filename)
@ -103,7 +85,10 @@ class DatalabMarkerLoader:
form_data['additional_config'] = self.additional_config form_data['additional_config'] = self.additional_config
log.info( log.info(
f"Datalab Marker POST request parameters: {{'filename': '{filename}', 'mime_type': '{mime_type}', **{form_data}}}" "Datalab Marker POST request parameters: {'filename': '%s', 'mime_type': '%s', **%s}",
filename,
mime_type,
form_data,
) )
try: try:
@ -167,7 +152,7 @@ class DatalabMarkerLoader:
'total_cost', 'total_cost',
) )
} }
log.info(f'Marker processing completed successfully: {json.dumps(summary, indent=2)}') log.info('Marker processing completed successfully: %s', json.dumps(summary, indent=2))
break break
if status_val == 'failed' or success_val is False: if status_val == 'failed' or success_val is False:
@ -234,7 +219,7 @@ class DatalabMarkerLoader:
try: try:
with open(output_path, 'w', encoding='utf-8') as f: with open(output_path, 'w', encoding='utf-8') as f:
f.write(full_text) f.write(full_text)
log.info(f'Saved Marker output to: {output_path}') log.info('Saved Marker output to: %s', output_path)
except Exception as e: except Exception as e:
log.warning(f'Failed to write marker output to disk: {e}') log.warning(f'Failed to write marker output to disk: {e}')
@ -249,11 +234,11 @@ class DatalabMarkerLoader:
images = final_result.get('images', {}) images = final_result.get('images', {})
if images: if images:
metadata['image_count'] = len(images) metadata['image_count'] = len(images)
metadata['images'] = json.dumps(list(images.keys())) metadata['images'] = JSONCodec.dumps(list(images.keys()))
for k, v in metadata.items(): for k, v in metadata.items():
if isinstance(v, (dict, list)): if isinstance(v, (dict, list)):
metadata[k] = json.dumps(v) metadata[k] = JSONCodec.dumps(v)
elif v is None: elif v is None:
metadata[k] = '' metadata[k] = ''

View file

@ -30,6 +30,9 @@ class ExternalWebLoader(BaseLoader):
response = requests.post( response = requests.post(
self.external_url, self.external_url,
headers={ headers={
# LICENSE covers this Open WebUI user-agent identifier.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
'User-Agent': 'Open WebUI (https://github.com/open-webui/open-webui) External Web Loader', 'User-Agent': 'Open WebUI (https://github.com/open-webui/open-webui) External Web Loader',
'Authorization': f'Bearer {self.external_api_key}', 'Authorization': f'Bearer {self.external_api_key}',
}, },

View file

@ -0,0 +1,121 @@
from importlib import import_module
from pathlib import Path
from bs4 import BeautifulSoup
from langchain_core.documents import Document
class TextLoader:
def __init__(self, file_path, encoding=None):
self.file_path = str(file_path)
self.encoding = encoding
def load(self) -> list[Document]:
try:
text = Path(self.file_path).read_text(encoding=self.encoding)
except Exception as e:
raise RuntimeError(f'Error loading {self.file_path}') from e
return [
Document(
page_content=text,
metadata={'source': self.file_path},
)
]
class HTMLLoader(TextLoader):
def load(self) -> list[Document]:
with open(self.file_path, encoding=self.encoding) as file:
soup = BeautifulSoup(file, 'lxml')
return [
Document(
page_content=soup.get_text(),
metadata={'source': self.file_path, 'title': str(soup.title.string) if soup.title else ''},
)
]
class DocxLoader(TextLoader):
def load(self) -> list[Document]:
import docx2txt
return [
Document(
page_content=docx2txt.process(Path(self.file_path).expanduser()),
metadata={'source': self.file_path},
)
]
class UnstructuredLoader:
def __init__(self, file_path, file_format, mode='single', **kwargs):
# Match the optional-package check; format dependencies are loaded when parsing.
import_module('unstructured')
self.file_path = file_path
self.file_format = file_format
self.mode = mode
self.kwargs = kwargs
def load(self) -> list[Document]:
file_format = self.file_format
if file_format in ('doc', 'ppt', 'pptx'):
from unstructured.file_utils.filetype import detect_filetype
legacy_format = 'doc' if file_format == 'doc' else 'ppt'
try:
import_module('magic')
except ImportError:
is_legacy = Path(self.file_path).suffix == f'.{legacy_format}'
else:
is_legacy = detect_filetype(self.file_path).name.lower() == legacy_format
file_format = legacy_format if is_legacy else legacy_format + 'x'
elif file_format == 'msg':
from unstructured.file_utils.filetype import detect_filetype
detected = detect_filetype(self.file_path).name
if detected not in ('EML', 'MSG'):
raise ValueError(f'Unsupported email file type: {detected}')
file_format = 'email' if detected == 'EML' else 'msg'
module = import_module(f'unstructured.partition.{file_format}')
elements = getattr(module, f'partition_{file_format}')(filename=self.file_path, **self.kwargs)
metadata = {'source': str(self.file_path)}
if self.mode == 'elements':
return [
Document(
page_content=str(element),
metadata={
**metadata,
**element.metadata.to_dict(),
'category': element.category,
'element_id': element.id,
},
)
for element in elements
]
return [Document(page_content='\n\n'.join(map(str, elements)), metadata=metadata)]
class DocumentIntelligenceLoader:
def __init__(self, file_path, api_endpoint, api_key=None, azure_credential=None, api_model='prebuilt-layout'):
if (api_key is None) == (azure_credential is None):
raise ValueError('Provide exactly one of api_key or azure_credential.')
self.file_path = file_path
self.api_endpoint = api_endpoint
self.api_key = api_key
self.azure_credential = azure_credential
self.api_model = api_model
def load(self) -> list[Document]:
from azure.ai.documentintelligence import DocumentIntelligenceClient
from azure.core.credentials import AzureKeyCredential
credential = self.azure_credential if self.azure_credential is not None else AzureKeyCredential(self.api_key)
with DocumentIntelligenceClient(self.api_endpoint, credential) as client, open(self.file_path, 'rb') as file:
result = client.begin_analyze_document(
self.api_model,
body=file,
content_type='application/octet-stream',
output_content_format='markdown',
).result()
return [Document(page_content=result.content, metadata=result.as_dict())]

View file

@ -1,33 +1,37 @@
import asyncio import asyncio
import json import csv
import logging import logging
import os
import sys import sys
import zipfile
import ftfy import ftfy
import requests import requests
from fastapi import HTTPException
from azure.identity import DefaultAzureCredential from azure.identity import DefaultAzureCredential
from langchain_community.document_loaders import (
AzureAIDocumentIntelligenceLoader,
BSHTMLLoader,
CSVLoader,
Docx2txtLoader,
PyPDFLoader,
TextLoader,
YoutubeLoader,
)
from langchain_core.documents import Document from langchain_core.documents import Document
from open_webui.env import ( from open_webui.env import (
AIOHTTP_CLIENT_SESSION_SSL, AIOHTTP_CLIENT_SESSION_SSL,
GLOBAL_LOG_LEVEL, GLOBAL_LOG_LEVEL,
USE_SLIM,
MINERU_MAX_MARKDOWN_BYTES, MINERU_MAX_MARKDOWN_BYTES,
REQUESTS_VERIFY, REQUESTS_VERIFY,
) )
from open_webui.retrieval.loaders.datalab_marker import DatalabMarkerLoader from open_webui.retrieval.loaders.datalab_marker import DatalabMarkerLoader
from open_webui.retrieval.loaders.external_document import ExternalDocumentLoader from open_webui.retrieval.loaders.external_document import ExternalDocumentLoader
from open_webui.retrieval.loaders.local import (
DocumentIntelligenceLoader,
DocxLoader,
HTMLLoader,
TextLoader,
UnstructuredLoader,
)
from open_webui.retrieval.loaders.mineru import MinerULoader from open_webui.retrieval.loaders.mineru import MinerULoader
from open_webui.retrieval.loaders.mistral import MistralLoader from open_webui.retrieval.loaders.mistral import MistralLoader
from open_webui.retrieval.loaders.paddleocr_vl import PADDLEOCR_VL_SUPPORTED_EXTENSIONS, PaddleOCRVLLoader from open_webui.retrieval.loaders.paddleocr_vl import PADDLEOCR_VL_SUPPORTED_EXTENSIONS, PaddleOCRVLLoader
from open_webui.retrieval.loaders.pdf import PDFLoader
from open_webui.utils.headers import get_user_groups_for_custom_headers from open_webui.utils.headers import get_user_groups_for_custom_headers
from open_webui.utils.json_codec import JSONCodec
logging.basicConfig(stream=sys.stdout, level=GLOBAL_LOG_LEVEL) logging.basicConfig(stream=sys.stdout, level=GLOBAL_LOG_LEVEL)
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@ -48,6 +52,7 @@ known_source_ext = [
'h', 'h',
'c', 'c',
'cs', 'cs',
'ino',
'sql', 'sql',
'log', 'log',
'ini', 'ini',
@ -86,8 +91,18 @@ known_source_ext = [
'yaml', 'yaml',
'yml', 'yml',
'toml', 'toml',
'svg',
] ]
known_archive_ext = {'docx', 'epub', 'odt', 'pptx', 'xlsx'}
known_archive_content_types = {
'application/epub+zip',
'application/vnd.oasis.opendocument.text',
'application/vnd.openxmlformats-officedocument.presentationml.presentation',
'application/vnd.openxmlformats-officedocument.spreadsheetml.sheet',
'application/vnd.openxmlformats-officedocument.wordprocessingml.document',
}
class ExcelLoader: class ExcelLoader:
"""Fallback Excel loader using pandas when unstructured is not installed.""" """Fallback Excel loader using pandas when unstructured is not installed."""
@ -111,6 +126,66 @@ class ExcelLoader:
] ]
def get_csv_summary(filename: str, file_path: str, encoding: str) -> str | None:
try:
with open(file_path, newline='', encoding=encoding) as f:
sample = f.read(4096)
f.seek(0)
try:
dialect = csv.Sniffer().sniff(sample)
except csv.Error:
dialect = csv.excel
total_rows = 0
max_columns = 0
headers = []
for row in csv.reader(f, dialect):
total_rows += 1
max_columns = max(max_columns, len(row))
if total_rows == 1:
headers = [header.lstrip('\ufeff') for header in row]
except Exception:
return None
if total_rows == 0:
return None
return (
f'Table: {total_rows} rows incl. header; '
f'{max(total_rows - 1, 0)} data rows; '
f'{max_columns} columns: {", ".join(headers)}.'
)
class CSVLoaderWithSummary:
def __init__(self, file_path: str, filename: str, encoding: str):
self.file_path = file_path
self.filename = filename
self.encoding = encoding
def load(self) -> list[Document]:
docs = []
try:
with open(self.file_path, newline='', encoding=self.encoding) as file:
for index, row in enumerate(csv.DictReader(file)):
fields = []
for key, value in row.items():
if isinstance(value, str):
value = value.strip()
elif isinstance(value, list):
value = ','.join(v.strip() for v in value)
fields.append(f'{key.strip() if key is not None else key}: {value}')
content = '\n'.join(fields)
docs.append(Document(page_content=content, metadata={'source': self.file_path, 'row': index}))
except Exception as e:
raise RuntimeError(f'Error loading {self.file_path}') from e
if os.getenv('ENABLE_RAG_CSV_SUMMARY', 'False').lower() == 'true':
summary = get_csv_summary(self.filename, self.file_path, self.encoding)
if summary:
docs.insert(0, Document(page_content=summary, metadata={'source': self.file_path, 'row': -1}))
return docs
class PptxLoader: class PptxLoader:
"""Fallback PowerPoint loader using python-pptx when unstructured is not installed.""" """Fallback PowerPoint loader using python-pptx when unstructured is not installed."""
@ -138,10 +213,11 @@ class PptxLoader:
class TikaLoader: class TikaLoader:
def __init__(self, url, file_path, mime_type=None, extract_images=None): def __init__(self, url, file_path, mime_type=None, extract_images=None, server_version='3'):
self.url = url self.url = url
self.file_path = file_path self.file_path = file_path
self.mime_type = mime_type self.mime_type = mime_type
self.server_version = str(server_version or '3')
self.extract_images = extract_images self.extract_images = extract_images
@ -157,16 +233,15 @@ class TikaLoader:
if self.extract_images == True: if self.extract_images == True:
headers['X-Tika-PDFextractInlineImages'] = 'true' headers['X-Tika-PDFextractInlineImages'] = 'true'
endpoint = self.url endpoint_path = 'tika/json/md' if self.server_version == '4' else 'tika/text'
if not endpoint.endswith('/'): content_key = 'tk:content' if self.server_version == '4' else 'X-TIKA:content'
endpoint += '/' endpoint = f'{self.url.rstrip("/")}/{endpoint_path}'
endpoint += 'tika/text'
r = requests.put(endpoint, data=data, headers=headers, verify=REQUESTS_VERIFY) r = requests.put(endpoint, data=data, headers=headers, verify=REQUESTS_VERIFY)
if r.ok: if r.ok:
raw_metadata = r.json() raw_metadata = r.json()
text = raw_metadata.get('X-TIKA:content', '<No text content found>').strip() text = raw_metadata.get(content_key, '<No text content found>').strip()
if 'Content-Type' in raw_metadata: if 'Content-Type' in raw_metadata:
headers['Content-Type'] = raw_metadata['Content-Type'] headers['Content-Type'] = raw_metadata['Content-Type']
@ -216,8 +291,17 @@ class DoclingLoader:
) )
if r.ok: if r.ok:
result = r.json() result = r.json()
# Docling reports failed and skipped conversions inside HTTP 200 responses.
conversion_status = result.get('status')
if conversion_status in ['failure', 'skipped']:
error_details = (
'; '.join(filter(None, (error.get('error_message') for error in result.get('errors', []))))
or 'no error message provided'
)
raise Exception(f'Error calling Docling: conversion status {conversion_status} - {error_details}')
document_data = result.get('document', {}) document_data = result.get('document', {})
md_content = document_data.get('md_content', '') md_content = document_data.get('md_content') or ''
text = md_content or '<No text content found>' text = md_content or '<No text content found>'
metadata = {'Content-Type': self.mime_type} if self.mime_type else {} metadata = {'Content-Type': self.mime_type} if self.mime_type else {}
@ -256,7 +340,11 @@ class Loader:
def load(self, filename: str, file_content_type: str, file_path: str) -> list[Document]: def load(self, filename: str, file_content_type: str, file_path: str) -> list[Document]:
loader = self._get_loader(filename, file_content_type, file_path) loader = self._get_loader(filename, file_content_type, file_path)
docs = loader.load() docs = loader.load()
return [Document(page_content=ftfy.fix_text(doc.page_content), metadata=doc.metadata) for doc in docs] # ftfy's auto mode unescapes entities on every line before the first literal '<', rewriting the document.
return [
Document(page_content=ftfy.fix_text(doc.page_content, unescape_html=False), metadata=doc.metadata)
for doc in docs
]
async def aload(self, filename: str, file_content_type: str, file_path: str) -> list[Document]: async def aload(self, filename: str, file_content_type: str, file_path: str) -> list[Document]:
""" """
@ -429,6 +517,27 @@ class Loader:
def _get_loader(self, filename: str, file_content_type: str, file_path: str): def _get_loader(self, filename: str, file_content_type: str, file_path: str):
file_ext = filename.split('.')[-1].lower() file_ext = filename.split('.')[-1].lower()
if file_ext in known_archive_ext or file_content_type in known_archive_content_types:
max_file_size = self.kwargs.get('FILE_MAX_SIZE')
try:
max_file_size_bytes = int(max_file_size) * 1024 * 1024 if max_file_size else 100 * 1024 * 1024
except (TypeError, ValueError):
max_file_size_bytes = 100 * 1024 * 1024
if max_file_size_bytes > 0:
try:
with zipfile.ZipFile(file_path) as archive:
uncompressed_size = sum(entry.file_size for entry in archive.infolist())
except (zipfile.BadZipFile, OSError):
pass
else:
max_bytes = min(
max(10 * 1024 * 1024, os.path.getsize(file_path) * 100),
max_file_size_bytes,
)
if uncompressed_size > max_bytes:
raise ValueError('Document archive is too large after decompression')
if ( if (
self.engine == 'external' self.engine == 'external'
and self.kwargs.get('EXTERNAL_DOCUMENT_LOADER_URL') and self.kwargs.get('EXTERNAL_DOCUMENT_LOADER_URL')
@ -455,6 +564,7 @@ class Loader:
loader = TikaLoader( loader = TikaLoader(
url=self.kwargs.get('TIKA_SERVER_URL'), url=self.kwargs.get('TIKA_SERVER_URL'),
file_path=file_path, file_path=file_path,
server_version=self.kwargs.get('TIKA_SERVER_VERSION'),
extract_images=self.kwargs.get('PDF_EXTRACT_IMAGES'), extract_images=self.kwargs.get('PDF_EXTRACT_IMAGES'),
) )
elif ( elif (
@ -508,8 +618,8 @@ class Loader:
params = self.kwargs.get('DOCLING_PARAMS', {}) params = self.kwargs.get('DOCLING_PARAMS', {})
if not isinstance(params, dict): if not isinstance(params, dict):
try: try:
params = json.loads(params) params = JSONCodec.loads(params)
except json.JSONDecodeError: except JSONCodec.JSONDecodeError:
log.error('Invalid DOCLING_PARAMS format, expected JSON object') log.error('Invalid DOCLING_PARAMS format, expected JSON object')
params = {} params = {}
@ -534,14 +644,14 @@ class Loader:
) )
): ):
if self.kwargs.get('DOCUMENT_INTELLIGENCE_KEY') != '': if self.kwargs.get('DOCUMENT_INTELLIGENCE_KEY') != '':
loader = AzureAIDocumentIntelligenceLoader( loader = DocumentIntelligenceLoader(
file_path=file_path, file_path=file_path,
api_endpoint=self.kwargs.get('DOCUMENT_INTELLIGENCE_ENDPOINT'), api_endpoint=self.kwargs.get('DOCUMENT_INTELLIGENCE_ENDPOINT'),
api_key=self.kwargs.get('DOCUMENT_INTELLIGENCE_KEY'), api_key=self.kwargs.get('DOCUMENT_INTELLIGENCE_KEY'),
api_model=self.kwargs.get('DOCUMENT_INTELLIGENCE_MODEL'), api_model=self.kwargs.get('DOCUMENT_INTELLIGENCE_MODEL'),
) )
else: else:
loader = AzureAIDocumentIntelligenceLoader( loader = DocumentIntelligenceLoader(
file_path=file_path, file_path=file_path,
api_endpoint=self.kwargs.get('DOCUMENT_INTELLIGENCE_ENDPOINT'), api_endpoint=self.kwargs.get('DOCUMENT_INTELLIGENCE_ENDPOINT'),
azure_credential=DefaultAzureCredential(), azure_credential=DefaultAzureCredential(),
@ -587,19 +697,34 @@ class Loader:
file_path=file_path, file_path=file_path,
) )
else: else:
if USE_SLIM:
if file_ext == 'csv':
return CSVLoaderWithSummary(file_path, filename, self._detect_text_encoding(file_path))
if file_ext in ['htm', 'html']:
return HTMLLoader(file_path, encoding=self._detect_text_encoding(file_path))
if file_ext in ['txt', 'md', 'markdown', 'rst', 'xml'] or self._is_text_file(
file_ext, file_content_type
):
return TextLoader(file_path, encoding=self._detect_text_encoding(file_path))
raise HTTPException(
503,
'This file type requires an external document extractor in slim. Configure one that supports it.',
)
if file_ext == 'pdf': if file_ext == 'pdf':
loader = PyPDFLoader( loader = PDFLoader(
file_path, file_path,
extract_images=self.kwargs.get('PDF_EXTRACT_IMAGES'), extract_images=self.kwargs.get('PDF_EXTRACT_IMAGES'),
mode=self.kwargs.get('PDF_LOADER_MODE', 'page'), mode=self.kwargs.get('PDF_LOADER_MODE', 'page'),
) )
elif file_ext == 'csv': elif file_ext == 'csv':
loader = CSVLoader(file_path, encoding=self._detect_text_encoding(file_path)) loader = CSVLoaderWithSummary(
file_path,
filename,
self._detect_text_encoding(file_path),
)
elif file_ext == 'rst': elif file_ext == 'rst':
try: try:
from langchain_community.document_loaders import UnstructuredRSTLoader loader = UnstructuredLoader(file_path, 'rst', mode='elements')
loader = UnstructuredRSTLoader(file_path, mode='elements')
except ImportError: except ImportError:
log.warning( log.warning(
"The 'unstructured' package is not installed. " "The 'unstructured' package is not installed. "
@ -609,9 +734,7 @@ class Loader:
loader = TextLoader(file_path, encoding=self._detect_text_encoding(file_path)) loader = TextLoader(file_path, encoding=self._detect_text_encoding(file_path))
elif file_ext == 'xml': elif file_ext == 'xml':
try: try:
from langchain_community.document_loaders import UnstructuredXMLLoader loader = UnstructuredLoader(file_path, 'xml')
loader = UnstructuredXMLLoader(file_path)
except ImportError: except ImportError:
log.warning( log.warning(
"The 'unstructured' package is not installed. " "The 'unstructured' package is not installed. "
@ -620,14 +743,12 @@ class Loader:
) )
loader = TextLoader(file_path, encoding=self._detect_text_encoding(file_path)) loader = TextLoader(file_path, encoding=self._detect_text_encoding(file_path))
elif file_ext in ['htm', 'html']: elif file_ext in ['htm', 'html']:
loader = BSHTMLLoader(file_path, open_encoding='unicode_escape') loader = HTMLLoader(file_path, encoding='unicode_escape')
elif file_ext == 'md': elif file_ext == 'md':
loader = TextLoader(file_path, encoding=self._detect_text_encoding(file_path)) loader = TextLoader(file_path, encoding=self._detect_text_encoding(file_path))
elif file_content_type == 'application/epub+zip': elif file_content_type == 'application/epub+zip':
try: try:
from langchain_community.document_loaders import UnstructuredEPubLoader loader = UnstructuredLoader(file_path, 'epub')
loader = UnstructuredEPubLoader(file_path)
except ImportError: except ImportError:
raise ValueError( raise ValueError(
"Processing .epub files requires the 'unstructured' package. " "Processing .epub files requires the 'unstructured' package. "
@ -637,12 +758,10 @@ class Loader:
file_content_type == 'application/vnd.openxmlformats-officedocument.wordprocessingml.document' file_content_type == 'application/vnd.openxmlformats-officedocument.wordprocessingml.document'
or file_ext == 'docx' or file_ext == 'docx'
): ):
loader = Docx2txtLoader(file_path) loader = DocxLoader(file_path)
elif file_ext == 'doc' or file_content_type == 'application/msword': elif file_ext == 'doc' or file_content_type == 'application/msword':
try: try:
from langchain_community.document_loaders import UnstructuredWordDocumentLoader loader = UnstructuredLoader(file_path, 'doc')
loader = UnstructuredWordDocumentLoader(file_path)
except ImportError: except ImportError:
raise ValueError( raise ValueError(
"Processing .doc files requires the 'unstructured' package. " "Processing .doc files requires the 'unstructured' package. "
@ -653,9 +772,7 @@ class Loader:
'application/vnd.openxmlformats-officedocument.spreadsheetml.sheet', 'application/vnd.openxmlformats-officedocument.spreadsheetml.sheet',
] or file_ext in ['xls', 'xlsx']: ] or file_ext in ['xls', 'xlsx']:
try: try:
from langchain_community.document_loaders import UnstructuredExcelLoader loader = UnstructuredLoader(file_path, 'xlsx')
loader = UnstructuredExcelLoader(file_path)
except ImportError: except ImportError:
log.warning( log.warning(
"The 'unstructured' package is not installed. " "The 'unstructured' package is not installed. "
@ -668,9 +785,7 @@ class Loader:
'application/vnd.openxmlformats-officedocument.presentationml.presentation', 'application/vnd.openxmlformats-officedocument.presentationml.presentation',
] or file_ext in ['ppt', 'pptx']: ] or file_ext in ['ppt', 'pptx']:
try: try:
from langchain_community.document_loaders import UnstructuredPowerPointLoader loader = UnstructuredLoader(file_path, 'ppt' if file_ext == 'ppt' else 'pptx')
loader = UnstructuredPowerPointLoader(file_path)
except ImportError: except ImportError:
log.warning( log.warning(
"The 'unstructured' package is not installed. " "The 'unstructured' package is not installed. "
@ -680,12 +795,8 @@ class Loader:
loader = PptxLoader(file_path) loader = PptxLoader(file_path)
elif file_ext == 'msg': elif file_ext == 'msg':
try: try:
from langchain_community.document_loaders import (
UnstructuredEmailLoader,
)
# unstructured parses .msg via python-oxmsg; avoids extract_msg's beautifulsoup4<4.14 conflict # unstructured parses .msg via python-oxmsg; avoids extract_msg's beautifulsoup4<4.14 conflict
loader = UnstructuredEmailLoader(file_path, process_attachments=False) loader = UnstructuredLoader(file_path, 'msg', process_attachments=False)
except ImportError: except ImportError:
raise ValueError( raise ValueError(
"Processing .msg files requires the 'unstructured' package. " "Processing .msg files requires the 'unstructured' package. "
@ -693,9 +804,7 @@ class Loader:
) )
elif file_ext == 'odt': elif file_ext == 'odt':
try: try:
from langchain_community.document_loaders import UnstructuredODTLoader loader = UnstructuredLoader(file_path, 'odt')
loader = UnstructuredODTLoader(file_path)
except ImportError: except ImportError:
raise ValueError( raise ValueError(
"Processing .odt files requires the 'unstructured' package. " "Processing .odt files requires the 'unstructured' package. "

View file

@ -10,7 +10,6 @@ from langchain_core.documents import Document
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
DEFAULT_MICROSOFT_WEB_IQ_API_BASE_URL = 'https://api.microsoft.ai/v3'
MICROSOFT_BROWSE_RETRY_STATUS_CODES = {202, 429, 500, 502, 503, 504} MICROSOFT_BROWSE_RETRY_STATUS_CODES = {202, 429, 500, 502, 503, 504}
MICROSOFT_BROWSE_MAX_RETRIES = 2 MICROSOFT_BROWSE_MAX_RETRIES = 2
@ -27,7 +26,7 @@ class MicrosoftWebIQLoader(BaseLoader):
continue_on_failure: bool = True, continue_on_failure: bool = True,
) -> None: ) -> None:
self.urls = urls if isinstance(urls, list) else [urls] self.urls = urls if isinstance(urls, list) else [urls]
self.api_base_url = (api_base_url or DEFAULT_MICROSOFT_WEB_IQ_API_BASE_URL).rstrip('/') self.api_base_url = api_base_url.rstrip('/')
self.api_key = api_key self.api_key = api_key
self.language = language self.language = language
self.verify_ssl = verify_ssl self.verify_ssl = verify_ssl

View file

@ -74,7 +74,7 @@ class MinerULoader:
Load document using Local API (synchronous). Load document using Local API (synchronous).
Posts file to /file_parse endpoint and gets immediate response. Posts file to /file_parse endpoint and gets immediate response.
""" """
log.info(f'Using MinerU Local API at {self.api_url}') log.info('Using MinerU Local API at %s', self.api_url)
filename = os.path.basename(self.file_path) filename = os.path.basename(self.file_path)
@ -97,8 +97,8 @@ class MinerULoader:
with open(self.file_path, 'rb') as f: with open(self.file_path, 'rb') as f:
files = {'files': (filename, f, 'application/octet-stream')} files = {'files': (filename, f, 'application/octet-stream')}
log.info(f'Sending file to MinerU Local API: {filename}') log.info('Sending file to MinerU Local API: %s', filename)
log.debug(f'Local API parameters: {form_data}') log.debug('Local API parameters: %s', form_data)
response = requests.post( response = requests.post(
f'{self.api_url}/file_parse', f'{self.api_url}/file_parse',
@ -163,7 +163,7 @@ class MinerULoader:
detail='MinerU returned empty markdown content', detail='MinerU returned empty markdown content',
) )
log.info(f'Successfully parsed document with MinerU Local API: {filename}') log.info('Successfully parsed document with MinerU Local API: %s', filename)
# Create metadata # Create metadata
metadata = { metadata = {
@ -180,7 +180,7 @@ class MinerULoader:
Load document using Cloud API (asynchronous). Load document using Cloud API (asynchronous).
Uses batch upload endpoint to avoid need for public file URLs. Uses batch upload endpoint to avoid need for public file URLs.
""" """
log.info(f'Using MinerU Cloud API at {self.api_url}') log.info('Using MinerU Cloud API at %s', self.api_url)
filename = os.path.basename(self.file_path) filename = os.path.basename(self.file_path)
@ -196,7 +196,7 @@ class MinerULoader:
# Step 4: Download and extract markdown from ZIP # Step 4: Download and extract markdown from ZIP
markdown_content = self._download_and_extract_zip(result['full_zip_url'], filename) markdown_content = self._download_and_extract_zip(result['full_zip_url'], filename)
log.info(f'Successfully parsed document with MinerU Cloud API: {filename}') log.info('Successfully parsed document with MinerU Cloud API: %s', filename)
# Create metadata # Create metadata
metadata = { metadata = {
@ -232,8 +232,8 @@ class MinerULoader:
if self.page_ranges: if self.page_ranges:
request_body['files'][0]['page_ranges'] = self.page_ranges request_body['files'][0]['page_ranges'] = self.page_ranges
log.info(f'Requesting upload URL for: {filename}') log.info('Requesting upload URL for: %s', filename)
log.debug(f'Cloud API request body: {request_body}') log.debug('Cloud API request body: %s', request_body)
try: try:
response = requests.post( response = requests.post(
@ -284,7 +284,7 @@ class MinerULoader:
) )
upload_url = file_urls[0] upload_url = file_urls[0]
log.info(f'Received upload URL for batch: {batch_id}') log.info('Received upload URL for batch: %s', batch_id)
return batch_id, upload_url return batch_id, upload_url
@ -334,7 +334,7 @@ class MinerULoader:
max_iterations = 300 # 10 minutes max (2 seconds per iteration) max_iterations = 300 # 10 minutes max (2 seconds per iteration)
poll_interval = 2 # seconds poll_interval = 2 # seconds
log.info(f'Polling batch status: {batch_id}') log.info('Polling batch status: %s', batch_id)
for iteration in range(max_iterations): for iteration in range(max_iterations):
try: try:
@ -393,7 +393,7 @@ class MinerULoader:
state = file_result.get('state') state = file_result.get('state')
if state == 'done': if state == 'done':
log.info(f'Processing complete for {filename}') log.info('Processing complete for %s', filename)
return file_result return file_result
elif state == 'failed': elif state == 'failed':
error_msg = file_result.get('err_msg', 'Unknown error') error_msg = file_result.get('err_msg', 'Unknown error')
@ -404,7 +404,7 @@ class MinerULoader:
elif state in ['waiting-file', 'pending', 'running', 'converting']: elif state in ['waiting-file', 'pending', 'running', 'converting']:
# Still processing # Still processing
if iteration % 10 == 0: # Log every 20 seconds if iteration % 10 == 0: # Log every 20 seconds
log.info(f'Processing status: {state} (iteration {iteration + 1}/{max_iterations})') log.info('Processing status: %s (iteration %s/%s)', state, iteration + 1, max_iterations)
time.sleep(poll_interval) time.sleep(poll_interval)
else: else:
log.warning(f'Unknown state: {state}') log.warning(f'Unknown state: {state}')
@ -421,7 +421,7 @@ class MinerULoader:
Download ZIP file from CDN and extract markdown content. Download ZIP file from CDN and extract markdown content.
Returns the markdown content as a string. Returns the markdown content as a string.
""" """
log.info(f'Downloading results from: {zip_url}') log.info('Downloading results from: %s', zip_url)
try: try:
response = requests.get(zip_url, timeout=60) response = requests.get(zip_url, timeout=60)
@ -452,7 +452,7 @@ class MinerULoader:
read_errors = [] read_errors = []
for member in md_members: for member in md_members:
log.info(f'Found markdown file in ZIP: {member.filename}') log.info('Found markdown file in ZIP: %s', member.filename)
try: try:
with zip_ref.open(member, 'r') as f: with zip_ref.open(member, 'r') as f:
if self.max_markdown_bytes is None: if self.max_markdown_bytes is None:
@ -515,5 +515,5 @@ class MinerULoader:
detail='Extracted markdown content is empty', detail='Extracted markdown content is empty',
) )
log.info(f'Successfully extracted markdown content ({len(markdown_content)} characters)') log.info('Successfully extracted markdown content (%s characters)', len(markdown_content))
return markdown_content return markdown_content

View file

@ -1,16 +1,13 @@
import asyncio
import base64 import base64
import logging import logging
import os import os
import sys import sys
import time import time
from contextlib import asynccontextmanager
from typing import Any, Dict, List, Optional from typing import Any, Dict, List, Optional
import aiohttp
import requests import requests
from langchain_core.documents import Document from langchain_core.documents import Document
from open_webui.env import AIOHTTP_CLIENT_SESSION_SSL, ENABLE_FORWARD_USER_INFO_HEADERS, GLOBAL_LOG_LEVEL from open_webui.env import ENABLE_FORWARD_USER_INFO_HEADERS, GLOBAL_LOG_LEVEL
from open_webui.utils.headers import include_user_info_headers from open_webui.utils.headers import include_user_info_headers
logging.basicConfig(stream=sys.stdout, level=GLOBAL_LOG_LEVEL) logging.basicConfig(stream=sys.stdout, level=GLOBAL_LOG_LEVEL)
@ -19,15 +16,12 @@ log = logging.getLogger(__name__)
class MistralLoader: class MistralLoader:
""" """
Enhanced Mistral OCR loader with both sync and async support. Enhanced Mistral OCR loader.
Loads documents by processing them through the Mistral OCR API. Loads documents by processing them through the Mistral OCR API.
Performance Optimizations: Performance Optimizations:
- Differentiated timeouts for different operations - Differentiated timeouts for different operations
- Intelligent retry logic with exponential backoff - Intelligent retry logic with exponential backoff
- Memory-efficient file streaming for large files
- Connection pooling and keepalive optimization
- Semaphore-based concurrency control for batch processing
- Enhanced error handling with retryable error classification - Enhanced error handling with retryable error classification
""" """
@ -63,7 +57,6 @@ class MistralLoader:
self.base_url = base_url.rstrip('/') if base_url else 'https://api.mistral.ai/v1' self.base_url = base_url.rstrip('/') if base_url else 'https://api.mistral.ai/v1'
self.api_key = api_key self.api_key = api_key
self.file_path = file_path self.file_path = file_path
self.timeout = timeout
self.max_retries = max_retries self.max_retries = max_retries
self.debug = enable_debug_logging self.debug = enable_debug_logging
self.use_base64 = use_base64 self.use_base64 = use_base64
@ -118,32 +111,6 @@ class MistralLoader:
log.error(f'JSON decode error: {json_err} - Response: {response.text}') log.error(f'JSON decode error: {json_err} - Response: {response.text}')
raise # Re-raise after logging raise # Re-raise after logging
async def _handle_response_async(self, response: aiohttp.ClientResponse) -> Dict[str, Any]:
"""Async version of response handling with better error info."""
try:
response.raise_for_status()
# Check content type
content_type = response.headers.get('content-type', '')
if 'application/json' not in content_type:
if response.status == 204:
return {}
text = await response.text()
raise ValueError(f'Unexpected content type: {content_type}, body: {text[:200]}...')
return await response.json()
except aiohttp.ClientResponseError as e:
error_text = await response.text() if response else 'No response'
log.error(f'HTTP {e.status}: {e.message} - Response: {error_text[:500]}')
raise
except aiohttp.ClientError as e:
log.error(f'Client error: {e}')
raise
except Exception as e:
log.error(f'Unexpected error processing response: {e}')
raise
def _is_retryable_error(self, error: Exception) -> bool: def _is_retryable_error(self, error: Exception) -> bool:
""" """
ENHANCEMENT: Intelligent error classification for retry logic. ENHANCEMENT: Intelligent error classification for retry logic.
@ -173,10 +140,6 @@ class MistralLoader:
status_code = error.response.status_code status_code = error.response.status_code
return status_code >= 500 or status_code == 429 return status_code >= 500 or status_code == 429
return False return False
if isinstance(error, (aiohttp.ClientConnectionError, aiohttp.ServerTimeoutError)):
return True # Async network/timeout errors are retryable
if isinstance(error, aiohttp.ClientResponseError):
return error.status >= 500 or error.status == 429
return False # All other errors are non-retryable return False # All other errors are non-retryable
def _retry_request_sync(self, request_func, *args, **kwargs): def _retry_request_sync(self, request_func, *args, **kwargs):
@ -203,32 +166,11 @@ class MistralLoader:
) )
time.sleep(wait_time) time.sleep(wait_time)
async def _retry_request_async(self, request_func, *args, **kwargs):
"""
ENHANCEMENT: Async retry logic with intelligent error classification.
Async version of retry logic that doesn't block the event loop during
wait periods. Uses the same exponential backoff strategy as sync version.
"""
for attempt in range(self.max_retries):
try:
return await request_func(*args, **kwargs)
except Exception as e:
if attempt == self.max_retries - 1 or not self._is_retryable_error(e):
raise
# PERFORMANCE OPTIMIZATION: Non-blocking exponential backoff
wait_time = min((2**attempt) + 0.5, 30) # Cap at 30 seconds
log.warning(
f'Retryable error (attempt {attempt + 1}/{self.max_retries}): {e}. Retrying in {wait_time}s...'
)
await asyncio.sleep(wait_time) # Non-blocking wait
def _upload_file(self) -> str: def _upload_file(self) -> str:
""" """
PERFORMANCE OPTIMIZATION: Enhanced file upload with streaming consideration. PERFORMANCE OPTIMIZATION: Enhanced file upload with streaming consideration.
Uploads the file to Mistral for OCR processing (sync version). Uploads the file to Mistral for OCR processing.
Uses context manager for file handling to ensure proper resource cleanup. Uses context manager for file handling to ensure proper resource cleanup.
Although streaming is not enabled for this endpoint, the file is opened Although streaming is not enabled for this endpoint, the file is opened
in a context manager to minimize memory usage duration. in a context manager to minimize memory usage duration.
@ -261,56 +203,15 @@ class MistralLoader:
file_id = response_data.get('id') file_id = response_data.get('id')
if not file_id: if not file_id:
raise ValueError('File ID not found in upload response.') raise ValueError('File ID not found in upload response.')
log.info(f'File uploaded successfully. File ID: {file_id}') log.info('File uploaded successfully. File ID: %s', file_id)
return file_id return file_id
except Exception as e: except Exception as e:
log.error(f'Failed to upload file: {e}') log.error(f'Failed to upload file: {e}')
raise raise
async def _upload_file_async(self, session: aiohttp.ClientSession) -> str:
"""Async file upload with streaming for better memory efficiency."""
url = f'{self.base_url}/files'
async def upload_request():
# Open inside the request so the handle stays valid for the whole
# streamed POST and is closed right after.
with open(self.file_path, 'rb') as f:
writer = aiohttp.MultipartWriter('form-data')
# Add purpose field
purpose_part = writer.append('ocr')
purpose_part.set_content_disposition('form-data', name='purpose')
# Stream the file. aiohttp builds a payload from the file object;
# the previous aiohttp.streams.FilePayload was removed upstream
# (payloads live in aiohttp.payload and there is no FilePayload),
# so this path raised AttributeError on every async OCR upload.
file_part = writer.append(f, {'Content-Type': 'application/pdf'})
file_part.set_content_disposition('form-data', name='file', filename=self.file_name)
self._debug_log(f'Uploading file: {self.file_name} ({self.file_size:,} bytes)')
async with session.post(
url,
data=writer,
headers=self.headers,
timeout=aiohttp.ClientTimeout(total=self.upload_timeout),
ssl=AIOHTTP_CLIENT_SESSION_SSL,
) as response:
return await self._handle_response_async(response)
response_data = await self._retry_request_async(upload_request)
file_id = response_data.get('id')
if not file_id:
raise ValueError('File ID not found in upload response.')
log.info(f'File uploaded successfully. File ID: {file_id}')
return file_id
def _get_signed_url(self, file_id: str) -> str: def _get_signed_url(self, file_id: str) -> str:
"""Retrieves a temporary signed URL for the uploaded file (sync version).""" """Retrieves a temporary signed URL for the uploaded file."""
log.info(f'Getting signed URL for file ID: {file_id}') log.info('Getting signed URL for file ID: %s', file_id)
url = f'{self.base_url}/files/{file_id}/url' url = f'{self.base_url}/files/{file_id}/url'
params = {'expiry': 1} params = {'expiry': 1}
signed_url_headers = {**self.headers, 'Accept': 'application/json'} signed_url_headers = {**self.headers, 'Accept': 'application/json'}
@ -330,35 +231,8 @@ class MistralLoader:
log.error(f'Failed to get signed URL: {e}') log.error(f'Failed to get signed URL: {e}')
raise raise
async def _get_signed_url_async(self, session: aiohttp.ClientSession, file_id: str) -> str:
"""Async signed URL retrieval."""
url = f'{self.base_url}/files/{file_id}/url'
params = {'expiry': 1}
headers = {**self.headers, 'Accept': 'application/json'}
async def url_request():
self._debug_log(f'Getting signed URL for file ID: {file_id}')
async with session.get(
url,
headers=headers,
params=params,
timeout=aiohttp.ClientTimeout(total=self.url_timeout),
ssl=AIOHTTP_CLIENT_SESSION_SSL,
) as response:
return await self._handle_response_async(response)
response_data = await self._retry_request_async(url_request)
signed_url = response_data.get('url')
if not signed_url:
raise ValueError('Signed URL not found in response.')
self._debug_log('Signed URL received successfully')
return signed_url
def _process_ocr(self, signed_url: str) -> Dict[str, Any]: def _process_ocr(self, signed_url: str) -> Dict[str, Any]:
"""Sends the signed URL to the OCR endpoint for processing (sync version).""" """Sends the signed URL to the OCR endpoint for processing."""
log.info('Processing OCR via Mistral API') log.info('Processing OCR via Mistral API')
url = f'{self.base_url}/ocr' url = f'{self.base_url}/ocr'
ocr_headers = { ocr_headers = {
@ -388,113 +262,24 @@ class MistralLoader:
log.error(f'Failed during OCR processing: {e}') log.error(f'Failed during OCR processing: {e}')
raise raise
async def _process_ocr_async(self, session: aiohttp.ClientSession, signed_url: str) -> Dict[str, Any]:
"""Async OCR processing with timing metrics."""
url = f'{self.base_url}/ocr'
headers = {
**self.headers,
'Content-Type': 'application/json',
'Accept': 'application/json',
}
payload = {
'model': 'mistral-ocr-latest',
'document': {
'type': 'document_url',
'document_url': signed_url,
},
'include_image_base64': False,
}
async def ocr_request():
log.info('Starting OCR processing via Mistral API')
start_time = time.time()
async with session.post(
url,
json=payload,
headers=headers,
timeout=aiohttp.ClientTimeout(total=self.ocr_timeout),
ssl=AIOHTTP_CLIENT_SESSION_SSL,
) as response:
ocr_response = await self._handle_response_async(response)
processing_time = time.time() - start_time
log.info(f'OCR processing completed in {processing_time:.2f}s')
return ocr_response
return await self._retry_request_async(ocr_request)
def _get_file_data_url(self) -> str: def _get_file_data_url(self) -> str:
with open(self.file_path, 'rb') as f: with open(self.file_path, 'rb') as f:
encoded_file = base64.b64encode(f.read()).decode('utf-8') encoded_file = base64.b64encode(f.read()).decode('utf-8')
return f'data:application/pdf;base64,{encoded_file}' return f'data:application/pdf;base64,{encoded_file}'
def _delete_file(self, file_id: str) -> None: def _delete_file(self, file_id: str) -> None:
"""Deletes the file from Mistral storage (sync version).""" """Deletes the file from Mistral storage."""
log.info(f'Deleting uploaded file ID: {file_id}') log.info('Deleting uploaded file ID: %s', file_id)
url = f'{self.base_url}/files/{file_id}' url = f'{self.base_url}/files/{file_id}'
try: try:
response = requests.delete(url, headers=self.headers, timeout=self.cleanup_timeout) response = requests.delete(url, headers=self.headers, timeout=self.cleanup_timeout)
delete_response = self._handle_response(response) delete_response = self._handle_response(response)
log.info(f'File deleted successfully: {delete_response}') log.info('File deleted successfully: %s', delete_response)
except Exception as e: except Exception as e:
# Log error but don't necessarily halt execution if deletion fails # Log error but don't necessarily halt execution if deletion fails
log.error(f'Failed to delete file ID {file_id}: {e}') log.error(f'Failed to delete file ID {file_id}: {e}')
async def _delete_file_async(self, session: aiohttp.ClientSession, file_id: str) -> None:
"""Async file deletion with error tolerance."""
try:
async def delete_request():
self._debug_log(f'Deleting file ID: {file_id}')
async with session.delete(
url=f'{self.base_url}/files/{file_id}',
headers=self.headers,
timeout=aiohttp.ClientTimeout(total=self.cleanup_timeout),
ssl=AIOHTTP_CLIENT_SESSION_SSL,
) as response:
return await self._handle_response_async(response)
await self._retry_request_async(delete_request)
self._debug_log(f'File {file_id} deleted successfully')
except Exception as e:
# Don't fail the entire process if cleanup fails
log.warning(f'Failed to delete file ID {file_id}: {e}')
@asynccontextmanager
async def _get_session(self):
"""Context manager for HTTP session with optimized settings."""
connector = aiohttp.TCPConnector(
limit=20, # Increased total connection limit for better throughput
limit_per_host=10, # Increased per-host limit for API endpoints
ttl_dns_cache=600, # Longer DNS cache TTL (10 minutes)
use_dns_cache=True,
keepalive_timeout=60, # Increased keepalive for connection reuse
enable_cleanup_closed=True,
force_close=False, # Allow connection reuse
resolver=aiohttp.AsyncResolver(), # Use async DNS resolver
)
timeout = aiohttp.ClientTimeout(
total=self.timeout,
connect=30, # Connection timeout
sock_read=60, # Socket read timeout
)
async with aiohttp.ClientSession(
connector=connector,
timeout=timeout,
headers={'User-Agent': 'OpenWebUI-MistralLoader/2.0'},
raise_for_status=False, # We handle status codes manually
trust_env=True,
) as session:
yield session
def _process_results(self, ocr_response: Dict[str, Any]) -> List[Document]: def _process_results(self, ocr_response: Dict[str, Any]) -> List[Document]:
"""Process OCR results into Document objects with enhanced metadata and memory efficiency.""" """Process OCR results into Document objects with enhanced metadata and memory efficiency."""
pages_data = ocr_response.get('pages') pages_data = ocr_response.get('pages')
@ -519,7 +304,7 @@ class MistralLoader:
if page_content is None or page_index is None: if page_content is None or page_index is None:
skipped_pages += 1 skipped_pages += 1
self._debug_log( self._debug_log(
f"Skipping page due to missing 'markdown' or 'index'. Data keys: {list(page_data.keys())}" "Skipping page due to missing 'markdown' or 'index'. Data keys: %s", list(page_data.keys())
) )
continue continue
@ -531,7 +316,7 @@ class MistralLoader:
if not cleaned_content: if not cleaned_content:
skipped_pages += 1 skipped_pages += 1
self._debug_log(f'Skipping empty page {page_index}') self._debug_log('Skipping empty page %s', page_index)
continue continue
# Create document with optimized metadata # Create document with optimized metadata
@ -551,7 +336,7 @@ class MistralLoader:
) )
if skipped_pages > 0: if skipped_pages > 0:
log.info(f'Processed {len(documents)} pages, skipped {skipped_pages} empty/invalid pages') log.info('Processed %s pages, skipped %s empty/invalid pages', len(documents), skipped_pages)
if not documents: if not documents:
# Case where pages existed but none had valid markdown/index # Case where pages existed but none had valid markdown/index
@ -572,7 +357,6 @@ class MistralLoader:
def load(self) -> List[Document]: def load(self) -> List[Document]:
""" """
Executes the full OCR workflow: upload, get URL, process OCR, delete file. Executes the full OCR workflow: upload, get URL, process OCR, delete file.
Synchronous version for backward compatibility.
Returns: Returns:
A list of Document objects, one for each page processed. A list of Document objects, one for each page processed.
@ -584,7 +368,7 @@ class MistralLoader:
if self.use_base64: if self.use_base64:
documents = self._process_results(self._process_ocr(self._get_file_data_url())) documents = self._process_results(self._process_ocr(self._get_file_data_url()))
total_time = time.time() - start_time total_time = time.time() - start_time
log.info(f'Sync OCR workflow completed in {total_time:.2f}s, produced {len(documents)} documents') log.info('Sync OCR workflow completed in %.2fs, produced %s documents', total_time, len(documents))
return documents return documents
# 1. Upload file # 1. Upload file
@ -600,7 +384,7 @@ class MistralLoader:
documents = self._process_results(ocr_response) documents = self._process_results(ocr_response)
total_time = time.time() - start_time total_time = time.time() - start_time
log.info(f'Sync OCR workflow completed in {total_time:.2f}s, produced {len(documents)} documents') log.info('Sync OCR workflow completed in %.2fs, produced %s documents', total_time, len(documents))
return documents return documents
@ -625,125 +409,3 @@ class MistralLoader:
except Exception as del_e: except Exception as del_e:
# Log deletion error, but don't overwrite original error if one occurred # Log deletion error, but don't overwrite original error if one occurred
log.error(f'Cleanup error: Could not delete file ID {file_id}. Reason: {del_e}') log.error(f'Cleanup error: Could not delete file ID {file_id}. Reason: {del_e}')
async def load_async(self) -> List[Document]:
"""
Asynchronous OCR workflow execution with optimized performance.
Returns:
A list of Document objects, one for each page processed.
"""
file_id = None
start_time = time.time()
try:
async with self._get_session() as session:
if self.use_base64:
ocr_response = await self._process_ocr_async(session, self._get_file_data_url())
documents = self._process_results(ocr_response)
total_time = time.time() - start_time
log.info(f'Async OCR workflow completed in {total_time:.2f}s, produced {len(documents)} documents')
return documents
# 1. Upload file with streaming
file_id = await self._upload_file_async(session)
# 2. Get signed URL
signed_url = await self._get_signed_url_async(session, file_id)
# 3. Process OCR
ocr_response = await self._process_ocr_async(session, signed_url)
# 4. Process results
documents = self._process_results(ocr_response)
total_time = time.time() - start_time
log.info(f'Async OCR workflow completed in {total_time:.2f}s, produced {len(documents)} documents')
return documents
except Exception as e:
total_time = time.time() - start_time
log.error(f'Async OCR workflow failed after {total_time:.2f}s: {e}')
return [
Document(
page_content=f'Error during OCR processing: {e}',
metadata={
'error': 'processing_failed',
'file_name': self.file_name,
},
)
]
finally:
# 5. Cleanup - always attempt file deletion
if file_id:
try:
async with self._get_session() as session:
await self._delete_file_async(session, file_id)
except Exception as cleanup_error:
log.error(f'Cleanup failed for file ID {file_id}: {cleanup_error}')
@staticmethod
async def load_multiple_async(
loaders: List['MistralLoader'],
max_concurrent: int = 5, # Limit concurrent requests
) -> List[List[Document]]:
"""
Process multiple files concurrently with controlled concurrency.
Args:
loaders: List of MistralLoader instances
max_concurrent: Maximum number of concurrent requests
Returns:
List of document lists, one for each loader
"""
if not loaders:
return []
log.info(f'Starting concurrent processing of {len(loaders)} files with max {max_concurrent} concurrent')
start_time = time.time()
# Use semaphore to control concurrency
semaphore = asyncio.Semaphore(max_concurrent)
async def process_with_semaphore(loader: 'MistralLoader') -> List[Document]:
async with semaphore:
return await loader.load_async()
# Process all files with controlled concurrency
tasks = [process_with_semaphore(loader) for loader in loaders]
results = await asyncio.gather(*tasks, return_exceptions=True)
# Handle any exceptions in results
processed_results = []
for i, result in enumerate(results):
if isinstance(result, Exception):
log.error(f'File {i} failed: {result}')
processed_results.append(
[
Document(
page_content=f'Error processing file: {result}',
metadata={
'error': 'batch_processing_failed',
'file_index': i,
},
)
]
)
else:
processed_results.append(result)
# MONITORING: Log comprehensive batch processing statistics
total_time = time.time() - start_time
total_docs = sum(len(docs) for docs in processed_results)
success_count = sum(1 for result in results if not isinstance(result, Exception))
failure_count = len(results) - success_count
log.info(
f'Batch processing completed in {total_time:.2f}s: '
f'{success_count} files succeeded, {failure_count} files failed, '
f'produced {total_docs} total documents'
)
return processed_results

View file

@ -35,7 +35,7 @@ class PaddleOCRVLLoader:
self.file_name = os.path.basename(file_path) self.file_name = os.path.basename(file_path)
def load(self) -> List[Document]: def load(self) -> List[Document]:
log.info(f'Processing with PaddleOCR-vl: {self.file_path}') log.info('Processing with PaddleOCR-vl: %s', self.file_path)
try: try:
with open(self.file_path, 'rb') as file: with open(self.file_path, 'rb') as file:
@ -96,7 +96,7 @@ class PaddleOCRVLLoader:
) )
if skipped_pages > 0: if skipped_pages > 0:
log.info(f'PaddleOCR-vl: Processed {len(documents)} pages, skipped {skipped_pages} empty pages.') log.info('PaddleOCR-vl: Processed %s pages, skipped %s empty pages.', len(documents), skipped_pages)
if not documents: if not documents:
log.warning('No valid text content found by PaddleOCR-vl.') log.warning('No valid text content found by PaddleOCR-vl.')

View file

@ -0,0 +1,101 @@
import datetime as dt
import io
import logging
from pathlib import Path
from langchain_core.document_loaders import BaseLoader
from langchain_core.documents import Document
log = logging.getLogger(__name__)
class PDFLoader(BaseLoader):
def __init__(self, file_path, *, extract_images=False, mode='page'):
if mode not in ('single', 'page'):
raise ValueError("PDF mode must be 'single' or 'page'")
self.file_path = str(Path(file_path).expanduser())
self.extract_images = extract_images
self.mode = mode
self.ocr = None
def lazy_load(self):
from pypdf import PdfReader
with open(self.file_path, 'rb') as file:
reader = PdfReader(file)
metadata = {'producer': 'PyPDF', 'creator': 'PyPDF', 'creationdate': ''}
for key, value in (reader.metadata or {}).items():
key = key.removeprefix('/').lower()
value = value if type(value) in (str, int) else str(value)
if key in ('creationdate', 'moddate') and isinstance(value, str):
try:
value = dt.datetime.strptime(value.replace("'", ''), 'D:%Y%m%d%H%M%S%z').isoformat()
except ValueError:
pass
metadata[key] = (
value.strip()
if isinstance(value, str) and key not in ('creationdate', 'moddate', 'page_count', 'file_path')
else value
)
metadata.update(source=self.file_path, total_pages=len(reader.pages))
labels = reader.page_labels if self.mode == 'page' else None
texts = []
for index, page in enumerate(reader.pages):
text = page.extract_text()
if self.extract_images:
image_text = self._extract_images(page)
if image_text:
text = self._merge_image_text(text, image_text)
text = text.strip()
if self.mode == 'page':
yield Document(page_content=text, metadata={**metadata, 'page': index, 'page_label': labels[index]})
else:
texts.append(text)
if self.mode == 'single':
yield Document(page_content='\n\f'.join(texts), metadata=metadata)
@staticmethod
def _merge_image_text(text, image_text):
# Insert before the final paragraphs/footer where possible, matching existing chunks.
position, separator = len(text), '\n\n'
for _ in range(2):
for delimiter in ('\n\n\n', '\n\n'):
found = text.rfind(delimiter, 0, position)
if found >= 0:
position, separator = found, delimiter
break
else:
break
return text[:position] + separator + image_text + text[position:]
def _extract_images(self, page):
import numpy as np
from PIL import Image, UnidentifiedImageError
if '/Resources' not in page or '/XObject' not in page['/Resources']:
return ''
texts = []
xobjects = page['/Resources']['/XObject']
for name in xobjects:
try:
stream = xobjects[name]
if stream.get('/Subtype') != '/Image':
continue
try:
# Encoded images, including CMYK JPEGs, can go straight to Pillow.
image = Image.open(io.BytesIO(stream.get_data()))
except UnidentifiedImageError:
image = stream.decode_as_image()
pixels = np.array(image.convert('RGB'))
except Exception as e:
log.warning('Skipping unreadable PDF image %s: %s', name, e)
continue
if self.ocr is None:
from rapidocr import RapidOCR
self.ocr = RapidOCR()
result = self.ocr(pixels)
if result and result.txts:
texts.append('\n'.join(result.txts).strip())
return '\n\n' + '\n'.join(filter(None, texts)) + '\n\n' if any(texts) else ''

View file

@ -4,6 +4,7 @@ from typing import Iterator, List, Literal, Union
import requests import requests
from langchain_core.document_loaders import BaseLoader from langchain_core.document_loaders import BaseLoader
from langchain_core.documents import Document from langchain_core.documents import Document
from open_webui.env import TAVILY_API_BASE_URL
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@ -48,7 +49,7 @@ class TavilyLoader(BaseLoader):
self.urls = urls if isinstance(urls, list) else [urls] self.urls = urls if isinstance(urls, list) else [urls]
self.extract_depth = extract_depth self.extract_depth = extract_depth
self.continue_on_failure = continue_on_failure self.continue_on_failure = continue_on_failure
self.api_url = 'https://api.tavily.com/extract' self.api_url = f'{TAVILY_API_BASE_URL}/extract'
def lazy_load(self) -> Iterator[Document]: def lazy_load(self) -> Iterator[Document]:
"""Extract and yield documents from the URLs using Tavily Extract API.""" """Extract and yield documents from the URLs using Tavily Extract API."""

View file

@ -18,6 +18,32 @@ ALLOWED_NETLOCS = {
} }
class YoutubeTranscriptError(Exception):
"""A YouTube transcript could not be retrieved."""
def _transcript_error_message(error: Exception, video_id: str) -> str:
name = type(error).__name__
if name in {'RequestBlocked', 'IpBlocked'}:
return (
f'YouTube blocked the transcript request for {video_id} from this server. '
'This usually means the server address is rate limited or belongs to a cloud '
'provider. A proxy for these requests can be configured under Admin Settings, '
'Web Search, Youtube Proxy URL.'
)
if name == 'TranscriptsDisabled':
return f'Transcripts are disabled for the YouTube video {video_id}.'
if name == 'AgeRestricted':
return f'The YouTube video {video_id} is age restricted, so its transcript cannot be retrieved.'
if name in {'VideoUnavailable', 'VideoUnplayable', 'InvalidVideoId'}:
return f'The YouTube video {video_id} is unavailable.'
if name == 'PoTokenRequired':
return f'YouTube requires additional verification to return the transcript for {video_id}.'
return f'Could not retrieve a transcript for the YouTube video {video_id}.'
def _parse_video_id(url: str) -> Optional[str]: def _parse_video_id(url: str) -> Optional[str]:
"""Parse a YouTube URL and return the video ID if valid, otherwise None.""" """Parse a YouTube URL and return the video ID if valid, otherwise None."""
parsed_url = urlparse(url) parsed_url = urlparse(url)
@ -90,7 +116,7 @@ class YoutubeLoader:
if self.proxy_url: if self.proxy_url:
youtube_proxies = GenericProxyConfig(http_url=self.proxy_url, https_url=self.proxy_url) youtube_proxies = GenericProxyConfig(http_url=self.proxy_url, https_url=self.proxy_url)
log.debug(f'Using proxy URL: {self.proxy_url[:14]}...') log.debug('Using proxy URL: %s...', self.proxy_url[:14])
else: else:
youtube_proxies = None youtube_proxies = None
@ -98,31 +124,31 @@ class YoutubeLoader:
try: try:
transcript_list = transcript_api.list(self.video_id) transcript_list = transcript_api.list(self.video_id)
except Exception as e: except Exception as e:
log.warning(f'Loading YouTube transcript failed: {e}') log.warning('Loading YouTube transcript failed: %s', e)
return [] raise YoutubeTranscriptError(_transcript_error_message(e, self.video_id)) from e
# Try each language in order of priority # Try each language in order of priority
for lang in self.language: for lang in self.language:
try: try:
transcript = transcript_list.find_transcript([lang]) transcript = transcript_list.find_transcript([lang])
if transcript.is_generated: if transcript.is_generated:
log.debug(f"Found generated transcript for language '{lang}'") log.debug("Found generated transcript for language '%s'", lang)
try: try:
transcript = transcript_list.find_manually_created_transcript([lang]) transcript = transcript_list.find_manually_created_transcript([lang])
log.debug(f"Found manual transcript for language '{lang}'") log.debug("Found manual transcript for language '%s'", lang)
except NoTranscriptFound: except NoTranscriptFound:
log.debug(f"No manual transcript found for language '{lang}', using generated") log.debug("No manual transcript found for language '%s', using generated", lang)
pass pass
log.debug(f"Found transcript for language '{lang}'") log.debug("Found transcript for language '%s'", lang)
try: try:
transcript_pieces: List[Dict[str, Any]] = transcript.fetch() transcript_pieces: List[Dict[str, Any]] = transcript.fetch()
except ParseError: except ParseError:
log.debug(f"Empty or invalid transcript for language '{lang}'") log.debug("Empty or invalid transcript for language '%s'", lang)
continue continue
if not transcript_pieces: if not transcript_pieces:
log.debug(f"Empty transcript for language '{lang}'") log.debug("Empty transcript for language '%s'", lang)
continue continue
transcript_text = ' '.join( transcript_text = ' '.join(
@ -135,18 +161,20 @@ class YoutubeLoader:
) )
return [Document(page_content=transcript_text, metadata=self._metadata)] return [Document(page_content=transcript_text, metadata=self._metadata)]
except NoTranscriptFound: except NoTranscriptFound:
log.debug(f"No transcript found for language '{lang}'") log.debug("No transcript found for language '%s'", lang)
continue continue
except Exception as e: except Exception as e:
log.info(f"Error finding transcript for language '{lang}'") log.info("Error finding transcript for language '%s'", lang)
raise e raise YoutubeTranscriptError(_transcript_error_message(e, self.video_id)) from e
# If we get here, all languages failed # If we get here, all languages failed
languages_tried = ', '.join(self.language) languages_tried = ', '.join(self.language)
log.warning( log.warning(
f'No transcript found for any of the specified languages: {languages_tried}. Verify if the video has transcripts, add more languages if needed.' f'No transcript found for any of the specified languages: {languages_tried}. Verify if the video has transcripts, add more languages if needed.'
) )
raise NoTranscriptFound(self.video_id, self.language, list(transcript_list)) raise YoutubeTranscriptError(
f'No transcript found for the YouTube video {self.video_id} in these languages: {languages_tried}.'
)
async def aload(self) -> Generator[Document, None, None]: async def aload(self) -> Generator[Document, None, None]:
"""Asynchronously load YouTube transcripts into `Document` objects.""" """Asynchronously load YouTube transcripts into `Document` objects."""

View file

@ -12,7 +12,7 @@ log = logging.getLogger(__name__)
class ColBERT(BaseReranker): class ColBERT(BaseReranker):
def __init__(self, name, **kwargs) -> None: def __init__(self, name, **kwargs) -> None:
log.info('ColBERT: Loading model', name) log.info('ColBERT: Loading model %s', name)
self.device = 'cuda' if torch.cuda.is_available() else 'cpu' self.device = 'cuda' if torch.cuda.is_available() else 'cpu'
DOCKER = kwargs.get('env') == 'docker' DOCKER = kwargs.get('env') == 'docker'

View file

@ -35,8 +35,8 @@ class ExternalReranker(BaseReranker):
} }
try: try:
log.info(f'ExternalReranker:predict:model {self.model}') log.info('ExternalReranker:predict:model %s', self.model)
log.info(f'ExternalReranker:predict:query {query}') log.info('ExternalReranker:predict:query %s', query)
headers = { headers = {
'Content-Type': 'application/json', 'Content-Type': 'application/json',

View file

@ -6,18 +6,17 @@ import logging
import os import os
import re import re
import time import time
from concurrent.futures import ThreadPoolExecutor
from typing import Awaitable, Optional, Union from typing import Awaitable, Optional, Union
from urllib.parse import quote from urllib.parse import quote
import aiohttp import aiohttp
import numpy as np
import requests import requests
from huggingface_hub import snapshot_download from fastapi import HTTPException
from langchain_classic.retrievers import ( from langchain_classic.retrievers import (
ContextualCompressionRetriever, ContextualCompressionRetriever,
EnsembleRetriever, EnsembleRetriever,
) )
from langchain_community.retrievers import BM25Retriever
from langchain_core.documents import Document from langchain_core.documents import Document
from open_webui.config import ( from open_webui.config import (
RAG_EMBEDDING_CONTENT_PREFIX, RAG_EMBEDDING_CONTENT_PREFIX,
@ -25,6 +24,7 @@ from open_webui.config import (
RAG_EMBEDDING_QUERY_PREFIX, RAG_EMBEDDING_QUERY_PREFIX,
VECTOR_DB, VECTOR_DB,
) )
from open_webui.constants import ERROR_MESSAGES
from open_webui.env import ( from open_webui.env import (
AIOHTTP_CLIENT_ALLOW_REDIRECTS, AIOHTTP_CLIENT_ALLOW_REDIRECTS,
AIOHTTP_CLIENT_SESSION_SSL, AIOHTTP_CLIENT_SESSION_SSL,
@ -32,7 +32,10 @@ from open_webui.env import (
BYPASS_RETRIEVAL_ACCESS_CONTROL, BYPASS_RETRIEVAL_ACCESS_CONTROL,
ENABLE_FORWARD_USER_INFO_HEADERS, ENABLE_FORWARD_USER_INFO_HEADERS,
ENABLE_RETRIEVAL_UNSCOPED_COLLECTIONS, ENABLE_RETRIEVAL_UNSCOPED_COLLECTIONS,
MPS_INFERENCE_LOCK,
OFFLINE_MODE, OFFLINE_MODE,
RAG_SOURCE_METADATA_KEYS,
USE_SLIM,
) )
from open_webui.models.access_grants import AccessGrants from open_webui.models.access_grants import AccessGrants
from open_webui.models.chats import Chats from open_webui.models.chats import Chats
@ -45,12 +48,12 @@ from open_webui.models.users import UserModel
from open_webui.retrieval.loaders.youtube import YoutubeLoader from open_webui.retrieval.loaders.youtube import YoutubeLoader
from open_webui.retrieval.vector.async_client import ASYNC_VECTOR_DB_CLIENT from open_webui.retrieval.vector.async_client import ASYNC_VECTOR_DB_CLIENT
from open_webui.retrieval.external import retrieve_external_knowledge from open_webui.retrieval.external import retrieve_external_knowledge
from open_webui.retrieval.vector.factory import VECTOR_DB_CLIENT from open_webui.retrieval.vector.factory import get_vector_db_client
from open_webui.retrieval.vector.main import GetResult, SearchResult from open_webui.retrieval.vector.main import GetResult, SearchResult
from open_webui.retrieval.web.utils import get_web_loader from open_webui.retrieval.web.utils import get_web_loader
from open_webui.utils.access_control.files import get_owner_accessible_folder_files, has_access_to_file from open_webui.utils.access_control.files import get_owner_accessible_folder_files, has_access_to_file
from open_webui.utils.access_control.folders import has_folder_access from open_webui.utils.access_control.folders import has_folder_access
from open_webui.utils.headers import include_user_info_headers from open_webui.utils.headers import get_json_bearer_headers, include_user_info_headers
from open_webui.utils.misc import get_content_from_message, get_message_list from open_webui.utils.misc import get_content_from_message, get_message_list
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@ -62,12 +65,22 @@ from langchain_core.callbacks import CallbackManagerForRetrieverRun
from langchain_core.retrievers import BaseRetriever from langchain_core.retrievers import BaseRetriever
class BM25Retriever(BaseRetriever):
docs: list[Document]
vectorizer: Any
k: int
def _get_relevant_documents(self, query: str, *, run_manager: CallbackManagerForRetrieverRun) -> list[Document]:
return self.vectorizer.get_top_n(query.split(), self.docs, n=self.k)
def is_youtube_url(url: str) -> bool: def is_youtube_url(url: str) -> bool:
youtube_regex = r'^(https?://)?(www\.)?(youtube\.com|youtu\.be)/.+$' youtube_regex = r'^(https?://)?(www\.)?(youtube\.com|youtu\.be)/.+$'
return re.match(youtube_regex, url) is not None return re.match(youtube_regex, url) is not None
LOADER_CONFIG_KEYS = { LOADER_CONFIG_KEYS = {
'file_max_size': 'rag.file.max_size',
'youtube_language': 'rag.youtube_loader_language', 'youtube_language': 'rag.youtube_loader_language',
'youtube_proxy_url': 'rag.youtube_loader_proxy_url', 'youtube_proxy_url': 'rag.youtube_loader_proxy_url',
'web_loader_ssl_verification': 'web.loader.ssl_verification', 'web_loader_ssl_verification': 'web.loader.ssl_verification',
@ -103,6 +116,7 @@ LOADER_CONFIG_KEYS = {
'EXTERNAL_DOCUMENT_LOADER_API_KEY': 'rag.external_document_loader_api_key', 'EXTERNAL_DOCUMENT_LOADER_API_KEY': 'rag.external_document_loader_api_key',
'EXTERNAL_DOCUMENT_LOADER_HEADERS': 'rag.external_document_loader_headers', 'EXTERNAL_DOCUMENT_LOADER_HEADERS': 'rag.external_document_loader_headers',
'TIKA_SERVER_URL': 'rag.tika_server_url', 'TIKA_SERVER_URL': 'rag.tika_server_url',
'TIKA_SERVER_VERSION': 'rag.tika_server_version',
'DOCLING_SERVER_URL': 'rag.docling_server_url', 'DOCLING_SERVER_URL': 'rag.docling_server_url',
'DOCLING_API_KEY': 'rag.docling_api_key', 'DOCLING_API_KEY': 'rag.docling_api_key',
'DOCLING_PARAMS': 'rag.docling_params', 'DOCLING_PARAMS': 'rag.docling_params',
@ -151,6 +165,7 @@ def build_loader_from_config(request, config: dict):
from open_webui.retrieval.loaders.main import Loader from open_webui.retrieval.loaders.main import Loader
loader_config = {key: config.get(key) for key in LOADER_CONFIG_KEYS if key.isupper()} loader_config = {key: config.get(key) for key in LOADER_CONFIG_KEYS if key.isupper()}
loader_config['FILE_MAX_SIZE'] = config.get('file_max_size')
return Loader( return Loader(
engine=loader_config['CONTENT_EXTRACTION_ENGINE'], engine=loader_config['CONTENT_EXTRACTION_ENGINE'],
**{key: value for key, value in loader_config.items() if key != 'CONTENT_EXTRACTION_ENGINE'}, **{key: value for key, value in loader_config.items() if key != 'CONTENT_EXTRACTION_ENGINE'},
@ -183,11 +198,20 @@ def _extract_text_from_binary_response(
suffix = '.' + filename.split('.')[-1].lower() if '.' in filename else '' suffix = '.' + filename.split('.')[-1].lower() if '.' in filename else ''
with tempfile.NamedTemporaryFile(suffix=suffix, delete=False) as tmp: max_size = loader_config.get('file_max_size')
tmp.write(response.content) max_bytes = int(max_size) * 1024 * 1024 if max_size else 0
tmp_path = tmp.name
tmp_fd, tmp_path = tempfile.mkstemp(suffix=suffix)
try: try:
downloaded = 0
# Stream to disk; response.content buffers the whole body in memory first.
with os.fdopen(tmp_fd, 'wb') as tmp:
for chunk in response.iter_content(64 * 1024):
downloaded += len(chunk)
if max_bytes and downloaded > max_bytes:
raise ValueError(ERROR_MESSAGES.FILE_TOO_LARGE(size=f'{max_size} MB'))
tmp.write(chunk)
loader = build_loader_from_config(request, loader_config) loader = build_loader_from_config(request, loader_config)
docs = loader.load(filename, content_type, tmp_path) docs = loader.load(filename, content_type, tmp_path)
for doc in docs: for doc in docs:
@ -198,14 +222,24 @@ def _extract_text_from_binary_response(
os.remove(tmp_path) os.remove(tmp_path)
TEXT_APPLICATION_CONTENT_TYPES = {
'application/javascript',
'application/json',
'application/xml',
'application/x-javascript',
}
def _is_text_content_type(content_type: str) -> bool: def _is_text_content_type(content_type: str) -> bool:
"""Return True if the content type should be handled by the web loader.""" """Return True if the content type should be handled by the web loader."""
ct = content_type.split(';')[0].strip().lower() ct = content_type.split(';')[0].strip().lower()
if not ct:
return True
if ct.startswith('text/'): if ct.startswith('text/'):
return True return True
if any(t in ct for t in ['xml', 'json', 'javascript']): if ct in TEXT_APPLICATION_CONTENT_TYPES:
return True return True
return not ct # empty / missing → assume HTML return ct.endswith(('+xml', '+json'))
async def get_content_from_url(request, url: str) -> str: async def get_content_from_url(request, url: str) -> str:
@ -218,7 +252,7 @@ async def get_content_from_url(request, url: str) -> str:
def _get_content_from_url_sync(request, url: str, loader_config): def _get_content_from_url_sync(request, url: str, loader_config):
from open_webui.retrieval.web.utils import validate_url, _SSRFSafeAdapter from open_webui.retrieval.web.utils import validate_url, get_ssrf_safe_requests_session
# Validate URL before making any request (blocks private IPs, non-HTTP, filter list) # Validate URL before making any request (blocks private IPs, non-HTTP, filter list)
validate_url(url) validate_url(url)
@ -242,9 +276,7 @@ def _get_content_from_url_sync(request, url: str, loader_config):
# cloud-metadata 169.254.169.254) via a public host that redirects internally. # cloud-metadata 169.254.169.254) via a public host that redirects internally.
try: try:
# Probe through the connect-time SSRF guard; bare requests.get re-resolves (DNS-rebinding gap). # Probe through the connect-time SSRF guard; bare requests.get re-resolves (DNS-rebinding gap).
session = requests.Session() session = get_ssrf_safe_requests_session()
session.mount('http://', _SSRFSafeAdapter())
session.mount('https://', _SSRFSafeAdapter())
response = session.get(url, stream=True, timeout=30, allow_redirects=AIOHTTP_CLIENT_ALLOW_REDIRECTS) response = session.get(url, stream=True, timeout=30, allow_redirects=AIOHTTP_CLIENT_ALLOW_REDIRECTS)
response.raise_for_status() response.raise_for_status()
content_type = response.headers.get('Content-Type', '') content_type = response.headers.get('Content-Type', '')
@ -310,30 +342,26 @@ class VectorSearchRetriever(BaseRetriever):
def query_doc(collection_name: str, query_embedding: list[float], k: int, user: UserModel = None): def query_doc(collection_name: str, query_embedding: list[float], k: int, user: UserModel = None):
try: log.debug('query_doc:doc %s', collection_name)
log.debug(f'query_doc:doc {collection_name}') result = get_vector_db_client().search(
result = VECTOR_DB_CLIENT.search( collection_name=collection_name,
collection_name=collection_name, vectors=[query_embedding],
vectors=[query_embedding], limit=k,
limit=k, )
)
if result: if result:
log.info(f'query_doc:result {result.ids} {result.metadatas}') log.info('query_doc:result %s %s', result.ids, result.metadatas)
return result return result
except Exception as e:
log.exception(f'Error querying doc {collection_name} with limit {k}: {e}')
raise e
def get_doc(collection_name: str, user: UserModel = None): def get_doc(collection_name: str, user: UserModel = None):
try: try:
log.debug(f'get_doc:doc {collection_name}') log.debug('get_doc:doc %s', collection_name)
result = VECTOR_DB_CLIENT.get(collection_name=collection_name) result = get_vector_db_client().get(collection_name=collection_name)
if result: if result:
log.info(f'query_doc:result {result.ids} {result.metadatas}') log.info('query_doc:result %s %s', result.ids, result.metadatas)
return result return result
except Exception as e: except Exception as e:
@ -458,7 +486,7 @@ async def query_doc_with_native_hybrid_search(
'metadatas': [metadatas], 'metadatas': [metadatas],
} }
except Exception as e: except Exception as e:
log.debug(f'Native hybrid search failed for {collection_name}, falling back to legacy hybrid search: {e}') log.debug('Native hybrid search failed for %s, falling back to legacy hybrid search: %s', collection_name, e)
return None return None
@ -475,122 +503,116 @@ async def query_doc_with_hybrid_search(
enable_enriched_texts: bool = False, enable_enriched_texts: bool = False,
native_hybrid_search: bool = True, native_hybrid_search: bool = True,
) -> dict: ) -> dict:
try: if native_hybrid_search and not enable_enriched_texts:
if native_hybrid_search and not enable_enriched_texts: native_result = await query_doc_with_native_hybrid_search(
native_result = await query_doc_with_native_hybrid_search(
collection_name=collection_name,
query=query,
embedding_function=embedding_function,
k=k,
reranking_function=reranking_function,
k_reranker=k_reranker,
r=r,
hybrid_bm25_weight=hybrid_bm25_weight,
)
if native_result is not None:
return native_result
if collection_result is None:
collection_result = await ASYNC_VECTOR_DB_CLIENT.get(collection_name=collection_name)
# First check if collection_result has the required attributes
if (
not collection_result
or not hasattr(collection_result, 'documents')
or not hasattr(collection_result, 'metadatas')
):
log.warning(f'query_doc_with_hybrid_search:no_docs {collection_name}')
return {'documents': [], 'metadatas': [], 'distances': []}
# Now safely check the documents content after confirming attributes exist
if (
not collection_result.documents
or len(collection_result.documents) == 0
or not collection_result.documents[0]
):
log.warning(f'query_doc_with_hybrid_search:no_docs {collection_name}')
return {'documents': [], 'metadatas': [], 'distances': []}
log.debug(f'query_doc_with_hybrid_search:doc {collection_name}')
original_texts = collection_result.documents[0]
bm25_metadatas = [
{**meta, CHUNK_HASH_KEY: _content_hash(original_texts[idx])}
for idx, meta in enumerate(collection_result.metadatas[0])
]
bm25_texts = get_enriched_texts(collection_result) if enable_enriched_texts else original_texts
bm25_retriever = BM25Retriever.from_texts(
texts=bm25_texts,
metadatas=bm25_metadatas,
)
bm25_retriever.k = k
vector_search_retriever = VectorSearchRetriever(
collection_name=collection_name, collection_name=collection_name,
query=query,
embedding_function=embedding_function, embedding_function=embedding_function,
top_k=k, k=k,
)
# Use CHUNK_HASH_KEY for dedup so enriched BM25 texts don't defeat RRF
if hybrid_bm25_weight <= 0:
ensemble_retriever = EnsembleRetriever(
retrievers=[vector_search_retriever],
weights=[1.0],
id_key=CHUNK_HASH_KEY,
)
elif hybrid_bm25_weight >= 1:
ensemble_retriever = EnsembleRetriever(
retrievers=[bm25_retriever],
weights=[1.0],
id_key=CHUNK_HASH_KEY,
)
else:
ensemble_retriever = EnsembleRetriever(
retrievers=[bm25_retriever, vector_search_retriever],
weights=[hybrid_bm25_weight, 1.0 - hybrid_bm25_weight],
id_key=CHUNK_HASH_KEY,
)
compressor = RerankCompressor(
embedding_function=embedding_function,
top_n=k_reranker,
reranking_function=reranking_function, reranking_function=reranking_function,
r_score=r, k_reranker=k_reranker,
r=r,
hybrid_bm25_weight=hybrid_bm25_weight,
)
if native_result is not None:
return native_result
if collection_result is None:
collection_result = await ASYNC_VECTOR_DB_CLIENT.get(collection_name=collection_name)
# First check if collection_result has the required attributes
if (
not collection_result
or not hasattr(collection_result, 'documents')
or not hasattr(collection_result, 'metadatas')
):
log.warning(f'query_doc_with_hybrid_search:no_docs {collection_name}')
return {'documents': [], 'metadatas': [], 'distances': []}
# Now safely check the documents content after confirming attributes exist
if not collection_result.documents or len(collection_result.documents) == 0 or not collection_result.documents[0]:
log.warning(f'query_doc_with_hybrid_search:no_docs {collection_name}')
return {'documents': [], 'metadatas': [], 'distances': []}
log.debug('query_doc_with_hybrid_search:doc %s', collection_name)
original_texts = collection_result.documents[0]
bm25_metadatas = [
{**meta, CHUNK_HASH_KEY: _content_hash(original_texts[idx])}
for idx, meta in enumerate(collection_result.metadatas[0])
]
bm25_texts = get_enriched_texts(collection_result) if enable_enriched_texts else original_texts
from rank_bm25 import BM25Okapi
bm25_retriever = BM25Retriever(
docs=[Document(page_content=text, metadata=meta) for text, meta in zip(bm25_texts, bm25_metadatas)],
vectorizer=BM25Okapi([text.split() for text in bm25_texts]),
k=k,
)
vector_search_retriever = VectorSearchRetriever(
collection_name=collection_name,
embedding_function=embedding_function,
top_k=k,
)
# Use CHUNK_HASH_KEY for dedup so enriched BM25 texts don't defeat RRF
if hybrid_bm25_weight <= 0:
ensemble_retriever = EnsembleRetriever(
retrievers=[vector_search_retriever],
weights=[1.0],
id_key=CHUNK_HASH_KEY,
)
elif hybrid_bm25_weight >= 1:
ensemble_retriever = EnsembleRetriever(
retrievers=[bm25_retriever],
weights=[1.0],
id_key=CHUNK_HASH_KEY,
)
else:
ensemble_retriever = EnsembleRetriever(
retrievers=[bm25_retriever, vector_search_retriever],
weights=[hybrid_bm25_weight, 1.0 - hybrid_bm25_weight],
id_key=CHUNK_HASH_KEY,
) )
compression_retriever = ContextualCompressionRetriever( compressor = RerankCompressor(
base_compressor=compressor, base_retriever=ensemble_retriever embedding_function=embedding_function,
) top_n=k_reranker,
reranking_function=reranking_function,
r_score=r,
)
result = await compression_retriever.ainvoke(query) compression_retriever = ContextualCompressionRetriever(
base_compressor=compressor, base_retriever=ensemble_retriever
)
distances = [d.metadata.get('score') for d in result] result = await compression_retriever.ainvoke(query)
documents = [d.page_content for d in result]
metadatas = [d.metadata for d in result]
# retrieve only min(k, k_reranker) items, sort and cut by distance if k < k_reranker distances = [d.metadata.get('score') for d in result]
if k < k_reranker: documents = [d.page_content for d in result]
sorted_items = sorted(zip(distances, documents, metadatas), key=lambda x: x[0], reverse=True) metadatas = [d.metadata for d in result]
sorted_items = sorted_items[:k]
if sorted_items: # retrieve only min(k, k_reranker) items, sort and cut by distance if k < k_reranker
distances, documents, metadatas = map(list, zip(*sorted_items)) if k < k_reranker:
else: sorted_items = sorted(zip(distances, documents, metadatas), key=lambda x: x[0], reverse=True)
distances, documents, metadatas = [], [], [] sorted_items = sorted_items[:k]
result = { if sorted_items:
'distances': [distances], distances, documents, metadatas = map(list, zip(*sorted_items))
'documents': [documents], else:
'metadatas': [metadatas], distances, documents, metadatas = [], [], []
}
log.info('query_doc_with_hybrid_search:result ' + f'{result["metadatas"]} {result["distances"]}') result = {
return result 'distances': [distances],
except Exception as e: 'documents': [documents],
log.exception(f'Error querying doc {collection_name} with hybrid search: {e}') 'metadatas': [metadatas],
raise e }
log.info('query_doc_with_hybrid_search:result %s %s', result['metadatas'], result['distances'])
return result
def merge_get_results(get_results: list[dict]) -> dict: def merge_get_results(get_results: list[dict]) -> dict:
@ -634,7 +656,7 @@ def merge_and_sort_query_results(query_results: list[dict], k: int) -> dict:
if isinstance(document, str): if isinstance(document, str):
doc_hash = (metadata or {}).get(CHUNK_HASH_KEY) or _content_hash(document) doc_hash = (metadata or {}).get(CHUNK_HASH_KEY) or _content_hash(document)
if doc_hash not in combined.keys(): if doc_hash not in combined:
combined[doc_hash] = (distance, document, metadata) combined[doc_hash] = (distance, document, metadata)
continue # if doc is new, no further comparison is needed continue # if doc is new, no further comparison is needed
@ -708,10 +730,11 @@ async def query_collection(
enable_enriched_texts=config.get('rag.enable_hybrid_search_enriched_texts'), enable_enriched_texts=config.get('rag.enable_hybrid_search_enriched_texts'),
) )
except Exception as e: except Exception as e:
log.debug(f'Hybrid search failed, falling back to vector search: {e}') log.debug('Hybrid search failed, falling back to vector search: %s', e)
results = [] results = []
error = False last_error = None
failed_collection_names = set()
def process_query_collection(collection_name, query_embedding): def process_query_collection(collection_name, query_embedding):
try: try:
@ -722,11 +745,10 @@ async def query_collection(
query_embedding=query_embedding, query_embedding=query_embedding,
) )
if result is not None: if result is not None:
return result.model_dump(), None return result.model_dump(), None, collection_name
return None, None return None, None, collection_name
except Exception as e: except Exception as e:
log.exception(f'Error when querying the collection: {e}') return None, e, collection_name
return None, e
# Sanitize: filter out None/empty queries to prevent embedding crashes # Sanitize: filter out None/empty queries to prevent embedding crashes
# (e.g. when get_last_user_message returns None) # (e.g. when get_last_user_message returns None)
@ -737,24 +759,30 @@ async def query_collection(
# Generate all query embeddings (in one call) # Generate all query embeddings (in one call)
query_embeddings = await embedding_function(queries, prefix=RAG_EMBEDDING_QUERY_PREFIX) query_embeddings = await embedding_function(queries, prefix=RAG_EMBEDDING_QUERY_PREFIX)
log.debug(f'query_collection: processing {len(queries)} queries across {len(collection_names)} collections') log.debug('query_collection: processing %s queries across %s collections', len(queries), len(collection_names))
with ThreadPoolExecutor() as executor: task_results = await asyncio.gather(
future_results = [] *[
for query_embedding in query_embeddings: asyncio.to_thread(process_query_collection, collection_name, query_embedding)
for collection_name in collection_names: for query_embedding in query_embeddings
result = executor.submit(process_query_collection, collection_name, query_embedding) for collection_name in collection_names
future_results.append(result) ]
task_results = [future.result() for future in future_results] )
for result, err in task_results: for result, err, collection_name in task_results:
if err is not None: if err is not None:
error = True last_error = err
failed_collection_names.add(collection_name)
elif result is not None: elif result is not None:
results.append(result) results.append(result)
if error and not results: if failed_collection_names:
log.warning('All collection queries failed. No results returned.') log.error(
'query_collection: %s collection(s) had failing queries: %s',
len(failed_collection_names),
', '.join(sorted(failed_collection_names)),
exc_info=last_error,
)
return merge_and_sort_query_results(results, k=k) return merge_and_sort_query_results(results, k=k)
@ -771,7 +799,8 @@ async def query_collection_with_hybrid_search(
enable_enriched_texts: bool = False, enable_enriched_texts: bool = False,
) -> dict: ) -> dict:
results = [] results = []
error = False last_error = None
failed_collection_names = set()
if not enable_enriched_texts: if not enable_enriched_texts:
@ -813,7 +842,7 @@ async def query_collection_with_hybrid_search(
collection_results = dict(await asyncio.gather(*(_fetch_collection(name) for name in collection_names))) collection_results = dict(await asyncio.gather(*(_fetch_collection(name) for name in collection_names)))
log.info(f'Starting hybrid search for {len(queries)} queries in {len(collection_names)} collections...') log.info('Starting hybrid search for %s queries in %s collections...', len(queries), len(collection_names))
async def process_query(collection_name, query): async def process_query(collection_name, query):
try: try:
@ -830,10 +859,9 @@ async def query_collection_with_hybrid_search(
enable_enriched_texts=enable_enriched_texts, enable_enriched_texts=enable_enriched_texts,
native_hybrid_search=False, native_hybrid_search=False,
) )
return result, None return result, None, collection_name
except Exception as e: except Exception as e:
log.exception(f'Error when querying the collection with hybrid_search: {e}') return None, e, collection_name
return None, e
# Prepare tasks for all collections and queries # Prepare tasks for all collections and queries
# Avoid running any tasks for collections that failed to fetch data (have assigned None) # Avoid running any tasks for collections that failed to fetch data (have assigned None)
@ -847,13 +875,22 @@ async def query_collection_with_hybrid_search(
# Run all queries in parallel using asyncio.gather # Run all queries in parallel using asyncio.gather
task_results = await asyncio.gather(*[process_query(collection_name, query) for collection_name, query in tasks]) task_results = await asyncio.gather(*[process_query(collection_name, query) for collection_name, query in tasks])
for result, err in task_results: for result, err, collection_name in task_results:
if err is not None: if err is not None:
error = True last_error = err
failed_collection_names.add(collection_name)
elif result is not None: elif result is not None:
results.append(result) results.append(result)
if error and not results: if failed_collection_names:
log.error(
'query_collection_with_hybrid_search: %s collection(s) had failing queries: %s',
len(failed_collection_names),
', '.join(sorted(failed_collection_names)),
exc_info=last_error,
)
if failed_collection_names and not results:
raise Exception('Hybrid search failed for all collections. Using Non-hybrid search as fallback.') raise Exception('Hybrid search failed for all collections. Using Non-hybrid search as fallback.')
return merge_and_sort_query_results(results, k=k) return merge_and_sort_query_results(results, k=k)
@ -867,15 +904,12 @@ def generate_openai_batch_embeddings(
prefix: str = None, prefix: str = None,
user: UserModel = None, user: UserModel = None,
) -> list[list[float]]: ) -> list[list[float]]:
log.debug(f'generate_openai_batch_embeddings:model {model} batch size: {len(texts)}') log.debug('generate_openai_batch_embeddings:model %s batch size: %s', model, len(texts))
json_data = {'input': texts, 'model': model} json_data = {'input': texts, 'model': model}
if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str): if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str):
json_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix json_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix
headers = { headers = get_json_bearer_headers(key)
'Content-Type': 'application/json',
'Authorization': f'Bearer {key}',
}
if ENABLE_FORWARD_USER_INFO_HEADERS and user: if ENABLE_FORWARD_USER_INFO_HEADERS and user:
headers = include_user_info_headers(headers, user) headers = include_user_info_headers(headers, user)
@ -900,15 +934,12 @@ async def agenerate_openai_batch_embeddings(
prefix: str = None, prefix: str = None,
user: UserModel = None, user: UserModel = None,
) -> list[list[float]]: ) -> list[list[float]]:
log.debug(f'agenerate_openai_batch_embeddings:model {model} batch size: {len(texts)}') log.debug('agenerate_openai_batch_embeddings:model %s batch size: %s', model, len(texts))
form_data = {'input': texts, 'model': model} form_data = {'input': texts, 'model': model}
if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str): if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str):
form_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix form_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix
headers = { headers = get_json_bearer_headers(key)
'Content-Type': 'application/json',
'Authorization': f'Bearer {key}',
}
if ENABLE_FORWARD_USER_INFO_HEADERS and user: if ENABLE_FORWARD_USER_INFO_HEADERS and user:
headers = include_user_info_headers(headers, user) headers = include_user_info_headers(headers, user)
@ -938,7 +969,7 @@ def generate_azure_openai_batch_embeddings(
prefix: str = None, prefix: str = None,
user: UserModel = None, user: UserModel = None,
) -> list[list[float]]: ) -> list[list[float]]:
log.debug(f'generate_azure_openai_batch_embeddings:deployment {model} batch size: {len(texts)}') log.debug('generate_azure_openai_batch_embeddings:deployment %s batch size: %s', model, len(texts))
json_data = {'input': texts} json_data = {'input': texts}
if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str): if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str):
json_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix json_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix
@ -980,7 +1011,7 @@ async def agenerate_azure_openai_batch_embeddings(
prefix: str = None, prefix: str = None,
user: UserModel = None, user: UserModel = None,
) -> list[list[float]]: ) -> list[list[float]]:
log.debug(f'agenerate_azure_openai_batch_embeddings:deployment {model} batch size: {len(texts)}') log.debug('agenerate_azure_openai_batch_embeddings:deployment %s batch size: %s', model, len(texts))
form_data = {'input': texts} form_data = {'input': texts}
if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str): if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str):
form_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix form_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix
@ -1019,15 +1050,12 @@ def generate_ollama_batch_embeddings(
prefix: str = None, prefix: str = None,
user: UserModel = None, user: UserModel = None,
) -> list[list[float]]: ) -> list[list[float]]:
log.debug(f'generate_ollama_batch_embeddings:model {model} batch size: {len(texts)}') log.debug('generate_ollama_batch_embeddings:model %s batch size: %s', model, len(texts))
json_data = {'input': texts, 'model': model, 'truncate': True} json_data = {'input': texts, 'model': model, 'truncate': True}
if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str): if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str):
json_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix json_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix
headers = { headers = get_json_bearer_headers(key)
'Content-Type': 'application/json',
'Authorization': f'Bearer {key}',
}
if ENABLE_FORWARD_USER_INFO_HEADERS and user: if ENABLE_FORWARD_USER_INFO_HEADERS and user:
headers = include_user_info_headers(headers, user) headers = include_user_info_headers(headers, user)
@ -1055,15 +1083,12 @@ async def agenerate_ollama_batch_embeddings(
prefix: str = None, prefix: str = None,
user: UserModel = None, user: UserModel = None,
) -> list[list[float]]: ) -> list[list[float]]:
log.debug(f'agenerate_ollama_batch_embeddings:model {model} batch size: {len(texts)}') log.debug('agenerate_ollama_batch_embeddings:model %s batch size: %s', model, len(texts))
form_data = {'input': texts, 'model': model, 'truncate': True} form_data = {'input': texts, 'model': model, 'truncate': True}
if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str): if isinstance(RAG_EMBEDDING_PREFIX_FIELD_NAME, str) and isinstance(prefix, str):
form_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix form_data[RAG_EMBEDDING_PREFIX_FIELD_NAME] = prefix
headers = { headers = get_json_bearer_headers(key)
'Content-Type': 'application/json',
'Authorization': f'Bearer {key}',
}
if ENABLE_FORWARD_USER_INFO_HEADERS and user: if ENABLE_FORWARD_USER_INFO_HEADERS and user:
headers = include_user_info_headers(headers, user) headers = include_user_info_headers(headers, user)
@ -1101,6 +1126,8 @@ def get_embedding_function(
if embedding_engine == '': if embedding_engine == '':
# Sentence transformers: CPU-bound sync operation # Sentence transformers: CPU-bound sync operation
async def async_embedding_function(query, prefix=None, user=None): async def async_embedding_function(query, prefix=None, user=None):
if USE_SLIM:
raise HTTPException(503, 'Configure an external embedding engine (openai, ollama, azure_openai).')
# Deferred so a missing local model degrades RAG instead of crashing boot. # Deferred so a missing local model degrades RAG instead of crashing boot.
if embedding_function is None: if embedding_function is None:
raise ValueError( raise ValueError(
@ -1108,17 +1135,16 @@ def get_embedding_function(
'SentenceTransformer model name, or configure an external ' 'SentenceTransformer model name, or configure an external '
'RAG_EMBEDDING_ENGINE (ollama, openai, azure_openai).' 'RAG_EMBEDDING_ENGINE (ollama, openai, azure_openai).'
) )
return await asyncio.to_thread(
( def encode():
lambda query, prefix=None: embedding_function.encode( with MPS_INFERENCE_LOCK:
return embedding_function.encode(
query, query,
batch_size=int(embedding_batch_size), batch_size=int(embedding_batch_size),
**({'prompt': prefix} if prefix else {}), **({'prompt': prefix} if prefix else {}),
).tolist() ).tolist()
),
query, return await asyncio.to_thread(encode)
prefix,
)
return async_embedding_function return async_embedding_function
elif embedding_engine in ['ollama', 'openai', 'azure_openai']: elif embedding_engine in ['ollama', 'openai', 'azure_openai']:
@ -1139,7 +1165,7 @@ def get_embedding_function(
batches = [query[i : i + embedding_batch_size] for i in range(0, len(query), embedding_batch_size)] batches = [query[i : i + embedding_batch_size] for i in range(0, len(query), embedding_batch_size)]
if enable_async: if enable_async:
log.debug(f'generate_multiple_async: Processing {len(batches)} batches in parallel') log.debug('generate_multiple_async: Processing %s batches in parallel', len(batches))
# Use semaphore to limit concurrent embedding API requests # Use semaphore to limit concurrent embedding API requests
# 0 = unlimited (no semaphore) # 0 = unlimited (no semaphore)
if concurrent_requests: if concurrent_requests:
@ -1154,7 +1180,7 @@ def get_embedding_function(
tasks = [embedding_function(batch, prefix=prefix, user=user) for batch in batches] tasks = [embedding_function(batch, prefix=prefix, user=user) for batch in batches]
batch_results = await asyncio.gather(*tasks) batch_results = await asyncio.gather(*tasks)
else: else:
log.debug(f'generate_multiple_async: Processing {len(batches)} batches sequentially') log.debug('generate_multiple_async: Processing %s batches sequentially', len(batches))
batch_results = [] batch_results = []
for batch in batches: for batch in batches:
batch_results.append(await embedding_function(batch, prefix=prefix, user=user)) batch_results.append(await embedding_function(batch, prefix=prefix, user=user))
@ -1167,7 +1193,9 @@ def get_embedding_function(
embeddings.extend(batch_embeddings) embeddings.extend(batch_embeddings)
log.debug( log.debug(
f'generate_multiple_async: Generated {len(embeddings)} embeddings from {len(batches)} parallel batches' 'generate_multiple_async: Generated %s embeddings from %s parallel batches',
len(embeddings),
len(batches),
) )
return embeddings return embeddings
else: else:
@ -1233,6 +1261,14 @@ async def generate_embeddings(
def get_reranking_function(reranking_engine, reranking_model, reranking_function, reranking_batch_size=32): def get_reranking_function(reranking_engine, reranking_model, reranking_function, reranking_batch_size=32):
if USE_SLIM and reranking_model and reranking_engine != 'external':
def unavailable(query, documents, user=None):
raise HTTPException(
503, 'Configure an external reranker, or clear the reranking model to use cosine scoring.'
)
return unavailable
if reranking_function is None: if reranking_function is None:
return None return None
if reranking_engine == 'external': if reranking_engine == 'external':
@ -1240,9 +1276,14 @@ def get_reranking_function(reranking_engine, reranking_model, reranking_function
[(query, doc.page_content) for doc in documents], user=user [(query, doc.page_content) for doc in documents], user=user
) )
else: else:
return lambda query, documents, user=None: reranking_function.predict(
[(query, doc.page_content) for doc in documents], batch_size=int(reranking_batch_size) def predict(query, documents, user=None):
) with MPS_INFERENCE_LOCK:
return reranking_function.predict(
[(query, doc.page_content) for doc in documents], batch_size=int(reranking_batch_size)
)
return predict
# UUIDs, SHA-256 digests, and prefixed variants thereof all fit [A-Za-z0-9_-]. # UUIDs, SHA-256 digests, and prefixed variants thereof all fit [A-Za-z0-9_-].
@ -1321,6 +1362,11 @@ async def filter_accessible_collections(
return validated return validated
def filter_source_metadata(metadata: dict) -> dict:
"""Keep only the chunk metadata keys the operator allowed the model to see."""
return {key: metadata[key] for key in RAG_SOURCE_METADATA_KEYS if metadata.get(key) is not None}
async def get_sources_from_items( async def get_sources_from_items(
request, request,
items, items,
@ -1415,8 +1461,21 @@ async def get_sources_from_items(
elif item.get('type') == 'chat': elif item.get('type') == 'chat':
# Chat Attached # Chat Attached
chat = await Chats.get_chat_by_id(item.get('id')) chat = await Chats.get_chat_by_id(item.get('id'))
has_read_access = bool(chat and (user.role == 'admin' or chat.user_id == user.id))
if chat and (user.role == 'admin' or chat.user_id == user.id): if chat and not has_read_access:
has_read_access = await AccessGrants.has_access(
user_id=user.id,
resource_type='shared_chat',
resource_id=chat.id,
permission='read',
)
if chat and not has_read_access and chat.folder_id:
folder = await Folders.get_folder_by_id(chat.folder_id)
has_read_access = folder and await has_folder_access(user.id, folder, 'read', db=None)
if has_read_access:
messages_map = chat.chat.get('history', {}).get('messages', {}) messages_map = chat.chat.get('history', {}).get('messages', {})
message_id = chat.chat.get('history', {}).get('currentId') message_id = chat.chat.get('history', {}).get('currentId')
@ -1607,14 +1666,14 @@ async def get_sources_from_items(
if query_result is None and collection_names: if query_result is None and collection_names:
collection_names = set(collection_names).difference(extracted_collections) collection_names = set(collection_names).difference(extracted_collections)
if not collection_names: if not collection_names:
log.debug(f'skipping {item} as it has already been extracted') log.debug('skipping %s as it has already been extracted', item)
continue continue
# Filter out collections the user cannot read # Filter out collections the user cannot read
if user and (item.get('type'), item.get('id')) not in folder_items: if user and (item.get('type'), item.get('id')) not in folder_items:
collection_names = await filter_accessible_collections(collection_names, user) collection_names = await filter_accessible_collections(collection_names, user)
if not collection_names: if not collection_names:
log.debug(f'access denied for all collections in item {item}') log.debug('access denied for all collections in item %s', item)
continue continue
try: try:
@ -1660,6 +1719,8 @@ async def get_sources_from_items(
def get_model_path(model: str, update_model: bool = False): def get_model_path(model: str, update_model: bool = False):
from huggingface_hub import snapshot_download
# Construct huggingface_hub kwargs with local_files_only to return the snapshot path # Construct huggingface_hub kwargs with local_files_only to return the snapshot path
cache_dir = os.getenv('SENTENCE_TRANSFORMERS_HOME') cache_dir = os.getenv('SENTENCE_TRANSFORMERS_HOME')
@ -1673,8 +1734,8 @@ def get_model_path(model: str, update_model: bool = False):
'local_files_only': local_files_only, 'local_files_only': local_files_only,
} }
log.debug(f'model: {model}') log.debug('model: %s', model)
log.debug(f'snapshot_kwargs: {snapshot_kwargs}') log.debug('snapshot_kwargs: %s', snapshot_kwargs)
# Inspiration from upstream sentence_transformers # Inspiration from upstream sentence_transformers
if os.path.exists(model) or ('\\' in model or model.count('/') > 1) and local_files_only: if os.path.exists(model) or ('\\' in model or model.count('/') > 1) and local_files_only:
@ -1689,7 +1750,7 @@ def get_model_path(model: str, update_model: bool = False):
# Attempt to query the huggingface_hub library to determine the local path and/or to update # Attempt to query the huggingface_hub library to determine the local path and/or to update
try: try:
model_repo_path = snapshot_download(**snapshot_kwargs) model_repo_path = snapshot_download(**snapshot_kwargs)
log.debug(f'model_repo_path: {model_repo_path}') log.debug('model_repo_path: %s', model_repo_path)
return model_repo_path return model_repo_path
except Exception as e: except Exception as e:
log.exception(f'Cannot determine model snapshot path: {e}') log.exception(f'Cannot determine model snapshot path: {e}')
@ -1705,6 +1766,17 @@ from langchain_core.callbacks import Callbacks
from langchain_core.documents import BaseDocumentCompressor, Document from langchain_core.documents import BaseDocumentCompressor, Document
def cosine_similarity(query, documents) -> np.ndarray:
"""Score one query against documents without loading a model runtime."""
if len(documents) == 0:
return np.array([], dtype=float)
query = np.asarray(query, dtype=float).reshape(-1)
documents = np.asarray(documents, dtype=float)
query = query / max(np.linalg.norm(query), 1e-12)
documents = documents / np.maximum(np.linalg.norm(documents, axis=1, keepdims=True), 1e-12)
return documents @ query
class RerankCompressor(BaseDocumentCompressor): class RerankCompressor(BaseDocumentCompressor):
embedding_function: Any embedding_function: Any
top_n: int top_n: int
@ -1740,18 +1812,18 @@ class RerankCompressor(BaseDocumentCompressor):
query: str, query: str,
callbacks: Callbacks | None = None, callbacks: Callbacks | None = None,
) -> Sequence[Document]: ) -> Sequence[Document]:
if not documents:
return []
reranking = self.reranking_function is not None reranking = self.reranking_function is not None
scores = None scores = None
if reranking: if reranking:
scores = await asyncio.to_thread(self.reranking_function, query, documents) scores = await asyncio.to_thread(self.reranking_function, query, documents)
else: else:
from sentence_transformers import util as st_util
query_embedding = await self.embedding_function(query, RAG_EMBEDDING_QUERY_PREFIX) query_embedding = await self.embedding_function(query, RAG_EMBEDDING_QUERY_PREFIX)
doc_texts = [doc.page_content for doc in documents] doc_texts = [doc.page_content for doc in documents]
document_embedding = await self.embedding_function(doc_texts, RAG_EMBEDDING_CONTENT_PREFIX) document_embedding = await self.embedding_function(doc_texts, RAG_EMBEDDING_CONTENT_PREFIX)
scores = st_util.cos_sim(query_embedding, document_embedding)[0] scores = cosine_similarity(query_embedding, document_embedding)
if scores is not None: if scores is not None:
docs_with_scores = list( docs_with_scores = list(

View file

@ -15,9 +15,8 @@ transparently dispatches each call to a worker thread via
`asyncio.to_thread`. Async callers can `await ASYNC_VECTOR_DB_CLIENT.x(...)` `asyncio.to_thread`. Async callers can `await ASYNC_VECTOR_DB_CLIENT.x(...)`
in place of `VECTOR_DB_CLIENT.x(...)` and the loop stays responsive. in place of `VECTOR_DB_CLIENT.x(...)` and the loop stays responsive.
The original `VECTOR_DB_CLIENT` is unchanged, so callers already running Client initialization and calls run in the worker thread. Synchronous callers
inside `run_in_threadpool` (e.g. `save_docs_to_vector_db`) are not already inside `run_in_threadpool` use `get_vector_db_client()` directly.
affected.
Thread-safety expectations Thread-safety expectations
-------------------------- --------------------------
@ -55,7 +54,7 @@ from __future__ import annotations
import asyncio import asyncio
from typing import Dict, List, Optional, Union from typing import Dict, List, Optional, Union
from open_webui.retrieval.vector.factory import VECTOR_DB_CLIENT from open_webui.retrieval.vector.factory import get_vector_db_client
from open_webui.retrieval.vector.main import ( from open_webui.retrieval.vector.main import (
GetResult, GetResult,
SearchResult, SearchResult,
@ -73,30 +72,30 @@ class AsyncVectorDBClient:
typically swallowed by surrounding ``try/except``). typically swallowed by surrounding ``try/except``).
""" """
def __init__(self, sync_client: VectorDBBase) -> None: def __init__(self, sync_client: Optional[VectorDBBase] = None) -> None:
self._sync = sync_client self._sync = sync_client
@property @property
def sync(self) -> VectorDBBase: def sync(self) -> VectorDBBase:
"""Escape hatch for code that must call the sync client directly """Escape hatch for code that must call the sync client directly
(e.g. already inside a worker thread).""" (e.g. already inside a worker thread)."""
return self._sync return self._sync if self._sync is not None else get_vector_db_client()
@property @property
def supports_hybrid_search(self) -> bool: def supports_hybrid_search(self) -> bool:
return type(self._sync).hybrid_search is not VectorDBBase.hybrid_search return type(self.sync).hybrid_search is not VectorDBBase.hybrid_search
async def has_collection(self, collection_name: str) -> bool: async def has_collection(self, collection_name: str) -> bool:
return await asyncio.to_thread(self._sync.has_collection, collection_name) return await asyncio.to_thread(lambda: self.sync.has_collection(collection_name))
async def delete_collection(self, collection_name: str) -> None: async def delete_collection(self, collection_name: str) -> None:
return await asyncio.to_thread(self._sync.delete_collection, collection_name) return await asyncio.to_thread(lambda: self.sync.delete_collection(collection_name))
async def insert(self, collection_name: str, items: List[VectorItem]) -> None: async def insert(self, collection_name: str, items: List[VectorItem]) -> None:
return await asyncio.to_thread(self._sync.insert, collection_name, items) return await asyncio.to_thread(lambda: self.sync.insert(collection_name, items))
async def upsert(self, collection_name: str, items: List[VectorItem]) -> None: async def upsert(self, collection_name: str, items: List[VectorItem]) -> None:
return await asyncio.to_thread(self._sync.upsert, collection_name, items) return await asyncio.to_thread(lambda: self.sync.upsert(collection_name, items))
async def search( async def search(
self, self,
@ -105,7 +104,7 @@ class AsyncVectorDBClient:
filter: Optional[Dict] = None, filter: Optional[Dict] = None,
limit: int = 10, limit: int = 10,
) -> Optional[SearchResult]: ) -> Optional[SearchResult]:
return await asyncio.to_thread(self._sync.search, collection_name, vectors, filter, limit) return await asyncio.to_thread(lambda: self.sync.search(collection_name, vectors, filter, limit))
async def hybrid_search( async def hybrid_search(
self, self,
@ -117,13 +116,7 @@ class AsyncVectorDBClient:
hybrid_bm25_weight: float = 0.5, hybrid_bm25_weight: float = 0.5,
) -> Optional[SearchResult]: ) -> Optional[SearchResult]:
return await asyncio.to_thread( return await asyncio.to_thread(
self._sync.hybrid_search, lambda: self.sync.hybrid_search(collection_name, query, vectors, filter, limit, hybrid_bm25_weight)
collection_name,
query,
vectors,
filter,
limit,
hybrid_bm25_weight,
) )
async def query( async def query(
@ -132,10 +125,10 @@ class AsyncVectorDBClient:
filter: Dict, filter: Dict,
limit: Optional[int] = None, limit: Optional[int] = None,
) -> Optional[GetResult]: ) -> Optional[GetResult]:
return await asyncio.to_thread(self._sync.query, collection_name, filter, limit) return await asyncio.to_thread(lambda: self.sync.query(collection_name, filter, limit))
async def get(self, collection_name: str) -> Optional[GetResult]: async def get(self, collection_name: str) -> Optional[GetResult]:
return await asyncio.to_thread(self._sync.get, collection_name) return await asyncio.to_thread(lambda: self.sync.get(collection_name))
async def delete( async def delete(
self, self,
@ -143,10 +136,10 @@ class AsyncVectorDBClient:
ids: Optional[List[str]] = None, ids: Optional[List[str]] = None,
filter: Optional[Dict] = None, filter: Optional[Dict] = None,
) -> None: ) -> None:
return await asyncio.to_thread(self._sync.delete, collection_name, ids, filter) return await asyncio.to_thread(lambda: self.sync.delete(collection_name, ids, filter))
async def reset(self) -> None: async def reset(self) -> None:
return await asyncio.to_thread(self._sync.reset) return await asyncio.to_thread(lambda: self.sync.reset())
ASYNC_VECTOR_DB_CLIENT = AsyncVectorDBClient(VECTOR_DB_CLIENT) ASYNC_VECTOR_DB_CLIENT = AsyncVectorDBClient()

View file

@ -16,6 +16,8 @@ from open_webui.config import (
CHROMA_HTTP_SSL, CHROMA_HTTP_SSL,
CHROMA_TENANT, CHROMA_TENANT,
) )
from open_webui.env import USE_SLIM
from fastapi import HTTPException
from open_webui.retrieval.vector.main import ( from open_webui.retrieval.vector.main import (
GetResult, GetResult,
SearchResult, SearchResult,
@ -29,6 +31,8 @@ log = logging.getLogger(__name__)
class ChromaClient(VectorDBBase): class ChromaClient(VectorDBBase):
def __init__(self): def __init__(self):
if USE_SLIM and not CHROMA_HTTP_HOST:
raise HTTPException(503, 'Configure CHROMA_HTTP_HOST: embedded Chroma is unavailable in slim.')
settings_dict = { settings_dict = {
'allow_reset': True, 'allow_reset': True,
'anonymized_telemetry': False, 'anonymized_telemetry': False,
@ -58,7 +62,7 @@ class ChromaClient(VectorDBBase):
def has_collection(self, collection_name: str) -> bool: def has_collection(self, collection_name: str) -> bool:
try: try:
self.client.get_collection(name=collection_name) self.client.get_collection(name=collection_name, embedding_function=None)
return True return True
except NotFoundError: except NotFoundError:
return False return False
@ -76,7 +80,7 @@ class ChromaClient(VectorDBBase):
) -> Optional[SearchResult]: ) -> Optional[SearchResult]:
# Search for the nearest neighbor items based on the vectors and return 'limit' number of results. # Search for the nearest neighbor items based on the vectors and return 'limit' number of results.
try: try:
collection = self.client.get_collection(name=collection_name) collection = self.client.get_collection(name=collection_name, embedding_function=None)
if collection: if collection:
result = collection.query( result = collection.query(
query_embeddings=vectors, query_embeddings=vectors,
@ -105,7 +109,7 @@ class ChromaClient(VectorDBBase):
def query(self, collection_name: str, filter: dict, limit: Optional[int] = None) -> Optional[GetResult]: def query(self, collection_name: str, filter: dict, limit: Optional[int] = None) -> Optional[GetResult]:
# Query the items from the collection based on the filter. # Query the items from the collection based on the filter.
try: try:
collection = self.client.get_collection(name=collection_name) collection = self.client.get_collection(name=collection_name, embedding_function=None)
if collection: if collection:
result = collection.get( result = collection.get(
where=filter, where=filter,
@ -125,7 +129,7 @@ class ChromaClient(VectorDBBase):
def get(self, collection_name: str) -> Optional[GetResult]: def get(self, collection_name: str) -> Optional[GetResult]:
# Get all the items in the collection. # Get all the items in the collection.
collection = self.client.get_collection(name=collection_name) collection = self.client.get_collection(name=collection_name, embedding_function=None)
if collection: if collection:
result = collection.get() result = collection.get()
return GetResult( return GetResult(
@ -139,7 +143,9 @@ class ChromaClient(VectorDBBase):
def insert(self, collection_name: str, items: list[VectorItem]): def insert(self, collection_name: str, items: list[VectorItem]):
# Insert the items into the collection, if the collection does not exist, it will be created. # Insert the items into the collection, if the collection does not exist, it will be created.
collection = self.client.get_or_create_collection(name=collection_name, metadata={'hnsw:space': 'cosine'}) collection = self.client.get_or_create_collection(
name=collection_name, metadata={'hnsw:space': 'cosine'}, embedding_function=None
)
ids = [item['id'] for item in items] ids = [item['id'] for item in items]
documents = [item['text'] for item in items] documents = [item['text'] for item in items]
@ -157,7 +163,9 @@ class ChromaClient(VectorDBBase):
def upsert(self, collection_name: str, items: list[VectorItem]): def upsert(self, collection_name: str, items: list[VectorItem]):
# Update the items in the collection, if the items are not present, insert them. If the collection does not exist, it will be created. # Update the items in the collection, if the items are not present, insert them. If the collection does not exist, it will be created.
collection = self.client.get_or_create_collection(name=collection_name, metadata={'hnsw:space': 'cosine'}) collection = self.client.get_or_create_collection(
name=collection_name, metadata={'hnsw:space': 'cosine'}, embedding_function=None
)
ids = [item['id'] for item in items] ids = [item['id'] for item in items]
documents = [item['text'] for item in items] documents = [item['text'] for item in items]
@ -174,7 +182,7 @@ class ChromaClient(VectorDBBase):
): ):
# Delete the items from the collection based on the ids. # Delete the items from the collection based on the ids.
try: try:
collection = self.client.get_collection(name=collection_name) collection = self.client.get_collection(name=collection_name, embedding_function=None)
if collection: if collection:
if ids: if ids:
collection.delete(ids=ids) collection.delete(ids=ids)
@ -182,7 +190,7 @@ class ChromaClient(VectorDBBase):
collection.delete(where=filter) collection.delete(where=filter)
except Exception as e: except Exception as e:
# If collection doesn't exist, that's fine - nothing to delete # If collection doesn't exist, that's fine - nothing to delete
log.debug(f'Attempted to delete from non-existent collection {collection_name}. Ignoring.') log.debug('Attempted to delete from non-existent collection %s. Ignoring.', collection_name)
pass pass
def reset(self): def reset(self):

View file

@ -3,7 +3,7 @@ NOTE: This vector database integration is community-supported and maintained on
""" """
import ssl import ssl
from typing import Optional from typing import Any, Optional
from elasticsearch import BadRequestError, Elasticsearch from elasticsearch import BadRequestError, Elasticsearch
from elasticsearch.helpers import bulk, scan from elasticsearch.helpers import bulk, scan
@ -23,7 +23,13 @@ from open_webui.retrieval.vector.main import (
VectorDBBase, VectorDBBase,
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import process_metadata from open_webui.retrieval.vector.utils import iter_filter_conditions, process_metadata
def _metadata_filter(key: str, op: str, value: Any) -> dict:
if op == '$in':
return {'terms': {f'metadata.{key}': value}}
return {'term': {f'metadata.{key}': value}}
class ElasticsearchClient(VectorDBBase): class ElasticsearchClient(VectorDBBase):
@ -161,12 +167,16 @@ class ElasticsearchClient(VectorDBBase):
filter: Optional[dict] = None, filter: Optional[dict] = None,
limit: int = 10, limit: int = 10,
) -> Optional[SearchResult]: ) -> Optional[SearchResult]:
filters = [{'term': {'collection': collection_name}}]
if filter:
filters.extend(_metadata_filter(key, op, value) for key, op, value in iter_filter_conditions(filter))
query = { query = {
'size': limit, 'size': limit,
'_source': ['text', 'metadata'], '_source': ['text', 'metadata'],
'query': { 'query': {
'script_score': { 'script_score': {
'query': {'bool': {'filter': [{'term': {'collection': collection_name}}]}}, 'query': {'bool': {'filter': filters}},
'script': { 'script': {
'source': "cosineSimilarity(params.vector, 'vector') + 1.0", 'source': "cosineSimilarity(params.vector, 'vector') + 1.0",
'params': {'vector': vectors[0]}, # Assuming single query vector 'params': {'vector': vectors[0]}, # Assuming single query vector

View file

@ -5,7 +5,6 @@ NOTE: This vector database integration is community-supported and maintained on
from __future__ import annotations from __future__ import annotations
import array import array
import json
import logging import logging
import math import math
import re import re
@ -30,6 +29,7 @@ from open_webui.retrieval.vector.main import (
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import process_metadata from open_webui.retrieval.vector.utils import process_metadata
from open_webui.utils.json_codec import JSONCodec
from sqlalchemy import create_engine from sqlalchemy import create_engine
from sqlalchemy.pool import NullPool, QueuePool from sqlalchemy.pool import NullPool, QueuePool
@ -72,7 +72,7 @@ def _safe_json(v: Any) -> Dict[str, Any]:
return {} return {}
if isinstance(v, str): if isinstance(v, str):
try: try:
j = json.loads(v) j = JSONCodec.loads(v)
return j if isinstance(j, dict) else {} return j if isinstance(j, dict) else {}
except Exception: except Exception:
return {} return {}
@ -324,7 +324,7 @@ class MariaDBVectorClient(VectorDBBase):
emb, emb,
collection_name, collection_name,
item.get('text'), item.get('text'),
json.dumps(meta), JSONCodec.dumps(meta),
) )
) )
cur.executemany(sql, params) cur.executemany(sql, params)
@ -367,7 +367,7 @@ class MariaDBVectorClient(VectorDBBase):
emb, emb,
collection_name, collection_name,
item.get('text'), item.get('text'),
json.dumps(meta), JSONCodec.dumps(meta),
) )
) )
cur.executemany(sql, params) cur.executemany(sql, params)

View file

@ -2,9 +2,9 @@
NOTE: This vector database integration is community-supported and maintained on a best-effort basis. NOTE: This vector database integration is community-supported and maintained on a best-effort basis.
""" """
import json
import logging import logging
from typing import Optional import re
from typing import Any, Optional
from open_webui.config import ( from open_webui.config import (
MILVUS_DB, MILVUS_DB,
@ -24,7 +24,8 @@ from open_webui.retrieval.vector.main import (
VectorDBBase, VectorDBBase,
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import process_metadata from open_webui.retrieval.vector.utils import iter_filter_conditions, process_metadata
from open_webui.utils.json_codec import JSONCodec
from pymilvus import DataType from pymilvus import DataType
from pymilvus import MilvusClient as Client from pymilvus import MilvusClient as Client
from pymilvus.exceptions import MilvusException from pymilvus.exceptions import MilvusException
@ -35,6 +36,36 @@ log = logging.getLogger(__name__)
# field). Clamp long chunks before insert so one oversized chunk can't fail the # field). Clamp long chunks before insert so one oversized chunk can't fail the
# whole batch and leave the file with zero embeddings. # whole batch and leave the file with zero embeddings.
MILVUS_TEXT_MAX_LENGTH = 65535 MILVUS_TEXT_MAX_LENGTH = 65535
_SAFE_METADATA_KEY_RE = re.compile(r'^[A-Za-z_][A-Za-z0-9_]{0,63}$')
def _escape_milvus_string(value: str) -> str:
if not isinstance(value, str):
raise TypeError(f'Expected str, got {type(value).__name__}')
return value.replace('\\', '\\\\').replace("'", "\\'")
def _milvus_literal(value: Any) -> str:
if isinstance(value, str):
return f"'{_escape_milvus_string(value)}'"
if isinstance(value, bool):
return str(value).lower()
if isinstance(value, (int, float)):
return str(value)
raise TypeError(f'Unsupported Milvus filter value type: {type(value).__name__}')
def _metadata_exprs(filter: Optional[dict]) -> list[str]:
exprs = []
for key, op, value in iter_filter_conditions(filter):
if not isinstance(key, str) or not _SAFE_METADATA_KEY_RE.fullmatch(key):
raise ValueError(f'Invalid Milvus metadata filter key: {key!r}')
if op == '$in':
items = [f"metadata['{key}'] == {_milvus_literal(item)}" for item in value]
exprs.append(f'({" or ".join(items)})' if items else 'false')
else:
exprs.append(f"metadata['{key}'] == {_milvus_literal(value)}")
return exprs
class MilvusClient(VectorDBBase): class MilvusClient(VectorDBBase):
@ -125,7 +156,7 @@ class MilvusClient(VectorDBBase):
index_type = MILVUS_INDEX_TYPE.upper() index_type = MILVUS_INDEX_TYPE.upper()
metric_type = MILVUS_METRIC_TYPE.upper() metric_type = MILVUS_METRIC_TYPE.upper()
log.info(f'Using Milvus index type: {index_type}, metric type: {metric_type}') log.info('Using Milvus index type: %s, metric type: %s', index_type, metric_type)
index_creation_params = {} index_creation_params = {}
if index_type == 'HNSW': if index_type == 'HNSW':
@ -133,18 +164,18 @@ class MilvusClient(VectorDBBase):
'M': MILVUS_HNSW_M, 'M': MILVUS_HNSW_M,
'efConstruction': MILVUS_HNSW_EFCONSTRUCTION, 'efConstruction': MILVUS_HNSW_EFCONSTRUCTION,
} }
log.info(f'HNSW params: {index_creation_params}') log.info('HNSW params: %s', index_creation_params)
elif index_type == 'IVF_FLAT': elif index_type == 'IVF_FLAT':
index_creation_params = {'nlist': MILVUS_IVF_FLAT_NLIST} index_creation_params = {'nlist': MILVUS_IVF_FLAT_NLIST}
log.info(f'IVF_FLAT params: {index_creation_params}') log.info('IVF_FLAT params: %s', index_creation_params)
elif index_type == 'DISKANN': elif index_type == 'DISKANN':
index_creation_params = { index_creation_params = {
'max_degree': MILVUS_DISKANN_MAX_DEGREE, 'max_degree': MILVUS_DISKANN_MAX_DEGREE,
'search_list_size': MILVUS_DISKANN_SEARCH_LIST_SIZE, 'search_list_size': MILVUS_DISKANN_SEARCH_LIST_SIZE,
} }
log.info(f'DISKANN params: {index_creation_params}') log.info('DISKANN params: %s', index_creation_params)
elif index_type in ['FLAT', 'AUTOINDEX']: elif index_type in ['FLAT', 'AUTOINDEX']:
log.info(f'Using {index_type} index with no specific build-time params.') log.info('Using %s index with no specific build-time params.', index_type)
else: else:
log.warning( log.warning(
f"Unsupported MILVUS_INDEX_TYPE: '{index_type}'. " f"Unsupported MILVUS_INDEX_TYPE: '{index_type}'. "
@ -167,7 +198,11 @@ class MilvusClient(VectorDBBase):
index_params=index_params, index_params=index_params,
) )
log.info( log.info(
f"Successfully created collection '{self.collection_prefix}_{collection_name}' with index type '{index_type}' and metric '{metric_type}'." "Successfully created collection '%s_%s' with index type '%s' and metric '%s'.",
self.collection_prefix,
collection_name,
index_type,
metric_type,
) )
def has_collection(self, collection_name: str) -> bool: def has_collection(self, collection_name: str) -> bool:
@ -189,6 +224,9 @@ class MilvusClient(VectorDBBase):
) -> Optional[SearchResult]: ) -> Optional[SearchResult]:
# Search for the nearest neighbor items based on the vectors and return 'limit' number of results. # Search for the nearest neighbor items based on the vectors and return 'limit' number of results.
collection_name = collection_name.replace('-', '_') collection_name = collection_name.replace('-', '_')
kwargs = {}
if filter:
kwargs['filter'] = ' and '.join(_metadata_exprs(filter))
# For some index types like IVF_FLAT, search params like nprobe can be set. # For some index types like IVF_FLAT, search params like nprobe can be set.
# Example: search_params = {"nprobe": 10} if using IVF_FLAT # Example: search_params = {"nprobe": 10} if using IVF_FLAT
# For simplicity, not adding configurable search_params here, but could be extended. # For simplicity, not adding configurable search_params here, but could be extended.
@ -197,6 +235,7 @@ class MilvusClient(VectorDBBase):
data=vectors, data=vectors,
limit=limit, limit=limit,
output_fields=['data', 'metadata'], output_fields=['data', 'metadata'],
**kwargs,
# search_params=search_params # Potentially add later if needed # search_params=search_params # Potentially add later if needed
) )
return self._result_to_search_result(result) return self._result_to_search_result(result)
@ -220,7 +259,11 @@ class MilvusClient(VectorDBBase):
try: try:
log.info( log.info(
f"Querying collection {self.collection_prefix}_{collection_name} with filter: '{filter_string}', limit: {limit}" "Querying collection %s_%s with filter: '%s', limit: %s",
self.collection_prefix,
collection_name,
filter_string,
limit,
) )
iterator = self.client.query_iterator( iterator = self.client.query_iterator(
@ -242,7 +285,7 @@ class MilvusClient(VectorDBBase):
break break
all_results.extend(batch) all_results.extend(batch)
log.debug(f'Total results from query: {len(all_results)}') log.debug('Total results from query: %s', len(all_results))
return self._result_to_get_result([all_results] if all_results else [[]]) return self._result_to_get_result([all_results] if all_results else [[]])
except Exception as e: except Exception as e:
@ -265,7 +308,7 @@ class MilvusClient(VectorDBBase):
# Insert the items into the collection, if the collection does not exist, it will be created. # Insert the items into the collection, if the collection does not exist, it will be created.
collection_name = collection_name.replace('-', '_') collection_name = collection_name.replace('-', '_')
if not self.client.has_collection(collection_name=f'{self.collection_prefix}_{collection_name}'): if not self.client.has_collection(collection_name=f'{self.collection_prefix}_{collection_name}'):
log.info(f'Collection {self.collection_prefix}_{collection_name} does not exist. Creating now.') log.info('Collection %s_%s does not exist. Creating now.', self.collection_prefix, collection_name)
if not items: if not items:
log.error( log.error(
f'Cannot create collection {self.collection_prefix}_{collection_name} without items to determine dimension.' f'Cannot create collection {self.collection_prefix}_{collection_name} without items to determine dimension.'
@ -273,7 +316,7 @@ class MilvusClient(VectorDBBase):
raise ValueError('Cannot create Milvus collection without items to determine vector dimension.') raise ValueError('Cannot create Milvus collection without items to determine vector dimension.')
self._create_collection(collection_name=collection_name, dimension=len(items[0]['vector'])) self._create_collection(collection_name=collection_name, dimension=len(items[0]['vector']))
log.info(f'Inserting {len(items)} items into collection {self.collection_prefix}_{collection_name}.') log.info('Inserting %s items into collection %s_%s.', len(items), self.collection_prefix, collection_name)
data = [] data = []
for item in items: for item in items:
text = item['text'] or '' text = item['text'] or ''
@ -301,7 +344,9 @@ class MilvusClient(VectorDBBase):
# Update the items in the collection, if the items are not present, insert them. If the collection does not exist, it will be created. # Update the items in the collection, if the items are not present, insert them. If the collection does not exist, it will be created.
collection_name = collection_name.replace('-', '_') collection_name = collection_name.replace('-', '_')
if not self.client.has_collection(collection_name=f'{self.collection_prefix}_{collection_name}'): if not self.client.has_collection(collection_name=f'{self.collection_prefix}_{collection_name}'):
log.info(f'Collection {self.collection_prefix}_{collection_name} does not exist for upsert. Creating now.') log.info(
'Collection %s_%s does not exist for upsert. Creating now.', self.collection_prefix, collection_name
)
if not items: if not items:
log.error( log.error(
f'Cannot create collection {self.collection_prefix}_{collection_name} for upsert without items to determine dimension.' f'Cannot create collection {self.collection_prefix}_{collection_name} for upsert without items to determine dimension.'
@ -311,7 +356,7 @@ class MilvusClient(VectorDBBase):
) )
self._create_collection(collection_name=collection_name, dimension=len(items[0]['vector'])) self._create_collection(collection_name=collection_name, dimension=len(items[0]['vector']))
log.info(f'Upserting {len(items)} items into collection {self.collection_prefix}_{collection_name}.') log.info('Upserting %s items into collection %s_%s.', len(items), self.collection_prefix, collection_name)
data = [] data = []
for item in items: for item in items:
text = item['text'] or '' text = item['text'] or ''
@ -348,15 +393,20 @@ class MilvusClient(VectorDBBase):
return None return None
if ids: if ids:
log.info(f'Deleting items by IDs from {self.collection_prefix}_{collection_name}. IDs: {ids}') log.info('Deleting items by IDs from %s_%s. IDs: %s', self.collection_prefix, collection_name, ids)
return self.client.delete( return self.client.delete(
collection_name=f'{self.collection_prefix}_{collection_name}', collection_name=f'{self.collection_prefix}_{collection_name}',
ids=ids, ids=ids,
) )
elif filter: elif filter:
filter_string = ' && '.join([f'metadata["{key}"] == {json.dumps(value)}' for key, value in filter.items()]) filter_string = ' && '.join(
[f'metadata["{key}"] == {JSONCodec.dumps(value)}' for key, value in filter.items()]
)
log.info( log.info(
f'Deleting items by filter from {self.collection_prefix}_{collection_name}. Filter: {filter_string}' 'Deleting items by filter from %s_%s. Filter: %s',
self.collection_prefix,
collection_name,
filter_string,
) )
return self.client.delete( return self.client.delete(
collection_name=f'{self.collection_prefix}_{collection_name}', collection_name=f'{self.collection_prefix}_{collection_name}',
@ -378,7 +428,7 @@ class MilvusClient(VectorDBBase):
try: try:
self.client.drop_collection(collection_name=collection_name_full) self.client.drop_collection(collection_name=collection_name_full)
deleted_collections.append(collection_name_full) deleted_collections.append(collection_name_full)
log.info(f'Deleted collection: {collection_name_full}') log.info('Deleted collection: %s', collection_name_full)
except Exception as e: except Exception as e:
log.error(f'Error deleting collection {collection_name_full}: {e}') log.error(f'Error deleting collection {collection_name_full}: {e}')
log.info(f'Milvus reset complete. Deleted collections: {deleted_collections}') log.info('Milvus reset complete. Deleted collections: %s', deleted_collections)

View file

@ -17,12 +17,14 @@ from open_webui.config import (
MILVUS_TOKEN, MILVUS_TOKEN,
MILVUS_URI, MILVUS_URI,
) )
from open_webui.retrieval.vector.dbs.milvus import _metadata_exprs
from open_webui.retrieval.vector.main import ( from open_webui.retrieval.vector.main import (
GetResult, GetResult,
SearchResult, SearchResult,
VectorDBBase, VectorDBBase,
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import process_metadata
from pymilvus import DataType from pymilvus import DataType
from pymilvus import MilvusClient as Client from pymilvus import MilvusClient as Client
from pymilvus.exceptions import MilvusException from pymilvus.exceptions import MilvusException
@ -146,7 +148,7 @@ class MilvusClient(VectorDBBase):
# The index only accelerates resource_id filters; never fail # The index only accelerates resource_id filters; never fail
# collection creation over it. # collection creation over it.
log.warning(f'Could not create {RESOURCE_ID_FIELD} index on {mt_collection_name}: {e}') log.warning(f'Could not create {RESOURCE_ID_FIELD} index on {mt_collection_name}: {e}')
log.info(f'Created shared collection: {mt_collection_name}') log.info('Created shared collection: %s', mt_collection_name)
def _ensure_collection(self, mt_collection_name: str, dimension: int): def _ensure_collection(self, mt_collection_name: str, dimension: int):
if not self.client.has_collection(mt_collection_name): if not self.client.has_collection(mt_collection_name):
@ -190,7 +192,7 @@ class MilvusClient(VectorDBBase):
'id': item['id'], 'id': item['id'],
'vector': item['vector'], 'vector': item['vector'],
'text': text, 'text': text,
'metadata': item['metadata'], 'metadata': process_metadata(item['metadata']),
RESOURCE_ID_FIELD: resource_id, RESOURCE_ID_FIELD: resource_id,
} }
) )
@ -221,13 +223,14 @@ class MilvusClient(VectorDBBase):
self.client.load_collection(mt_collection) self.client.load_collection(mt_collection)
expr = [f"{RESOURCE_ID_FIELD} == '{resource_id}'", *_metadata_exprs(filter)]
results = self.client.search( results = self.client.search(
collection_name=mt_collection, collection_name=mt_collection,
data=vectors, data=vectors,
anns_field='vector', anns_field='vector',
search_params={'metric_type': MILVUS_METRIC_TYPE, 'params': {}}, search_params={'metric_type': MILVUS_METRIC_TYPE, 'params': {}},
limit=limit, limit=limit,
filter=f"{RESOURCE_ID_FIELD} == '{resource_id}'", filter=' and '.join(expr),
output_fields=['id', 'text', 'metadata'], output_fields=['id', 'text', 'metadata'],
) )

View file

@ -2,7 +2,6 @@
NOTE: This vector database integration is community-supported and maintained on a best-effort basis. NOTE: This vector database integration is community-supported and maintained on a best-effort basis.
""" """
import json
import logging import logging
import re import re
from typing import Any, Dict, List, Optional from typing import Any, Dict, List, Optional
@ -63,20 +62,18 @@ from open_webui.config import (
OPENGAUSS_POOL_SIZE, OPENGAUSS_POOL_SIZE,
OPENGAUSS_POOL_TIMEOUT, OPENGAUSS_POOL_TIMEOUT,
) )
from open_webui.env import SRC_LOG_LEVELS
from open_webui.retrieval.vector.main import ( from open_webui.retrieval.vector.main import (
GetResult, GetResult,
SearchResult, SearchResult,
VectorDBBase, VectorDBBase,
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import process_metadata from open_webui.retrieval.vector.utils import iter_filter_conditions, process_metadata
VECTOR_LENGTH = OPENGAUSS_INITIALIZE_MAX_VECTOR_LENGTH VECTOR_LENGTH = OPENGAUSS_INITIALIZE_MAX_VECTOR_LENGTH
Base = declarative_base() Base = declarative_base()
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
log.setLevel(SRC_LOG_LEVELS['RAG'])
class DocumentChunk(Base): class DocumentChunk(Base):
@ -89,6 +86,12 @@ class DocumentChunk(Base):
vmetadata = Column(MutableDict.as_mutable(JSONB), nullable=True) vmetadata = Column(MutableDict.as_mutable(JSONB), nullable=True)
def _metadata_clause(key: str, op: str, value: Any):
if op == '$in':
return DocumentChunk.vmetadata[key].astext.in_([str(v) for v in value])
return DocumentChunk.vmetadata[key].astext == str(value)
class OpenGaussClient(VectorDBBase): class OpenGaussClient(VectorDBBase):
def __init__(self) -> None: def __init__(self) -> None:
if not OPENGAUSS_DB_URL: if not OPENGAUSS_DB_URL:
@ -182,7 +185,7 @@ class OpenGaussClient(VectorDBBase):
new_items.append(new_chunk) new_items.append(new_chunk)
self.session.bulk_save_objects(new_items) self.session.bulk_save_objects(new_items)
self.session.commit() self.session.commit()
log.info(f"Inserting {len(new_items)} items into collection '{collection_name}'.") log.info("Inserting %s items into collection '%s'.", len(new_items), collection_name)
except Exception as e: except Exception as e:
self.session.rollback() self.session.rollback()
log.exception(f'Failed to insert data: {e}') log.exception(f'Failed to insert data: {e}')
@ -208,7 +211,7 @@ class OpenGaussClient(VectorDBBase):
) )
self.session.add(new_chunk) self.session.add(new_chunk)
self.session.commit() self.session.commit()
log.info(f"Inserting/updating {len(items)} items in collection '{collection_name}'.") log.info("Inserting/updating %s items in collection '%s'.", len(items), collection_name)
except Exception as e: except Exception as e:
self.session.rollback() self.session.rollback()
log.exception(f'Failed to insert or update data.: {e}') log.exception(f'Failed to insert or update data.: {e}')
@ -245,10 +248,15 @@ class OpenGaussClient(VectorDBBase):
DocumentChunk.vmetadata, DocumentChunk.vmetadata,
(DocumentChunk.vector.cosine_distance(query_vectors.c.q_vector)).label('distance'), (DocumentChunk.vector.cosine_distance(query_vectors.c.q_vector)).label('distance'),
] ]
where_clauses = [DocumentChunk.collection_name == collection_name]
if filter:
where_clauses.extend(
_metadata_clause(key, op, value) for key, op, value in iter_filter_conditions(filter)
)
subq = ( subq = (
select(*result_fields) select(*result_fields)
.where(DocumentChunk.collection_name == collection_name) .where(*where_clauses)
.order_by(DocumentChunk.vector.cosine_distance(query_vectors.c.q_vector)) .order_by(DocumentChunk.vector.cosine_distance(query_vectors.c.q_vector))
) )
if limit is not None: if limit is not None:
@ -303,6 +311,7 @@ class OpenGaussClient(VectorDBBase):
results = query.all() results = query.all()
if not results: if not results:
self.session.rollback()
return None return None
ids = [[result.id for result in results]] ids = [[result.id for result in results]]
@ -325,6 +334,7 @@ class OpenGaussClient(VectorDBBase):
results = query.all() results = query.all()
if not results: if not results:
self.session.rollback()
return None return None
ids = [[result.id for result in results]] ids = [[result.id for result in results]]
@ -353,7 +363,7 @@ class OpenGaussClient(VectorDBBase):
query = query.filter(DocumentChunk.vmetadata[key].astext == str(value)) query = query.filter(DocumentChunk.vmetadata[key].astext == str(value))
deleted = query.delete(synchronize_session=False) deleted = query.delete(synchronize_session=False)
self.session.commit() self.session.commit()
log.info(f"Deleted {deleted} items from collection '{collection_name}'") log.info("Deleted %s items from collection '%s'", deleted, collection_name)
except Exception as e: except Exception as e:
self.session.rollback() self.session.rollback()
log.exception(f'Failed to delete data: {e}') log.exception(f'Failed to delete data: {e}')
@ -363,7 +373,7 @@ class OpenGaussClient(VectorDBBase):
try: try:
deleted = self.session.query(DocumentChunk).delete() deleted = self.session.query(DocumentChunk).delete()
self.session.commit() self.session.commit()
log.info(f'Reset completed. Deleted {deleted} items') log.info('Reset completed. Deleted %s items', deleted)
except Exception as e: except Exception as e:
self.session.rollback() self.session.rollback()
log.exception(f'Reset failed: {e}') log.exception(f'Reset failed: {e}')
@ -387,4 +397,4 @@ class OpenGaussClient(VectorDBBase):
def delete_collection(self, collection_name: str) -> None: def delete_collection(self, collection_name: str) -> None:
self.delete(collection_name) self.delete(collection_name)
log.info(f"Collection '{collection_name}' has been deleted") log.info("Collection '%s' has been deleted", collection_name)

View file

@ -2,7 +2,7 @@
NOTE: This vector database integration is community-supported and maintained on a best-effort basis. NOTE: This vector database integration is community-supported and maintained on a best-effort basis.
""" """
from typing import Optional from typing import Any, Optional
from open_webui.config import ( from open_webui.config import (
OPENSEARCH_CERT_VERIFY, OPENSEARCH_CERT_VERIFY,
@ -17,11 +17,17 @@ from open_webui.retrieval.vector.main import (
VectorDBBase, VectorDBBase,
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import process_metadata from open_webui.retrieval.vector.utils import iter_filter_conditions, process_metadata
from opensearchpy import OpenSearch from opensearchpy import OpenSearch
from opensearchpy.helpers import bulk from opensearchpy.helpers import bulk
def _metadata_filter(key: str, op: str, value: Any) -> dict:
if op == '$in':
return {'terms': {f'metadata.{key}.keyword': value}}
return {'term': {f'metadata.{key}.keyword': value}}
class OpenSearchClient(VectorDBBase): class OpenSearchClient(VectorDBBase):
def __init__(self): def __init__(self):
self.index_prefix = 'open_webui' self.index_prefix = 'open_webui'
@ -121,6 +127,8 @@ class OpenSearchClient(VectorDBBase):
filter: Optional[dict] = None, filter: Optional[dict] = None,
limit: int = 10, limit: int = 10,
) -> Optional[SearchResult]: ) -> Optional[SearchResult]:
filter_clauses = [_metadata_filter(key, op, value) for key, op, value in iter_filter_conditions(filter)]
try: try:
if not self.has_collection(collection_name): if not self.has_collection(collection_name):
return None return None
@ -130,7 +138,7 @@ class OpenSearchClient(VectorDBBase):
'_source': ['text', 'metadata'], '_source': ['text', 'metadata'],
'query': { 'query': {
'script_score': { 'script_score': {
'query': {'match_all': {}}, 'query': {'bool': {'filter': filter_clauses}} if filter_clauses else {'match_all': {}},
'script': { 'script': {
'source': '(cosineSimilarity(params.query_value, doc[params.field]) + 1.0) / 2.0', 'source': '(cosineSimilarity(params.query_value, doc[params.field]) + 1.0) / 2.0',
'params': { 'params': {

View file

@ -32,6 +32,7 @@ import array
import json import json
import logging import logging
import os import os
import re
import threading import threading
import time import time
from decimal import Decimal from decimal import Decimal
@ -56,8 +57,29 @@ from open_webui.retrieval.vector.main import (
VectorDBBase, VectorDBBase,
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import iter_filter_conditions, process_metadata
from open_webui.utils.json_codec import JSONCodec
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
_SAFE_METADATA_KEY_RE = re.compile(r'^[A-Za-z_][A-Za-z0-9_]{0,63}$')
def _metadata_where(filter: Optional[dict]) -> tuple[str, dict[str, Any]]:
clause = ''
params: dict[str, Any] = {}
for i, (key, op, value) in enumerate(iter_filter_conditions(filter)):
if not isinstance(key, str) or not _SAFE_METADATA_KEY_RE.fullmatch(key):
raise ValueError(f'Invalid Oracle metadata filter key: {key!r}')
json_value = f"JSON_VALUE(dc.vmetadata, '$.{key}' RETURNING VARCHAR2(4096))"
if op == '$in':
names = [f'value_{i}_{j}' for j, _ in enumerate(value)]
clause += f' AND {json_value} IN ({", ".join(f":{name}" for name in names)})' if names else ' AND 1 = 0'
params.update({name: str(item) for name, item in zip(names, value)})
else:
name = f'value_{i}'
clause += f' AND {json_value} = :{name}'
params[name] = str(value)
return clause, params
class Oracle23aiClient(VectorDBBase): class Oracle23aiClient(VectorDBBase):
@ -93,10 +115,10 @@ class Oracle23aiClient(VectorDBBase):
self._create_dbcs_pool() self._create_dbcs_pool()
dsn = ORACLE_DB_DSN dsn = ORACLE_DB_DSN
log.info(f'Creating Connection Pool [{ORACLE_DB_USER}:**@{dsn}]') log.info('Creating Connection Pool [%s:**@%s]', ORACLE_DB_USER, dsn)
with self.get_connection() as connection: with self.get_connection() as connection:
log.info(f'Connection version: {connection.version}') log.info('Connection version: %s', connection.version)
self._initialize_database(connection) self._initialize_database(connection)
log.info('Oracle Vector Search initialization complete.') log.info('Oracle Vector Search initialization complete.')
@ -158,7 +180,7 @@ class Oracle23aiClient(VectorDBBase):
if attempt < max_retries - 1: if attempt < max_retries - 1:
wait_time = 2**attempt wait_time = 2**attempt
log.info(f'Retrying in {wait_time} seconds...') log.info('Retrying in %s seconds...', wait_time)
time.sleep(wait_time) time.sleep(wait_time)
else: else:
raise raise
@ -183,7 +205,7 @@ class Oracle23aiClient(VectorDBBase):
thread = threading.Thread(target=_monitor, daemon=True) thread = threading.Thread(target=_monitor, daemon=True)
thread.start() thread.start()
log.info(f'Started DB health monitor every {interval_seconds} seconds.') log.info('Started DB health monitor every %s seconds.', interval_seconds)
def _reconnect_pool(self): def _reconnect_pool(self):
""" """
@ -378,7 +400,7 @@ class Oracle23aiClient(VectorDBBase):
Returns: Returns:
str: JSON representation of metadata str: JSON representation of metadata
""" """
return json.dumps(metadata, default=self._decimal_handler) if metadata else '{}' return json.dumps(process_metadata(metadata), default=self._decimal_handler) if metadata else '{}'
def _json_to_metadata(self, json_str: str) -> Dict: def _json_to_metadata(self, json_str: str) -> Dict:
""" """
@ -390,7 +412,7 @@ class Oracle23aiClient(VectorDBBase):
Returns: Returns:
Dict: Metadata dictionary Dict: Metadata dictionary
""" """
return json.loads(json_str) if json_str else {} return JSONCodec.loads(json_str) if json_str else {}
def insert(self, collection_name: str, items: List[VectorItem]) -> None: def insert(self, collection_name: str, items: List[VectorItem]) -> None:
""" """
@ -411,7 +433,7 @@ class Oracle23aiClient(VectorDBBase):
... ] ... ]
>>> client.insert("my_collection", items) >>> client.insert("my_collection", items)
""" """
log.info(f"Inserting {len(items)} items into collection '{collection_name}'.") log.info("Inserting %s items into collection '%s'.", len(items), collection_name)
with self.get_connection() as connection: with self.get_connection() as connection:
try: try:
@ -436,7 +458,7 @@ class Oracle23aiClient(VectorDBBase):
) )
connection.commit() connection.commit()
log.info(f"Successfully inserted {len(items)} items into collection '{collection_name}'.") log.info("Successfully inserted %s items into collection '%s'.", len(items), collection_name)
except Exception as e: except Exception as e:
connection.rollback() connection.rollback()
@ -465,7 +487,7 @@ class Oracle23aiClient(VectorDBBase):
... ] ... ]
>>> client.upsert("my_collection", items) >>> client.upsert("my_collection", items)
""" """
log.info(f"Upserting {len(items)} items into collection '{collection_name}'.") log.info("Upserting %s items into collection '%s'.", len(items), collection_name)
with self.get_connection() as connection: with self.get_connection() as connection:
try: try:
@ -504,7 +526,7 @@ class Oracle23aiClient(VectorDBBase):
) )
connection.commit() connection.commit()
log.info(f"Successfully upserted {len(items)} items into collection '{collection_name}'.") log.info("Successfully upserted %s items into collection '%s'.", len(items), collection_name)
except Exception as e: except Exception as e:
connection.rollback() connection.rollback()
@ -540,7 +562,7 @@ class Oracle23aiClient(VectorDBBase):
... for i, (id, dist) in enumerate(zip(results.ids[0], results.distances[0])): ... for i, (id, dist) in enumerate(zip(results.ids[0], results.distances[0])):
... log.info(f"Match {i+1}: id={id}, distance={dist}") ... log.info(f"Match {i+1}: id={id}, distance={dist}")
""" """
log.info(f"Searching items from collection '{collection_name}' with limit {limit}.") log.info("Searching items from collection '%s' with limit %s.", collection_name, limit)
try: try:
if not vectors: if not vectors:
@ -548,6 +570,7 @@ class Oracle23aiClient(VectorDBBase):
return None return None
num_queries = len(vectors) num_queries = len(vectors)
filter_clause, filter_params = _metadata_where(filter)
ids = [[] for _ in range(num_queries)] ids = [[] for _ in range(num_queries)]
distances = [[] for _ in range(num_queries)] distances = [[] for _ in range(num_queries)]
@ -560,12 +583,12 @@ class Oracle23aiClient(VectorDBBase):
vector_blob = self._vector_to_blob(vector) vector_blob = self._vector_to_blob(vector)
cursor.execute( cursor.execute(
""" f"""
SELECT dc.id, dc.text, SELECT dc.id, dc.text,
JSON_SERIALIZE(dc.vmetadata RETURNING VARCHAR2(4096)) as vmetadata, JSON_SERIALIZE(dc.vmetadata RETURNING VARCHAR2(4096)) as vmetadata,
VECTOR_DISTANCE(dc.vector, :query_vector, COSINE) as distance VECTOR_DISTANCE(dc.vector, :query_vector, COSINE) as distance
FROM document_chunk dc FROM document_chunk dc
WHERE dc.collection_name = :collection_name WHERE dc.collection_name = :collection_name{filter_clause}
ORDER BY VECTOR_DISTANCE(dc.vector, :query_vector, COSINE) ORDER BY VECTOR_DISTANCE(dc.vector, :query_vector, COSINE)
FETCH APPROX FIRST :limit ROWS ONLY FETCH APPROX FIRST :limit ROWS ONLY
""", """,
@ -573,6 +596,7 @@ class Oracle23aiClient(VectorDBBase):
'query_vector': vector_blob, 'query_vector': vector_blob,
'collection_name': collection_name, 'collection_name': collection_name,
'limit': limit, 'limit': limit,
**filter_params,
}, },
) )
@ -586,7 +610,7 @@ class Oracle23aiClient(VectorDBBase):
metadatas[qid].append(self._json_to_metadata(metadata_str)) metadatas[qid].append(self._json_to_metadata(metadata_str))
distances[qid].append(float(row[3])) distances[qid].append(float(row[3]))
log.info(f'Search completed. Found {sum(len(ids[i]) for i in range(num_queries))} total results.') log.info('Search completed. Found %s total results.', sum(len(ids[i]) for i in range(num_queries)))
return SearchResult(ids=ids, distances=distances, documents=documents, metadatas=metadatas) return SearchResult(ids=ids, distances=distances, documents=documents, metadatas=metadatas)
@ -615,7 +639,7 @@ class Oracle23aiClient(VectorDBBase):
>>> if results: >>> if results:
... print(f"Found {len(results.ids[0])} matching documents") ... print(f"Found {len(results.ids[0])} matching documents")
""" """
log.info(f"Querying items from collection '{collection_name}' with filters.") log.info("Querying items from collection '%s' with filters.", collection_name)
try: try:
limit = limit or 100 limit = limit or 100
@ -655,7 +679,7 @@ class Oracle23aiClient(VectorDBBase):
] ]
] ]
log.info(f'Query completed. Found {len(results)} results.') log.info('Query completed. Found %s results.', len(results))
return GetResult(ids=ids, documents=documents, metadatas=metadatas) return GetResult(ids=ids, documents=documents, metadatas=metadatas)
@ -746,7 +770,7 @@ class Oracle23aiClient(VectorDBBase):
>>> # Or delete by metadata filter >>> # Or delete by metadata filter
>>> client.delete("my_collection", filter={"source": "deprecated_source"}) >>> client.delete("my_collection", filter={"source": "deprecated_source"})
""" """
log.info(f"Deleting items from collection '{collection_name}'.") log.info("Deleting items from collection '%s'.", collection_name)
try: try:
query = 'DELETE FROM document_chunk WHERE collection_name = :collection_name' query = 'DELETE FROM document_chunk WHERE collection_name = :collection_name'
@ -771,7 +795,7 @@ class Oracle23aiClient(VectorDBBase):
deleted = cursor.rowcount deleted = cursor.rowcount
connection.commit() connection.commit()
log.info(f"Deleted {deleted} items from collection '{collection_name}'.") log.info("Deleted %s items from collection '%s'.", deleted, collection_name)
except Exception as e: except Exception as e:
log.exception(f'Error during delete: {e}') log.exception(f'Error during delete: {e}')
@ -799,7 +823,7 @@ class Oracle23aiClient(VectorDBBase):
deleted = cursor.rowcount deleted = cursor.rowcount
connection.commit() connection.commit()
log.info(f"Reset complete. Deleted {deleted} items from 'document_chunk' table.") log.info("Reset complete. Deleted %s items from 'document_chunk' table.", deleted)
except Exception as e: except Exception as e:
log.exception(f'Error during reset: {e}') log.exception(f'Error during reset: {e}')
@ -874,7 +898,7 @@ class Oracle23aiClient(VectorDBBase):
>>> client = Oracle23aiClient() >>> client = Oracle23aiClient()
>>> client.delete_collection("obsolete_collection") >>> client.delete_collection("obsolete_collection")
""" """
log.info(f"Deleting collection '{collection_name}'.") log.info("Deleting collection '%s'.", collection_name)
try: try:
with self.get_connection() as connection: with self.get_connection() as connection:
@ -890,7 +914,7 @@ class Oracle23aiClient(VectorDBBase):
deleted = cursor.rowcount deleted = cursor.rowcount
connection.commit() connection.commit()
log.info(f"Collection '{collection_name}' deleted. Removed {deleted} items.") log.info("Collection '%s' deleted. Removed %s items.", collection_name, deleted)
except Exception as e: except Exception as e:
log.exception(f"Error deleting collection '{collection_name}': {e}") log.exception(f"Error deleting collection '{collection_name}': {e}")

View file

@ -1,4 +1,3 @@
import json
import logging import logging
from typing import Any, Dict, List, Optional, Tuple from typing import Any, Dict, List, Optional, Tuple
@ -9,6 +8,7 @@ from open_webui.config import (
PGVECTOR_HNSW_M, PGVECTOR_HNSW_M,
PGVECTOR_INDEX_METHOD, PGVECTOR_INDEX_METHOD,
PGVECTOR_INITIALIZE_MAX_VECTOR_LENGTH, PGVECTOR_INITIALIZE_MAX_VECTOR_LENGTH,
PGVECTOR_ITERATIVE_SCAN,
PGVECTOR_IVFFLAT_LISTS, PGVECTOR_IVFFLAT_LISTS,
PGVECTOR_PGCRYPTO, PGVECTOR_PGCRYPTO,
PGVECTOR_PGCRYPTO_KEY, PGVECTOR_PGCRYPTO_KEY,
@ -18,6 +18,7 @@ from open_webui.config import (
PGVECTOR_POOL_TIMEOUT, PGVECTOR_POOL_TIMEOUT,
PGVECTOR_USE_HALFVEC, PGVECTOR_USE_HALFVEC,
) )
from open_webui.internal.db import ScopedSession, enable_iam_token_auth
from open_webui.retrieval.vector.main import ( from open_webui.retrieval.vector.main import (
GetResult, GetResult,
SearchResult, SearchResult,
@ -25,6 +26,7 @@ from open_webui.retrieval.vector.main import (
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import merge_hybrid_search_results, process_metadata from open_webui.retrieval.vector.utils import merge_hybrid_search_results, process_metadata
from open_webui.utils.json_codec import JSONCodec
from open_webui.utils.misc import sanitize_text_for_db from open_webui.utils.misc import sanitize_text_for_db
from pgvector.sqlalchemy import HALFVEC, Vector from pgvector.sqlalchemy import HALFVEC, Vector
from sqlalchemy import ( from sqlalchemy import (
@ -87,8 +89,6 @@ class PgvectorClient(VectorDBBase):
def __init__(self) -> None: def __init__(self) -> None:
# if no pgvector uri, use the existing database connection # if no pgvector uri, use the existing database connection
if not PGVECTOR_DB_URL: if not PGVECTOR_DB_URL:
from open_webui.internal.db import ScopedSession
self.session = ScopedSession self.session = ScopedSession
else: else:
if isinstance(PGVECTOR_POOL_SIZE, int): if isinstance(PGVECTOR_POOL_SIZE, int):
@ -107,6 +107,7 @@ class PgvectorClient(VectorDBBase):
else: else:
engine = create_engine(PGVECTOR_DB_URL, pool_pre_ping=True) engine = create_engine(PGVECTOR_DB_URL, pool_pre_ping=True)
enable_iam_token_auth(engine)
SessionLocal = sessionmaker(autocommit=False, autoflush=False, bind=engine, expire_on_commit=False) SessionLocal = sessionmaker(autocommit=False, autoflush=False, bind=engine, expire_on_commit=False)
self.session = scoped_session(SessionLocal) self.session = scoped_session(SessionLocal)
@ -154,6 +155,7 @@ class PgvectorClient(VectorDBBase):
index_method, index_options = self._vector_index_configuration() index_method, index_options = self._vector_index_configuration()
self._ensure_vector_index(index_method, index_options) self._ensure_vector_index(index_method, index_options)
self._ensure_text_search_index() self._ensure_text_search_index()
self.iterative_scan_sql = self._iterative_scan_setting(index_method)
self.session.execute( self.session.execute(
text( text(
@ -223,6 +225,9 @@ class PgvectorClient(VectorDBBase):
) )
if not existing_index_def: if not existing_index_def:
if index_method == 'ivfflat' and not self._has_enough_ivfflat_training_rows():
return
index_sql = ( index_sql = (
f'CREATE INDEX IF NOT EXISTS {index_name} ' f'CREATE INDEX IF NOT EXISTS {index_name} '
f'ON document_chunk USING {index_method} (vector {VECTOR_OPCLASS})' f'ON document_chunk USING {index_method} (vector {VECTOR_OPCLASS})'
@ -237,6 +242,38 @@ class PgvectorClient(VectorDBBase):
f' {index_options}' if index_options else '', f' {index_options}' if index_options else '',
) )
def _has_enough_ivfflat_training_rows(self) -> bool:
# ivfflat samples 50 rows per list to place its centroids, so recall stays poor until the table holds that many
min_training_rows = 50 * PGVECTOR_IVFFLAT_LISTS
row_count = self.session.execute(
text('SELECT count(*) FROM (SELECT 1 FROM document_chunk LIMIT :min_training_rows) AS sample'),
{'min_training_rows': min_training_rows},
).scalar()
if row_count < min_training_rows:
log.info(
"Deferring vector index 'idx_document_chunk_vector' until document_chunk holds %s rows to cluster on, "
'it has %s. Searches run as an exact scan until then.',
min_training_rows,
row_count,
)
return False
return True
def _iterative_scan_setting(self, index_method: str) -> Optional[str]:
if PGVECTOR_ITERATIVE_SCAN == 'off':
return None
version = self.session.execute(text("SELECT extversion FROM pg_extension WHERE extname = 'vector'")).scalar()
version_parts = [int(part) for part in (version or '').split('.') if part.isdigit()]
if version_parts[:2] < [0, 8]:
log.info('Iterative scan needs pgvector 0.8 or newer, the server has %s.', version or 'none')
return None
# ivfflat only accepts relaxed_order
mode = 'relaxed_order' if index_method == 'ivfflat' else PGVECTOR_ITERATIVE_SCAN
return f'SET LOCAL {index_method}.iterative_scan = {mode}'
def _ensure_text_search_index(self) -> None: def _ensure_text_search_index(self) -> None:
if PGVECTOR_PGCRYPTO: if PGVECTOR_PGCRYPTO:
return return
@ -303,7 +340,7 @@ class PgvectorClient(VectorDBBase):
# Use raw SQL for BYTEA/pgcrypto # Use raw SQL for BYTEA/pgcrypto
# Ensure metadata is converted to its JSON text representation # Ensure metadata is converted to its JSON text representation
# Sanitize to strip null bytes / surrogates that PostgreSQL cannot store # Sanitize to strip null bytes / surrogates that PostgreSQL cannot store
json_metadata = sanitize_text_for_db(json.dumps(item['metadata'])) json_metadata = sanitize_text_for_db(JSONCodec.dumps(item['metadata']))
item_text = sanitize_text_for_db(item['text']) item_text = sanitize_text_for_db(item['text'])
self.session.execute( self.session.execute(
text(""" text("""
@ -326,7 +363,7 @@ class PgvectorClient(VectorDBBase):
}, },
) )
self.session.commit() self.session.commit()
log.info(f"Encrypted & inserted {len(items)} into '{collection_name}'") log.info("Encrypted & inserted %s into '%s'", len(items), collection_name)
else: else:
new_items = [] new_items = []
@ -342,7 +379,7 @@ class PgvectorClient(VectorDBBase):
new_items.append(new_chunk) new_items.append(new_chunk)
self.session.bulk_save_objects(new_items) self.session.bulk_save_objects(new_items)
self.session.commit() self.session.commit()
log.info(f"Inserted {len(new_items)} items into collection '{collection_name}'.") log.info("Inserted %s items into collection '%s'.", len(new_items), collection_name)
except Exception as e: except Exception as e:
self.session.rollback() self.session.rollback()
log.exception(f'Error during insert: {e}') log.exception(f'Error during insert: {e}')
@ -354,7 +391,7 @@ class PgvectorClient(VectorDBBase):
for item in items: for item in items:
vector = self.adjust_vector_length(item['vector']) vector = self.adjust_vector_length(item['vector'])
# Sanitize to strip null bytes / surrogates that PostgreSQL cannot store # Sanitize to strip null bytes / surrogates that PostgreSQL cannot store
json_metadata = sanitize_text_for_db(json.dumps(item['metadata'])) json_metadata = sanitize_text_for_db(JSONCodec.dumps(item['metadata']))
item_text = sanitize_text_for_db(item['text']) item_text = sanitize_text_for_db(item['text'])
self.session.execute( self.session.execute(
text(""" text("""
@ -381,7 +418,7 @@ class PgvectorClient(VectorDBBase):
}, },
) )
self.session.commit() self.session.commit()
log.info(f"Encrypted & upserted {len(items)} into '{collection_name}'") log.info("Encrypted & upserted %s into '%s'", len(items), collection_name)
else: else:
for item in items: for item in items:
vector = self.adjust_vector_length(item['vector']) vector = self.adjust_vector_length(item['vector'])
@ -401,7 +438,7 @@ class PgvectorClient(VectorDBBase):
) )
self.session.add(new_chunk) self.session.add(new_chunk)
self.session.commit() self.session.commit()
log.info(f"Upserted {len(items)} items into collection '{collection_name}'.") log.info("Upserted %s items into collection '%s'.", len(items), collection_name)
except Exception as e: except Exception as e:
self.session.rollback() self.session.rollback()
log.exception(f'Error during upsert: {e}') log.exception(f'Error during upsert: {e}')
@ -503,6 +540,9 @@ class PgvectorClient(VectorDBBase):
.order_by(query_vectors.c.qid, subq.c.distance) .order_by(query_vectors.c.qid, subq.c.distance)
) )
if self.iterative_scan_sql:
self.session.execute(text(self.iterative_scan_sql))
result_proxy = self.session.execute(stmt) result_proxy = self.session.execute(stmt)
results = result_proxy.all() results = result_proxy.all()
@ -512,6 +552,7 @@ class PgvectorClient(VectorDBBase):
metadatas = [[] for _ in range(num_queries)] metadatas = [[] for _ in range(num_queries)]
if not results: if not results:
self.session.rollback()
return SearchResult( return SearchResult(
ids=ids, ids=ids,
distances=distances, distances=distances,
@ -631,6 +672,7 @@ class PgvectorClient(VectorDBBase):
results = query.all() results = query.all()
if not results: if not results:
self.session.rollback()
return None return None
ids = [[result.id for result in results]] ids = [[result.id for result in results]]
@ -670,6 +712,7 @@ class PgvectorClient(VectorDBBase):
results = query.all() results = query.all()
if not results: if not results:
self.session.rollback()
return None return None
ids = [[result.id for result in results]] ids = [[result.id for result in results]]
@ -712,7 +755,7 @@ class PgvectorClient(VectorDBBase):
query = query.filter(DocumentChunk.vmetadata[key].astext == str(value)) query = query.filter(DocumentChunk.vmetadata[key].astext == str(value))
deleted = query.delete(synchronize_session=False) deleted = query.delete(synchronize_session=False)
self.session.commit() self.session.commit()
log.info(f"Deleted {deleted} items from collection '{collection_name}'.") log.info("Deleted %s items from collection '%s'.", deleted, collection_name)
except Exception as e: except Exception as e:
self.session.rollback() self.session.rollback()
log.exception(f'Error during delete: {e}') log.exception(f'Error during delete: {e}')
@ -722,7 +765,7 @@ class PgvectorClient(VectorDBBase):
try: try:
deleted = self.session.query(DocumentChunk).delete() deleted = self.session.query(DocumentChunk).delete()
self.session.commit() self.session.commit()
log.info(f"Reset complete. Deleted {deleted} items from 'document_chunk' table.") log.info("Reset complete. Deleted %s items from 'document_chunk' table.", deleted)
except Exception as e: except Exception as e:
self.session.rollback() self.session.rollback()
log.exception(f'Error during reset: {e}') log.exception(f'Error during reset: {e}')
@ -746,4 +789,4 @@ class PgvectorClient(VectorDBBase):
def delete_collection(self, collection_name: str) -> None: def delete_collection(self, collection_name: str) -> None:
self.delete(collection_name) self.delete(collection_name)
log.info(f"Collection '{collection_name}' deleted.") log.info("Collection '%s' deleted.", collection_name)

View file

@ -35,7 +35,7 @@ from open_webui.retrieval.vector.main import (
VectorDBBase, VectorDBBase,
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import process_metadata from open_webui.retrieval.vector.utils import normalize_filter, process_metadata
NO_LIMIT = 10000 # Reasonable limit to avoid overwhelming the system NO_LIMIT = 10000 # Reasonable limit to avoid overwhelming the system
BATCH_SIZE = 100 # Recommended batch size for Pinecone operations BATCH_SIZE = 100 # Recommended batch size for Pinecone operations
@ -106,16 +106,16 @@ class PineconeClient(VectorDBBase):
try: try:
# Check if index exists # Check if index exists
if self.index_name not in self.client.list_indexes().names(): if self.index_name not in self.client.list_indexes().names():
log.info(f"Creating Pinecone index '{self.index_name}'...") log.info("Creating Pinecone index '%s'...", self.index_name)
self.client.create_index( self.client.create_index(
name=self.index_name, name=self.index_name,
dimension=self.dimension, dimension=self.dimension,
metric=self.metric, metric=self.metric,
spec=ServerlessSpec(cloud=self.cloud, region=self.environment), spec=ServerlessSpec(cloud=self.cloud, region=self.environment),
) )
log.info(f"Successfully created Pinecone index '{self.index_name}'") log.info("Successfully created Pinecone index '%s'", self.index_name)
else: else:
log.info(f"Using existing Pinecone index '{self.index_name}'") log.info("Using existing Pinecone index '%s'", self.index_name)
# Connect to the index # Connect to the index
self.index = self.client.Index( self.index = self.client.Index(
@ -245,7 +245,7 @@ class PineconeClient(VectorDBBase):
collection_name_with_prefix = self._get_collection_name_with_prefix(collection_name) collection_name_with_prefix = self._get_collection_name_with_prefix(collection_name)
try: try:
self.index.delete(filter={'collection_name': collection_name_with_prefix}) self.index.delete(filter={'collection_name': collection_name_with_prefix})
log.info(f"Collection '{collection_name_with_prefix}' deleted (all vectors removed).") log.info("Collection '%s' deleted (all vectors removed).", collection_name_with_prefix)
except Exception as e: except Exception as e:
log.warning(f"Failed to delete collection '{collection_name_with_prefix}': {e}") log.warning(f"Failed to delete collection '{collection_name_with_prefix}': {e}")
raise raise
@ -274,9 +274,9 @@ class PineconeClient(VectorDBBase):
log.error(f'Error inserting batch: {e}') log.error(f'Error inserting batch: {e}')
raise raise
elapsed = time.time() - start_time elapsed = time.time() - start_time
log.debug(f'Insert of {len(points)} vectors took {elapsed:.2f} seconds') log.debug('Insert of %s vectors took %.2f seconds', len(points), elapsed)
log.info( log.info(
f"Successfully inserted {len(points)} vectors in parallel batches into '{collection_name_with_prefix}'" "Successfully inserted %s vectors in parallel batches into '%s'", len(points), collection_name_with_prefix
) )
def upsert(self, collection_name: str, items: List[VectorItem]) -> None: def upsert(self, collection_name: str, items: List[VectorItem]) -> None:
@ -303,9 +303,9 @@ class PineconeClient(VectorDBBase):
log.error(f'Error upserting batch: {e}') log.error(f'Error upserting batch: {e}')
raise raise
elapsed = time.time() - start_time elapsed = time.time() - start_time
log.debug(f'Upsert of {len(points)} vectors took {elapsed:.2f} seconds') log.debug('Upsert of %s vectors took %.2f seconds', len(points), elapsed)
log.info( log.info(
f"Successfully upserted {len(points)} vectors in parallel batches into '{collection_name_with_prefix}'" "Successfully upserted %s vectors in parallel batches into '%s'", len(points), collection_name_with_prefix
) )
async def insert_async(self, collection_name: str, items: List[VectorItem]) -> None: async def insert_async(self, collection_name: str, items: List[VectorItem]) -> None:
@ -326,7 +326,9 @@ class PineconeClient(VectorDBBase):
if isinstance(result, Exception): if isinstance(result, Exception):
log.error(f'Error in async insert batch: {result}') log.error(f'Error in async insert batch: {result}')
raise result raise result
log.info(f"Successfully async inserted {len(points)} vectors in batches into '{collection_name_with_prefix}'") log.info(
"Successfully async inserted %s vectors in batches into '%s'", len(points), collection_name_with_prefix
)
async def upsert_async(self, collection_name: str, items: List[VectorItem]) -> None: async def upsert_async(self, collection_name: str, items: List[VectorItem]) -> None:
"""Async version of upsert using asyncio and run_in_executor for improved performance.""" """Async version of upsert using asyncio and run_in_executor for improved performance."""
@ -346,7 +348,9 @@ class PineconeClient(VectorDBBase):
if isinstance(result, Exception): if isinstance(result, Exception):
log.error(f'Error in async upsert batch: {result}') log.error(f'Error in async upsert batch: {result}')
raise result raise result
log.info(f"Successfully async upserted {len(points)} vectors in batches into '{collection_name_with_prefix}'") log.info(
"Successfully async upserted %s vectors in batches into '%s'", len(points), collection_name_with_prefix
)
def search( def search(
self, self,
@ -368,13 +372,15 @@ class PineconeClient(VectorDBBase):
try: try:
# Search using the first vector (assuming this is the intended behavior) # Search using the first vector (assuming this is the intended behavior)
query_vector = vectors[0] query_vector = vectors[0]
pinecone_filter = normalize_filter(filter)
pinecone_filter['collection_name'] = collection_name_with_prefix
# Perform the search # Perform the search
query_response = self.index.query( query_response = self.index.query(
vector=query_vector, vector=query_vector,
top_k=limit, top_k=limit,
include_metadata=True, include_metadata=True,
filter={'collection_name': collection_name_with_prefix}, filter=pinecone_filter,
) )
matches = getattr(query_response, 'matches', []) or [] matches = getattr(query_response, 'matches', []) or []
@ -474,8 +480,10 @@ class PineconeClient(VectorDBBase):
# Note: When deleting by ID, we can't filter by collection_name # Note: When deleting by ID, we can't filter by collection_name
# This is a limitation of Pinecone - be careful with ID uniqueness # This is a limitation of Pinecone - be careful with ID uniqueness
self.index.delete(ids=batch_ids) self.index.delete(ids=batch_ids)
log.debug(f"Deleted batch of {len(batch_ids)} vectors by ID from '{collection_name_with_prefix}'") log.debug(
log.info(f"Successfully deleted {len(ids)} vectors by ID from '{collection_name_with_prefix}'") "Deleted batch of %s vectors by ID from '%s'", len(batch_ids), collection_name_with_prefix
)
log.info("Successfully deleted %s vectors by ID from '%s'", len(ids), collection_name_with_prefix)
elif filter: elif filter:
# Combine user filter with collection_name # Combine user filter with collection_name
@ -484,7 +492,7 @@ class PineconeClient(VectorDBBase):
pinecone_filter.update(filter) pinecone_filter.update(filter)
# Delete by metadata filter # Delete by metadata filter
self.index.delete(filter=pinecone_filter) self.index.delete(filter=pinecone_filter)
log.info(f"Successfully deleted vectors by filter from '{collection_name_with_prefix}'") log.info("Successfully deleted vectors by filter from '%s'", collection_name_with_prefix)
else: else:
log.warning('No ids or filter provided for delete operation') log.warning('No ids or filter provided for delete operation')

View file

@ -3,7 +3,7 @@ NOTE: This vector database integration is community-supported and maintained on
""" """
import logging import logging
from typing import Optional from typing import Any, Optional
from urllib.parse import urlparse from urllib.parse import urlparse
from open_webui.config import ( from open_webui.config import (
@ -22,6 +22,7 @@ from open_webui.retrieval.vector.main import (
VectorDBBase, VectorDBBase,
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import iter_filter_conditions, process_metadata
from qdrant_client import QdrantClient as Qclient from qdrant_client import QdrantClient as Qclient
from qdrant_client.http.models import PointStruct from qdrant_client.http.models import PointStruct
from qdrant_client.models import models from qdrant_client.models import models
@ -31,6 +32,11 @@ NO_LIMIT = 999999999
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
def _metadata_filter(key: str, op: str, value: Any) -> models.FieldCondition:
match = models.MatchAny(any=value) if op == '$in' else models.MatchValue(value=value)
return models.FieldCondition(key=f'metadata.{key}', match=match)
class QdrantClient(VectorDBBase): class QdrantClient(VectorDBBase):
def __init__(self): def __init__(self):
self.collection_prefix = QDRANT_COLLECTION_PREFIX self.collection_prefix = QDRANT_COLLECTION_PREFIX
@ -119,7 +125,7 @@ class QdrantClient(VectorDBBase):
on_disk=self.QDRANT_ON_DISK, on_disk=self.QDRANT_ON_DISK,
), ),
) )
log.info(f'collection {collection_name_with_prefix} successfully created!') log.info('collection %s successfully created!', collection_name_with_prefix)
def _create_collection_if_not_exists(self, collection_name, dimension): def _create_collection_if_not_exists(self, collection_name, dimension):
if not self.has_collection(collection_name=collection_name): if not self.has_collection(collection_name=collection_name):
@ -130,7 +136,7 @@ class QdrantClient(VectorDBBase):
PointStruct( PointStruct(
id=item['id'], id=item['id'],
vector=item['vector'], vector=item['vector'],
payload={'text': item['text'], 'metadata': item['metadata']}, payload={'text': item['text'], 'metadata': process_metadata(item['metadata'])},
) )
for item in items for item in items
] ]
@ -152,10 +158,13 @@ class QdrantClient(VectorDBBase):
if limit is None: if limit is None:
limit = NO_LIMIT # otherwise qdrant would set limit to 10! limit = NO_LIMIT # otherwise qdrant would set limit to 10!
conditions = [_metadata_filter(key, op, value) for key, op, value in iter_filter_conditions(filter)]
query_filter = models.Filter(must=conditions) if conditions else None
query_response = self.client.query_points( query_response = self.client.query_points(
collection_name=f'{self.collection_prefix}_{collection_name}', collection_name=f'{self.collection_prefix}_{collection_name}',
query=vectors[0], query=vectors[0],
limit=limit, limit=limit,
query_filter=query_filter,
) )
get_result = self._result_to_get_result(query_response.points) get_result = self._result_to_get_result(query_response.points)
return SearchResult( return SearchResult(

View file

@ -23,6 +23,7 @@ from open_webui.retrieval.vector.main import (
VectorDBBase, VectorDBBase,
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import iter_filter_conditions, process_metadata
from qdrant_client import QdrantClient as Qclient from qdrant_client import QdrantClient as Qclient
from qdrant_client.http.exceptions import UnexpectedResponse from qdrant_client.http.exceptions import UnexpectedResponse
from qdrant_client.http.models import PointStruct from qdrant_client.http.models import PointStruct
@ -39,8 +40,9 @@ def _tenant_filter(tenant_id: str) -> models.FieldCondition:
return models.FieldCondition(key=TENANT_ID_FIELD, match=models.MatchValue(value=tenant_id)) return models.FieldCondition(key=TENANT_ID_FIELD, match=models.MatchValue(value=tenant_id))
def _metadata_filter(key: str, value: Any) -> models.FieldCondition: def _metadata_filter(key: str, op: str, value: Any) -> models.FieldCondition:
return models.FieldCondition(key=f'metadata.{key}', match=models.MatchValue(value=value)) match = models.MatchAny(any=value) if op == '$in' else models.MatchValue(value=value)
return models.FieldCondition(key=f'metadata.{key}', match=match)
class QdrantClient(VectorDBBase): class QdrantClient(VectorDBBase):
@ -148,7 +150,7 @@ class QdrantClient(VectorDBBase):
m=0, m=0,
), ),
) )
log.info(f'Multi-tenant collection {mt_collection_name} created with dimension {dimension}!') log.info('Multi-tenant collection %s created with dimension %s!', mt_collection_name, dimension)
self.client.create_payload_index( self.client.create_payload_index(
collection_name=mt_collection_name, collection_name=mt_collection_name,
@ -180,7 +182,7 @@ class QdrantClient(VectorDBBase):
vector=item['vector'], vector=item['vector'],
payload={ payload={
'text': item['text'], 'text': item['text'],
'metadata': item['metadata'], 'metadata': process_metadata(item['metadata']),
TENANT_ID_FIELD: tenant_id, TENANT_ID_FIELD: tenant_id,
}, },
) )
@ -224,7 +226,7 @@ class QdrantClient(VectorDBBase):
mt_collection, tenant_id = self._get_collection_and_tenant_id(collection_name) mt_collection, tenant_id = self._get_collection_and_tenant_id(collection_name)
if not self.client.collection_exists(collection_name=mt_collection): if not self.client.collection_exists(collection_name=mt_collection):
log.debug(f"Collection {mt_collection} doesn't exist, nothing to delete") log.debug("Collection %s doesn't exist, nothing to delete", mt_collection)
return None return None
must_conditions = [_tenant_filter(tenant_id)] must_conditions = [_tenant_filter(tenant_id)]
@ -234,7 +236,7 @@ class QdrantClient(VectorDBBase):
# whose payload omits an id (e.g. memories), leaving orphaned vectors. # whose payload omits an id (e.g. memories), leaving orphaned vectors.
must_conditions.append(models.HasIdCondition(has_id=ids)) must_conditions.append(models.HasIdCondition(has_id=ids))
elif filter: elif filter:
must_conditions += [_metadata_filter(k, v) for k, v in filter.items()] must_conditions += [_metadata_filter(k, '$eq', v) for k, v in filter.items()]
return self.client.delete( return self.client.delete(
collection_name=mt_collection, collection_name=mt_collection,
@ -255,15 +257,17 @@ class QdrantClient(VectorDBBase):
return None return None
mt_collection, tenant_id = self._get_collection_and_tenant_id(collection_name) mt_collection, tenant_id = self._get_collection_and_tenant_id(collection_name)
if not self.client.collection_exists(collection_name=mt_collection): if not self.client.collection_exists(collection_name=mt_collection):
log.debug(f"Collection {mt_collection} doesn't exist, search returns None") log.debug("Collection %s doesn't exist, search returns None", mt_collection)
return None return None
tenant_filter = _tenant_filter(tenant_id) conditions = [_tenant_filter(tenant_id)]
if filter:
conditions.extend(_metadata_filter(key, op, value) for key, op, value in iter_filter_conditions(filter))
query_response = self.client.query_points( query_response = self.client.query_points(
collection_name=mt_collection, collection_name=mt_collection,
query=vectors[0], query=vectors[0],
limit=limit, limit=limit,
query_filter=models.Filter(must=[tenant_filter]), query_filter=models.Filter(must=conditions),
) )
get_result = self._result_to_get_result(query_response.points) get_result = self._result_to_get_result(query_response.points)
return SearchResult( return SearchResult(
@ -281,12 +285,12 @@ class QdrantClient(VectorDBBase):
return None return None
mt_collection, tenant_id = self._get_collection_and_tenant_id(collection_name) mt_collection, tenant_id = self._get_collection_and_tenant_id(collection_name)
if not self.client.collection_exists(collection_name=mt_collection): if not self.client.collection_exists(collection_name=mt_collection):
log.debug(f"Collection {mt_collection} doesn't exist, query returns None") log.debug("Collection %s doesn't exist, query returns None", mt_collection)
return None return None
if limit is None: if limit is None:
limit = NO_LIMIT limit = NO_LIMIT
tenant_filter = _tenant_filter(tenant_id) tenant_filter = _tenant_filter(tenant_id)
field_conditions = [_metadata_filter(k, v) for k, v in filter.items()] field_conditions = [_metadata_filter(k, '$eq', v) for k, v in filter.items()]
combined_filter = models.Filter(must=[tenant_filter, *field_conditions]) combined_filter = models.Filter(must=[tenant_filter, *field_conditions])
points = self.client.scroll( points = self.client.scroll(
collection_name=mt_collection, collection_name=mt_collection,
@ -303,7 +307,7 @@ class QdrantClient(VectorDBBase):
return None return None
mt_collection, tenant_id = self._get_collection_and_tenant_id(collection_name) mt_collection, tenant_id = self._get_collection_and_tenant_id(collection_name)
if not self.client.collection_exists(collection_name=mt_collection): if not self.client.collection_exists(collection_name=mt_collection):
log.debug(f"Collection {mt_collection} doesn't exist, get returns None") log.debug("Collection %s doesn't exist, get returns None", mt_collection)
return None return None
tenant_filter = _tenant_filter(tenant_id) tenant_filter = _tenant_filter(tenant_id)
points = self.client.scroll( points = self.client.scroll(
@ -350,7 +354,7 @@ class QdrantClient(VectorDBBase):
return None return None
mt_collection, tenant_id = self._get_collection_and_tenant_id(collection_name) mt_collection, tenant_id = self._get_collection_and_tenant_id(collection_name)
if not self.client.collection_exists(collection_name=mt_collection): if not self.client.collection_exists(collection_name=mt_collection):
log.debug(f"Collection {mt_collection} doesn't exist, nothing to delete") log.debug("Collection %s doesn't exist, nothing to delete", mt_collection)
return None return None
self.client.delete( self.client.delete(
collection_name=mt_collection, collection_name=mt_collection,

View file

@ -13,7 +13,7 @@ from open_webui.retrieval.vector.main import (
VectorDBBase, VectorDBBase,
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import process_metadata from open_webui.retrieval.vector.utils import metadata_matches_filter, normalize_filter, process_metadata
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@ -36,7 +36,7 @@ class S3VectorClient(VectorDBBase):
if self.bucket_name and self.region: if self.bucket_name and self.region:
try: try:
self.client = boto3.client('s3vectors', region_name=self.region) self.client = boto3.client('s3vectors', region_name=self.region)
log.info(f"S3Vector client initialized for bucket '{self.bucket_name}' in region '{self.region}'") log.info("S3Vector client initialized for bucket '%s' in region '%s'", self.bucket_name, self.region)
except Exception as e: except Exception as e:
log.error(f'Failed to initialize S3Vector client: {e}') log.error(f'Failed to initialize S3Vector client: {e}')
self.client = None self.client = None
@ -54,7 +54,7 @@ class S3VectorClient(VectorDBBase):
Create a new index in the S3 vector bucket for the given collection if it does not exist. Create a new index in the S3 vector bucket for the given collection if it does not exist.
""" """
if self.has_collection(index_name): if self.has_collection(index_name):
log.debug(f"Index '{index_name}' already exists, skipping creation") log.debug("Index '%s' already exists, skipping creation", index_name)
return return
try: try:
@ -70,7 +70,9 @@ class S3VectorClient(VectorDBBase):
] ]
}, },
) )
log.info(f'Created S3 index: {index_name} (dim={dimension}, type={data_type}, metric={distance_metric})') log.info(
'Created S3 index: %s (dim=%s, type=%s, metric=%s)', index_name, dimension, data_type, distance_metric
)
except Exception as e: except Exception as e:
log.error(f"Error creating S3 index '{index_name}': {e}") log.error(f"Error creating S3 index '{index_name}': {e}")
raise raise
@ -137,9 +139,9 @@ class S3VectorClient(VectorDBBase):
return return
try: try:
log.info(f"Deleting collection '{collection_name}'") log.info("Deleting collection '%s'", collection_name)
self.client.delete_index(vectorBucketName=self.bucket_name, indexName=collection_name) self.client.delete_index(vectorBucketName=self.bucket_name, indexName=collection_name)
log.info(f"Successfully deleted collection '{collection_name}'") log.info("Successfully deleted collection '%s'", collection_name)
except Exception as e: except Exception as e:
log.error(f"Error deleting collection '{collection_name}': {e}") log.error(f"Error deleting collection '{collection_name}': {e}")
raise raise
@ -156,7 +158,7 @@ class S3VectorClient(VectorDBBase):
try: try:
if not self.has_collection(collection_name): if not self.has_collection(collection_name):
log.info(f"Index '{collection_name}' does not exist. Creating index.") log.info("Index '%s' does not exist. Creating index.", collection_name)
self._create_index( self._create_index(
index_name=collection_name, index_name=collection_name,
dimension=dimension, dimension=dimension,
@ -202,9 +204,11 @@ class S3VectorClient(VectorDBBase):
indexName=collection_name, indexName=collection_name,
vectors=batch, vectors=batch,
) )
log.info(f"Inserted batch {i // batch_size + 1}: {len(batch)} vectors into index '{collection_name}'.") log.info(
"Inserted batch %s: %s vectors into index '%s'.", i // batch_size + 1, len(batch), collection_name
)
log.info(f"Completed insertion of {len(vectors)} vectors into index '{collection_name}'.") log.info("Completed insertion of %s vectors into index '%s'.", len(vectors), collection_name)
except Exception as e: except Exception as e:
log.error(f'Error inserting vectors: {e}') log.error(f'Error inserting vectors: {e}')
raise raise
@ -218,11 +222,11 @@ class S3VectorClient(VectorDBBase):
return return
dimension = len(items[0]['vector']) dimension = len(items[0]['vector'])
log.info(f'Upsert dimension: {dimension}') log.info('Upsert dimension: %s', dimension)
try: try:
if not self.has_collection(collection_name): if not self.has_collection(collection_name):
log.info(f"Index '{collection_name}' does not exist. Creating index for upsert.") log.info("Index '%s' does not exist. Creating index for upsert.", collection_name)
self._create_index( self._create_index(
index_name=collection_name, index_name=collection_name,
dimension=dimension, dimension=dimension,
@ -264,10 +268,14 @@ class S3VectorClient(VectorDBBase):
batch = vectors[i : i + batch_size] batch = vectors[i : i + batch_size]
if i == 0: # Log sample info for first batch only if i == 0: # Log sample info for first batch only
log.info( log.info(
f'Upserting batch 1: {len(batch)} vectors. First vector sample: key={batch[0]["key"]}, data_type={type(batch[0]["data"]["float32"])}, data_len={len(batch[0]["data"]["float32"])}' 'Upserting batch 1: %s vectors. First vector sample: key=%s, data_type=%s, data_len=%s',
len(batch),
batch[0]['key'],
type(batch[0]['data']['float32']),
len(batch[0]['data']['float32']),
) )
else: else:
log.info(f'Upserting batch {i // batch_size + 1}: {len(batch)} vectors.') log.info('Upserting batch %s: %s vectors.', i // batch_size + 1, len(batch))
self.client.put_vectors( self.client.put_vectors(
vectorBucketName=self.bucket_name, vectorBucketName=self.bucket_name,
@ -275,7 +283,7 @@ class S3VectorClient(VectorDBBase):
vectors=batch, vectors=batch,
) )
log.info(f"Completed upsert of {len(vectors)} vectors into index '{collection_name}'.") log.info("Completed upsert of %s vectors into index '%s'.", len(vectors), collection_name)
except Exception as e: except Exception as e:
log.error(f'Error upserting vectors: {e}') log.error(f'Error upserting vectors: {e}')
raise raise
@ -300,7 +308,8 @@ class S3VectorClient(VectorDBBase):
return None return None
try: try:
log.info(f"Searching collection '{collection_name}' with {len(vectors)} query vectors, limit={limit}") log.info("Searching collection '%s' with %s query vectors, limit=%s", collection_name, len(vectors), limit)
vector_filter = normalize_filter(filter)
# Initialize result lists # Initialize result lists
all_ids = [] all_ids = []
@ -310,20 +319,23 @@ class S3VectorClient(VectorDBBase):
# Process each query vector # Process each query vector
for i, query_vector in enumerate(vectors): for i, query_vector in enumerate(vectors):
log.debug(f'Processing query vector {i + 1}/{len(vectors)}') log.debug('Processing query vector %s/%s', i + 1, len(vectors))
# Prepare the query vector in S3 Vector format # Prepare the query vector in S3 Vector format
query_vector_dict = {'float32': [float(x) for x in query_vector]} query_vector_dict = {'float32': [float(x) for x in query_vector]}
# Call S3 Vector query API request_params = {
response = self.client.query_vectors( 'vectorBucketName': self.bucket_name,
vectorBucketName=self.bucket_name, 'indexName': collection_name,
indexName=collection_name, 'topK': limit,
topK=limit, 'queryVector': query_vector_dict,
queryVector=query_vector_dict, 'returnMetadata': True,
returnMetadata=True, 'returnDistance': True,
returnDistance=True, }
) if vector_filter:
request_params['filter'] = vector_filter
response = self.client.query_vectors(**request_params)
# Process results for this query # Process results for this query
query_ids = [] query_ids = []
@ -338,6 +350,9 @@ class S3VectorClient(VectorDBBase):
vector_metadata = vector.get('metadata', {}) vector_metadata = vector.get('metadata', {})
vector_distance = vector.get('distance', 0.0) vector_distance = vector.get('distance', 0.0)
if vector_filter and not metadata_matches_filter(vector_metadata, vector_filter):
continue
# Extract document text from metadata # Extract document text from metadata
document_text = '' document_text = ''
if isinstance(vector_metadata, dict): if isinstance(vector_metadata, dict):
@ -362,7 +377,7 @@ class S3VectorClient(VectorDBBase):
all_metadatas.append(query_metadatas) all_metadatas.append(query_metadatas)
all_distances.append(query_distances) all_distances.append(query_distances)
log.info(f'Search completed. Found results for {len(all_ids)} queries') log.info('Search completed. Found results for %s queries', len(all_ids))
# Return SearchResult format # Return SearchResult format
return SearchResult( return SearchResult(
@ -402,7 +417,7 @@ class S3VectorClient(VectorDBBase):
return self.get(collection_name) return self.get(collection_name)
try: try:
log.info(f"Querying collection '{collection_name}' with filter: {filter}") log.info("Querying collection '%s' with filter: %s", collection_name, filter)
# For S3 Vector, we need to use list_vectors and then filter results # For S3 Vector, we need to use list_vectors and then filter results
# Since S3 Vector may not support complex server-side filtering, # Since S3 Vector may not support complex server-side filtering,
@ -437,7 +452,7 @@ class S3VectorClient(VectorDBBase):
if limit and len(filtered_ids) >= limit: if limit and len(filtered_ids) >= limit:
break break
log.info(f'Filter applied: {len(filtered_ids)} vectors match out of {len(all_ids)} total') log.info('Filter applied: %s vectors match out of %s total', len(filtered_ids), len(all_ids))
# Return GetResult format # Return GetResult format
if filtered_ids: if filtered_ids:
@ -472,7 +487,7 @@ class S3VectorClient(VectorDBBase):
return GetResult(ids=[[]], documents=[[]], metadatas=[[]]) return GetResult(ids=[[]], documents=[[]], metadatas=[[]])
try: try:
log.info(f"Retrieving all vectors from collection '{collection_name}'") log.info("Retrieving all vectors from collection '%s'", collection_name)
# Initialize result lists # Initialize result lists
all_ids = [] all_ids = []
@ -521,7 +536,7 @@ class S3VectorClient(VectorDBBase):
) )
# Log the actual content for debugging # Log the actual content for debugging
log.debug(f'Document text preview (first 200 chars): {str(document_text)[:200]}') log.debug('Document text preview (first 200 chars): %s', str(document_text)[:200])
else: else:
document_text = vector_id document_text = vector_id
@ -534,7 +549,7 @@ class S3VectorClient(VectorDBBase):
if not next_token: if not next_token:
break break
log.info(f"Retrieved {len(all_ids)} vectors from collection '{collection_name}'") log.info("Retrieved %s vectors from collection '%s'", len(all_ids), collection_name)
# Return in GetResult format # Return in GetResult format
# The Open WebUI GetResult expects lists of lists, so we wrap each list # The Open WebUI GetResult expects lists of lists, so we wrap each list
@ -576,17 +591,17 @@ class S3VectorClient(VectorDBBase):
try: try:
if ids: if ids:
# Delete by specific vector IDs/keys # Delete by specific vector IDs/keys
log.info(f"Deleting {len(ids)} vectors by IDs from collection '{collection_name}'") log.info("Deleting %s vectors by IDs from collection '%s'", len(ids), collection_name)
self.client.delete_vectors( self.client.delete_vectors(
vectorBucketName=self.bucket_name, vectorBucketName=self.bucket_name,
indexName=collection_name, indexName=collection_name,
keys=ids, keys=ids,
) )
log.info(f"Deleted {len(ids)} vectors from index '{collection_name}'") log.info("Deleted %s vectors from index '%s'", len(ids), collection_name)
elif filter: elif filter:
# Handle filter-based deletion # Handle filter-based deletion
log.info(f"Deleting vectors by filter from collection '{collection_name}': {filter}") log.info("Deleting vectors by filter from collection '%s': %s", collection_name, filter)
# If this is a knowledge collection and we have a file_id filter, # If this is a knowledge collection and we have a file_id filter,
# also clean up the corresponding file-specific collection # also clean up the corresponding file-specific collection
@ -595,7 +610,8 @@ class S3VectorClient(VectorDBBase):
file_collection_name = f'file-{file_id}' file_collection_name = f'file-{file_id}'
if self.has_collection(file_collection_name): if self.has_collection(file_collection_name):
log.info( log.info(
f"Found related file-specific collection '{file_collection_name}', deleting it to prevent duplicates" "Found related file-specific collection '%s', deleting it to prevent duplicates",
file_collection_name,
) )
self.delete_collection(file_collection_name) self.delete_collection(file_collection_name)
@ -604,7 +620,7 @@ class S3VectorClient(VectorDBBase):
query_result = self.query(collection_name, filter) query_result = self.query(collection_name, filter)
if query_result and query_result.ids and query_result.ids[0]: if query_result and query_result.ids and query_result.ids[0]:
matching_ids = query_result.ids[0] matching_ids = query_result.ids[0]
log.info(f'Found {len(matching_ids)} vectors matching filter, deleting them') log.info('Found %s vectors matching filter, deleting them', len(matching_ids))
# Delete the matching vectors by ID # Delete the matching vectors by ID
self.client.delete_vectors( self.client.delete_vectors(
@ -612,7 +628,7 @@ class S3VectorClient(VectorDBBase):
indexName=collection_name, indexName=collection_name,
keys=matching_ids, keys=matching_ids,
) )
log.info(f"Deleted {len(matching_ids)} vectors from index '{collection_name}' using filter") log.info("Deleted %s vectors from index '%s' using filter", len(matching_ids), collection_name)
else: else:
log.warning('No vectors found matching the filter criteria') log.warning('No vectors found matching the filter criteria')
else: else:
@ -645,11 +661,11 @@ class S3VectorClient(VectorDBBase):
try: try:
self.client.delete_index(vectorBucketName=self.bucket_name, indexName=index_name) self.client.delete_index(vectorBucketName=self.bucket_name, indexName=index_name)
deleted_count += 1 deleted_count += 1
log.info(f'Deleted index: {index_name}') log.info('Deleted index: %s', index_name)
except Exception as e: except Exception as e:
log.error(f"Error deleting index '{index_name}': {e}") log.error(f"Error deleting index '{index_name}': {e}")
log.info(f'Reset completed: deleted {deleted_count} indexes') log.info('Reset completed: deleted %s indexes', deleted_count)
except Exception as e: except Exception as e:
log.error(f'Error during reset: {e}') log.error(f'Error during reset: {e}')

View file

@ -2,7 +2,6 @@
# Requires Valkey core >= 9.0.1 with the valkey-search module >= 1.2.0 loaded. # Requires Valkey core >= 9.0.1 with the valkey-search module >= 1.2.0 loaded.
import atexit import atexit
import json
import logging import logging
import re import re
import struct import struct
@ -24,6 +23,7 @@ from open_webui.retrieval.vector.main import (
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import process_metadata from open_webui.retrieval.vector.utils import process_metadata
from open_webui.utils.json_codec import JSONCodec
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@ -279,7 +279,7 @@ class ValkeyClient(VectorDBBase):
f'{self._format_version(MIN_VALKEY_VERSION)}. valkey-search 1.2.0 requires Valkey core ' f'{self._format_version(MIN_VALKEY_VERSION)}. valkey-search 1.2.0 requires Valkey core '
'9.0.1 or later. Upgrade your server or use valkey-bundle:9.1.0-rc2+.' '9.0.1 or later. Upgrade your server or use valkey-bundle:9.1.0-rc2+.'
) )
log.info(f'Valkey core version: {self._format_version(version) if version else "unknown"}') log.info('Valkey core version: %s', self._format_version(version) if version else 'unknown')
def _check_search_module(self) -> None: def _check_search_module(self) -> None:
try: try:
@ -331,7 +331,7 @@ class ValkeyClient(VectorDBBase):
'TEXT field type and filter-only FT.SEARCH support required by this backend. ' 'TEXT field type and filter-only FT.SEARCH support required by this backend. '
'Upgrade to valkey-bundle:9.1.0-rc2+ or load valkey-search 1.2.0+ as a module.' 'Upgrade to valkey-bundle:9.1.0-rc2+ or load valkey-search 1.2.0+ as a module.'
) )
log.info(f'valkey-search version: {self._format_version(search_version) if search_version else "unknown"}') log.info('valkey-search version: %s', self._format_version(search_version) if search_version else 'unknown')
def _index_name(self, collection_name: str) -> str: def _index_name(self, collection_name: str) -> str:
return f'idx:{self.collection_prefix}:{collection_name}' return f'idx:{self.collection_prefix}:{collection_name}'
@ -385,12 +385,15 @@ class ValkeyClient(VectorDBBase):
try: try:
g['glide_ft'].create(self.client, index_name, schema, options) g['glide_ft'].create(self.client, index_name, schema, options)
log.info( log.info(
f'Created Valkey index {index_name} with dimension={dimension}, ' 'Created Valkey index %s with dimension=%s, type=%s, metric=%s',
f'type={self.index_type}, metric={self.distance_metric}' index_name,
dimension,
self.index_type,
self.distance_metric,
) )
except g['RequestError'] as e: except g['RequestError'] as e:
if 'already exists' in str(e).lower(): if 'already exists' in str(e).lower():
log.debug(f'Index {index_name} already exists, skipping creation.') log.debug('Index %s already exists, skipping creation.', index_name)
else: else:
raise raise
@ -456,9 +459,9 @@ class ValkeyClient(VectorDBBase):
index_name = self._index_name(collection_name) index_name = self._index_name(collection_name)
try: try:
self._g['glide_ft'].dropindex(self.client, index_name) self._g['glide_ft'].dropindex(self.client, index_name)
log.info(f'Dropped index {index_name}') log.info('Dropped index %s', index_name)
except self._g['RequestError'] as e: except self._g['RequestError'] as e:
log.debug(f'Could not drop index {index_name}: {e}') log.debug('Could not drop index %s: %s', index_name, e)
self._delete_keys_by_prefix(self._key_prefix(collection_name)) self._delete_keys_by_prefix(self._key_prefix(collection_name))
@ -482,7 +485,7 @@ class ValkeyClient(VectorDBBase):
'id': item['id'], 'id': item['id'],
'vector': _vector_to_bytes(item['vector']), 'vector': _vector_to_bytes(item['vector']),
'text': item['text'], 'text': item['text'],
'metadata_json': json.dumps(metadata), 'metadata_json': JSONCodec.dumps(metadata),
# `or ''` prevents indexing literal 'None' as a TAG value, which would # `or ''` prevents indexing literal 'None' as a TAG value, which would
# poison $ne / equality queries. # poison $ne / equality queries.
'hash': str(metadata.get('hash') or ''), 'hash': str(metadata.get('hash') or ''),
@ -492,7 +495,7 @@ class ValkeyClient(VectorDBBase):
} }
self.batch_client.hset(self._item_key(collection_name, item['id']), mapping) self.batch_client.hset(self._item_key(collection_name, item['id']), mapping)
log.debug(f'Inserted {len(items)} items into collection {collection_name}') log.debug('Inserted %s items into collection %s', len(items), collection_name)
def upsert(self, collection_name: str, items: list[VectorItem]): def upsert(self, collection_name: str, items: list[VectorItem]):
self.insert(collection_name, items) self.insert(collection_name, items)
@ -588,8 +591,8 @@ class ValkeyClient(VectorDBBase):
ids.append(_decode(fields.get(b'id', b''))) ids.append(_decode(fields.get(b'id', b'')))
documents.append(_decode(fields.get(b'text', b''))) documents.append(_decode(fields.get(b'text', b'')))
try: try:
metadatas.append(json.loads(_decode(fields.get(b'metadata_json', b'{}')))) metadatas.append(JSONCodec.loads(_decode(fields.get(b'metadata_json', b'{}'))))
except (json.JSONDecodeError, TypeError): except (ValueError, TypeError):
metadatas.append({}) metadatas.append({})
if limit is not None and limit > 0 and len(ids) >= limit: if limit is not None and limit > 0 and len(ids) >= limit:
return GetResult(ids=[ids], documents=[documents], metadatas=[metadatas]) return GetResult(ids=[ids], documents=[documents], metadatas=[metadatas])
@ -656,7 +659,7 @@ class ValkeyClient(VectorDBBase):
collections.append(name[len(idx_prefix) :]) collections.append(name[len(idx_prefix) :])
try: try:
glide_ft.dropindex(self.client, idx) glide_ft.dropindex(self.client, idx)
log.info(f'Dropped index: {name}') log.info('Dropped index: %s', name)
except Exception as e: except Exception as e:
log.error(f'Error dropping index {name}: {e}') log.error(f'Error dropping index {name}: {e}')
except Exception as e: except Exception as e:
@ -664,7 +667,7 @@ class ValkeyClient(VectorDBBase):
for collection in collections: for collection in collections:
self._delete_keys_by_prefix(self._key_prefix(collection)) self._delete_keys_by_prefix(self._key_prefix(collection))
log.info(f'Valkey vector store reset complete (prefix: {self.collection_prefix})') log.info('Valkey vector store reset complete (prefix: %s)', self.collection_prefix)
def _delete_keys_by_prefix(self, prefix: str) -> None: def _delete_keys_by_prefix(self, prefix: str) -> None:
cursor = '0' cursor = '0'
@ -734,8 +737,8 @@ class ValkeyClient(VectorDBBase):
ids.append(_decode(fields.get(b'id', b''))) ids.append(_decode(fields.get(b'id', b'')))
documents.append(_decode(fields.get(b'text', b''))) documents.append(_decode(fields.get(b'text', b'')))
try: try:
metadatas.append(json.loads(_decode(fields.get(b'metadata_json', b'{}')))) metadatas.append(JSONCodec.loads(_decode(fields.get(b'metadata_json', b'{}'))))
except (json.JSONDecodeError, TypeError): except (ValueError, TypeError):
metadatas.append({}) metadatas.append({})
if include_score: if include_score:

View file

@ -23,7 +23,7 @@ from open_webui.retrieval.vector.main import (
VectorDBBase, VectorDBBase,
VectorItem, VectorItem,
) )
from open_webui.retrieval.vector.utils import process_metadata from open_webui.retrieval.vector.utils import iter_filter_conditions, process_metadata
def _convert_uuids_to_strings(obj: Any) -> Any: def _convert_uuids_to_strings(obj: Any) -> Any:
@ -54,6 +54,20 @@ def _convert_uuids_to_strings(obj: Any) -> Any:
return obj return obj
def _metadata_filter(filter: Optional[dict]) -> Any:
clauses = []
for key, op, value in iter_filter_conditions(filter):
if op == '$in':
clauses.append(
weaviate.classes.query.Filter.any_of(
[weaviate.classes.query.Filter.by_property(name=key).equal(item) for item in value]
)
)
else:
clauses.append(weaviate.classes.query.Filter.by_property(name=key).equal(value))
return weaviate.classes.query.Filter.all_of(clauses) if len(clauses) > 1 else (clauses[0] if clauses else None)
class WeaviateClient(VectorDBBase): class WeaviateClient(VectorDBBase):
def __init__(self): def __init__(self):
self.url = WEAVIATE_HTTP_HOST self.url = WEAVIATE_HTTP_HOST
@ -168,6 +182,7 @@ class WeaviateClient(VectorDBBase):
return None return None
collection = self.client.collections.get(sane_collection_name) collection = self.client.collections.get(sane_collection_name)
weaviate_filter = _metadata_filter(filter)
result_ids, result_documents, result_metadatas, result_distances = ( result_ids, result_documents, result_metadatas, result_distances = (
[], [],
@ -181,6 +196,7 @@ class WeaviateClient(VectorDBBase):
response = collection.query.near_vector( response = collection.query.near_vector(
near_vector=vector_embedding, near_vector=vector_embedding,
limit=limit, limit=limit,
filters=weaviate_filter,
return_metadata=weaviate.classes.query.MetadataQuery(distance=True), return_metadata=weaviate.classes.query.MetadataQuery(distance=True),
) )

View file

@ -1,8 +1,12 @@
from threading import Lock
from fastapi import HTTPException
from open_webui.config import ( from open_webui.config import (
ENABLE_MILVUS_MULTITENANCY_MODE, ENABLE_MILVUS_MULTITENANCY_MODE,
ENABLE_QDRANT_MULTITENANCY_MODE, ENABLE_QDRANT_MULTITENANCY_MODE,
VECTOR_DB, VECTOR_DB,
) )
from open_webui.env import USE_SLIM
from open_webui.retrieval.vector.main import VectorDBBase from open_webui.retrieval.vector.main import VectorDBBase
from open_webui.retrieval.vector.type import VectorType from open_webui.retrieval.vector.type import VectorType
@ -13,6 +17,11 @@ class Vector:
""" """
get vector db instance by vector type get vector db instance by vector type
""" """
if USE_SLIM and vector_type != VectorType.PGVECTOR:
raise HTTPException(
503,
'Slim requires PostgreSQL/pgvector for vector storage. Set VECTOR_DB=pgvector and PGVECTOR_DB_URL, or use the standard image.',
)
match vector_type: match vector_type:
case VectorType.MILVUS: case VectorType.MILVUS:
if ENABLE_MILVUS_MULTITENANCY_MODE: if ENABLE_MILVUS_MULTITENANCY_MODE:
@ -88,4 +97,27 @@ class Vector:
raise ValueError(f'Unsupported vector type: {vector_type}') raise ValueError(f'Unsupported vector type: {vector_type}')
VECTOR_DB_CLIENT = Vector.get_vector(VECTOR_DB) VECTOR_DB_CLIENT = None if USE_SLIM else Vector.get_vector(VECTOR_DB)
_vector_client_lock = Lock()
def get_vector_db_client() -> VectorDBBase:
"""Initialize slim's remote client on first use so chat can start without it."""
global VECTOR_DB_CLIENT
if VECTOR_DB_CLIENT is not None:
return VECTOR_DB_CLIENT
with _vector_client_lock:
if VECTOR_DB_CLIENT is None:
from open_webui import config
if VECTOR_DB == VectorType.PGVECTOR and not config.PGVECTOR_DB_URL.startswith('postgres'):
raise HTTPException(503, 'Configure PGVECTOR_DB_URL for remote vector storage.')
try:
VECTOR_DB_CLIENT = Vector.get_vector(VECTOR_DB)
except HTTPException:
raise
except Exception as exc:
raise HTTPException(
503, f'Unable to connect to configured vector database ({VECTOR_DB}): {exc}'
) from exc
return VECTOR_DB_CLIENT

View file

@ -1,16 +1,38 @@
import datetime as dt import datetime as dt
from typing import Any from typing import Any
from open_webui.env import RAG_METADATA_MAX_VALUE_CHARS
from open_webui.retrieval.vector.main import SearchResult from open_webui.retrieval.vector.main import SearchResult
from open_webui.utils.misc import sanitize_text_for_db from open_webui.utils.misc import sanitize_text_for_db
KEYS_TO_EXCLUDE = ['content', 'pages', 'tables', 'paragraphs', 'sections', 'figures'] KEYS_TO_EXCLUDE = [
'content',
'pages',
'tables',
'paragraphs',
'sections',
'figures',
'documents',
'keyValuePairs',
'styles',
'languages',
]
def filter_metadata(metadata: dict[str, any]) -> dict[str, any]: def filter_metadata(metadata: dict[str, any]) -> dict[str, any]:
# Removes large/redundant fields from metadata dict. # Removes large/redundant fields from metadata dict.
metadata = {key: value for key, value in metadata.items() if key not in KEYS_TO_EXCLUDE} result = {}
return metadata for key, value in metadata.items():
if key in KEYS_TO_EXCLUDE:
continue
if RAG_METADATA_MAX_VALUE_CHARS is not None and isinstance(value, (list, dict)):
try:
if len(str(value)) > RAG_METADATA_MAX_VALUE_CHARS:
continue
except (MemoryError, RecursionError, ValueError):
continue
result[key] = value
return result
def process_metadata( def process_metadata(
@ -25,6 +47,12 @@ def process_metadata(
continue continue
if value is None: if value is None:
continue continue
if RAG_METADATA_MAX_VALUE_CHARS is not None and isinstance(value, (list, dict)):
try:
if len(str(value)) > RAG_METADATA_MAX_VALUE_CHARS:
continue
except (MemoryError, RecursionError, ValueError):
continue
# Convert non-serializable fields to strings # Convert non-serializable fields to strings
if isinstance(value, (dt.datetime, list, dict)): if isinstance(value, (dt.datetime, list, dict)):
result[key] = sanitize_text_for_db(str(value)) result[key] = sanitize_text_for_db(str(value))
@ -33,6 +61,33 @@ def process_metadata(
return result return result
def iter_filter_conditions(filter: dict[str, Any] | None):
for key, value in (filter or {}).items():
if isinstance(value, dict):
if set(value) != {'$in'}:
raise ValueError(f"Unsupported metadata filter for '{key}': {value}")
yield key, '$in', list(value['$in'])
else:
yield key, '$eq', value
def normalize_filter(filter: dict[str, Any] | None) -> dict[str, Any]:
return {key: {'$in': value} if op == '$in' else value for key, op, value in iter_filter_conditions(filter)}
def metadata_matches_filter(metadata: dict[str, Any], filter: dict[str, Any] | None) -> bool:
if not isinstance(metadata, dict):
return False
for key, op, value in iter_filter_conditions(filter):
actual = metadata.get(key)
if op == '$in':
if actual not in value:
return False
elif actual != value:
return False
return True
def merge_hybrid_search_results( def merge_hybrid_search_results(
vector_result: SearchResult | None, vector_result: SearchResult | None,
fts_results: list[dict[str, Any]], fts_results: list[dict[str, Any]],

View file

@ -1,9 +1,9 @@
import json
import logging import logging
from typing import Optional from typing import Optional
import requests import requests
from open_webui.retrieval.web.main import SearchResult, get_filtered_results from open_webui.retrieval.web.main import SearchResult, get_filtered_results
from open_webui.utils.json_codec import JSONCodec
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@ -41,7 +41,7 @@ def search_bocha(api_key: str, query: str, count: int, filter_list: Optional[lis
url = 'https://api.bochaai.com/v1/web-search?utm_source=ollama' url = 'https://api.bochaai.com/v1/web-search?utm_source=ollama'
headers = {'Authorization': f'Bearer {api_key}', 'Content-Type': 'application/json'} headers = {'Authorization': f'Bearer {api_key}', 'Content-Type': 'application/json'}
payload = json.dumps({'query': query, 'summary': True, 'freshness': 'noLimit', 'count': count}) payload = JSONCodec.dumps({'query': query, 'summary': True, 'freshness': 'noLimit', 'count': count})
response = requests.post(url, headers=headers, data=payload, timeout=5) response = requests.post(url, headers=headers, data=payload, timeout=5)
response.raise_for_status() response.raise_for_status()

View file

@ -16,7 +16,7 @@ async def search_brave(
api_key: str, api_key: str,
query: str, query: str,
count: int, count: int,
filter_list: list[str | None] | None = None, filter_list: list[str] | None = None,
) -> list[SearchResult]: ) -> list[SearchResult]:
"""Query the Brave Web Search API and return normalised results. """Query the Brave Web Search API and return normalised results.

View file

@ -3,8 +3,7 @@ from __future__ import annotations
import logging import logging
import urllib.request import urllib.request
from ddgs import DDGS from open_webui.env import USE_SLIM
from ddgs.exceptions import RatelimitException
from open_webui.retrieval.web.main import SearchResult, get_filtered_results from open_webui.retrieval.web.main import SearchResult, get_filtered_results
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@ -13,7 +12,7 @@ log = logging.getLogger(__name__)
def search_duckduckgo( def search_duckduckgo(
query: str, query: str,
count: int, count: int,
filter_list: list[str | None] = None, filter_list: list[str] | None = None,
concurrent_requests: int | None = None, concurrent_requests: int | None = None,
backend: str | None = 'auto', backend: str | None = 'auto',
) -> list[SearchResult]: ) -> list[SearchResult]:
@ -27,25 +26,22 @@ def search_duckduckgo(
Returns: Returns:
list[SearchResult]: A list of search results list[SearchResult]: A list of search results
""" """
if USE_SLIM:
raise ValueError(
'DDGS is unavailable in slim. Configure another web search provider in Admin Settings > Web Search.'
)
from ddgs import DDGS
# The ddgs library (primp-based) does not auto-detect proxy env vars. # The ddgs library (primp-based) does not auto-detect proxy env vars.
# Resolve via stdlib getproxies() — same pattern as the other loaders. # Resolve via stdlib getproxies() — same pattern as the other loaders.
env_proxies = urllib.request.getproxies() env_proxies = urllib.request.getproxies()
proxy = env_proxies.get('https') or env_proxies.get('http') proxy = env_proxies.get('https') or env_proxies.get('http')
search_results = []
with DDGS(proxy=proxy) as ddgs: with DDGS(proxy=proxy) as ddgs:
if concurrent_requests: if concurrent_requests:
ddgs.threads = concurrent_requests ddgs.threads = concurrent_requests
# Use the ddgs.text() method to perform the search search_results = ddgs.text(query, safesearch='moderate', max_results=count, backend=backend or 'auto')
try:
kwargs = {'safesearch': 'moderate', 'max_results': count}
if backend and backend != 'auto':
kwargs['backend'] = backend
results = ddgs.text(query, **kwargs)
search_results = results if results is not None else []
except RatelimitException as e:
log.error(f'RatelimitException: {e}')
search_results = []
if filter_list: if filter_list:
search_results = get_filtered_results(search_results, filter_list) search_results = get_filtered_results(search_results, filter_list)

View file

@ -1,6 +1,4 @@
import logging import logging
from dataclasses import dataclass
from typing import Optional
import requests import requests
from open_webui.retrieval.web.main import SearchResult from open_webui.retrieval.web.main import SearchResult
@ -10,18 +8,12 @@ log = logging.getLogger(__name__)
EXA_API_BASE = 'https://api.exa.ai' EXA_API_BASE = 'https://api.exa.ai'
@dataclass
class ExaResult:
url: str
title: str
text: str
def search_exa( def search_exa(
api_key: str, api_key: str,
query: str, query: str,
count: int, count: int,
filter_list: Optional[list[str]] = None, filter_list: list[str] | None = None,
max_content_length: int | None = None,
) -> list[SearchResult]: ) -> list[SearchResult]:
"""Search using Exa Search API and return the results as a list of SearchResult objects. """Search using Exa Search API and return the results as a list of SearchResult objects.
@ -29,9 +21,10 @@ def search_exa(
api_key (str): A Exa Search API key api_key (str): A Exa Search API key
query (str): The query to search for query (str): The query to search for
count (int): Number of results to return count (int): Number of results to return
filter_list (Optional[list[str]]): List of domains to filter results by filter_list (list[str] | None): List of domains to filter results by
max_content_length (int | None): Maximum characters per result; None leaves text unlimited.
""" """
log.info(f'Searching with Exa for query: {query}') log.info('Searching with Exa for query: %s', query)
headers = {'Authorization': f'Bearer {api_key}', 'Content-Type': 'application/json'} headers = {'Authorization': f'Bearer {api_key}', 'Content-Type': 'application/json'}
@ -39,7 +32,7 @@ def search_exa(
'query': query, 'query': query,
'numResults': count or 5, 'numResults': count or 5,
'includeDomains': filter_list, 'includeDomains': filter_list,
'contents': {'text': True, 'highlights': True}, 'contents': {'text': {'maxCharacters': max_content_length} if max_content_length is not None else True},
'type': 'auto', # Use the auto search type (keyword or neural) 'type': 'auto', # Use the auto search type (keyword or neural)
} }
@ -48,22 +41,13 @@ def search_exa(
response.raise_for_status() response.raise_for_status()
data = response.json() data = response.json()
results = [] results = data['results']
for result in data['results']: log.info('Found %s results', len(results))
results.append(
ExaResult(
url=result['url'],
title=result['title'],
text=result['text'],
)
)
log.info(f'Found {len(results)} results')
return [ return [
SearchResult( SearchResult(
link=result.url, link=result['url'],
title=result.title, title=result['title'],
snippet=result.text, snippet=(result.get('text') or '')[:max_content_length],
) )
for result in results for result in results
] ]

View file

@ -21,6 +21,9 @@ def search_external(
) -> List[SearchResult]: ) -> List[SearchResult]:
try: try:
headers = { headers = {
# LICENSE covers this Open WebUI user-agent identifier.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
'User-Agent': 'Open WebUI (https://github.com/open-webui/open-webui) RAG Bot', 'User-Agent': 'Open WebUI (https://github.com/open-webui/open-webui) RAG Bot',
'Authorization': f'Bearer {external_api_key}', 'Authorization': f'Bearer {external_api_key}',
} }
@ -50,7 +53,7 @@ def search_external(
) )
for result in results[:count] for result in results[:count]
] ]
log.info(f'External search results: {results}') log.info('External search results: %s', results)
return results return results
except Exception as e: except Exception as e:
log.error(f'Error in External search: {e}') log.error(f'Error in External search: {e}')

View file

@ -225,7 +225,7 @@ def search_firecrawl(
) )
) )
log.info(f'FireCrawl search results: {search_results}') log.info('FireCrawl search results: %s', search_results)
return search_results return search_results
except Exception as e: except Exception as e:
log.error(f'Error in FireCrawl search: {e}') log.error(f'Error in FireCrawl search: {e}')

View file

@ -13,7 +13,7 @@ async def search_google_pse(
search_engine_id: str, search_engine_id: str,
query: str, query: str,
count: int, count: int,
filter_list: list[str | None] | None = None, filter_list: list[str] | None = None,
referer: str | None = None, referer: str | None = None,
) -> list[SearchResult]: ) -> list[SearchResult]:
"""Query Google Programmable Search Engine with automatic pagination. """Query Google Programmable Search Engine with automatic pagination.

View file

@ -1,11 +1,10 @@
from __future__ import annotations from __future__ import annotations
import ipaddress
from urllib.parse import urlparse from urllib.parse import urlparse
import validators import validators
from open_webui.retrieval.web.utils import resolve_hostname from open_webui.retrieval.web.utils import resolve_hostname
from open_webui.utils.misc import get_allow_block_lists, is_host_allowed from open_webui.utils.misc import as_network, get_allow_block_lists, is_host_allowed
from pydantic import BaseModel from pydantic import BaseModel
@ -14,14 +13,8 @@ def get_filtered_results(results, filter_list):
return results return results
allow_list, block_list = get_allow_block_lists(filter_list) allow_list, block_list = get_allow_block_lists(filter_list)
resolve_ips = False # Only worth a lookup when an entry names an address, since a hostname entry matches by name.
for entry in allow_list + block_list: resolve_ips = any(as_network(entry) is not None for entry in allow_list + block_list)
try:
ipaddress.ip_address(entry)
except ValueError:
continue
resolve_ips = True
break
filtered_results = [] filtered_results = []

View file

@ -9,20 +9,18 @@ from open_webui.utils.headers import include_user_info_headers
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
DEFAULT_MICROSOFT_WEB_IQ_API_BASE_URL = 'https://api.microsoft.ai/v3'
def search_microsoft_web_iq( def search_microsoft_web_iq(
api_base_url: str, api_base_url: str,
api_key: str, api_key: str,
query: str, query: str,
count: int, count: int,
filter_list: list[str | None] | None = None, filter_list: list[str] | None = None,
language: str = 'en', language: str = 'en',
user=None, user=None,
) -> list[SearchResult]: ) -> list[SearchResult]:
try: try:
api_base_url = (api_base_url or DEFAULT_MICROSOFT_WEB_IQ_API_BASE_URL).rstrip('/') api_base_url = api_base_url.rstrip('/')
headers = { headers = {
'host': urlparse(api_base_url).netloc or 'api.microsoft.ai', 'host': urlparse(api_base_url).netloc or 'api.microsoft.ai',
'x-apikey': api_key, 'x-apikey': api_key,

View file

@ -8,7 +8,7 @@ from open_webui.retrieval.web.main import SearchResult, get_filtered_results
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
def search_mojeek(api_key: str, query: str, count: int, filter_list: list[str | None] = None) -> list[SearchResult]: def search_mojeek(api_key: str, query: str, count: int, filter_list: list[str] | None = None) -> list[SearchResult]:
"""Search using Mojeek's Search API and return the results as a list of SearchResult objects. """Search using Mojeek's Search API and return the results as a list of SearchResult objects.
Args: Args:

View file

@ -23,7 +23,7 @@ def search_ollama_cloud(
count (int): Number of results to return count (int): Number of results to return
filter_list (Optional[list[str]]): List of domains to filter results by filter_list (Optional[list[str]]): List of domains to filter results by
""" """
log.info(f'Searching with Ollama for query: {query}') log.info('Searching with Ollama for query: %s', query)
headers = {'Authorization': f'Bearer {api_key}', 'Content-Type': 'application/json'} headers = {'Authorization': f'Bearer {api_key}', 'Content-Type': 'application/json'}
payload = {'query': query, 'max_results': count} payload = {'query': query, 'max_results': count}
@ -34,7 +34,7 @@ def search_ollama_cloud(
data = response.json() data = response.json()
results = data.get('results', []) results = data.get('results', [])
log.info(f'Found {len(results)} results') log.info('Found %s results', len(results))
if filter_list: if filter_list:
results = get_filtered_results(results, filter_list) results = get_filtered_results(results, filter_list)

View file

@ -12,7 +12,7 @@ async def search_openserp(
base_url: str, base_url: str,
query: str, query: str,
count: int, count: int,
filter_list: list[str | None] | None = None, filter_list: list[str] | None = None,
) -> list[SearchResult]: ) -> list[SearchResult]:
"""Query an OpenSERP instance and return normalised results. """Query an OpenSERP instance and return normalised results.

View file

@ -26,19 +26,26 @@ def search_searchapi(
engine = engine or 'google' engine = engine or 'google'
payload = {'engine': engine, 'q': query, 'api_key': api_key} payload = {'engine': engine, 'q': query, 'api_key': api_key}
if engine.startswith('google'):
payload['link'] = 'resolved'
url = f'{url}?{urlencode(payload)}' url = f'{url}?{urlencode(payload)}'
response = requests.request('GET', url) response = requests.request('GET', url, timeout=30)
response.raise_for_status()
json_response = response.json() json_response = response.json()
log.info(f'results from searchapi search: {json_response}') log.debug('results from searchapi search: %s', json_response)
results = sorted(json_response.get('organic_results', []), key=lambda x: x.get('position', 0)) # top_stories entries carry no position, so the merged list keeps API order
results = [
*json_response.get('organic_results', []),
*json_response.get('top_stories', []),
]
if filter_list: if filter_list:
results = get_filtered_results(results, filter_list) results = get_filtered_results(results, filter_list)
return [ return [
SearchResult( SearchResult(
link=result['link'], link=result.get('link', ''),
title=result.get('title'), title=result.get('title'),
snippet=result.get('snippet'), snippet=result.get('snippet'),
) )

View file

@ -12,6 +12,9 @@ log = logging.getLogger(__name__)
# SearXNG request headers — identifies the bot to instance operators. # SearXNG request headers — identifies the bot to instance operators.
_SEARXNG_HEADERS = { _SEARXNG_HEADERS = {
# LICENSE covers this Open WebUI user-agent identifier.
# Do not alter, remove, obscure, or replace it except as LICENSE permits:
# https://docs.openwebui.com/license.
'User-Agent': 'Open WebUI (https://github.com/open-webui/open-webui) RAG Bot', 'User-Agent': 'Open WebUI (https://github.com/open-webui/open-webui) RAG Bot',
'Accept': 'text/html', 'Accept': 'text/html',
'Accept-Encoding': 'gzip, deflate', 'Accept-Encoding': 'gzip, deflate',
@ -37,7 +40,7 @@ async def search_searxng(
query_url: str, query_url: str,
query: str, query: str,
count: int, count: int,
filter_list: list[str | None] | None = None, filter_list: list[str] | None = None,
**kwargs, **kwargs,
) -> list[SearchResult]: ) -> list[SearchResult]:
"""Query a SearXNG instance and return results sorted by relevance score. """Query a SearXNG instance and return results sorted by relevance score.

View file

@ -31,7 +31,7 @@ def search_serpapi(
response = requests.request('GET', url) response = requests.request('GET', url)
json_response = response.json() json_response = response.json()
log.info(f'results from serpapi search: {json_response}') log.info('results from serpapi search: %s', json_response)
results = sorted(json_response.get('organic_results', []), key=lambda x: x.get('position', 0)) results = sorted(json_response.get('organic_results', []), key=lambda x: x.get('position', 0))
if filter_list: if filter_list:

View file

@ -1,9 +1,9 @@
from __future__ import annotations from __future__ import annotations
import json
import logging import logging
from open_webui.retrieval.web.main import SearchResult, get_filtered_results from open_webui.retrieval.web.main import SearchResult, get_filtered_results
from open_webui.utils.json_codec import JSONCodec
from open_webui.utils.session_pool import get_session from open_webui.utils.session_pool import get_session
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@ -13,7 +13,7 @@ async def search_serper(
api_key: str, api_key: str,
query: str, query: str,
count: int, count: int,
filter_list: list[str | None] | None = None, filter_list: list[str] | None = None,
) -> list[SearchResult]: ) -> list[SearchResult]:
"""Query the serper.dev Google Search API and return normalised results. """Query the serper.dev Google Search API and return normalised results.
@ -23,7 +23,7 @@ async def search_serper(
headers = {'X-API-KEY': api_key, 'Content-Type': 'application/json'} headers = {'X-API-KEY': api_key, 'Content-Type': 'application/json'}
session = await get_session() session = await get_session()
async with session.post(url, headers=headers, data=json.dumps({'q': query})) as response: async with session.post(url, headers=headers, data=JSONCodec.dumps({'q': query})) as response:
response.raise_for_status() response.raise_for_status()
payload = await response.json() payload = await response.json()

View file

@ -9,7 +9,7 @@ async def search_serphouse(
domain: str, domain: str,
query: str, query: str,
count: int, count: int,
filter_list: list[str | None] | None = None, filter_list: list[str] | None = None,
) -> list[SearchResult]: ) -> list[SearchResult]:
"""Query SERPHouse and return normalised organic results.""" """Query SERPHouse and return normalised organic results."""
session = await get_session() session = await get_session()

View file

@ -17,7 +17,7 @@ def search_serply(
limit: int = 10, limit: int = 10,
device_type: str = 'desktop', device_type: str = 'desktop',
proxy_location: str = 'US', proxy_location: str = 'US',
filter_list: list[str | None] = None, filter_list: list[str] | None = None,
) -> list[SearchResult]: ) -> list[SearchResult]:
"""Search using serper.dev's API and return the results as a list of SearchResult objects. """Search using serper.dev's API and return the results as a list of SearchResult objects.
@ -51,7 +51,7 @@ def search_serply(
response.raise_for_status() response.raise_for_status()
json_response = response.json() json_response = response.json()
log.info(f'results from serply search: {json_response}') log.info('results from serply search: %s', json_response)
results = sorted(json_response.get('results', []), key=lambda x: x.get('realPosition', 0)) results = sorted(json_response.get('results', []), key=lambda x: x.get('realPosition', 0))
if filter_list: if filter_list:

View file

@ -12,7 +12,7 @@ async def search_serpstack(
api_key: str, api_key: str,
query: str, query: str,
count: int, count: int,
filter_list: list[str | None] | None = None, filter_list: list[str] | None = None,
https_enabled: bool = True, https_enabled: bool = True,
) -> list[SearchResult]: ) -> list[SearchResult]:
"""Query the serpstack.com API and return normalised results. """Query the serpstack.com API and return normalised results.

View file

@ -1,8 +1,8 @@
import json
import logging import logging
from typing import List, Optional from typing import List, Optional
from open_webui.retrieval.web.main import SearchResult, get_filtered_results from open_webui.retrieval.web.main import SearchResult, get_filtered_results
from open_webui.utils.json_codec import JSONCodec
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@ -28,10 +28,11 @@ def search_sougou(
http_profile.endpoint = 'tms.tencentcloudapi.com' http_profile.endpoint = 'tms.tencentcloudapi.com'
client_profile = ClientProfile() client_profile = ClientProfile()
client_profile.http_profile = http_profile client_profile.http_profile = http_profile
params = json.dumps({'Query': query, 'Cnt': 20}) params = JSONCodec.dumps({'Query': query, 'Cnt': 20})
common_client = CommonClient('tms', '2020-12-29', cred, '', profile=client_profile) common_client = CommonClient('tms', '2020-12-29', cred, '', profile=client_profile)
results = [ results = [
json.loads(page) for page in common_client.call_json('SearchPro', json.loads(params))['Response']['Pages'] JSONCodec.loads(page)
for page in common_client.call_json('SearchPro', JSONCodec.loads(params))['Response']['Pages']
] ]
sorted_results = sorted(results, key=lambda x: x.get('scour', 0.0), reverse=True) sorted_results = sorted(results, key=lambda x: x.get('scour', 0.0), reverse=True)
if filter_list: if filter_list:

View file

@ -0,0 +1,60 @@
from __future__ import annotations
import requests
from open_webui.retrieval.web.main import SearchResult, get_filtered_results
def search_staan(
api_key: str,
query: str,
count: int,
filter_list: list[str] | None = None,
market: str | None = None,
max_snippets: int | None = None,
) -> list[SearchResult]:
"""Search using Staan's Web Search API and return the results as a list of SearchResult objects.
Args:
api_key (str): A Staan API key
query (str): The query to search for
count (int): The maximum number of results to return
filter_list (list[str] | None): The domains to allow or block
market (str | None): The market to search in, e.g. 'en-us'
max_snippets (int | None): The maximum extra snippets to request per result
Returns:
A list of SearchResult objects.
"""
url = 'https://api.staan.ai/v2/search/web'
headers = {
'Accept': 'application/json',
'Authorization': f'Bearer {api_key}',
}
params = {'q': query, 'market': market}
if max_snippets:
params['extra_snippets'] = 'true'
params['max_snippets'] = max_snippets
response = requests.get(url, headers=headers, params=params)
response.raise_for_status()
results = response.json().get('web', {}).get('results', [])
if filter_list:
results = get_filtered_results(results, filter_list)
return [
SearchResult(
link=result.get('url', ''),
title=result.get('title'),
snippet=_build_snippet(result),
)
for result in results[:count]
]
def _build_snippet(result: dict) -> str:
"""Combine the snippet and the extra snippets list into a single string."""
parts = [result.get('snippet')]
parts.extend(extra.get('chunk') for extra in result.get('extra_snippets', []))
return '\n\n'.join(part for part in parts if part)

View file

@ -3,6 +3,7 @@ from __future__ import annotations
import logging import logging
import requests import requests
from open_webui.env import TAVILY_API_BASE_URL
from open_webui.retrieval.web.main import SearchResult, get_filtered_results from open_webui.retrieval.web.main import SearchResult, get_filtered_results
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@ -12,7 +13,7 @@ def search_tavily(
api_key: str, api_key: str,
query: str, query: str,
count: int, count: int,
filter_list: list[str | None] = None, filter_list: list[str] | None = None,
# **kwargs, # **kwargs,
) -> list[SearchResult]: ) -> list[SearchResult]:
"""Search using Tavily's Search API and return the results as a list of SearchResult objects. """Search using Tavily's Search API and return the results as a list of SearchResult objects.
@ -25,7 +26,7 @@ def search_tavily(
Returns: Returns:
A list of SearchResult objects. A list of SearchResult objects.
""" """
url = 'https://api.tavily.com/search' url = f'{TAVILY_API_BASE_URL}/search'
headers = { headers = {
'Content-Type': 'application/json', 'Content-Type': 'application/json',
'Authorization': f'Bearer {api_key}', 'Authorization': f'Bearer {api_key}',

Some files were not shown because too many files have changed in this diff Show more