Outlet filters see the finished reply with its content and token usage, but had no way to tell whether the model ended on its own, ran into the token limit or stopped to call a tool, short of reading every chunk in a stream filter. For OpenAI-compatible and Ollama models, the assistant message handed to outlet now carries finish_reason with the value the provider reported on the last model call of the turn, for streaming and non-streaming replies. Ollama replies cut off by the token limit now report length as well, where they always said stop before.
Fill the 2020 empty strings in it-IT and use "Skill" consistently.
Co-authored-by: Daniele Giovanardi <177720659+danielegiovanardi@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
With ENABLE_FORWARD_USER_INFO_HEADERS on and hybrid search enabled, rerank requests went out without the user headers when a chat searched attached knowledge, when a model used the built-in knowledge search tool, and when the collection query API was called with hybrid turned off for that one request. The embedding requests of the same search did carry them, so external rerankers that use these headers for per-user auth, rate limits or auditing saw anonymous calls. Rerank requests now carry the signed-in user, the same as embedding requests.
Fixes#32060
When a model streams its images in separate chunks, each new chunk saved all earlier images again, so the copies doubled with every image: four images were stored as fifteen and the chat showed duplicates. New replies now keep each image once. Chats that already hold duplicates stay as they are.
Fixes#32053
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Waiting to run
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Waiting to run
The full security policy now lives on the documentation site, so it only has to be kept up to date in one place. This file keeps a link to it and the one reporting channel we accept, since GitHub shows this file to reporters on the Security tab.
Good-faith reports that turn out not to be vulnerabilities are now explicitly welcomed, and the line about barring reporters is softened so it only targets repeated or deliberate rule breaking. The section on foreign CNAs moves near the end, dashes are gone, a few redundant sentences are cut and the last-updated date is today.
* fix: Direct Connection replies are saved scrambled or empty
With a Direct Connection the browser tab forwards the model's reply to the server piece by piece. Since the server started checking the tab's session again for every one of those pieces, the pieces could overtake each other while the checks ran, so the saved reply came out in the wrong order, or empty when the end of the stream arrived first. Pieces from one tab are now handled one after another in the order they arrived, and each one is still checked.
Fixes#31953
* fix: check Direct Connection reply pieces side by side while keeping them in order
Handling one tab's reply pieces strictly one after another also made each piece wait for the previous piece's session check, so a fast 2000-piece reply took 12 to 22% longer to save than without the ordering. Each piece's check now starts as soon as it arrives and only the hand-over waits its turn, so replies save as fast as before and still in order. Once a check fails, for example after a sign-out, no later piece from that tab gets through either.