open-webui/backend/open_webui
Nexory 19222eff68 fix: retain handles + done_callback + shutdown-cancel for pool cleanup tasks
Issue #25216 named two root causes for unbounded SESSION_POOL/USAGE_POOL
growth. The first — `periodic_usage_pool_cleanup` exiting permanently on
Redis lock acquire/renew failure — is being addressed by PR #21798's
`run_with_lock` helper. This commit closes the second: the
`asyncio.create_task(...)` calls at the cleanup-daemon spawn sites
discard their return value, so if either coroutine ever does exit for a
reason that escapes the inner loop (notably `asyncio.CancelledError`,
which is `BaseException` and is not caught by the `except Exception`
guards in `run_with_lock`), there is no real-time operator signal — the
asyncio default exception handler emits only a delayed "Task exception
was never retrieved" WARNING at GC time. By that point both pools have
already been growing unbounded for an unknown interval.

Three changes against `backend/open_webui/main.py`, mirroring the
existing `app.state.redis_task_command_listener` pattern at lines 676
and 754:

1. Store each cleanup-task handle on `app.state` and give it a `name=`
   so log lines identify which daemon died.
2. Attach a `_daemon_done_cb` that logs the exception at `ERROR` with
   `exc_info` at the moment of task death, not at GC time. The callback
   filters out `task.cancelled()` so it produces no noise during normal
   shutdown.
3. Cancel both tasks cooperatively in the lifespan shutdown block.
   Without this, `CancelledError` is never delivered, the tasks die
   while pending, and asyncio emits "Task was destroyed but it is
   pending!" warnings on every clean restart — which trains operators
   to ignore that warning and masks real failures.

The change is orthogonal to PR #21798 (zero file overlap; that PR
modifies `socket/main.py` and `socket/utils.py`, this one touches only
`main.py`). The two PRs can be reviewed, merged, and reverted
independently. With both in place, the cleanup loop is internally
resilient (it will not die from Redis events) AND the spawn site is
observable (if it dies anyway, the operator knows immediately).

Note: line 687 spawns `scheduler_worker_loop` with the same
fire-and-forget pattern. It is out of scope for this PR — issue #25216
specifically tracks the pool cleanup daemons — but the same treatment
would apply.

Refs: #25216, #21798
2026-05-30 17:37:37 +02:00
..
data refac: mv backend files to /open_webui dir 2024-09-04 16:54:48 +02:00
internal refac 2026-05-21 16:25:25 +04:00
migrations refac 2026-05-21 15:29:49 +04:00
models fix: gate chat-file links by caller access + repair insert_chat_files db arg (#25054) 2026-05-28 17:42:17 -05:00
retrieval refac 2026-05-21 16:44:36 +04:00
routers fix: harden model profile image against SVG stored XSS (#25060) 2026-05-28 17:41:55 -05:00
socket fix: resolve NameError for redis_sentinels in session_cleanup_lock 2026-05-21 13:41:21 +04:00
static refac 2026-05-12 03:04:35 +09:00
storage refac: modernize type annotations (PEP 604 / PEP 585) 2026-05-12 17:10:15 +09:00
tools fix: add knowledge_id access check in search_knowledge_files (BOLA) (#25113) 2026-05-28 16:23:49 -05:00
utils fix(auth): use request.scope["path"] to prevent CVE-2026-48710 (BadHost) (#25123) 2026-05-28 16:41:56 -05:00
__init__.py refac 2026-05-09 02:38:08 +09:00
alembic.ini fix: Alembic CLI commands from failing 2025-08-15 04:17:47 -04:00
config.py refac 2026-05-21 16:56:56 +04:00
constants.py refac: modernize type annotations (PEP 604 / PEP 585) 2026-05-12 17:10:15 +09:00
env.py refac 2026-05-20 00:22:27 +04:00
functions.py fix: emit [DONE] for AsyncGenerator pipe returns (#24763) 2026-05-19 22:07:56 +04:00
main.py fix: retain handles + done_callback + shutdown-cancel for pool cleanup tasks 2026-05-30 17:37:37 +02:00
tasks.py refac: modernize type annotations (PEP 604 / PEP 585) 2026-05-12 17:10:15 +09:00