fix: stop the memory block leaking into replies after a trailing prompt instruction

With Memory on, a system prompt ending in an instruction like "At the end of every response, append exactly: HELLO, WORLD!" makes some models print the whole injected memory block, stored memories included, right after that string. The block was joined to the prompt with a single newline, so it sat directly under the last instruction and the model read it as part of the text to append.

A blank line now separates the block from an existing system prompt. The block stays at the end of the system message, so the static prompt ahead of it is still a stable prefix for KV caching; moving it first would put per-turn memory content ahead of the prompt.

Measured over OpenRouter, 12 trials each, same system message as Open WebUI builds it:

| Model | Before | After |
|---|---|---|
| gpt-5.6-luna | 12/12 leaked | 0/12 |
| gpt-5-mini | 12/12 leaked | 0/12 |

gpt-4.1-mini and gemini-2.5-flash never leaked on either. With the fix, models still answer from the memory when asked and still end with the requested string.

Fixes #29557
This commit is contained in:
Classic298 2026-09-24 10:53:00 +02:00
parent bbfa876afd
commit a3f39d908b

View file

@ -403,6 +403,9 @@ async def add_memory_context(request, form_data: dict, user, model: dict | None
return form_data
memory_context = f'{MEMORY_CONTEXT_OPEN}\n{rendered}\n{MEMORY_CONTEXT_CLOSE}'
if messages and messages[0].get('role') == 'system':
# Without a blank line, models echo the block as part of the prompt
memory_context = f'\n{memory_context}'
form_data['messages'] = add_or_update_system_message(memory_context, messages, append=True)
return form_data