* fix(execution): only strip images on image-rejection errors
The image-strip recovery fired on any 400/404/422, wiping a session's screenshots and replaying the turn on unrelated errors (e.g. an upstream incomplete-tool-call 400, a context overflow). Require an image-related message too.
* fix(execution): match only replies that mean the model takes no images
* fix(execution): match DeepSeek's image rejection
* fix(execution): guess at Vercel AI Gateway's image rejection
* feat(llm): scrub images from every request to a text-only model
Replace images at the MCP call before any result_transform, give load_skill the
same agent_browser screenshot swap as the system prompt, and add a
call_model_input_filter that turns any image left in the input into a tool error
while chaining any filter already set.
* chore: fix the ruff and bandit failures on main
* feat(llm): keep images away from text-only models
Leave view_image out of the tools, replace MCP image blocks with a text note,
and swap agent_browser's screenshot section for text-only guidance when the
run's model does not accept images.
* refactor(llm): disable view_image via is_enabled and gate skill image guidance by modality
* refactor(skills): render opt-in Jinja skills instead of comment markers
* Tweaked skill
* feat(llm): detect whether a model accepts image input
Add model_supports_images(): OpenRouter models are checked against architecture.input_modalities from OpenRouter's public model list (fetched once per process), others against LiteLLM's supports_vision. Unknown models and lookup failures count as image-capable. Not wired in yet.
* refactor(llm): read image support from litellm's catalog, assuming it for unknown models
* style: drop a trailing space in the image-support docstring
A stop that lands between the silent-wait check and mark_running was
overwritten with running. resume_silent_user_wait does the check and the
state change under the coordinator lock and leaves anything else alone.
Plain text is now the only channel to the user. wait_for_user takes no
arguments and only parks the agent, so a reply can no longer be written
twice (once as text, once as a tool argument). The driver bounces a
wait_for_user call that follows no assistant text back into recovery.
TUI and viewer render wait_for_user as the waiting marker alone and keep
rendering recorded respond_to_user events with their message.
docker.from_env() ignores the current docker context, so Docker Desktop on
macOS (socket under ~/.docker/run unless the default-socket option is on),
OrbStack and Colima reported DOCKER NOT AVAILABLE while `docker ps` worked.
Every failure also printed the same "ensure Docker Desktop is running" text
followed by a RuntimeError traceback.
- resolve the endpoint as the CLI does: DOCKER_HOST, then DOCKER_CONTEXT or
the current context, then the default socket; the sandbox backend uses the
same resolution so startup and scan talk to the same daemon
- classify the SDK error (socket missing, permission denied, connection
refused, Windows named pipe) and print the fix for the current platform,
the endpoint that was tried and the underlying error
- exit 1 cleanly instead of raising after the panel
- telemetry reports docker_unavailable_<reason>
parse_arguments() rejected --fail-on without -n before main() could switch
to headless, so a CI job running `strix -t x --fail-on high` without -n
stopped at an argument error instead of scanning. The terminal check now
lives in cli_args as terminal_attached(); the parser only enforces the
headless-only flag when a terminal is attached and the TUI would open.
Without a tty on stdin and stdout (CI, nohup, pipes, cron, TERM=dumb) the Go
TUI handshakes fine and then exits 1 the moment Bubble Tea tries to take over
the screen. Because the backend was already activated, that exit escaped the
pre-activation mapping and surfaced as a raw traceback:
RuntimeError: Bubble Tea TUI exited with status 1.
Now main() checks for a terminal right after argument parsing. With a target
it switches to the headless path (same as -n) and prints one dim notice.
Without a target the start screen is the only way to enter one, so it prints
a panel telling the user to pass -t <target> -n and exits 1. A bare --resume
is left to the picker, which already lists runs when there is no terminal.
If the sidecar still dies after startup, check_return_code raises
TuiProcessExitedError, run_tui maps it to InteractiveInterfaceExitedError,
and main() prints an INTERACTIVE INTERFACE STOPPED panel with the -n hint
under its own telemetry name instead of re-raising.
preflight_request takes api_base_setting so a dedupe model that timed out
points the user at DEDUPE_LLM_API_BASE, not LLM_API_BASE. warm_up_llm is
now exercised with a dedicated dedupe model: own headers, same preflight
timeout, resolved through resolve_dedupe_model.
The warm-up request used settings.llm.timeout (LLM_TIMEOUT, 300s) both as
the request timeout and the wait_for bound, so a wrong LLM_API_BASE or a
dead proxy hung five minutes and then printed an empty Error line.
Add LlmSettings.preflight_timeout (LLM_PREFLIGHT_TIMEOUT, default 30) and
a shared preflight_request() used for the main and dedupe models. When it
expires the panel names the model, the limit and the setting.
Add uv lock --check, a wheel install smoke test, a PyInstaller dry run,
actionlint and zizmor on the workflows, CodeQL, and Dependabot for action
pins, uv, Go and npm dependencies. Release caches are disabled so zizmor's
cache-poisoning audit passes.
strix --resume with no run name now lists the runs in ./strix_runs inline
in the terminal (started, target, status, findings, run name), newest
first, with arrow-key selection, type-to-search and esc to cancel.
Picking a run continues through the same path as --resume <name>.
Headless or non-TTY launches error with the run list instead.
No install-once global, no hand-rolled traceback formatting, no SystemExit special case: one helper logs the report with exc_info on strix.telemetry when a handler exists, and both hooks call it.
Keep this PR to the logging change only: the finalizer reports are routed to strix.log by the unraisable hook, so the sandbox teardown stays as on main.
Completes the previous commit, which only carried the test rename. sys.unraisablehook (finalizers, __del__, GC and weakref callbacks, files closed at exit) and threading.excepthook are replaced for the life of the process with hooks that log the report at WARNING on strix.telemetry: strix.log, and stderr only when the stream handler runs at DEBUG (STRIX_DEBUG=1). Nothing is delegated to Python's default printers; a report arriving after the strix handlers are torn down is dropped instead of reaching logging.lastResort, and failures inside logging itself are swallowed.
Replaces the pattern-matched urllib3/http.client filter with a generic rule: sys.unraisablehook (finalizers, __del__, GC and weakref callbacks, files closed at exit) and threading.excepthook are replaced for the life of the process with hooks that log the report at WARNING on strix.telemetry. That goes to strix.log, and to stderr only when the stream handler runs at DEBUG (STRIX_DEBUG=1). Nothing is delegated to Python's default printers; a report arriving after the strix handlers are torn down is dropped instead of reaching logging.lastResort, and failures inside logging itself are swallowed so nothing can print mid-shutdown.
tests/test_hook_exceptions_logged.py covers both hooks in-process and end to end in a fresh interpreter (a __del__ raising at shutdown, a thread traceback, a urllib3 response whose file was closed first): stderr stays empty.
The object-shape match accepted any http.* class; it now takes urllib3 and http.client only, so a closed-file error from another module still reaches the previous unraisable hook. New tests build a real DockerSandboxSession holding a PTY exec stream and run the real SDK delete() with the container gone: the SDK alone leaves the socket and its pinned response open, StrixDockerSandboxClient.delete() closes both.
StrixDockerSandboxClient.delete() killed the container and then let the SDK's delete() run shutdown(), which only terminates the agent's PTY exec streams while the container still exists. Once the container was gone the hijacked exec sockets were left to the garbage collector and their HTTP responses failed to close at interpreter exit, which Python 3.14 reports as "Exception ignored while finalizing file <urllib3.response.HTTPResponse>" after the scan summary. delete() now awaits pty_terminate_all() first, whatever the container's state.
The unraisable filter now also matches Python 3.14's finalizer shape (object=None, repr in err_msg) and http.client responses, and logs the match at DEBUG on strix.telemetry instead of letting it reach stderr.
Every sidebar panel gets a header with a click-to-collapse chevron and a
click-to-zoom glyph; a one-row chip on the viewer line hides the whole
sidebar and brings it back from a narrow rail. Tab skips collapsed and
hidden panels, scrollbar hit zones follow the rendered bars, and
fillBackground also repaints after the bare ESC[m reset so no cell shows
the terminal background.