Commit graph

254 commits

Author SHA1 Message Date
Ahmed Allam
aa3144dac4 fix(resume-picker): cancel on ctrl+c during drawing too 2026-10-04 20:52:33 +03:00
Ahmed Allam
28fa7f7734 fix(resume-picker): treat ctrl+c as cancel instead of leaking a KeyboardInterrupt traceback 2026-10-04 20:52:33 +03:00
Ahmed Allam
7a31053348 test(pricing): accept any identically priced provider for bare grok-4.5 2026-10-04 20:38:45 +03:00
Ahmed Allam
cd3bc8011e fix(resume-picker): keep --instruction on a picked run, tolerate blank instructions, read UTF-8 keys, fit narrow terminals, cap the window at 8 rows 2026-10-04 20:07:47 +03:00
Ahmed Allam
e954301a09 feat(cli): pick a prior run interactively when --resume has no name
strix --resume with no run name now lists the runs in ./strix_runs inline
in the terminal (started, target, status, findings, run name), newest
first, with arrow-key selection, type-to-search and esc to cancel.
Picking a run continues through the same path as --resume <name>.
Headless or non-TTY launches error with the run list instead.
2026-10-04 20:07:47 +03:00
Ahmed Allam
deb6d81d79 refactor(telemetry): keep only the unraisable hook, no docstrings 2026-10-04 18:35:07 +03:00
Ahmed Allam
eaf16c6817 refactor(telemetry): collapse the exception hooks into one plain function
No install-once global, no hand-rolled traceback formatting, no SystemExit special case: one helper logs the report with exc_info on strix.telemetry when a handler exists, and both hooks call it.
2026-10-04 18:35:07 +03:00
Ahmed Allam
8ea2c99225 revert(runtime): drop the PTY teardown change in StrixDockerSandboxClient.delete()
Keep this PR to the logging change only: the finalizer reports are routed to strix.log by the unraisable hook, so the sandbox teardown stays as on main.
2026-10-04 18:35:07 +03:00
Ahmed Allam
fa5f99c4e0 fix(telemetry): route every unraisable and thread exception to strix.log, never stderr
Completes the previous commit, which only carried the test rename. sys.unraisablehook (finalizers, __del__, GC and weakref callbacks, files closed at exit) and threading.excepthook are replaced for the life of the process with hooks that log the report at WARNING on strix.telemetry: strix.log, and stderr only when the stream handler runs at DEBUG (STRIX_DEBUG=1). Nothing is delegated to Python's default printers; a report arriving after the strix handlers are torn down is dropped instead of reaching logging.lastResort, and failures inside logging itself are swallowed.
2026-10-04 18:35:07 +03:00
Ahmed Allam
121216096c fix(telemetry): route every unraisable and thread exception to strix.log, never stderr
Replaces the pattern-matched urllib3/http.client filter with a generic rule: sys.unraisablehook (finalizers, __del__, GC and weakref callbacks, files closed at exit) and threading.excepthook are replaced for the life of the process with hooks that log the report at WARNING on strix.telemetry. That goes to strix.log, and to stderr only when the stream handler runs at DEBUG (STRIX_DEBUG=1). Nothing is delegated to Python's default printers; a report arriving after the strix handlers are torn down is dropped instead of reaching logging.lastResort, and failures inside logging itself are swallowed so nothing can print mid-shutdown.

tests/test_hook_exceptions_logged.py covers both hooks in-process and end to end in a fresh interpreter (a __del__ raising at shutdown, a thread traceback, a urllib3 response whose file was closed first): stderr stays empty.
2026-10-04 18:35:07 +03:00
Ahmed Allam
d945eeee32 fix(telemetry): match only http.client responses in the finalizer filter; test PTY teardown through the real SDK session
The object-shape match accepted any http.* class; it now takes urllib3 and http.client only, so a closed-file error from another module still reaches the previous unraisable hook. New tests build a real DockerSandboxSession holding a PTY exec stream and run the real SDK delete() with the container gone: the SDK alone leaves the socket and its pinned response open, StrixDockerSandboxClient.delete() closes both.
2026-10-04 18:35:07 +03:00
Ahmed Allam
b4be726a3b fix(runtime): terminate PTY exec streams on teardown and silence 3.14 response finalizer noise
StrixDockerSandboxClient.delete() killed the container and then let the SDK's delete() run shutdown(), which only terminates the agent's PTY exec streams while the container still exists. Once the container was gone the hijacked exec sockets were left to the garbage collector and their HTTP responses failed to close at interpreter exit, which Python 3.14 reports as "Exception ignored while finalizing file <urllib3.response.HTTPResponse>" after the scan summary. delete() now awaits pty_terminate_all() first, whatever the container's state.

The unraisable filter now also matches Python 3.14's finalizer shape (object=None, repr in err_msg) and http.client responses, and logs the match at DEBUG on strix.telemetry instead of letting it reach stderr.
2026-10-04 18:35:07 +03:00
Ahmed Allam
25f3978804 fix(cli): install the stderr log handler at startup and prepare the run state before the TUI launches 2026-10-04 05:44:06 +03:00
Ahmed Allam
5325009f65 fix(cli): verify the model before the TUI opens on a direct launch 2026-10-04 05:44:06 +03:00
Ahmed Allam
b49ac2fd9e fix(preflight): check extra headers for subscription models and the dedupe key as sent 2026-10-04 01:48:35 +03:00
Ahmed Allam
08df05db61 fix(preflight): skip unused credentials for subscription models, check dedupe credentials too 2026-10-04 01:48:35 +03:00
Ahmed Allam
4ca15a6e60 style: drop docstrings from the header-value check 2026-10-04 01:48:35 +03:00
Ahmed Allam
539960d953 fix(preflight): name a non-ASCII character in the API key instead of raising UnicodeEncodeError
httpx encodes header values as ASCII, so a smart quote, non-breaking
space or byte-order mark pasted into LLM_API_KEY (or LLM_EXTRA_HEADERS)
surfaced as a bare UnicodeEncodeError from inside the client, wrapped in
ModelConnectionError. The preflight now checks these settings first and
fails with the setting name, the code point, its Unicode name and its
position. Nothing is trimmed or rewritten and the key itself is never
printed.
2026-10-04 01:48:35 +03:00
Ahmed Allam
7280c40caa style: drop docstrings from the UTF-8 stream helpers 2026-10-04 01:19:48 +03:00
Ahmed Allam
5c3476795a fix(cli): force UTF-8 stdout/stderr on Windows so Rich output never raises
Windows hands a redirected or legacy console stream the ANSI code page
(cp1252), which cannot encode Rich's panels or the model's text: one
check mark in a finding ended a headless run with UnicodeEncodeError, and
the error handler raised again rendering its own panel. Reconfigure both
streams to UTF-8 at startup on win32, and open the Go TUI's os.devnull
sink as UTF-8 so logging into it stops producing --- Logging error ---
reports under the same code page.
2026-10-04 01:19:48 +03:00
Ahmed Allam
45b775dcd5 fix(cli): apply the --fail-on threshold regardless of run status
A stopped run with no findings already exits 0, so skipping the threshold
only when low findings exist made the gate depend on the wrong thing.
The threshold now always applies; reporting an unfinished run is a separate
exit-code concern.
2026-10-04 00:03:23 +03:00
siundu254
b11228192d Resolve PR Comments 2026-10-04 00:03:23 +03:00
siundu254
3cd6c93fa0 Fail on Severity 2026-10-04 00:03:23 +03:00
Ahmed Allam
0107a15295 test: reject a leading heading in any finish_scan example field 2026-10-03 23:07:56 +03:00
Ahmed Allam
c0258b25fa fix(finish_scan): stop asking for a section heading in every report field
Every renderer of the final report already titles each section, so the
heading the docstring asked for printed twice. Describe the fields as
section bodies and drop the example headings.
2026-10-03 23:07:56 +03:00
Ahmed Allam
99c0711687 fix(config): use the Responses API whenever the model's catalog entry lists /v1/responses 2026-10-02 18:41:53 +03:00
Ahmed Allam
8ab5e39c06 fix(inputs): send reasoning_effort as configured; no route-specific handling 2026-10-02 18:17:27 +03:00
Ahmed Allam
e1ec259ac2 fix(inputs): send reasoning_effort=none explicitly on chat completions; hint at the Responses API when tools+effort are rejected 2026-10-02 18:17:27 +03:00
Ahmed Allam
066bd60a03 fix(runner): pick the SDK route from the resolved model override
configure_sdk_model_defaults only sees STRIX_LLM; a model= override to
run_strix_scan re-applies the route for the model that actually runs.
2026-10-02 18:17:27 +03:00
Ahmed Allam
6ab123484d fix(config): choose Responses vs chat completions from the model, not the base URL
A base URL no longer forces chat completions. resolve_api_type() keeps an
explicit STRIX_API_TYPE, uses Responses without a base URL or for
api.openai.com, uses Responses for models whose LiteLLM catalog entry has
no /v1/chat/completions endpoint, and chat completions for other gateways.

On the chat completions route, reasoning_effort is omitted for models
whose LiteLLM parameter map does not list it there instead of failing the
request with function tools. STRIX_REASONING_EFFORT and STRIX_API_TYPE
are matched case-insensitively.
2026-10-02 18:17:27 +03:00
alex s
007ed1a94e
fix(reporting): restore create_vulnerability_report parameter descrip… (#1391)
* fix(reporting): restore create_vulnerability_report parameter descriptions

A docstring line beginning with a backtick fence example opened a markdown
code block that griffe's Google-style parser never saw closed, so the Args
section was parsed as plain text and the generated tool schema carried no
per-parameter descriptions. Reword the example, move Args after the trailing
notes so nothing after it is dropped from the tool description, and add a
test asserting every scan-agent tool parameter has a description.

* test(reporting): cover respond_to_user and reject null parameter descriptions
2026-09-30 13:26:48 -04:00
Ahmed Allam
ef272b8e0d chore(models): remove the model quality warning and its allowlists 2026-09-30 10:04:28 +03:00
Ian
9b72488c92 docs + unit test fix 2026-09-30 04:46:01 +03:00
Ian
c814f6bf30 Made session IDs optional 2026-09-30 04:46:01 +03:00
Ian
3fbccc6b1d Non-streaming path 2026-09-30 04:46:01 +03:00
Ian
c9aebc6c87 larger default block size 2026-09-30 04:46:01 +03:00
Ian
95fbd8d687 Openrouter sticky sessions for caching, with telemetry 2026-09-30 04:46:01 +03:00
Ahmed Allam
463b149bdb fix(budget): parked agents count as active; park never overwrites a stop
active_agents_except (finish_scan, wait_for_message) treats budget_paused as
active, so a root cannot finish the scan over a parked child. park_for_budget
only transitions a running agent, and the wake back to running happens under
the coordinator lock.
2026-09-30 03:13:40 +03:00
Ahmed Allam
d355838ea0 feat(budget): budget_policy=pause parks every agent at the limit until the operator resumes
Adds budget_policy: stop | pause to run_strix_scan / ReportUsageHooks /
AgentCoordinator, independent of interactive mode. Under pause the agents
get no budget warnings and no sub-agent reserve; each agent parks before
its next LLM call once spent >= limit or the scan is paused, sessions and
sandbox stay alive, and coordinator.resume_budget(max_budget_usd=...)
replaces the limit and wakes every parked agent without adding anything
to any session. coordinator.pause_budget() parks a running scan the same
way. In-flight calls are never cancelled, so spent may end above the
limit. Parked agents count as active for stop_agent.
2026-09-30 03:13:40 +03:00
ian-at-strix
0ff9f8c324
feat(tui): animate the wait_for_agents indicator (#1383) 2026-09-29 14:39:07 -07:00
ian-at-strix
954bc0d527
perf(prompt): load requested skills after a cache point (#1382)
Siblings differ only in the skills they were spawned with, but those came
first in <specialized_knowledge>, so their prompts diverged at 39%. Shared
skills and the catalog now come first, and the requested skills follow a
cache point, so siblings share 93%.

The extra system message takes a fourth Claude breakpoint, so the Bedrock
tool_config one goes: the first system breakpoint already covers the tools.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 00:35:22 +03:00
ian-at-strix
e66c56c473
perf(llm): give Claude a cache point before the per-run scope (#1376)
* perf(llm): give Claude a cache point before the per-run scope

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(llm): split the system prompt at a generic <cache_point> marker

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 00:35:22 +03:00
ian-at-strix
50425c2c99
perf(prompt): put per-run scope at the end of the system prompt (#1375)
* perf(prompt): put per-run scope at the end of the system prompt

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(prompt): assert scope renders once, after the shared prefix

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 00:35:21 +03:00
devin-ai-integration[bot]
ae38fe70cd
Fill in blank tool-call ids so strict providers accept the history (#1355) 2026-09-23 18:34:22 -07:00
alex s
4c1be22150
Let agents delete a vulnerability report they filed (#1354) 2026-09-23 17:25:23 -07:00
devin-ai-integration[bot]
56f7d45388
feat(llm): structured per-attempt provider request log with provider request ids (#1353) 2026-09-22 20:49:41 -07:00
yoni-at-strix
e158eab3f8
feat(mcp): initialize connections lazily (#1347)
* feat(mcp): initialize connections lazily

* fix(mcp): replace terminally dead sessions

* fix(mcp): improve targeted tool discovery

* fix(mcp): limit active tool fallback
2026-09-22 14:02:02 -04:00
Ahmed Allam
56e9ae982c runtime: read_only local sources become :ro bind mounts
A local_code target can mark its tree read_only; collect_local_sources
forwards the flag and build_bind_mounts mounts the tree read-only instead
of relying on host mode bits, skipping the per-metadata remounts since the
whole tree is already immutable. Used for pulled container image layouts.
2026-09-20 06:10:13 +03:00
Ahmed Allam
355a8bb437 fix(reporting): move the git blame hint to the end of the tool description 2026-09-18 21:39:35 +03:00
Ahmed Allam
77a0cf839b fix(reporting): make the git blame hint a casual inline note 2026-09-18 21:39:35 +03:00