Commit graph

236 commits

Author SHA1 Message Date
Ahmed Allam
7280c40caa style: drop docstrings from the UTF-8 stream helpers 2026-10-04 01:19:48 +03:00
Ahmed Allam
5c3476795a fix(cli): force UTF-8 stdout/stderr on Windows so Rich output never raises
Windows hands a redirected or legacy console stream the ANSI code page
(cp1252), which cannot encode Rich's panels or the model's text: one
check mark in a finding ended a headless run with UnicodeEncodeError, and
the error handler raised again rendering its own panel. Reconfigure both
streams to UTF-8 at startup on win32, and open the Go TUI's os.devnull
sink as UTF-8 so logging into it stops producing --- Logging error ---
reports under the same code page.
2026-10-04 01:19:48 +03:00
Ahmed Allam
45b775dcd5 fix(cli): apply the --fail-on threshold regardless of run status
A stopped run with no findings already exits 0, so skipping the threshold
only when low findings exist made the gate depend on the wrong thing.
The threshold now always applies; reporting an unfinished run is a separate
exit-code concern.
2026-10-04 00:03:23 +03:00
siundu254
b11228192d Resolve PR Comments 2026-10-04 00:03:23 +03:00
siundu254
3cd6c93fa0 Fail on Severity 2026-10-04 00:03:23 +03:00
Ahmed Allam
0107a15295 test: reject a leading heading in any finish_scan example field 2026-10-03 23:07:56 +03:00
Ahmed Allam
c0258b25fa fix(finish_scan): stop asking for a section heading in every report field
Every renderer of the final report already titles each section, so the
heading the docstring asked for printed twice. Describe the fields as
section bodies and drop the example headings.
2026-10-03 23:07:56 +03:00
Ahmed Allam
99c0711687 fix(config): use the Responses API whenever the model's catalog entry lists /v1/responses 2026-10-02 18:41:53 +03:00
Ahmed Allam
8ab5e39c06 fix(inputs): send reasoning_effort as configured; no route-specific handling 2026-10-02 18:17:27 +03:00
Ahmed Allam
e1ec259ac2 fix(inputs): send reasoning_effort=none explicitly on chat completions; hint at the Responses API when tools+effort are rejected 2026-10-02 18:17:27 +03:00
Ahmed Allam
066bd60a03 fix(runner): pick the SDK route from the resolved model override
configure_sdk_model_defaults only sees STRIX_LLM; a model= override to
run_strix_scan re-applies the route for the model that actually runs.
2026-10-02 18:17:27 +03:00
Ahmed Allam
6ab123484d fix(config): choose Responses vs chat completions from the model, not the base URL
A base URL no longer forces chat completions. resolve_api_type() keeps an
explicit STRIX_API_TYPE, uses Responses without a base URL or for
api.openai.com, uses Responses for models whose LiteLLM catalog entry has
no /v1/chat/completions endpoint, and chat completions for other gateways.

On the chat completions route, reasoning_effort is omitted for models
whose LiteLLM parameter map does not list it there instead of failing the
request with function tools. STRIX_REASONING_EFFORT and STRIX_API_TYPE
are matched case-insensitively.
2026-10-02 18:17:27 +03:00
alex s
007ed1a94e
fix(reporting): restore create_vulnerability_report parameter descrip… (#1391)
* fix(reporting): restore create_vulnerability_report parameter descriptions

A docstring line beginning with a backtick fence example opened a markdown
code block that griffe's Google-style parser never saw closed, so the Args
section was parsed as plain text and the generated tool schema carried no
per-parameter descriptions. Reword the example, move Args after the trailing
notes so nothing after it is dropped from the tool description, and add a
test asserting every scan-agent tool parameter has a description.

* test(reporting): cover respond_to_user and reject null parameter descriptions
2026-09-30 13:26:48 -04:00
Ahmed Allam
ef272b8e0d chore(models): remove the model quality warning and its allowlists 2026-09-30 10:04:28 +03:00
Ian
9b72488c92 docs + unit test fix 2026-09-30 04:46:01 +03:00
Ian
c814f6bf30 Made session IDs optional 2026-09-30 04:46:01 +03:00
Ian
3fbccc6b1d Non-streaming path 2026-09-30 04:46:01 +03:00
Ian
c9aebc6c87 larger default block size 2026-09-30 04:46:01 +03:00
Ian
95fbd8d687 Openrouter sticky sessions for caching, with telemetry 2026-09-30 04:46:01 +03:00
Ahmed Allam
463b149bdb fix(budget): parked agents count as active; park never overwrites a stop
active_agents_except (finish_scan, wait_for_message) treats budget_paused as
active, so a root cannot finish the scan over a parked child. park_for_budget
only transitions a running agent, and the wake back to running happens under
the coordinator lock.
2026-09-30 03:13:40 +03:00
Ahmed Allam
d355838ea0 feat(budget): budget_policy=pause parks every agent at the limit until the operator resumes
Adds budget_policy: stop | pause to run_strix_scan / ReportUsageHooks /
AgentCoordinator, independent of interactive mode. Under pause the agents
get no budget warnings and no sub-agent reserve; each agent parks before
its next LLM call once spent >= limit or the scan is paused, sessions and
sandbox stay alive, and coordinator.resume_budget(max_budget_usd=...)
replaces the limit and wakes every parked agent without adding anything
to any session. coordinator.pause_budget() parks a running scan the same
way. In-flight calls are never cancelled, so spent may end above the
limit. Parked agents count as active for stop_agent.
2026-09-30 03:13:40 +03:00
ian-at-strix
0ff9f8c324
feat(tui): animate the wait_for_agents indicator (#1383) 2026-09-29 14:39:07 -07:00
ian-at-strix
954bc0d527
perf(prompt): load requested skills after a cache point (#1382)
Siblings differ only in the skills they were spawned with, but those came
first in <specialized_knowledge>, so their prompts diverged at 39%. Shared
skills and the catalog now come first, and the requested skills follow a
cache point, so siblings share 93%.

The extra system message takes a fourth Claude breakpoint, so the Bedrock
tool_config one goes: the first system breakpoint already covers the tools.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 00:35:22 +03:00
ian-at-strix
e66c56c473
perf(llm): give Claude a cache point before the per-run scope (#1376)
* perf(llm): give Claude a cache point before the per-run scope

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(llm): split the system prompt at a generic <cache_point> marker

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 00:35:22 +03:00
ian-at-strix
50425c2c99
perf(prompt): put per-run scope at the end of the system prompt (#1375)
* perf(prompt): put per-run scope at the end of the system prompt

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(prompt): assert scope renders once, after the shared prefix

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 00:35:21 +03:00
devin-ai-integration[bot]
ae38fe70cd
Fill in blank tool-call ids so strict providers accept the history (#1355) 2026-09-23 18:34:22 -07:00
alex s
4c1be22150
Let agents delete a vulnerability report they filed (#1354) 2026-09-23 17:25:23 -07:00
devin-ai-integration[bot]
56f7d45388
feat(llm): structured per-attempt provider request log with provider request ids (#1353) 2026-09-22 20:49:41 -07:00
yoni-at-strix
e158eab3f8
feat(mcp): initialize connections lazily (#1347)
* feat(mcp): initialize connections lazily

* fix(mcp): replace terminally dead sessions

* fix(mcp): improve targeted tool discovery

* fix(mcp): limit active tool fallback
2026-09-22 14:02:02 -04:00
Ahmed Allam
56e9ae982c runtime: read_only local sources become :ro bind mounts
A local_code target can mark its tree read_only; collect_local_sources
forwards the flag and build_bind_mounts mounts the tree read-only instead
of relying on host mode bits, skipping the per-metadata remounts since the
whole tree is already immutable. Used for pulled container image layouts.
2026-09-20 06:10:13 +03:00
Ahmed Allam
355a8bb437 fix(reporting): move the git blame hint to the end of the tool description 2026-09-18 21:39:35 +03:00
Ahmed Allam
77a0cf839b fix(reporting): make the git blame hint a casual inline note 2026-09-18 21:39:35 +03:00
Ahmed Allam
cafa4b19fd fix(reporting): keep git blame guidance to the technical_analysis field 2026-09-18 21:39:35 +03:00
alex s
976835194d
Prompt agents to include local Git blame in technical details (#1329)
* Enrich issue technical details with local Git blame

* Bound report history enrichment and require unambiguous repository identity

* test(history): drive attribution through the CLI scan setup and isolate git config

* Simplify Git blame attribution to existing reporting instructions

* Make local blame guidance reliable in live reporting
2026-09-18 13:05:53 -04:00
Ahmed Allam
4c1f00d1ee fix(runtime): tear the sandbox down when staging is cancelled
CancelledError is not an Exception, so a run cancelled during the extra-file
upload or unpack left a created-but-uncached sandbox running.
2026-09-17 21:53:35 +03:00
Ahmed Allam
46d7bdb290 fix(runtime): place extra files as agent-writable sandbox files on every backend
Extra files (knowledge trees, workspace files) reached the docker sandbox as
per-file read-only bind mounts whose parent directories docker created as
root, so the sandbox user could neither edit them nor create siblings. They
now travel as one tar archive uploaded after bring-up and unpacked as the
sandbox user, on every backend.
2026-09-17 21:53:35 +03:00
alex s
65d495bb7f
feat(config): STRIX_API_TYPE forces responses vs chat completions (#1324)
* Add api_type field to LlmSettings

Added 'api_type' field to LlmSettings for API path selection.

* Refactor API type handling in models.py

* Implement test for LlmSettings API type

Add test for API type override settings in LlmSettings.

* fix(tests): lint api_type test, cover the api_base override route, document STRIX_API_TYPE

* fix(models): keep LiteLLM chat-completions tool schema when STRIX_API_TYPE=responses

---------

Co-authored-by: RAJVARDHAN <95933896+vardhans07@users.noreply.github.com>
2026-09-16 11:13:30 -04:00
Ahmed Allam
84f4108195 fix(web_search): send only the agent's query to Exa search
Exa /search is a neural search endpoint, not a chat model, so prepending
the Perplexity system prompt made Exa match the prompt's own vocabulary
(Kali, OWASP, apt, NIST) instead of the query. The system prompt stays
on the Perplexity path where it is a chat system message; the Exa
summary instruction is unchanged.
2026-09-13 19:40:50 +03:00
alex s
22959a7ba6
feat(reporting): link HTTP exchange evidence (#1281)
Co-authored-by: Ahmed Allam <ahmed39652003@gmail.com>
2026-09-09 07:50:23 -07:00
devin-ai-integration[bot]
52b1923347
fix(models): frontier model check matches the model name only, never the provider route (#1280)
Co-authored-by: Ahmed Allam <ahmed39652003@gmail.com>
2026-09-06 11:09:24 -07:00
Ahmed Allam
2e1db25786 feat(telemetry): classify error beacons by phase and exception class
error events now carry phase (startup/preflight/sandbox_init/agent_setup/
agent_loop) and the exception class name (plus its cause), never the message
or trace. Startup and preflight failures that exit(1) before the scan starts
are beaconed with a stable error_type instead of vanishing. scan_ended
distinguishes budget_exceeded, rate_limited, and headless agent_stopped
from user_exit.
2026-09-05 04:08:09 +03:00
Ahmed Allam
a3bf864e1e test(warmup): assert wait_for_import_warmup blocks until the thread finishes 2026-09-05 02:45:34 +03:00
Ahmed Allam
e60fd83931 refactor(warmup): drop the orphan purge and join the warm-up once before the engine imports 2026-09-05 02:45:34 +03:00
Ahmed Allam
7f46dd17d3 fix(cli): wait for the import warm-up before importing the agents SDK on the main thread
The warm-up thread imports strix.core.runner while warm_up_llm and
preflight_model_connection import agents.models.interface. Both walk the
agents SDK graph from different entry points, CPython fails one side to
break the import-lock cycle, and the orphan purge then removes agents.*
from sys.modules while the main thread is still importing it, crashing
strix -n with KeyError: 'agents.models'.
2026-09-05 02:45:34 +03:00
devin-ai-integration[bot]
afa7c4a77f
feat(web_search): add Exa as a web search provider alongside Perplexity (#1270) 2026-09-04 10:34:28 -07:00
oyasumi
f6d9790ecb fix(viewer): show stopped run status 2026-09-04 01:10:10 +03:00
alex s
5d015df6b1
fix(cloud): print top-up instructions on 402 and guide oversize or archive --source (#1242)
- Every payment-required error now ends with a "Next step" line: the
  platform hint when one is sent, else the topup command and the billing
  URL for the configured platform. JSON output gets the same text as
  next_step. The platform hint is no longer repeated inside the error.
- An archive file passed to --source is rejected with guidance to pass
  the directory instead, which packs and excludes deps/build output.
- An oversize archive names its largest files and points to --exclude
  and --dry-run --show-files.
- uploads request help points to scans start --source for local code.
2026-09-02 15:26:37 -04:00
Ahmed Allam
1edafd3e80 fix(agents): stop parents waiting on finished non-interactive children
A non-interactive agent's loop returns after its terminal state, yet
send_message_to_agent kept reporting messages to it as delivered and the
parent then waited out wait_for_agents on a reply that could never come.

- AgentRuntime.resumable records whether the loop parks for wake-ups after a
  terminal state; run_agent_loop / _start_child_runner set it from interactive.
- AgentCoordinator.send returns False (nothing queued) for a terminal agent
  that is not resumable; send_message_to_agent surfaces target_status and
  delivery_status=not_delivered with a pointer to list_reports / get_report.
- wait_for_agents returns wait_outcome=no_active_agents at once when no other
  agent is running or waiting in a non-interactive run.
- agent_finish reads the reports the finishing agent filed from the report
  state and puts their ids in the completion report, the parent message
  (filed_report_ids) and its own return payload, so parents no longer have to
  infer what was filed from prose.
2026-09-02 22:11:12 +03:00
Ahmed Allam
0ab7244807 feat(models): refresh the recommended model list and docs examples
Add Claude Fable 5.1, Gemini 3.7 Flash, and Z.ai GLM-5.3 / GLM-5.3-Flash
to RECOMMENDED_MODEL_NAMES, add a Z.ai GLM frontier family so GLM-5.x is
accepted through OpenRouter and Novita routes, and drop the superseded
GPT-5.4, GPT-5.3-codex, Opus 4.8, Sonnet 4.6, Gemini 3.6 Flash, and
Qwen3.7 entries. Update the README, docs provider pages, quickstart, and
CLI hint strings to the same current models, including DeepSeek V4,
Kimi K3, and GLM-5.3.
2026-09-02 16:52:53 +03:00
Ahmed Allam
c514f712f4 fix(config): persist only the alias the runtime settings read
pydantic-settings takes the first alias present in the environment, even
when it is empty. persist_current() must save that same alias, so an empty
LLM_API_KEY does not let a non-empty OPENAI_API_KEY sibling land in the
file and restore a credential the run did not use.
2026-09-02 16:10:52 +03:00