Greptile caught a real logic bug in the initial patch: the `continue`
inside the transient-thinking branch advanced the outer
`for attempt in range(max_retries + 1)` counter, so each transient retry
consumed a generic-retry slot. With the default `max_retries=5`:
- Loop runs 6 iterations (attempts 0–5)
- Each transient retry burned one iteration
- On attempt 5, `attempt >= max_retries` fired `_raise_error` before the
6th sleep (160s) ever ran
- The documented "fall through to generic retry" path was unreachable
Fix: inner `while` loop that does its own `_stream()` retry without
advancing the outer `attempt` counter. The transient budget of 6 (5 / 10
/ 20 / 40 / 80 / 160 s) is now independent of `max_retries`, and once
the inner budget is exhausted the most recent exception falls through
to the generic retry path as the PR originally intended.
Also extracts the literal `6` into `max_transient_thinking_retries` for
readability — still unconfigurable since the envelope budget is based on
observed Bedrock transient durations, not something we want users tuning
blindly.
Bedrock occasionally returns HTTP 400 claiming the assistant's thinking
blocks have been modified, even when the payload is structurally valid.
The same payload replayed immediately succeeds, indicating a transient
server-side condition rather than a client bug.
Observed on bedrock/us.anthropic.claude-sonnet-4-6 with adaptive thinking
(reasoning_effort=high maps to {type: "adaptive"} on claude-4-6). The
error message is:
messages.N.content.M: `thinking` or `redacted_thinking` blocks in the
latest assistant message cannot be modified. These blocks must remain
as they were in the original response.
Reproduced during a large-scale multi-scan study (800+ Strix runs). Most
failures clear within seconds but we observed one case where the error
persisted across ~2 minutes of backoff, so the retry budget allows up to
~5 minutes total.
Fix:
- Add _is_transient_thinking_error() helper that matches on the status
code (400) AND the characteristic message ("thinking" + "cannot be
modified"), to avoid treating unrelated 400s as transient.
- On detection, retry with exponential backoff (5s, 10s, 20s, 40s, 80s,
160s — total ~5 min) before falling through to the generic retry path.
- Self-contained: the detector checks the exception directly, so it works
with or without other 400 handlers in place.
Testing:
- Replayed 33 captured payloads that appeared adjacent to a failure: all
succeeded, confirming the condition is transient and not a client-side
malformation.
- In a multi-repo scan study, 3-retry budget was insufficient in one case
(exhausted retries over ~30s). The 6-retry budget implemented here
succeeded on the same repo on a subsequent attempt.
* fix: --config flag now fully overrides ~/.strix/cli-config.json (fixes#377)
Previously, env vars applied from the default config at module import time
were not cleared when --config was later processed, causing settings from
~/.strix/cli-config.json to leak into runs that specified a custom config.
Track which vars were applied by the initial default-config load in
Config._applied_from_default. In apply_config_override, clear those vars
before applying the custom config so only the custom file's settings take effect.
* Add config override regression test
* Make config override test setup explicit
---------
Co-authored-by: octo-patch <octo-patch@github.com>
Co-authored-by: bearsyankees <bearsyankees@gmail.com>
* fix: wrap acompletion in asyncio.wait_for to prevent indefinite hangs
litellm's timeout parameter doesn't always propagate to the underlying
httpx transport for Bedrock converse streaming. When Bedrock accepts the
TCP connection but never starts streaming chunks, the acompletion call
hangs indefinitely with all connections in CLOSED state.
This wraps the acompletion call in asyncio.wait_for() using the
configured LLM_TIMEOUT (default 300s). TimeoutError is already retryable
via _should_retry (status_code=None), so the retry loop handles it.
Diagnosed via faulthandler thread dump showing the main asyncio event
loop blocked in selectors.select() with no pending callbacks.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: add per-chunk timeout to streaming loop
Addresses review feedback: the initial asyncio.wait_for only guards the
acompletion call. If Bedrock returns headers but stalls mid-stream, the
async for loop could still hang indefinitely.
Replaces async for with explicit __anext__ calls wrapped in
asyncio.wait_for, using the same configured timeout. Mid-stream stalls
now raise TimeoutError and trigger the existing retry logic.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Sean Turner <sean.turner@zerohash.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Models occasionally output text-only narration ("Planning the
assessment...") without a tool call, which halts the interactive agent
loop since the system interprets no-tool-call as "waiting for user
input." Rewrite both interactive and autonomous prompt sections to make
the tool-call requirement absolute with explicit warnings about the
system halt consequence.
- Change default model from gpt-5 to gpt-5.4 across docs, tests, and examples
- Remove Strix Router references from docs, quickstart, overview, and README
- Delete models.mdx (Strix Router page) and its nav entry
- Simplify install script to suggest openai/ prefix directly
- Keep strix/ model routing support intact in code
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Re-architects the agent loop to support interactive (chat-like) mode
where text-only responses pause execution and wait for user input,
while tool-call responses continue looping autonomously.
- Add `interactive` flag to LLMConfig (default False, no regression)
- Add configurable `waiting_timeout` to AgentState (0 = disabled)
- _process_iteration returns None for text-only → agent_loop pauses
- Conditional system prompt: interactive allows natural text responses
- Skip <meta>Continue the task.</meta> injection in interactive mode
- Sub-agents inherit interactive from parent (300s auto-resume timeout)
- Root interactive agents wait indefinitely for user input (timeout=0)
- TUI sets interactive=True; CLI unchanged (non_interactive=True)
The perplexity API key check in strix/tools/__init__.py used
Config.get() which only checks os.environ. At import time, the
config file (~/.strix/cli-config.json) hasn't been applied to
env vars yet, so the check always returned False.
Replace with _has_perplexity_api() that checks os.environ first
(fast path for SaaS/env var), then falls back to Config.load()
which reads the config file directly.
Users can now access the Caido web UI from their browser to inspect traffic,
replay requests, and perform manual testing alongside the automated scan.
- Map Caido port (48080) to a random host port in DockerRuntime
- Add caido_port to SandboxInfo and track across container lifecycle
- Display Caido URL in TUI sidebar stats panel with selectable text
- Bind Caido to 0.0.0.0 in entrypoint (requires image rebuild)
- Bump sandbox image to 0.1.12
- Restore discord link in exit screen
The badge image URL used invite code which is expired,
causing the badge to render 'Invalid invite' instead of the server info.
Updated to use the vanity URL which resolves correctly.
Fixes#313