Commit graph

17 commits

Author SHA1 Message Date
Hazemwaddah
1d05637b68
fix(cost): close three review gaps in the fan-out cost knobs
Three issues raised in review on #1217, all real:

- The STRIX_MAX_AGENTS check read agent_count() and then spawned, so two
  parents racing for the last slot both passed before either child
  registered and the graph overshot the cap. The coordinator now hands out
  slots atomically (try_reserve_agent_slot / release_agent_slot), counting
  outstanding reservations alongside live agents under the same lock. The
  slot is released once the spawner has registered the child or failed.
  Depth stays a plain check — a parent's depth cannot change mid-spawn.

- _trim_parent_history kept the newest item whole when that item alone
  exceeded the budget ("and kept"), so the cap bounded nothing on the
  child's first request and could overflow the provider context window. An
  oversized newest item is now rendered as a truncated text message. No
  tool-call pairing can break: nothing else survives that trim.

- --reasoning-effort exported STRIX_REASONING_EFFORT, which persist_current()
  then wrote into cli-config.json, turning a documented per-run flag into
  the default for every later run. Env vars can now be marked run-scoped and
  are skipped when persisting.

Tests cover the concurrent reservation, slot release, the oversized-item
trim, and the persist exemption.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 10:55:59 +03:00
Hazemwaddah
7712e65d51
feat(cost): opt-in agent fan-out caps, context trim, --reasoning-effort
Autonomous scans can spawn unbounded agents, each re-paying the full
system prompt on every turn and (by default) inheriting a full copy of
the parent's history — a large token-cost driver on a single target.
This adds knobs to bound it, all OFF by default so out-of-the-box
behavior is unchanged.

New (opt-in via env, default 0 = disabled):
- STRIX_MAX_AGENTS — cap total agents in the graph; create_agent refuses
  past the cap with a model-facing message to reuse/wait/self-serve.
- STRIX_MAX_AGENT_DEPTH — cap spawn depth (root = 1).
- STRIX_INHERIT_CONTEXT_MAX_TOKENS — trim inherited parent history to the
  most-recent tail within a token budget.

Also:
- New --reasoning-effort CLI flag (overrides STRIX_REASONING_EFFORT per
  run); pure addition, no default change.

Coordinator gains agent_count()/depth_of() helpers. Tests cover the caps
and the history trim. Docs updated. No default behavior changes.
2026-09-13 10:55:07 +03:00
devin-ai-integration[bot]
cf179d564e
fix(llm): bind dedupe credentials to a provider; send reasoning=max via extra_body (#1187)
Co-authored-by: Ahmed Allam <ahmed39652003@gmail.com>
2026-08-28 09:27:57 -07:00
devin-ai-integration[bot]
583af23d9a
fix(llm): only attach prompt-cache points on routes LiteLLM serves (#1186)
Co-authored-by: Ahmed Allam <ahmed39652003@gmail.com>
2026-08-28 09:26:57 -07:00
devin-ai-integration[bot]
391d81bea7
feat(agents): evidence discipline, and coverage as a first-class artifact (#961)
Co-authored-by: Ahmed Allam <ahmed39652003@gmail.com>
2026-08-24 03:34:09 -07:00
devin-ai-integration[bot]
174c16fa26
fix(llm): send OpenRouter app attribution on the request itself (#1045) 2026-08-10 11:24:02 -07:00
Ahmed Allam
649a2e2140 fix(llm): omit parallel_tool_calls on tool-less requests 2026-08-09 15:44:16 +03:00
oyasumi
5bb9fe896b
feat(tui): replace Textual with a Go/Bubble Tea interface (#941) 2026-08-03 19:23:07 -07:00
devin-ai-integration[bot]
2e7040240d
feat(config): accept STRIX_REASONING_EFFORT=max for providers that support it (#956)
Co-authored-by: Ahmed Allam <allam@usestrix.com>
2026-08-01 17:10:01 -07:00
devin-ai-integration[bot]
d4e58b2cd0
fix(llm): pass LLM_EXTRA_HEADERS through ModelSettings so they reach the agent loop (#937) 2026-07-29 19:38:06 -07:00
seanturner83
27f9750cdc
feat(llm): enable Bedrock/Anthropic prompt caching for Claude models (#772)
Co-authored-by: Sean Turner <sean.turner@zerohash.com>
Co-authored-by: Ahmed Allam <ahmed39652003@gmail.com>
2026-07-26 17:12:57 -07:00
Ahmed Allam
7d5a67d234 chore(llm): shorten timeout helper docstring; update tests 2026-07-17 19:45:32 -07:00
Ahmed Allam
cf7689e927 fix(llm): use httpx.Timeout read-inactivity for per-turn model timeout 2026-07-17 18:40:23 -07:00
Ahmed Allam
3bb95ab43d fix(llm): add per-turn model request timeout so stalled streams fail fast and retry 2026-07-17 18:40:23 -07:00
alex s
c13960ae01 Support routed OpenAI required tool choice (#732) 2026-07-10 18:43:07 -04:00
alex s
22d327d21f feat(settings): add force_required_tool_choice to LlmSettings (#730)
feat(inputs): implement logic for required tool choice based on model

test(inputs): add tests for force_required_tool_choice behavior

test(runner): update tests to include force_required_tool_choice in settings
2026-07-10 18:36:33 -04:00
Rome Thorstenson
8cdf0683a3
fix(core): collapse child agent initial input into a single user message (#589) 2026-06-29 06:51:47 -07:00