- Extract a shared fetch_run_events_page helper so the three client
paging loops (full list, until, tail) no longer repeat the request/
convert/has_more skeleton; fold the tail loop's two descending-order
checks into one and drop its redundant had_events flag.
- Skip the latest-seq lookup in list_events_before_with_limit when the
caller supplies a before_seq cursor, so a cold projection cache costs
at most one full history scan per pagination session instead of one
per page.
- Remove the dead before_seq max(1) clamp and the passthrough order()
accessor from EventListParams.
- Document the CLI --tail 0 --follow seeding trick and the reader
event_seq placeholder invariant.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Extract the quadruplicated watchdog check-and-clear logic in
schedule_worker_cancel_escalation into ManagedRun methods
(escalation_still_current, clear_escalation_for)
- Derive strum::IntoStaticStr for WorkerRef instead of a hand-written
variant-to-string match in kind()
- Use the generated AgentControlState constant instead of the raw
"waiting_for_steer" literal in run-detail.tsx
- Replace optimisticCancellationRunId state with a boolean; the
component is keyed by run id, so the stored id could only ever be
this run's own
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The crate reorganization renamed lib/crates/ to lib/apps|components|foundation/.
Git followed all modified files across the rename; the only conflict was the
newly added codec/cache.rs, now placed at lib/components/fabro-llm/src/codec/cache.rs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reuses run_multi_turn_cache_test — the same live cache verification the
anthropic, openai, and gemini routes already have. OpenRouter was the
one caching route with no live caller, which is exactly where the
missing-breakpoints bug hid: unit and wire tests prove we now send
cache_control, but only a live call proves OpenRouter forwards it to
Anthropic and cache reads actually appear.
Runs with: set -a && source .env && set +a && \
cargo nextest run -p fabro-llm --profile e2e --run-ignored only \
-E 'test(openrouter_claude_multi_turn_cache)'
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Anthropic prompt caching is opt-in per request: without explicit
ephemeral cache_control breakpoints in the body, no cache writes or
reads ever happen. The OpenAI-compatible codec never emitted them, so
every run on openrouter Claude models billed the full conversation at
the uncached input rate on every turn (0 cache tokens on the billing
page, confirmed by OpenRouter's activity portal).
- Add a `cache_control_breakpoints` model feature declaring that a
route only caches when the request marks the cacheable prefix; set it
on the builtin OpenRouter Claude rows. Catalog build rejects the flag
without `prompt_cache`.
- Teach the Chat Completions wire shape a parts-form content variant so
a message can carry the annotation; unmarked messages keep the
plain-string form for compatibility with strict servers.
- Mark the last system message (covers tools + system upstream) and the
second-to-last user turn, counting tool results as user turns —
mirroring the anthropic codec's placement so agent loops get
incremental cache hits.
- Extract the shared placement/opt-out policy into codec::cache and
refactor the anthropic codec onto it; anthropic wire snapshots are
unchanged.
- Honor `provider_options.<name>.auto_cache = false` as an opt-out and
consume the control key instead of merging it into the body.
- Mirror the new feature through settings (fabro-config), the OpenAPI
schema, and the generated TypeScript client.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Document the contract of read_last_file_routing_json (terminal JSON
extraction only; routing validation happens downstream), extract a
shared sandbox_with_file test helper, and drop the misleading
"standalone" wording from the fallback docs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Return 400 (not 500) for WorkflowError::ModelReference from run
creation, matching ModelSelection: an ambiguous model/provider token
is user input, not a server fault.
- Gate fabro-workflow's test_support module behind
cfg(any(test, feature = "test-support")) so the feature actually
controls exposure, per the repo's test-support boundary guidance.
Add the self dev-dependency so tests/it keeps compiling, and gate
the pipeline helpers that only test_support consumed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Tag every LLM request in a run with an x-session-id header carrying the
run ID, so gateways that understand session tracing (e.g. OpenRouter
broadcast) can group a run's requests into one session.
Adds ExtraHeadersCredentialSource to fabro-auth: a CredentialSource
decorator that appends fixed headers to every resolved credential,
leaving operator-configured extra_headers untouched. The run pipeline
wraps its vault/env source with it, so agent stages, prompt stages,
hooks, and PR-content generation all pick up the header through the
existing extra_headers plumbing with no fabro-llm changes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Both this branch and the remote qa branch fixed the same provider-pin
regression; the merge stacked the two implementations. Keep the remote's
semantics: pin the run's provider whenever it offers the model, otherwise
fall back to priority selection.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- rename resolve_route catalog-instance test to describe its actual
id-based resolution assertion
- use EnvVars::OPENAI_API_KEY instead of a raw string in the automation
scheduler test fixture
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Merging main brought in billing tests that construct ModelRef with String
model ids and an integration test that pins an OpenRouter run via the
backend's provider id. The ModelRef sites now use ModelId conversions.
The integration test also exposed a real regression: resolve_provider_context
ignored the persisted run provider whenever the model selector resolved
globally, re-routing pinned OpenRouter runs to a higher-priority provider for
nodes without explicit model/provider attrs. Request-time routing now treats
the run's selected provider as a pin with custom-model passthrough, matching
transform-time selection semantics.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>