Derive configured providers from env and vault when choosing default
models during run creation and materialization, and thread the resolved
run provider through execution handlers instead of recomputing it.
Also return a user-facing error when fabro-agent cannot infer a default
model for the selected provider.
When `disk_cache = true` in `[server.slatedb]`, Fabro enables SlateDB's
object-store cache at `<storage_root>/cache/slatedb`, caching raw S3
bytes on local disk to reduce read latency. All cache parameters use
SlateDB defaults (16 GB max, 4 MB parts). A warning is emitted if
enabled with `provider = "local"` since the cache adds overhead when
the object store is already on the local filesystem.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
No production deployments exist, so there's no need for migration shims.
Remove all six backwards-compat type aliases (AgentError, SdkError,
CoreError, GraphvizError, StoreError, FabroError) and migrate ~880
callsites to use the canonical Error name directly within each crate,
or qualified imports (e.g., `use fabro_llm::Error as LlmError`) for
cross-crate references. Also fix a pre-existing absolute-path clippy
lint in fabro-server error.rs.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add a shared MiniJinja-based template crate and migrate workflow prompts,
imports, hooks, and InterpString env references to the new {{ ... }}
syntax. This also threads typed run inputs through workflow rendering and
updates docs and tests to match the new templating model.
Cleanup pass on the events schema v2 work merged from origin/main.
Quality fixes:
- prompt.rs: drop dead `_visit` local; use stage_scope.visit at the
emit site (the value was being recomputed inline next to a scope
that already had it).
- llm/cli.rs: rename `_context` to `context` in CodergenBackend::run
(it's actually used now); delete the lingering `current_visit`
helper that was deleted from llm/api.rs in 49767a43f but missed
here; use stage_scope.visit at the emit site.
- llm/api.rs: rename `event_scope` to `stage_scope` for consistency
with every other handler.
- agent.rs, fan_in.rs, parallel.rs: same `visit_from_context` →
`stage_scope.visit` substitution at every event-emit site.
- parallel.rs: switch ParallelStarted/ParallelCompleted from `emit`
to `emit_scoped` so they carry stage_id in the envelope.
- event.rs: fix the StageScope::for_handler docstring — the lifecycle
hook is `before_node`, not `before_attempt`.
Reuse fixes:
- run_event/mod.rs: add `ActorRef::agent(session_id, display)` symmetric
with the existing `ActorRef::user`; use it from agent_actor_for_event
in workflow event.rs.
Correctness fixes:
- event.rs: introduce `StageScope::for_parallel_branch` to name the
"branch starts at visit 1" invariant the parallel handler was
hardcoding via a struct literal at parallel.rs:307. This makes
the assumption auditable and gives a single place to fix when
parallel nodes ever loop.
Efficiency fixes:
- stage_id.rs: switch StageId/ParallelBranchId Serialize impls from
`serializer.serialize_str(&self.to_string())` to `collect_str(self)`,
removing one transient String allocation per ID per emitted event.
Hardening:
- event.rs: add `#[must_use]` on `to_run_event`, `to_run_event_at`,
and `event_name`.
- store/types.rs: add a second wire-envelope round-trip test that
populates stage_id, parallel_group_id, parallel_branch_id,
session_id, parent_session_id, tool_call_id, and actor — the
existing test only exercised stage_id, so a regression in any of
the other envelope fields' #[serde(flatten)] interaction would
have been silent.
All 3810 workspace tests pass; clippy and fmt clean.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Mostly consolidation of code added in the recent schema v2 work:
- Share a single ActorRef::user() constructor between server control
actions and workflow provenance conversions.
- Share StageScope::from_context() between current_stage_scope and
StageScope::for_handler so the 4-field construction lives in one place.
- Collapse RunEvent::to_value's if-let chain into an insert_opt helper.
- Use Value::String(id.to_string()) instead of serde_json::to_value for
StageId/ParallelBranchId when seeding the parallel branch context.
- Share parse_event_envelopes via tests/it/support/mod.rs instead of
duplicating the parsing block in two CLI run_events helpers.
Also fix parallel-branch git.commit to emit via emit_scoped with a
branch-specific StageScope so it carries stage_id / parallel_group_id /
parallel_branch_id alongside the other stage-scoped events.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Populate stage_id / parallel_group_id / parallel_branch_id on every
event tied to a concrete stage execution, per the spec at
docs-internal/fabro-event-schema-v2-concrete-shape.md:223-279.
Before this commit, stored_event_fields() only set stage_id for the
four Event::Stage* variants and Event::Agent -- the only variants
that carried visit/parallel_group_id/parallel_branch_id in their
payload. Every other stage-scoped event (Checkpoint*, PromptCompleted,
Command*, AgentCli*, Prompt, Interview*, Failover, StallWatchdog,
GitCommit, ArtifactCaptured) fell through to node_stored_fields()
and left stage_id as None.
New approach: scope is carried alongside the event, not on the
variant.
- fabro-workflow/src/event.rs: new StageScope type
{ node_id, visit, parallel_group_id, parallel_branch_id }. New
Emitter::emit_scoped(&event, &scope) for stage-level emission.
to_run_event_at and stored_event_fields take an
Option<&StageScope> that merges into the returned envelope
fields. StageScope::for_handler(context, node_id) is the
canonical handler-side constructor -- prefers
context.current_stage_scope() set by the fidelity lifecycle,
falls back to a scope synthesized from the node_id + context
visit count for tests that don't go through the full lifecycle.
- fabro-workflow/src/context.rs: new
WorkflowContext::current_stage_scope() method reads CURRENT_NODE,
internal.node_visit_count, internal.parallel_group_id,
internal.parallel_branch_id from the context.
- Remove the now-redundant visit/parallel_group_id/parallel_branch_id
fields from Event::Stage{Started,Completed,Failed,Retrying} and
the parallel_* fields from Event::Agent. These existed only to
feed stored_event_fields() and are obsolete once scope is
threaded through the emitter.
Emission site migration (all stage-scoped handlers now use
emit_scoped):
- lifecycle/event.rs: StageStarted, StageCompleted, StageFailed,
StageRetrying, CheckpointCompleted, GitCommit (from on_checkpoint)
- lifecycle/git.rs: CheckpointFailed
- lifecycle/artifact.rs: ArtifactCaptured
- handler/command.rs: CommandStarted, CommandCompleted
- handler/prompt.rs: Prompt, PromptCompleted
- handler/agent.rs: Prompt, PromptCompleted
- handler/fan_in.rs: Prompt, PromptCompleted
- handler/human.rs: InterviewStarted, InterviewTimeout,
InterviewInterrupted, InterviewCompleted
- handler/llm/api.rs: Failover, Agent (via spawn_event_forwarder
which now carries a StageScope across the tokio::spawn boundary)
- handler/llm/cli.rs: AgentCliStarted, AgentCliCompleted
- handler/parallel.rs: ParallelBranchStarted, ParallelBranchCompleted
StallWatchdogTimeout stays on plain emit() because the watchdog
fires from an error path without a live stage context.
Deleted the local StageEventScope struct + current_stage_event_scope
helper from handler/llm/api.rs; it's generalized into StageScope.
Tests: two new unit tests in event.rs --
stage_scope_populates_stage_id_on_non_stage_events verifies
CommandStarted / Prompt / GitCommit all pick up stage_id from scope,
run_level_events_without_scope_leave_stage_id_absent confirms
run.* events still get no stage scope. Updated all test fixtures
across fabro-workflow, fabro-cli to drop the removed Event variant
fields. Accepted two insta snapshot updates in
fabro-cli/tests/it/cmd/{attach,run}.rs that now include the
formerly-missing stage_id fields on checkpoint and interview events.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Promote RunEvent.stage_id / parallel_group_id / parallel_branch_id
and the internal Event enum's matching fields from stringly-typed
Option<String> to Option<StageId> / Option<ParallelBranchId>. The
wire contract is now self-enforcing: malformed strings are rejected
at the serde seam, not quietly round-tripped, and the three
StageId::new(...).to_string() calls in stored_event_fields() just
drop the .to_string() since the newtypes flow straight through.
- fabro-types/src/stage_id.rs: new ParallelBranchId { group: StageId,
index: u32 } mirroring StageId's Display / FromStr / serde string
form. "{group}:{index}" (e.g. "fanout@2:0"). Tests for round-trip
and parse rejections.
- fabro-types/src/lib.rs: re-export ParallelBranchId.
- fabro-types/src/run_event/mod.rs: RunEvent, RunEventRaw, and
RunEventParts take Option<StageId> / Option<ParallelBranchId>.
from_ref gains a small generic opt_field<T: Deserialize> helper
that also replaces the bespoke actor null-handling branch. to_value
uses serde_json::to_value(value) for the three typed fields.
- fabro-workflow/src/event.rs: Event::Stage{Started,Completed,
Failed,Retrying} and Event::Agent take Option<StageId> /
Option<ParallelBranchId>. Event::ParallelBranch{Started,Completed}
take the required (non-Option) typed forms. StoredEventFields
and stored_event_fields() plumb the newtypes end-to-end.
- fabro-workflow/src/context.rs: WorkflowContext::parallel_group_id()
returns Option<StageId>, parallel_branch_id() returns
Option<ParallelBranchId>. Read via serde_json::from_value which
validates the shape on the way out.
- fabro-workflow/src/handler/parallel.rs: builds typed values
directly, stores in context via serde_json::to_value (still
produces a JSON string through the custom Serialize). BranchSetup
holds a ParallelBranchId.
- fabro-workflow/src/handler/llm/api.rs: StageEventScope holds
typed ids.
- fabro-workflow/src/lifecycle/event.rs: stage_parallel_ids returns
typed tuple.
Wire JSON is byte-identical before and after (StageId serializes as
"{node_id}@{visit}", ParallelBranchId as "{node_id}@{visit}:{index}",
matching the existing spec). Progenitor-generated types and OpenAPI
schema untouched. Existing None-only fixtures in runtime_store,
git, pipeline, error, run_state, rewind, pr_view, and store/dump
didn't need any edit because None fits any Option<T>.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Centralize flattened EventEnvelope conversion in fabro-store so the CLI,
server, and test helpers reuse one wire-shape path. Also thread parallel
group and branch ids through nested stage and agent events so the new
envelope fields stay populated inside parallel branches.
Adds visit: u32 to Event::StageStarted/Completed/Failed/Retrying so
stored_event_fields() can derive stage_id = "{node_id}@{visit}".
Adds parallel_group_id/parallel_branch_id to ParallelBranchStarted/
Completed Events, computed once in handler/parallel.rs from the
parent parallel node id + visit_from_context + branch index.
Emission sites in lifecycle/event.rs populate visit from
state.node_visits via a new stage_visit helper.
Stored_event_fields() still leaves stage_id and parallel ids None
pending the extraction pass in the next commit.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Move detached workers onto an HTTP-backed runtime store so the server
remains the only SlateDB owner. This replaces the worker's seeded local
RunDatabase with a canonical server-backed handle for state, events, and
blobs, and updates workflow runtime plumbing to use that abstraction.
Pass a real cancel signal from __run-worker through the workflow engine
into sandbox command execution so cancelled runs reap gated shell loops
instead of leaking slow.gate waiters.
- centralize FABRO_HOME and storage path resolution in fabro-config
- rename store types, extract ArtifactStore, and simplify run key layout
- switch run scratch to scratch/, remove RuntimeState, and refresh docs/clients
Move durable run metadata and path derivation onto RunId, simplify the
Slate catalog/index format, and carry the storage-specific run directory
through workflow creation so detached and lookup flows stay aligned.
Also update affected CLI snapshots and test helpers to match the new
run discovery behavior.
All data in these files is already stored in SlateDB via events and
projected into RunState. No production code reads them from disk.
Removed writes: prompt.md, response.md, stdout.log, stderr.log,
script_invocation.json, script_timing.json, parallel_results.json,
provider_used.json, retro/{prompt,response,status,session}, live.json,
detached_failure.json.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Follow-up to the prior commit that removed production callers. This commit:
- Removes InMemoryRunStore methods and dead key functions
- Rewrites store/workflow tests to use append_event + state() instead of removed methods
- Updates CLI snapshot tests for new event-projected output
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Make run lookup fail with RunNotFound instead of returning Option, thread a required RunStore through workflow and retro paths, and update CLI, server, and test callers to match. Also treat null optional event properties as absent during store-backed replay so event-sourced state stays robust.
Aligns naming with the convention that "Config" is for file-level configuration
while "Options" and "Settings" describe runtime parameters. Also applies
rustfmt formatting fixes in web_auth.rs.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add richer run, stage, prompt, command, retro, and agent session event
metadata so progress output and stored workflow events carry the context
needed by the new plan. Normalize event serialization and update CLI log
handling to prefer progress.jsonl with consistent redaction, and fix the
detached wait/log race covered by the updated integration and snapshot
tests.