Commit graph

3028 commits

Author SHA1 Message Date
Bryan Helmkamp
1166caed0b
Give up on an unreachable worker control stream and reap .ft- test daemons
A worker whose server is killed outright retried its control-stream
connect forever at a 5 s capped backoff, so it never finalized, never
stopped its sandbox, and lived until reboot. The worker now tracks the
start of each run of continuous connection failure and gives up after
60 s (10 s once its parent is pid 1), through the existing fatal
control-loss path that interrupts interviews and cancels the run. After
that fatal fires, the runner keeps driving the cancelled pipeline for a
bounded grace so `conclude` can stop the sandbox before the process
exits, instead of dropping the pipeline future mid-flight.

The test harness's stale-daemon reaper only matched `fabro server`
titles bound under `/tmp/.tmp*`, but `TestContext` roots are
`.ft-<label>-*` under `std::env::temp_dir()`, so daemons bound there
were never reaped. The regex now also matches sockets below a `.ft-`
directory component under any parent, and the unit test covers both
roots plus real-looking paths and `.ft-` outside a directory component.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:27:55 -04:00
Bryan Helmkamp
3e065b806b
Bump lithos-llm to 43a42ac and migrate catalogs to the codecs schema
Move the lithos-llm pin from 55add459 to 43a42ac28e9d9bcf40a91abc02be4f12ca274ebb,
and the three Pebble pins from a39f43e to 67c9f48, Pebble `main`, which pins
that same lithos-llm revision so Cargo holds one lithos-llm crate. lithos-llm
`main` (ca19fac) is one commit further; that commit touches only its nightly
workflow, so this pin stays on the revision Pebble unifies with.

The `openai`, `anthropic`, `gemini`, and `openai-compatible` features are
gone upstream; each expanded to `runtime`, which `bedrock` implies, so the
four names leave the fabro-llm feature list. Every other manifest already
names `runtime`.

The catalog schema now names one adapter and many codecs per provider.
`adapter` defaults to `http`, `codecs = [...]` replaces `codec` and defaults
to `["openai-chat"]`, and the loader rejects the old `codec` key and the four
protocol-named adapter ids. Every inline catalog in tests and docs moves to
the new shape: the `openai-compatible` + `openai-chat` pair is dropped as the
default, `adapter = "openai"` + `codec = "openai-responses"` becomes
`codecs = ["openai-responses"]`, and the one test that swaps in a custom
adapter id now adds the line instead of replacing one. The settings
reference, the API schema's `Provider.adapter` description, and the SDK page
describe the new fields; the three `docs/superpowers/plans/` files that show
the old shape are dated, unchecked historical plans and are left as they are.

The implied agent profile for an operator provider that declares none used
to read the removed protocol adapter ids; it now reads the provider's first
codec (Anthropic Messages and Gemini map to their harnesses, the `bedrock`
adapter to Anthropic, everything else to OpenAI), with a test for the codec
path.

Absorbing the rest of the range: OpenRouter and Fireworks now ship enabled,
so the two fabro-llm tests that used OpenRouter as the disabled fixture use
`bedrock-openai`, and the docs and comments that said the two ship disabled
are corrected. The built-in catalog grew past 100 enabled model rows
(Vercel, TypeSafe, and the enabled OpenRouter and Fireworks rosters), so the
pagination shape test walks `page[offset]` to the last page instead of
assuming one page fits.

`cargo update -p` on the four crates also re-resolved a few already-locked
edges to match the lithos-llm lockfile: `windows-sys` 0.61.2/0.60.2 ->
0.59.0 under dirs-sys, errno, nu-ansi-term, quinn-udp, rustix,
rustls-platform-verifier, tempfile, terminal_size, and winapi-util;
`windows-core` 0.61.2 -> 0.62.2 under iana-time-zone; `errno` 0.2.8 ->
0.3.14 under signal-hook-registry; and `indexmap` 2.13.0 as a new public
dependency of lithos-llm. No package version was added or removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 22:07:00 -04:00
Scott Werner
e5ba7e3a0f Keep approval help concise and document attaching to interviews 2026-09-17 15:15:42 -04:00
Scott Werner
1630d26354 Clarify interview guidance in approval command help 2026-09-17 14:54:00 -04:00
Bryan Helmkamp
27f16f89c4
Delete fabro's own pricing: lithos-llm prices every response once
lithos-llm attaches a cost to every response at the client: the codec
keeps a provider-reported cost when the provider supplies one, and the
resolver fills the catalog's price for the route when it does not.
Pebble records that priced usage on every assistant turn and sums it,
so each AssistantMessage on the stream, and the store fold's live stage
usage, already carries the cost. Fabro's catalog re-pricing of the same
tokens was redundant, and is gone.

model_usage_from_llm, with_reported_cost, and every estimate_cost call
in fabro are deleted. The pebble handler's stage_usage groups pebble's
accounts by route and sums them with Usage::saturating_add, keeping the
cost and source pebble carried, so the terminal stage.completed usage is
the live fold's sum; it no longer fails when the catalog does not know a
provider. A one-shot prompt stage records the response's own usage and
cost as lithos-llm returned it. The per-model price cards in fabro-llm's
API module stay.

Tests: the pebble handler sums Catalog and Provider costs per row and
leaves a row's and the total's cost unknown once an answer was unpriced;
the store fold shows the same tokens and cost live and at completion,
and None at both for an unpriced answer; the agent integration test
compares the whole completed Usage with the live fold, cost included;
a one-shot prompt stage on a mocked OpenAI-compatible provider reports a
Catalog cost that lithos-llm's resolver attached, with no fabro pricing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:46:44 -06:00
Bryan Helmkamp
7d5e696c18
Merge pull request #873 from fabro-sh/sandbox-driver-64c14b8
Pin sandbox-driver main 64c14b8 and surface dropped Daytona output
2026-09-14 16:31:07 -04:00
Bryan Helmkamp
c98785d84a
Pin pebble a39f43e and refuse an older stored session record clearly
Pebble main a39f43e26effdf99635eaf343f095c17157c9c93 (pebble #22) carries
an assistant turn's usage as Usage in the session record and moves the
record format to version 5. CodingRuntime::from_record refuses a record
in another format with UnsupportedRecord { version, supported } before
it reads the route. Fabro persists those records in SQLite for Ask Fabro
resume, and old runs get no migration, so a record written by an older
build is read back as stored and refused on the next turn.

Two tests pin that down. The store reads pebble's own version 4 fixture
back through get without a parse error and reports it unsupported. A
resumed Ask Fabro session whose stored record declares the previous
format fails its next turn with the agent_error code and the message
"session record format version 4 is not supported (this build requires
5)", runs no turn, and leaves the stored record in place.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:58:31 -06:00
Bryan Helmkamp
e9ee0aaeaa
Merge remote-tracking branch 'origin/main' into one-usage-type
# Conflicts:
#	Cargo.lock
#	Cargo.toml
2026-09-14 12:40:49 -06:00
Bryan Helmkamp
aca9bbbf97
Regenerate the TypeScript API client for the usage schemas
The generated models follow the spec: TokenCounts, Cost, Usage,
ModelUsage, UsageModelRef, UsageStageRef, Speed, RunUsage,
RunUsageStage, RunUsageTotals, UsageByModel, AggregateUsage, and
AggregateUsageTotals replace the billing models, and UsageApi replaces
BillingApi. The stale billing model files are deleted.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:31:34 -06:00
Bryan Helmkamp
ef86ad278a
Carry usage as lithos-llm's Usage and rename billing to usage
Re-pin lithos-llm to 55add4596b861a0623d00c3a54aa5c147c8d504b and
pebble to c91810fe51aece80359b9cd8efea971af0c46925, where token usage
and cost travel together as Usage { tokens: TokenCounts, cost:
Option<Cost> }. Fabro now carries that one type everywhere it used to
carry BilledTokenCounts, BilledModelUsage, UsdMicros, or a token count
beside a cost_usd_micros.

fabro-types: billing.rs is usage.rs with ModelRef, ModelUsage { model,
usage }, sum_usage, and usage_is_empty; billing_rollup.rs is
usage_rollup.rs with ProjectionUsageStage, ProjectionUsageByModel,
ProjectionUsageRollup, and usage_rollup_from_projection. Every usage
field is named usage: StageProjection.usage and usage_by_model,
Outcome<Option<ModelUsage>>, stage.completed and stage.failed usage and
usage_by_model, prompt.completed usage, run.completed and run.failed
usage (total_usd_micros is gone), Conclusion.usage, StageSummary.usage,
Run.usage. RunSize buckets by Cost.

fabro-workflow: model_usage_from_llm prices tokens from the catalog with
a Catalog cost source, with_reported_cost keeps a provider cost, and the
pebble handler's stage_usage groups pebble's accounts by model and sums
rows with Usage::saturating_add, so a total has a cost only when every
priced part was priced. The store fold's live usage is the agent's
usage plus its descendants'.

API: the OpenAPI spec deletes BilledTokenCounts, BilledModelUsage,
CompletionUsage, CompletionCost, TokenUsage, and RunBillingSummary,
adds TokenCounts, Cost, Usage, and ModelUsage, and renames every
billing schema, property, tag, path, and operation to usage. fabro-api
reuses lithos-llm's and fabro-types' types through with_replacement,
with a round-trip test per replacement.

Old stored runs get no migration: their pebble events in the old shape
read back with zero usage, and their rebuilt projections lose agent
usage.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:31:34 -06:00
Bryan Helmkamp
14f69f50cd
Surface dropped Daytona output to the agent as a stderr line
When the driver reports an output loss on a streaming exec, the pebble
Environment adapter now appends one line to the stderr it hands back:

    [sandbox] N output frame(s), M bytes dropped by the provider

Pebble renders the result's stderr into the tool output, so the model
and the run log both see that the command's output is incomplete rather
than reading a silently shortened stream. The adapter also logs one
`warn!` with `dropped_frames`, `dropped_bytes`, and the command's first
word, bounded, so the operator can find the event without the log
carrying the command itself.

The loss is not folded into pebble's per-stream capture stats: the
dropped frames' stream is unknown and the counts are of encoded bytes,
so attributing them to stdout or stderr would be a guess. The driver's
`truncated` flags on both captures already say the counts undercount.
`ExecOutputTail` is pebble-owned and mirrored in the OpenAPI spec, so
the run events are left alone.

Tests cover the appended line with and without existing stderr, a
lossless command over the scripted double staying unchanged, the
bounded program name, and a lossy command end to end through
`Environment::exec` over a scripted sandbox whose exec facet reports a
loss (the driver's scripted double has no knob for it).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:24:02 -06:00
Bryan Helmkamp
5264227ca8
Merge pull request #872 from fabro-sh/exec-verbose-rendering
Pin pebble main 6d802a9 and print tool calls and the transcript under fabro exec --verbose
2026-09-14 13:07:37 -04:00
Bryan Helmkamp
86219355b6
Merge pull request #871 from fabro-sh/daytona-sandbox-guard
Delete a live Daytona test's sandbox even when the test panics
2026-09-14 13:07:24 -04:00
Bryan Helmkamp
3b8d712edb
Delete a live Daytona test's sandbox even when the test panics
The live Daytona tests create a provider sandbox and delete it on their
last line, so any panic or failed assertion before that line leaks a
running, billed sandbox. Two leaked that way on 2026-09-14 when a
sandbox-driver decoder flake panicked daytona_playwright_mcp_sandbox_transport.

Add fabro_sandbox::test_support::DeletedOnDrop, a guard that owns the
RunSandbox (Deref keeps the tests reading unchanged), offers an explicit
delete(self) for the happy path, and deletes from Drop otherwise. The
drop-time delete runs on its own thread and runtime because the test's
runtime may be unwinding. It reconnects the provider through the
ProviderAccess the test built the sandbox with, because the sandbox's
own handle pools HTTP connections whose tasks live on the test's runtime;
a live check of that path timed out after 10s.

Every Daytona test that creates a sandbox now holds it through the guard.
Unit tests over the scripted double prove delete-on-drop runs once, an
explicit delete runs once, and a panic inside catch_unwind still deletes
with and without a runtime.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 11:01:24 -06:00
Bryan Helmkamp
b46c293ceb
Print tool calls and the transcript under fabro exec --verbose
Since #852, `fabro exec --verbose` only turned on the request/response
middleware on the LLM client and no longer printed tool calls, tool
results, or the transcript. Pebble #18 gives pebble-cli-core rendering
options, so `--verbose` now runs the prompt through
`run_prompt_with(..., RenderOptions::verbose())`: each tool call's
arguments and result in full under its `[tool]` and `[result]` lines,
plus the transcript. The middleware is enabled as before. Without the
flag the renderer gets the default options, so the output is unchanged.

The twin shell test now scripts the tool call and the final answer as
two turns, so the answer on stdout is the scripted one rather than the
twin's fallback echo, and it asserts that stderr carries no result,
reasoning, or verbose blocks. A new twin test runs the same prompt with
`--verbose` and asserts the tool and result blocks and the request dump.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 10:56:53 -06:00
Bryan Helmkamp
e3ea3ff2e6
Assert the graph name field in local_run_lifecycle
`ps --json` reports the digraph name as `workflow_graph_name` and reserves
`workflow_name` for an explicit `[workflow] name` (6a86ced77). That change
updated the ps tests but not this ignored e2e test, which still expected
the graph name under `workflow_name`. The test now asserts the contract
the ps tests assert: `workflow_name` is null for a bare graph file and
`workflow_graph_name` is the digraph name.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 09:41:05 -06:00
Bryan Helmkamp
b46ed0a429
Point twin_doctor's isolated server at the twin
The provider probe runs inside the isolated server, which never sees the
test process environment. The test used to store `OPENAI_BASE_URL` in the
vault, and cd74013d0 dropped that entry without replacing it, so the
server probed the real OpenAI API with the namespace as its key and the
doctor reported the provider as failed. The server settings now repoint
the `openai` provider at the twin through the operator `[llm]` overlay.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 09:41:05 -06:00
Bryan Helmkamp
325fe0c4bb
Fix the twin-mode hook and arc e2e tests
The four hook tests and arc_e2e_with_real_llm in workflow/hooks.rs failed
in twin mode for three reasons, all in the test fixtures.

The hooked workflows were written as `<name>.toml` beside `<name>.fabro`.
Version packaging accepts a config only as `workflow.toml` beside its
graph (44dccfa3d), so `fabro run` failed at collection. Each hooked
workflow now lives in its own `<name>/` directory as `workflow.toml`.

The isolated server never learned the twin's base URL. The CLI command
carried `OPENAI_BASE_URL`, but the run executes in the server, which does
not see the test process environment, so it called the real OpenAI API
with the namespace as its key. The twin-mode server settings now repoint
the `openai` provider at the twin through the operator `[llm]` overlay,
the same way `run_uses_vault_credentials_for_worker_execution` does.

With the server reaching the twin, the hook scenarios were consumed by
the wrong request: the server asks the model for a run title in the same
namespace before the hook fires, and the scenarios had no matcher. The
block test then saw the twin's default response and the hook failed open,
so the run succeeded. Hook scenarios now match on the `Hook prompt:`
prefix of the evaluator's user message.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 09:41:05 -06:00
Scott Werner
0dcec5af77 Drop redundant create tests and reuse the production client in fixtures 2026-09-13 09:16:31 -06:00
Scott Werner
012b556367 Remove tests and fixtures tied to retired manifest fields 2026-09-13 08:55:55 -06:00
Scott Werner
ea11538038 Retire legacy manifest run creation 2026-09-13 08:49:50 -06:00
Bryan Helmkamp
a2ace0888b
Drop the projection parity tests and fold agent_control into agent.activity
The session_projection_parity module pinned fabro's stage fold to pebble's
SessionProjection while both existed; the stage view reads the fold now,
and the one usage rule has its own tests. StageProjection.agent_control
and AgentControlState go too: pebble's fold carries the interrupted and
steered facts as agent.activity, and the stage's state says whether the
stage still runs, which is what the reset on fabro's own stage events was
for. The run-detail banner reads activity plus state.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:38:52 -06:00
Bryan Helmkamp
32dca54cd7
Show the agent sidebar's sections in demo mode
The demo agent stage's stored events now read as one pebble session: MCP
servers up and failed, skills, a subagent, a failover, a compaction, and a
written file, ending with ProcessingEnd. Demo mode serves the run state it
answered not_implemented to, with the agent stage carrying the coding
agent's fold of those events, so the stage sidebar renders them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:28:32 -06:00
Bryan Helmkamp
7a6eea0399
Delete the agent mirrors and move failover to prompt stages
Pebble's stream is the agent event contract. The run's own agent.mcp.ready,
agent.mcp.failed, and agent.mcp.disconnected events, which mirrored pebble's
McpServer* events, are gone with their props, the sink arms that emitted
them, and their conversion and naming entries; pebble's stored
agent.mcp.server.* events are the only record and feed the stage's fold.

The sink no longer mirrors RouteFailover onto agent.failover either: an
agent stage's moves are pebble's agent.route.failover. The event is now
prompt.failover, emitted only by a one-shot prompt stage that walks its
fallback plan itself, and its props are trimmed to the two routes, the
attempt, and the error; nothing read the rest.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:21:41 -06:00
Bryan Helmkamp
b5ee16b37a
Read the stage view from the agent's fold
StageProjection loses todos, subagents, skills, mcp_servers, and
context_window, the types behind them, their fold arms and helpers, and
their OpenAPI schemas: every one of those facts is pebble's fold in
StageProjection.agent now. The context-window endpoint reads the fold's
snapshot, whose event_seq is the agent's own sequence. The parity module
keeps its assertions on the surviving own fields, usage and model, and
checks that what the stage view reads from agent is the whole-session
fold's for the stage's events. The TypeScript client is regenerated and
its stale models removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:13:10 -06:00
Bryan Helmkamp
234953d5c4
Bill a failed agent stage what it spent
An agent stage that failed billed nothing: the backend returned a bare
error and the outcome built from it carried no usage. A terminal failure
now becomes the stage's failed outcome from the same fold that bills a
completed stage, with the tree's usage, the rows by model, the files it
wrote, and its active time; stage.failed carries billing and
billing_by_model and the store keeps both. Cancellation and retryable
failures still go up as the error.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:58:34 -06:00
Bryan Helmkamp
36e9259eb7
Describe billing_by_model on the API and the Billing tab
StageProjection.billing_by_model and a BilledModelUsage schema that reuses
fabro's type; the TypeScript client regenerated; the Billing tab's token
tooltip says subagent tokens are included and priced at each subagent's
model; the stage.completed docs describe the rows and the one usage rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:49:28 -06:00
Bryan Helmkamp
850d5cca52
Bill an agent stage's whole session tree from one fold
One usage rule: a stage's usage is its session tree's, the root and every
subagent, live and at completion. The worker's event sink folds pebble's
SessionProjection over the events it records and the stage's billing and
files come from that fold at stage end, so the completed values are what
the run showed live. The store's live usage is the fold's tree usage, and
completion brings the catalog's price for the same tokens instead of
resetting them to the root's.

Fabro keeps catalog pricing: the root at its route, each descendant at its
own route where the catalog knows it and at the root's otherwise, a
provider-reported cost standing in where pebble has one. The rows travel
as billing_by_model on stage.completed and the stage projection, and the
billing rollup splits by_model by them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:49:27 -06:00
Bryan Helmkamp
fa6902c448
Describe StageProjection.agent on the API as pebble's own types
AgentSessionProjection and the schemas nested in it reuse pebble's types
through with_replacement; the AgentSession prefix marks the projection's
own types where fabro already has a schema of that name, and pebble's
event-level types keep their names. The round-trip test builds a
projection over a scripted stream, validates it against the spec with the
spec as the root document, checks every serialized key is declared, and
validates every enum variant this build knows. The TypeScript client is
regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:25:38 -06:00
Bryan Helmkamp
424a6a1a47
Embed pebble's SessionProjection in StageProjection
StageProjection.agent is pebble's fold of the stage's agent events, fed
every stored agent event before the fabro-only arms run. Every existing
field and arm stays for now. The parity tests prove each old field is
derivable from the embedded fold: the tree's usage, the route as the model,
the context window without fabro's stamped seq, the root's todo list, the
subagent rows, the skills, and the MCP servers under the disconnected,
error, ready rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:25:38 -06:00
Bryan Helmkamp
217a5860b6
Store ProcessingEnd and the mirrored pebble events in the run log
Pebble's SessionProjection reads ProcessingEnd to complete a prompt and mark
the session idle, so a projection rebuilt from the run's log needs it: one
small event per prompt. The four pebble events the sink mirrored onto
fabro's own agent.failover and agent.mcp.* are now stored verbatim as well,
so the fold sees the route moves and the MCP outcomes; the mirrors stay
until every reader is on the projection.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:25:38 -06:00
Bryan Helmkamp
1d1cacbef3
Pin pebble main 6996942 and implement its PortRoutes trait
Pebble now owns the trait an application hands its MCP support to reach a
port inside the environment. SandboxPortRoutes implements PortRoutes over
the run sandbox's preview-URL facet: a missing facet is Unsupported, a
driver failure is Failed with the driver error as its source. The mcp
feature no longer pins sandbox-driver, so fabro's sandbox-driver pin moves
on its own from here.

The new pebble rev also puts the summary call's usage and cost on
CompactionCompleted, which the CLI progress test literal names.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:03:36 -06:00
Bryan Helmkamp
697d8622e9
Merge pull request #843 from fabro-sh/remove/run-metadata-branches
Some checks are pending
TypeScript / Build (push) Waiting to run
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox plugins (stdio) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
TypeScript / Typecheck (push) Waiting to run
TypeScript / Test (push) Waiting to run
Remove Git run metadata branches
2026-09-12 16:41:24 -06:00
Bryan Helmkamp
e5d5c534ab
Merge remote-tracking branch 'origin/main' into remove/run-metadata-branches
Resolve conflicts between the metadata-branch removal and the
sandbox-driver adoption on main:

- fabro-sandbox docker.rs, sandbox.rs, daytona/mod.rs: take main's driver
  rewrite. The Sandbox trait is gone, so the PR's push_token_source
  removal now applies to RunSandbox instead; drop that accessor and the
  RepoCredentials::source helper that only served it.
- run_metadata.rs: keep deleted. Main's edits there were adaptations to
  the driver API and the run git identity field.
- lifecycle/git.rs, finalize.rs: keep the PR's removal of metadata
  snapshots and write_finalize_commit; carry main's RunSandbox,
  GitRetryPolicy, git_identity, local_sandbox, and test catalog changes.
- sandbox_git.rs: take main's version and drop the shadow_sha parameter
  and Fabro-Checkpoint trailer.
- git_integration.rs: remove meta_branch from the new git identity test.
- Cargo.toml: main's dependency set with fabro-dump kept as a
  dev-dependency.
- checkpoints.mdx: keep both the git identity paragraph and the durable
  execution state section.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:28:01 -06:00
Bryan Helmkamp
9550ff0806
Cover the failover continuation and the stopped failover
The existing failover test asserts the backup continued the turn after
the primary committed a tool result. A new test exhausts a two-route
chain and checks that one agent.failover and one
agent.route.failover.stopped are stored on the work stage, the stop
after the error it reports, with the exhausted reason and the failing
route. The events catalog documents both.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:26:32 -06:00
Bryan Helmkamp
252e7e0643
Name pebble's stopped failover agent.route.failover.stopped
Pebble publishes RouteFailoverStopped when a model error ends a prompt on
its route although fallback routes were configured: the error is
ineligible or the chain is exhausted. Fabro has no event of its own for
that case, so the pebble event is stored as it is under a derived name
next to agent.route.failover instead of the generic agent.event.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:26:32 -06:00
Bryan Helmkamp
57a746235e
Pin pebble to main after lithoscomputer/pebble#12
Pebble's RouteFailover event now describes the route that failed and how
the new route carried the prompt on. Record the continuation on fabro's
agent.failover event as an optional string (replay_prompt or
continue_turn); events written before it existed, and one-shot prompt
stages that walk the plan themselves, read as absent. The failed route's
usage, cost, and timing are not mirrored: the stage's totals already
include them through the prompt report, and no fabro run event carries
per-route usage yet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:26:32 -06:00
Scott Werner
1aaca98ea4
Merge pull request #845 from fabro-sh/codex/cli-workflow-target-selection
Select CLI workflow sources and run targets independently
2026-09-12 16:38:01 -04:00
Bryan Helmkamp
7879e6d223
Merge pull request #856 from fabro-sh/brynary/run-git-identity
Resolve one Git identity per run and inject it into every workflow command
2026-09-12 14:23:30 -06:00
Scott Werner
86bcc8f128 Simplify CLI repository selectors and add workflow shorthand 2026-09-12 14:18:34 -06:00
Bryan Helmkamp
3b33730bac
Regenerate the TypeScript API client for McpServerStatusDisconnected
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 12:39:31 -06:00
Bryan Helmkamp
ff7821dcd2
Adopt pebble's MCP disconnect event and startup timing
Pebble reports an MCP server whose connection closed mid-session once,
as McpServerDisconnected, and carries startup_ms on McpServerReady and
McpServerFailed. The workflow event sink mirrors the disconnect onto a
new agent.mcp.disconnected run event shaped like agent.mcp.failed, and
passes startup_ms through on agent.mcp.ready and agent.mcp.failed. The
raw pebble event is not stored for these, so the timing would otherwise
be dropped at the boundary.

The stage projection's McpServerStatus gains a `disconnected` kind next
to `ready` and `failed`. The fold keeps the server's tool count and
sticky invoked flag and only moves the status. The OpenAPI
McpServerStatus oneOf gains McpServerStatusDisconnected, and the
fabro-api round-trip test covers its JSON shape. A new
session_projection_parity test folds the same MCP events through
pebble's SessionProjection and fabro's stage projection and compares
them, including pebble's `disconnected`.

ToolErrorKind::Timeout needs no fabro change: the kind is stored as
pebble serializes it and never matched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 12:39:31 -06:00
Bryan Helmkamp
1aee59585b
Pin pebble to mcp-embedder-events and accept startup_ms on MCP outcomes
Pebble's McpServerReady and McpServerFailed events now carry startup_ms.
The workflow event sink destructured both variants by name, so it stops
listing every field. The pin is temporary: it moves to pebble main once
lithoscomputer/pebble merges the branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 12:29:36 -06:00
Scott Werner
94c7159ef8 Preserve shared target inference with explicit CLI targets 2026-09-12 11:54:48 -06:00
Scott Werner
a9f28828a6 Adapt workflow target selection to sandbox provider kinds 2026-09-12 11:48:21 -06:00
Scott Werner
321a190315 Share one supervised process runner between server and CLI Git
The CLI's native Git runner and the server's run_git_plan each hand-rolled
the same mechanics: kill-on-drop, a wall-clock timeout, output capture,
and (only in the CLI) process-group teardown, bounded capture, and
cancellation. Add fabro_proc::SupervisedCommand, which owns stdio, the
process group, the timeout, cooperative cancellation, and bounded
capture, and put both runners on it. The server gains group teardown on
timeout, so helpers a stuck clone or fetch spawned no longer outlive it;
the CLI keeps discarding output on failure and gains nothing but less
code. The hardened -c overrides become one named list in the CLI.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 11:48:21 -06:00
Scott Werner
7c97ba1b70 Move the run-driven remote workflow test to cmd/run.rs
remote_workflow_run_starts_once_create_leaves_submitted_and_failures_do_not_refetch
lived in cmd/create.rs but drove fabro run in four of its five
iterations and asserted the start call, which is fabro run's contract.
Split it: cmd/create.rs keeps the single create invocation that must
leave the run submitted without starting it, cmd/run.rs owns the
run-driven success and failure iterations, and the workflow and remote
repository fixtures move to the shared command test support module.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 11:48:21 -06:00
Scott Werner
f15972f8f2 Keep Ctrl-C owned for the whole native Git command phase
owned() polled tokio's ctrl_c() only for the duration of one Git
acquisition. On Unix that listener permanently replaces the default
SIGINT disposition, so once acquisition finished nothing handled Ctrl-C
and fabro create, fabro run --detach, and fabro run before attach
silently ignored it while waiting on the server. Introduce an
Interruption handle that the command entry points create from the run
arguments: it installs a listener only when --workflow-git or
--target-git is in play, guards create (and start for fabro run) as one
phase, and tracks owned Git tasks so interruption waits for their
cleanup before returning. attach keeps installing its own listener.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 11:48:21 -06:00
Scott Werner
39694ec5e8 Acquire the remote workflow before observing a local Git target
For --workflow-git selections, target resolution ran before the remote
workflow ref was verified to exist. On Docker and Daytona environments a
path target that is a GitHub checkout is observed via
observe_git_run_target, which may silently push the attached branch, so
a typo in --workflow-ref produced a remote side effect with no run
created. Resolve the remote workflow after parent and environment
validation but before target observation, restoring the pre-existing
workflow-then-target order, and cover it with a caller checkout whose
unpushed branch must stay unpublished.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 11:48:21 -06:00
Scott Werner
0998f60624 Contain the selected workflow file instead of pre-walking the checkout
Remote acquisition walked the entire depth-1 checkout and failed on any
symlink that dangled or resolved outside the root, even when the link
was nowhere near the selected workflow. Submodule-style dangling links
and links into the host are common in workflow repositories and made
--workflow-git fail where the same commit collected fine locally. The
bundler already root-checks every file it opens; the only unchecked
reads were the selected TOML (or a graph selector's sibling TOML) during
location resolution. Check those in collect_workflow_versions and drop
the O(repo) walk. walkdir stays a dev-dependency for the dump tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 11:48:21 -06:00