Commit graph

664 commits

Author SHA1 Message Date
Bryan Helmkamp
bfc2abebab
Wait for the terminal lifecycle record before reading a stream's end
Fabro's terminal `run.lifecycle` record lands a moment after Petri's
`run.finished`: the worker exits, the server records the status, the
projector folds it. A CLI scenario that asserts on the end of the stream
now waits for that record instead of reading the stream as soon as the
runs row turns `succeeded`, which the projector writes from the engine's
finish alone.

The fabro-petri README names the projector's stream reader and commit
signal, the server's reconnect test with its fixture capture, and the
CLI scenarios that read a run back through the stream.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 01:25:40 -04:00
Bryan Helmkamp
523f831a7a
Serve a Petri run's events as one stream with a stream_seq cursor
`GET /runs/{id}/events` and `GET /runs/{id}/attach` serve a Petri run's
public events and Fabro's platform records as one ordered stream in a
Fabro envelope (`RunStreamItem`: `run_id`, `stream_seq`, `kind`, `id`,
`recorded_at`, `item`), read from the projector's `petri_stream` table.
The cursor is `stream_seq` (`?after=`); the item's own identity (the
Petri `EventId` as `<log>/<seq>/<index>`, or the platform record's seq)
travels beside it for deduplication. A legacy run keeps its envelope on
the same endpoints; the OpenAPI response is the union of the two lists,
and the stream list reports Petri's `EVENT_CONTRACT_VERSION`.

The attached stream follows the projector's commit signal (a wake-up,
with a poll as the fallback) and ends after the platform record of the
run's terminal lifecycle transition, the analog of the legacy stream's
`run.completed`, or a bounded grace after the projection went terminal.

`RunSpec.engine` (`RunEngine`, `PetriAdmission`, `PetriGraphRef`) is
named in the spec and reuses the Rust types. `fabro-client` matches the
union and adds `list_run_stream`, `list_run_stream_page` and
`attach_run_stream`.

A server test attaches to a two-branch parallel run, disconnects once
both branches started, records a platform notice while both branch
scripts run, reconnects from the last `stream_seq`, and checks the
union is the whole stream: every item once, in order, no gap, no
duplicate, the notice between the branch events, and the same as the
paged listing. The Petri scenarios capture their settled projection and
stream as JSON fixtures for the web app under
`FABRO_CAPTURE_PETRI_FIXTURES`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:47:26 -04:00
Bryan Helmkamp
a6ac3f120e
Give a Petri question one identity across the adapter and the projection
The interview adapter derived its own question id from Petri's identity
and posted it on `interview.started`, while the projection over Petri's
records serves the pending question under Petri's `Question.id` with the
firing's stage label. The answer endpoint validates against the
projection, so an answer under the projection's id never reached the
adapter's wait.

The adapter now waits under Petri's id and labels the question's stage
through the projection's own rule: `stage_label`, `is_shown` and
`visit_of` move out of `start_visit` into shared functions, and the
adapter's observer derives each firing's `visit.started` through Petri's
`Projection`, as the projector does, so the label matches by
construction. The full Petri identity stays on `AskedQuestion`.

The legacy `interview.*` events are still posted, under Petri's id, for
the readers that follow the event stream rather than the projection: the
Slack service, `run attach`, the web app's Q&A renderer and the server's
answer claim. The store already derives the `interview.answered` platform
record from `interview.completed` for a Petri run, so who answered is
recorded under Petri's id with the answering principal.

The gate scenarios assert the new identity and encode the id as one path
segment, as the generated clients do. Projection tests cover an expired
question and an auto-approved answer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:19:30 -04:00
Bryan Helmkamp
5dae891a98
Merge branch 'petri-integration-read' into petri-integration
# Conflicts:
#	Cargo.lock
#	lib/components/fabro-petri/Cargo.toml
#	lib/components/fabro-petri/README.md
#	lib/components/fabro-petri/src/lib.rs
2026-09-17 23:48:36 -04:00
Bryan Helmkamp
ee4850689e
Name the per-run pass lock's type through an import
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:58:31 -04:00
Bryan Helmkamp
4a20acafc7
Run one projection pass at a time per run
The startup pass called the pass directly while a signalled pass could
run for the same run, so both read one committed stream sequence and
the second insert into the stream failed on its primary key, which
stopped the restarted server. Passes now take a per-run lock, and a run
whose startup pass fails is logged and left for its next signal instead
of stopping the server. A test races four passes, a signal and the
startup pass over one run and checks the stream stays contiguous.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:56:30 -04:00
Bryan Helmkamp
e0b546d465
Read the view tables where the run summary store keeps them
The projector takes two pools: the one Petri's records live in and the
one the view tables live in. In the server both are the one database;
a test fixture keeps the runs row, the platform records and the
projection tables in the run summary store's own pool, which the
projector was not reading, so a run projected in a test server folded
its Petri events before its run.created record. The startup run-history
verification checks only a Petri run's identity and legacy guard, since
its row is the projector's. An agent stage's response is the
response.<node> its outcome wrote into the run context, as the prompt
step writes it. The scenario tests assert each branch's own index.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:49:15 -04:00
Bryan Helmkamp
53b16b591d
Merge branch 'petri-integration-adapters' into petri-integration 2026-09-17 22:40:30 -04:00
Bryan Helmkamp
e2bf05c0f0
Project a Petri run's records into Fabro's run view
The projection folds Petri's public events (replay_since over the run's
stored records) and Fabro's platform records into the RunProjection the
API serves, row by row as VIEWS.md maps them. The stage key is the
execution and firing; the StageId label is node@visit, made unique with
the execution when two child invocations would share one. A stage's
first_event_seq is the milliseconds from the run's creation to its
visit.started, so the view built live equals the view rebuilt from the
records whatever order two logs' records were committed in.

The projector is the view pass and its wake-up. Records first: an append
returns before any view work; a pass reads what is committed, folds the
items past the committed positions, and writes the projection document,
the ordered stream (one stream_seq per Petri event or platform record,
with the item's own identity beside it) and the narrowed runs row in one
later transaction. Signals coalesce per run, a lost signal costs only
latency, the startup pass folds every run the view trails, and a pass
that races a platform record leaves the view alone and runs again. A
torn tail holds the view where it stands and reports the run incomplete
with the replay's error; inspect_run decides completeness once the run
recorded its finish.

The tests build the view live for the hello bundle, a command workflow
and a two-branch parallel workflow and compare it with the rebuild; drop
every wake-up and catch up by a signal and by the startup pass; crash
between the record commit and the view transaction and apply only the
suffix; restart the projector over child executions; and hold at a torn
tail.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:32:23 -04:00
Bryan Helmkamp
162791979c
Wake the projector after a Petri run's first event too
The run's creation commits its first event on the create path, not the
append path, so the platform record hook never fired for run.created.
The store now notifies after that commit as well, and a test proves a
Petri run's lifecycle events leave platform records beside them with one
wake-up per record while a legacy run leaves none.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:32:23 -04:00
Bryan Helmkamp
3aadc21739
Test the interview, secret, blob and home adapters through the engine
Integration tests in `fabro-petri` run workflows through `engine::run` on
the host sandbox: a gate answered under the posted question id, two
parallel gates each bound to their own answer, an expired question
completed as a timeout with the gate's default, an auto-approved run, a
cancelled run; a secret resolved from a vault into a command and masked in
every `petri_records` row; a command's large output round-tripped through
the `blobs` table under `blob://sha256/<hex>`; and the `hello` bundle on
the OpenAI twin with a model client over a vault that holds the key,
whose skills step searched the configured home.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:31:35 -04:00
Bryan Helmkamp
223e10ea20
Install Petri's interview, secret, blob and home adapters in a Fabro run
A Petri run in the worker, and in the server under its test override, now
gets Fabro's platform adapters instead of the standalone defaults:

- `fabro_petri::interview`: Petri's `Interviewer` over the questions API
  and the worker's control channel. A human gate's question is posted as
  the `interview.started` event a legacy stage emits, keyed by an id
  derived from Petri's identity (node, execution, firing, occurrence,
  ask), so the API, the web app and Slack list it; the answer posted to
  the questions endpoint reaches the control interviewer the adapter waits
  on and is mapped onto Petri's answer. An expiry the gate reports is
  completed as `interview.timeout`, a cancel as `interview.interrupted`,
  and an auto-approved run answers itself. The hook points the read side
  takes over are marked.
- `fabro_petri::secrets`: Petri's `SecretProvider` over the vault's token
  entries, so `{{ secrets.NAME }}` resolves at spawn and is masked in every
  record; a sensitive answer registers as a dynamic secret.
- `fabro_petri::blobs`: Petri's `OutputStore` over Fabro's `blobs` table,
  through the server's blob store or the worker's client.
- The Fabro home the server resolved travels to the worker as
  `--fabro-home`, so the skills step reads it whatever the worker's
  environment says.

`engine::RunRequest` takes the interviewer, its observers, the secret
provider and the blob table from the caller; `interviewer::Unattended` is
gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:31:35 -04:00
Bryan Helmkamp
abbc7ca11d
Add platform records and the Petri projection tables
A Petri run's own Fabro facts (its lifecycle before and after the
engine, a checkpoint commit, a pull request, a notification, a pairing)
are platform records in a table beside Petri's records, one typed enum
of kinds tagged on the wire, each keyed to a Petri stage where it
belongs to one and carrying the operation identity of the effect it
records. The run summary store derives the lifecycle kinds from the
legacy run events a Petri run still appends, in the event's
transaction, and calls a hook after the commit so the run's projector
can wake up.

Two more tables serve the projection that follows: the per-run
projection document with its committed positions, and the ordered
stream of everything the view consumed. The run summary store reads the
Petri projection back for the API and writes the narrowed runs row from
it without touching the legacy concurrency guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:46:21 -04:00
Bryan Helmkamp
a745d0b75d
Check a workflow version in memory with Runtime::check_source
`fabro_petri::check` materialized the version's bundle into a temporary
directory because `Runtime::check` read the workflow and its settings
files from disk. Petri now has `Runtime::check_source`, which takes the
workflow's repository-relative path, its text and a `FileSource`, so the
bundle goes into a `frontend::MapFiles` map instead: every file at its
bundle-relative path, `workflow.toml` beside the workflow, and
`.fabro/project.toml` at the root when the caller has one. Nothing is
written to disk, and the diagnostics name the bundle-relative paths
directly, with no root to strip.

The compile inputs are unchanged: the intent's inputs, the launch model
and provider, and `petri.repository` bound by Fabro itself (the launch's
path, or `null`). An entrypoint that is not one of the bundle's files is
now `CheckError::MissingEntrypoint`; `CheckError::Materialize` goes away.
`tempfile` becomes a dev-dependency, as only the tests use it.

New tests: a version whose `workflow.toml` names `engine = "petri"` is
admitted (the new Petri pin knows the key), an unknown `[workflow]` key
is refused with `unsupported.workflow_toml.key` and named in
`workflow.toml`, the project settings are read from the map, and a
missing entrypoint is an error.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:31:04 -04:00
Bryan Helmkamp
a621fb72e1
Let the Petri engine start or resume a run over any store
`fabro_petri::engine` is now the one assembly the worker process and the
server share: `RunRequest` takes the run's store as `Arc<dyn RunStore>` and
an `Execution`, either `Start` with the admitted graphs or `Resume` from the
run's records through `host::resume_configured`, with the same interview
observer a start installs. A resume whose record has no root invocation is
refused with a named error instead of a panic in the host. The outcome is
mapped to a `Conclusion` (succeeded, or failed with Fabro's reason and a
message) so both callers record the same terminal event.

`admission::load_with` loads the admitted graphs through any blob read, so
a worker loads them through its client; `admission::load` over the server's
`BlobStore` delegates to it.

`HttpRunStore::for_worker` takes every lease for the worker's launch id,
whatever owner Petri minted for the run runtime, and the open logs the
owner. The module docs state the rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:20:22 -04:00
Bryan Helmkamp
832f39f704
Merge branch 'petri-integration-http' into petri-integration
# Conflicts:
#	Cargo.lock
#	lib/components/fabro-petri/Cargo.toml
#	lib/components/fabro-petri/README.md
#	lib/components/fabro-petri/src/lib.rs
2026-09-17 20:37:03 -04:00
Bryan Helmkamp
d2f0e70dca
Run a workflow through Petri in the server, behind the engine flag
When a run's engine is Petri, the create handler hands the bundle, inputs
and launch to Petri's check instead of the legacy compile, lint and model
pinning, refuses the run with the validation error the legacy validator
uses (Petri's codes as the rules, listed in the API detail), and records
the admission on the run spec. The Fabro graph the read side displays is
parsed without validation. The scheduler executes a Petri run in the server
process through fabro_petri::engine, appending only the run lifecycle
events the read side needs (run.starting, run.running, run.completed or
run.failed); no stage or agent event is projected yet.

Scenario tests run the hello bundle on the OpenAI twin under the version
flag and a command-only bundle under the server setting, check Petri's
record agrees, and cover the refusals for an unknown attribute, an
undeclared node and an unknown model.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:27:44 -04:00
Bryan Helmkamp
03309d4240
Let Petri compile and execute a Fabro run through fabro-petri
`check` materializes a workflow version's bundle into a temporary directory
(`Runtime::check` reads files from disk), lowers it with the run's inputs
and launch, and returns the admitted graphs or Petri's diagnostics in a
shape the server maps onto Fabro's. `admission` keeps the admitted graphs
in the blob store, named on the run spec and verified by digest on load.
`runtime` assembles the same Petri runtime at create and at execution: the
Fabro frontend with the server's settings layer, the Attractor step kinds,
the model client as the PebbleClient capability so admission pins every
model. `engine` runs the admitted graph in the server process over
SqliteRunStore under the Fabro run id, with the standalone defaults, an
interviewer that fails any question, and cancel on a token, and derives the
outcome from inspect_run over the run's record.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:27:44 -04:00
Bryan Helmkamp
25d47ebcd3
Implement Petri's run store over the server's API
`HttpRunStore` is Petri's `RunStore` and `RunLogs` as a worker process
reaches them: over `fabro_client::Client` with the worker's token,
against the server's SQLite store. A run key is a Fabro run id, the
`{id}` of every request, which is the plan's rule that Petri's run key
is Fabro's run id.

The lease rules are the store's. A same-owner reopen shares the live
handle in the process, and the server makes a same-owner reopen after
a lost reply the same lease. Dropping the last handle of an owner sends
`release` on the current Tokio runtime, and the store awaits every such
release before its next open, so a drop followed by an open observes
it. The server's worker-exit release is the backstop.

A reply that never arrives, a transport error or the client's request
timeout, is retried by resending the same request up to three times.
Every request is idempotent on the server, so that is safe; a reply
that did arrive is never retried. Each server error code maps back to
its `StoreError` variant, with the leased owner and the conflict
position read from `meta`.

`fabro_petri::petri` re-exports the store vocabulary for the server,
and the `test-support` feature re-exports Petri's test kit so the
server's tests can run the conformance suite over the wire. Both keep
this crate the one place that names a Petri package.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:12:20 -04:00
Bryan Helmkamp
c383b6a70b
Add the engine flag and record the engine on the run spec
A workflow version names its engine with `engine = "petri"` in the
`[workflow]` table of `workflow.toml`, and `[server.execution] engine`
(`FABRO_SERVER_ENGINE`, `--engine`) defaults it for every version that
names none. The choice, with what Petri admitted (the lowered root graph
and its children by blob and digest), is recorded on the run spec as
`RunEngine`, carried on `run.created`, and replayed into the projection.
A legacy run's spec omits the field, so existing specs decode unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:02:47 -04:00
Bryan Helmkamp
55a145f5a4
Add the Fabro-on-Petri view coverage matrix
Plan item F2.1: every Fabro view of a run, the Petri event or platform
record that supplies each fact, and the identity it is keyed on. Ends
with the two completeness checks (every EVENTS.md family, every Fabro
view) and the gaps table.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 19:38:25 -04:00
Bryan Helmkamp
3ffe7e00ce
Implement Petri's run store over Fabro's SQLite database
`SqliteRunStore` implements Petri's `RunStore` and `RunLogs` on the pool
Fabro's other stores share. A run's existence and writer lease live in
`petri_runs`; every record of every log lives in `petri_records`, keyed by
(run, log, seq) with the record stored as JSON and read back unchanged;
blobs share the `blobs` table with `BlobStore`. The lease is taken
idempotently per owner, ends when the last handle drops or when an
operator releases it, and never by timeout; every write checks it inside
its own transaction. An append is one `BEGIN IMMEDIATE` transaction per
batch: a repeated record is accepted, a different record at a taken seq or
a seq past the head is a conflict that stores nothing.

Petri's store conformance suite passes against it, with the operator
release, lease exclusivity, a crash between appends, and blob
interoperation checked beside it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 19:30:30 -04:00
Bryan Helmkamp
7eb5ca502c
Add the fabro-petri crate and pin the Petri packages
Fabro runs its workflows on Petri. The six Petri packages and the testkit
are pinned by revision in the workspace manifest under `petri_*` keys, and
`fabro-petri` is the one crate that depends on them. The crate's tests run
the `hello` bundle in memory on the stub registry and a command-only
workflow on the host sandbox; both skip without the sandbox-driver host
plugin, and the sandbox-plugins CI job requires it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 19:23:30 -04:00
Bryan Helmkamp
27f16f89c4
Delete fabro's own pricing: lithos-llm prices every response once
lithos-llm attaches a cost to every response at the client: the codec
keeps a provider-reported cost when the provider supplies one, and the
resolver fills the catalog's price for the route when it does not.
Pebble records that priced usage on every assistant turn and sums it,
so each AssistantMessage on the stream, and the store fold's live stage
usage, already carries the cost. Fabro's catalog re-pricing of the same
tokens was redundant, and is gone.

model_usage_from_llm, with_reported_cost, and every estimate_cost call
in fabro are deleted. The pebble handler's stage_usage groups pebble's
accounts by route and sums them with Usage::saturating_add, keeping the
cost and source pebble carried, so the terminal stage.completed usage is
the live fold's sum; it no longer fails when the catalog does not know a
provider. A one-shot prompt stage records the response's own usage and
cost as lithos-llm returned it. The per-model price cards in fabro-llm's
API module stay.

Tests: the pebble handler sums Catalog and Provider costs per row and
leaves a row's and the total's cost unknown once an answer was unpriced;
the store fold shows the same tokens and cost live and at completion,
and None at both for an unpriced answer; the agent integration test
compares the whole completed Usage with the live fold, cost included;
a one-shot prompt stage on a mocked OpenAI-compatible provider reports a
Catalog cost that lithos-llm's resolver attached, with no fabro pricing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:46:44 -06:00
Bryan Helmkamp
7d5e696c18
Merge pull request #873 from fabro-sh/sandbox-driver-64c14b8
Pin sandbox-driver main 64c14b8 and surface dropped Daytona output
2026-09-14 16:31:07 -04:00
Bryan Helmkamp
c98785d84a
Pin pebble a39f43e and refuse an older stored session record clearly
Pebble main a39f43e26effdf99635eaf343f095c17157c9c93 (pebble #22) carries
an assistant turn's usage as Usage in the session record and moves the
record format to version 5. CodingRuntime::from_record refuses a record
in another format with UnsupportedRecord { version, supported } before
it reads the route. Fabro persists those records in SQLite for Ask Fabro
resume, and old runs get no migration, so a record written by an older
build is read back as stored and refused on the next turn.

Two tests pin that down. The store reads pebble's own version 4 fixture
back through get without a parse error and reports it unsupported. A
resumed Ask Fabro session whose stored record declares the previous
format fails its next turn with the agent_error code and the message
"session record format version 4 is not supported (this build requires
5)", runs no turn, and leaves the stored record in place.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:58:31 -06:00
Bryan Helmkamp
e9ee0aaeaa
Merge remote-tracking branch 'origin/main' into one-usage-type
# Conflicts:
#	Cargo.lock
#	Cargo.toml
2026-09-14 12:40:49 -06:00
Bryan Helmkamp
ef86ad278a
Carry usage as lithos-llm's Usage and rename billing to usage
Re-pin lithos-llm to 55add4596b861a0623d00c3a54aa5c147c8d504b and
pebble to c91810fe51aece80359b9cd8efea971af0c46925, where token usage
and cost travel together as Usage { tokens: TokenCounts, cost:
Option<Cost> }. Fabro now carries that one type everywhere it used to
carry BilledTokenCounts, BilledModelUsage, UsdMicros, or a token count
beside a cost_usd_micros.

fabro-types: billing.rs is usage.rs with ModelRef, ModelUsage { model,
usage }, sum_usage, and usage_is_empty; billing_rollup.rs is
usage_rollup.rs with ProjectionUsageStage, ProjectionUsageByModel,
ProjectionUsageRollup, and usage_rollup_from_projection. Every usage
field is named usage: StageProjection.usage and usage_by_model,
Outcome<Option<ModelUsage>>, stage.completed and stage.failed usage and
usage_by_model, prompt.completed usage, run.completed and run.failed
usage (total_usd_micros is gone), Conclusion.usage, StageSummary.usage,
Run.usage. RunSize buckets by Cost.

fabro-workflow: model_usage_from_llm prices tokens from the catalog with
a Catalog cost source, with_reported_cost keeps a provider cost, and the
pebble handler's stage_usage groups pebble's accounts by model and sums
rows with Usage::saturating_add, so a total has a cost only when every
priced part was priced. The store fold's live usage is the agent's
usage plus its descendants'.

API: the OpenAPI spec deletes BilledTokenCounts, BilledModelUsage,
CompletionUsage, CompletionCost, TokenUsage, and RunBillingSummary,
adds TokenCounts, Cost, Usage, and ModelUsage, and renames every
billing schema, property, tag, path, and operation to usage. fabro-api
reuses lithos-llm's and fabro-types' types through with_replacement,
with a round-trip test per replacement.

Old stored runs get no migration: their pebble events in the old shape
read back with zero usage, and their rebuilt projections lose agent
usage.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:31:34 -06:00
Bryan Helmkamp
14f69f50cd
Surface dropped Daytona output to the agent as a stderr line
When the driver reports an output loss on a streaming exec, the pebble
Environment adapter now appends one line to the stderr it hands back:

    [sandbox] N output frame(s), M bytes dropped by the provider

Pebble renders the result's stderr into the tool output, so the model
and the run log both see that the command's output is incomplete rather
than reading a silently shortened stream. The adapter also logs one
`warn!` with `dropped_frames`, `dropped_bytes`, and the command's first
word, bounded, so the operator can find the event without the log
carrying the command itself.

The loss is not folded into pebble's per-stream capture stats: the
dropped frames' stream is unknown and the counts are of encoded bytes,
so attributing them to stdout or stderr would be a guess. The driver's
`truncated` flags on both captures already say the counts undercount.
`ExecOutputTail` is pebble-owned and mirrored in the OpenAPI spec, so
the run events are left alone.

Tests cover the appended line with and without existing stderr, a
lossless command over the scripted double staying unchanged, the
bounded program name, and a lossy command end to end through
`Environment::exec` over a scripted sandbox whose exec facet reports a
loss (the driver's scripted double has no knob for it).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:24:02 -06:00
Bryan Helmkamp
3b8d712edb
Delete a live Daytona test's sandbox even when the test panics
The live Daytona tests create a provider sandbox and delete it on their
last line, so any panic or failed assertion before that line leaks a
running, billed sandbox. Two leaked that way on 2026-09-14 when a
sandbox-driver decoder flake panicked daytona_playwright_mcp_sandbox_transport.

Add fabro_sandbox::test_support::DeletedOnDrop, a guard that owns the
RunSandbox (Deref keeps the tests reading unchanged), offers an explicit
delete(self) for the happy path, and deletes from Drop otherwise. The
drop-time delete runs on its own thread and runtime because the test's
runtime may be unwinding. It reconnects the provider through the
ProviderAccess the test built the sandbox with, because the sandbox's
own handle pools HTTP connections whose tasks live on the test's runtime;
a live check of that path timed out after 10s.

Every Daytona test that creates a sandbox now holds it through the guard.
Unit tests over the scripted double prove delete-on-drop runs once, an
explicit delete runs once, and a panic inside catch_unwind still deletes
with and without a runtime.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 11:01:24 -06:00
Scott Werner
012b556367 Remove tests and fixtures tied to retired manifest fields 2026-09-13 08:55:55 -06:00
Scott Werner
ea11538038 Retire legacy manifest run creation 2026-09-13 08:49:50 -06:00
Bryan Helmkamp
a2ace0888b
Drop the projection parity tests and fold agent_control into agent.activity
The session_projection_parity module pinned fabro's stage fold to pebble's
SessionProjection while both existed; the stage view reads the fold now,
and the one usage rule has its own tests. StageProjection.agent_control
and AgentControlState go too: pebble's fold carries the interrupted and
steered facts as agent.activity, and the stage's state says whether the
stage still runs, which is what the reset on fabro's own stage events was
for. The run-detail banner reads activity plus state.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:38:52 -06:00
Bryan Helmkamp
7a6eea0399
Delete the agent mirrors and move failover to prompt stages
Pebble's stream is the agent event contract. The run's own agent.mcp.ready,
agent.mcp.failed, and agent.mcp.disconnected events, which mirrored pebble's
McpServer* events, are gone with their props, the sink arms that emitted
them, and their conversion and naming entries; pebble's stored
agent.mcp.server.* events are the only record and feed the stage's fold.

The sink no longer mirrors RouteFailover onto agent.failover either: an
agent stage's moves are pebble's agent.route.failover. The event is now
prompt.failover, emitted only by a one-shot prompt stage that walks its
fallback plan itself, and its props are trimmed to the two routes, the
attempt, and the error; nothing read the rest.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:21:41 -06:00
Bryan Helmkamp
b5ee16b37a
Read the stage view from the agent's fold
StageProjection loses todos, subagents, skills, mcp_servers, and
context_window, the types behind them, their fold arms and helpers, and
their OpenAPI schemas: every one of those facts is pebble's fold in
StageProjection.agent now. The context-window endpoint reads the fold's
snapshot, whose event_seq is the agent's own sequence. The parity module
keeps its assertions on the surviving own fields, usage and model, and
checks that what the stage view reads from agent is the whole-session
fold's for the stage's events. The TypeScript client is regenerated and
its stale models removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:13:10 -06:00
Bryan Helmkamp
234953d5c4
Bill a failed agent stage what it spent
An agent stage that failed billed nothing: the backend returned a bare
error and the outcome built from it carried no usage. A terminal failure
now becomes the stage's failed outcome from the same fold that bills a
completed stage, with the tree's usage, the rows by model, the files it
wrote, and its active time; stage.failed carries billing and
billing_by_model and the store keeps both. Cancellation and retryable
failures still go up as the error.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:58:34 -06:00
Bryan Helmkamp
850d5cca52
Bill an agent stage's whole session tree from one fold
One usage rule: a stage's usage is its session tree's, the root and every
subagent, live and at completion. The worker's event sink folds pebble's
SessionProjection over the events it records and the stage's billing and
files come from that fold at stage end, so the completed values are what
the run showed live. The store's live usage is the fold's tree usage, and
completion brings the catalog's price for the same tokens instead of
resetting them to the root's.

Fabro keeps catalog pricing: the root at its route, each descendant at its
own route where the catalog knows it and at the root's otherwise, a
provider-reported cost standing in where pebble has one. The rows travel
as billing_by_model on stage.completed and the stage projection, and the
billing rollup splits by_model by them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:49:27 -06:00
Bryan Helmkamp
424a6a1a47
Embed pebble's SessionProjection in StageProjection
StageProjection.agent is pebble's fold of the stage's agent events, fed
every stored agent event before the fabro-only arms run. Every existing
field and arm stays for now. The parity tests prove each old field is
derivable from the embedded fold: the tree's usage, the route as the model,
the context window without fabro's stamped seq, the root's todo list, the
subagent rows, the skills, and the MCP servers under the disconnected,
error, ready rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:25:38 -06:00
Bryan Helmkamp
217a5860b6
Store ProcessingEnd and the mirrored pebble events in the run log
Pebble's SessionProjection reads ProcessingEnd to complete a prompt and mark
the session idle, so a projection rebuilt from the run's log needs it: one
small event per prompt. The four pebble events the sink mirrored onto
fabro's own agent.failover and agent.mcp.* are now stored verbatim as well,
so the fold sees the route moves and the MCP outcomes; the mirrors stay
until every reader is on the projection.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:25:38 -06:00
Bryan Helmkamp
1d1cacbef3
Pin pebble main 6996942 and implement its PortRoutes trait
Pebble now owns the trait an application hands its MCP support to reach a
port inside the environment. SandboxPortRoutes implements PortRoutes over
the run sandbox's preview-URL facet: a missing facet is Unsupported, a
driver failure is Failed with the driver error as its source. The mcp
feature no longer pins sandbox-driver, so fabro's sandbox-driver pin moves
on its own from here.

The new pebble rev also puts the summary call's usage and cost on
CompactionCompleted, which the CLI progress test literal names.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 07:03:36 -06:00
Bryan Helmkamp
697d8622e9
Merge pull request #843 from fabro-sh/remove/run-metadata-branches
Some checks are pending
TypeScript / Build (push) Waiting to run
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox plugins (stdio) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
TypeScript / Typecheck (push) Waiting to run
TypeScript / Test (push) Waiting to run
Remove Git run metadata branches
2026-09-12 16:41:24 -06:00
Bryan Helmkamp
e5d5c534ab
Merge remote-tracking branch 'origin/main' into remove/run-metadata-branches
Resolve conflicts between the metadata-branch removal and the
sandbox-driver adoption on main:

- fabro-sandbox docker.rs, sandbox.rs, daytona/mod.rs: take main's driver
  rewrite. The Sandbox trait is gone, so the PR's push_token_source
  removal now applies to RunSandbox instead; drop that accessor and the
  RepoCredentials::source helper that only served it.
- run_metadata.rs: keep deleted. Main's edits there were adaptations to
  the driver API and the run git identity field.
- lifecycle/git.rs, finalize.rs: keep the PR's removal of metadata
  snapshots and write_finalize_commit; carry main's RunSandbox,
  GitRetryPolicy, git_identity, local_sandbox, and test catalog changes.
- sandbox_git.rs: take main's version and drop the shadow_sha parameter
  and Fabro-Checkpoint trailer.
- git_integration.rs: remove meta_branch from the new git identity test.
- Cargo.toml: main's dependency set with fabro-dump kept as a
  dev-dependency.
- checkpoints.mdx: keep both the git identity paragraph and the durable
  execution state section.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:28:01 -06:00
Bryan Helmkamp
9550ff0806
Cover the failover continuation and the stopped failover
The existing failover test asserts the backup continued the turn after
the primary committed a tool result. A new test exhausts a two-route
chain and checks that one agent.failover and one
agent.route.failover.stopped are stored on the work stage, the stop
after the error it reports, with the exhausted reason and the failing
route. The events catalog documents both.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:26:32 -06:00
Bryan Helmkamp
57a746235e
Pin pebble to main after lithoscomputer/pebble#12
Pebble's RouteFailover event now describes the route that failed and how
the new route carried the prompt on. Record the continuation on fabro's
agent.failover event as an optional string (replay_prompt or
continue_turn); events written before it existed, and one-shot prompt
stages that walk the plan themselves, read as absent. The failed route's
usage, cost, and timing are not mirrored: the stage's totals already
include them through the prompt report, and no fabro run event carries
per-route usage yet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:26:32 -06:00
Scott Werner
1aaca98ea4
Merge pull request #845 from fabro-sh/codex/cli-workflow-target-selection
Select CLI workflow sources and run targets independently
2026-09-12 16:38:01 -04:00
Bryan Helmkamp
7879e6d223
Merge pull request #856 from fabro-sh/brynary/run-git-identity
Resolve one Git identity per run and inject it into every workflow command
2026-09-12 14:23:30 -06:00
Bryan Helmkamp
ff7821dcd2
Adopt pebble's MCP disconnect event and startup timing
Pebble reports an MCP server whose connection closed mid-session once,
as McpServerDisconnected, and carries startup_ms on McpServerReady and
McpServerFailed. The workflow event sink mirrors the disconnect onto a
new agent.mcp.disconnected run event shaped like agent.mcp.failed, and
passes startup_ms through on agent.mcp.ready and agent.mcp.failed. The
raw pebble event is not stored for these, so the timing would otherwise
be dropped at the boundary.

The stage projection's McpServerStatus gains a `disconnected` kind next
to `ready` and `failed`. The fold keeps the server's tool count and
sticky invoked flag and only moves the status. The OpenAPI
McpServerStatus oneOf gains McpServerStatusDisconnected, and the
fabro-api round-trip test covers its JSON shape. A new
session_projection_parity test folds the same MCP events through
pebble's SessionProjection and fabro's stage projection and compares
them, including pebble's `disconnected`.

ToolErrorKind::Timeout needs no fabro change: the kind is stored as
pebble serializes it and never matched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 12:39:31 -06:00
Bryan Helmkamp
1aee59585b
Pin pebble to mcp-embedder-events and accept startup_ms on MCP outcomes
Pebble's McpServerReady and McpServerFailed events now carry startup_ms.
The workflow event sink destructured both variants by name, so it stops
listing every field. The pin is temporary: it moves to pebble main once
lithoscomputer/pebble merges the branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 12:29:36 -06:00
Scott Werner
0998f60624 Contain the selected workflow file instead of pre-walking the checkout
Remote acquisition walked the entire depth-1 checkout and failed on any
symlink that dangled or resolved outside the root, even when the link
was nowhere near the selected workflow. Submodule-style dangling links
and links into the host are common in workflow repositories and made
--workflow-git fail where the same commit collected fine locally. The
bundler already root-checks every file it opens; the only unchecked
reads were the selected TOML (or a graph selector's sibling TOML) during
location resolution. Check those in collect_workflow_versions and drop
the O(repo) walk. walkdir stays a dev-dependency for the dump tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 11:48:21 -06:00
Bryan Helmkamp
ee6576cff7
Resolve one Git identity per run and inject it everywhere
A run now resolves a single author and committer identity once, after its
GitHub credentials are selected and before anything can commit, and uses it
for every commit it creates. Resolution order: a complete explicit
`run.git.author`; the run's GitHub App bot account
(`<slug>[bot] <id+slug[bot]@users.noreply.github.com>`); the authenticated
user of the run's PAT; the generic `Fabro <noreply@fabro.sh>`. A partial
explicit author overlays the fields it supplies. Only the selected
credential is consulted; a failed lookup is a setup error. A standalone
installation token falls back to the generic identity with a warning.

The resolved identity is carried on `RunOptions` and `EngineServices`,
recorded as a `git.identity.resolved` event and `RunProjection.git_identity`
so resume reuses it, and exposed through the run state API. Engine
checkpoints and metadata commits read it through `RunOptions::git_author`.
Every workflow execution path receives it as `GIT_AUTHOR_NAME`,
`GIT_AUTHOR_EMAIL`, `GIT_COMMITTER_NAME`, and `GIT_COMMITTER_EMAIL`, applied
last so it wins over inherited host variables and `[run.environment]`
entries: prepare steps, command stages, native agent shell tools, and ACP
launches. The identity is injected even without a Git origin, and the old
local `git config user.*` write is removed.

fabro-github gains `GET /user` and `/users/{slug}[bot]` lookups with mocked
tests for success, unauthorized, malformed, and transient cases. Real-Git
integration tests commit in the primary checkout, a clone, and a fresh
repository under conflicting local config, `[run.environment]`, and host
variables, and prove concurrent runs do not leak identities. CLI workflow
tests cover host script stages and ACP launch env through `fabro run`.

Docs and generated option metadata now describe the credential-derived
defaults instead of the stale `fabro`/`fabro@local` values.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 11:35:46 -06:00