Commit graph

553 commits

Author SHA1 Message Date
Bryan Helmkamp
d3d6697aa1
Merge origin/main and pin sandbox-driver 583a164, Pebble eb08b70, Petri 23f78a2
main (#891) upstreamed the Pebble sandbox adapter and deleted
fabro-pebble-sandbox, so Fabro now builds Pebble's sandbox-driver feature
beside Petri's crates and must hold one sandbox-driver copy. The three
pins move together: sandbox-driver to its main after #61 (host
attach-by-path, merged), Pebble to pebble#27's head after it moved its
own sandbox-driver pin to the same revision, and Petri to petri#36's
head, which carries both bumps.

Resolution: Cargo.toml keeps main's shape with the three revisions
rewritten; Cargo.lock regenerated from main's copy and holds one copy of
each library.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-22 07:50:57 -04:00
Bryan Helmkamp
baf32b70a7
Attach to a run's host sandbox with the worker's managed record
A run's host sandbox is a managed directory the worker recorded, with
Petri's labels, in the host registry inside the run's Petri directory.
The server reached it only by designating the directory again: a handle
with no record, so no labels and no ownership check, unlike Docker and
Daytona where `petri.run` is checked on every attach.

The server now observes the run's registry (`HostProvider::
observe_registry`, read and never written, so the worker stays its only
writer), resolves the directory to its record by path
(`attach_directory`), and runs the same `petri.run` ownership check as
on Docker (`OwnedProvider::check`). The handle refuses every lifecycle
change, so starting, stopping, and deleting stay the worker's, and a
stopped sandbox's retained workspace is usable through it; the access
paths therefore return a host handle as attached instead of activating
it. `ProviderAccess` carries the storage root the registry is found
under; a directory no registry records (a run older than this driver, a
caller without a storage root, a pruned Petri directory) is designated
again as before, without labels, for the read paths only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 17:44:48 -04:00
Bryan Helmkamp
45d94ce711
Use pebble's sandbox-driver adapter and delete fabro-pebble-sandbox
`lib/components/fabro-pebble-sandbox` moved into pebble as
`pebble_coding_agent::sandbox_driver` (lithoscomputer/pebble#27): the
`Environment` over a driver handle (`SandboxEnvironment`, was
`PebbleSandbox`), the `SandboxExec` policy, the port routes, and
`display_for_log`. The pebble pin moves to that branch head, 6d03b3b,
with the `sandbox-driver` feature on (`sandbox-driver-test-util` for the
server's tests, which take `MockSandbox` from pebble now). Nothing in the
crate was Fabro's by design; what was Fabro's stays: `SecretRedactor`
moves to `fabro-redact` as pebble's `Redactor` over `redact_string`, and
the log renderer takes it where a driver failure is rendered.

`fabro-petri` hands pebble types to Petri's crates, so Petri must pin the
same pebble revision: the petri pins move to lithoscomputer/petri#36
(9ee3f85), which pins pebble at the same head. Both re-pin to the pebble
merge commit together once #27 merges.

The 14 pebble commits between the pins fold the session projection's
lifetime tallies into `SessionProjection::totals` (and `PromptDelta`'s
into a flattened `totals`, which renames the prompt's `subagents` key to
`subagent_counts`, as the projection's already was). The stage progress
fold, the runs handler, the OpenAPI schema, the generated client model,
and the round-trip test follow.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 17:42:15 -04:00
Bryan Helmkamp
968c46d69b
Settle an in-process run at Petri's finish, not at its terminal record
The in-process Petri path settled the run's in-memory managed run once
Fabro's terminal lifecycle record was stored, after the engine returned.
GET /runs/{id} reads the stored summary, which the projector ends at
Petri's own `run.finished` record, a moment earlier, so a delete issued
the moment the run read as ended could reach the delete precheck, which
prefers the managed run, while it still said running, and was refused
with 409 "cannot remove active run". Against the real engine the window
hit eight times in thirty.

The worker path settles its run at the worker's records endpoint, ahead
of the store (#888). The in-process run now settles at the same record
through the run store it executes over: a store whose coordinator
appends settle the managed run at the `run.finished` record before the
record reaches the store and the projector's signal, with the same
finish mapping and settle the worker path uses. The settle is in memory
only. The terminal lifecycle record stored once the engine returns
refines the status and error and ends the run's live state as before,
and stays the settle of a run that ends without an engine finish.

The prune scenario no longer waits for the managed run to settle before
its delete: the wait guarded only this window.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 12:08:48 -04:00
Bryan Helmkamp
81b7427add
Merge pull request #888 from fabro-sh/settle-worker-run-on-terminal-record
Settle a worker's run at Petri's finish, not at the worker's exit
2026-09-21 08:08:31 -04:00
Bryan Helmkamp
2b30689e77
Settle a worker's run at Petri's finish, not at the worker's exit
A worker-backed run's in-memory managed run settled only when the
worker process exited. GET /runs/{id} reads the stored summary, which
the projector ends at Petri's own `run.finished` record, a moment
before the worker stores Fabro's terminal lifecycle record and exits.
The delete precheck prefers the managed run, so a delete issued in that
window was refused with 409 "cannot remove active run". Against a real
worker the window hit about six times in ten.

The server sees both records before they are stored: `run.finished` on
the coordinator log through the worker's records endpoint, and the
terminal lifecycle record through the platform-records endpoint. The
managed run now settles at either, ahead of the store, so the view
never reports the run ended while the managed run still says running.
The stream follower no longer reopens a settled run with the records
that precede its terminal one, and the worker's exit keeps the settled
status: it reaps the process, records a missing terminal record as it
did, and takes the store's status only when the store ended the run
differently. The mapping from Petri's finish to the run's status is the
projection's own, shared with its fold.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 03:49:44 -04:00
Bryan Helmkamp
1fda9633e4
Bind the run's resolved goal into Petri's check
An intent's goal override, and a `[run.goal]` layer, reached the run's
display graph and its settings but not Petri's check, so the agent stages
executed with the workflow's own goal while the run showed the override.
The launch now carries the run's resolved goal (the settings' inline
`run.goal`, layered as the create path layers it) as Petri's
`petri.launch_goal` compile variable, which Petri binds over the bundle's
`[run] goal` and the graph's own `goal`, so admission's frozen plan carries
the goal the run shows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 03:48:35 -04:00
Bryan Helmkamp
b5c63474e0
Qualify the terminal check: the scenario already imports the engine's RunStatus
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 16:27:02 -04:00
Bryan Helmkamp
c4086f1804
Name the terminal check directly in the managed settle wait
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 16:16:06 -04:00
Bryan Helmkamp
9a93bbcfbd
Wait for the managed run to settle before the prune scenario deletes it
The stored view reports a Petri run ended as soon as its own run.finished
record is folded, which is before the server stores the terminal
lifecycle record and settles the managed run in its map. The delete
precheck reads that map, so a delete sent as soon as the API reports the
run ended can be refused as active instead of by the held lease. The
scenario now waits for the managed run to settle, through a test-support
accessor for its status, before it deletes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 16:03:20 -04:00
Bryan Helmkamp
e010563ad5
Read the resume action off the unified restore log line
The Docker recovery scenarios scraped the worker log for "sandbox
workspace brought to its durable snapshot"; since the host and sandbox
restores share one path, the line reads "workspace brought to its
durable snapshot" with the site as a field.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 15:31:39 -04:00
Bryan Helmkamp
4216364e92
Split the projector into the pass, its ordering rule, its stream and its wake-up
projector.rs mixed the commit rule with things that are not the pass:
the ordering rule and its tests, the in-process signalling store, the
stream table's rows and reads, and six readers documented "for a test".
projector/mod.rs now holds the pass alone; order.rs, signalling.rs and
stream.rs hold the rest; the test readers (rebuild, stored_projection,
stored_stream, stored_platform_records, event_json) live in test_support
behind the test-support feature, as the test-support boundary rule asks,
and the two test callers reach them there. The fault injection and the
cache's test-only reader are gated the same way, so neither ships.

pass() reads as three steps: at_head, read_new_events (a NewEvents
struct in place of a tuple) and stream::stream_rows, which rebuild
shares instead of repeating the fold loop. PassReport::skipped and
PassReport::contended replace the zero-filled literals.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 15:10:03 -04:00
Bryan Helmkamp
1f6e53c949
Share one poison-tolerant lock helper from fabro-util
Seven crates' files each carried the same four-line lock function that
recovers a poisoned mutex. fabro_util::sync::lock is that function, once;
the copies in fabro-petri and fabro-server are gone. fabro-template's
helper panics on poison instead, a different policy, and is left as is.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 14:47:46 -04:00
Bryan Helmkamp
c264a3567c
Settle the managed run as soon as its terminal record is stored
The in-process Petri path persisted the run's terminal lifecycle
record, settled the projector, aggregated usage, and only then settled
the managed run. GET /runs/{id} reads the stored summary, so it reported
the run as ended while the delete precheck, which prefers the managed
run, still saw it running and refused the delete as active. The prune
scenario hit that window about once in thirty runs. Settle the managed
run right after the record is stored, before the view catches up.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:56:48 -04:00
Bryan Helmkamp
634891ded2
Filter the working-directory depth out of the partial include snapshot
The include error names the partial relative to the run's working
directory, so its `../` run is as long as that directory is deep: eight
on this machine's temp dir, three on the CI runner's. Collapse the run
to a token before the snapshot compares.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:39:18 -04:00
Bryan Helmkamp
fdf5141917
Wait for the terminal lifecycle record before reading the cancelled run
The detached cancel test waited for the run's status to read `failed`
and then asserted on the stored `run.lifecycle` record. The projection
concludes the run from Petri's `run.finished` coordinator record, and
the worker stores the platform's terminal lifecycle record a moment
later, so the read raced the write and the assertion failed about once
in thirty runs. Wait for the record itself, and print the stored events
and the run state when the assertion fails.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:31:49 -04:00
Bryan Helmkamp
beac547a1f
Run the workflow scenarios on the Docker provider instead of the stdio plugins
The host_plugin_ and docker_plugin_ variants ran each scenario under
Fabro's old plugin transport with the provider kinds `host` and
`docker-plugin`, which Petri's Fabro frontend rejects. Under Petri every
provider is already served by a sandbox-driver plugin, so those variants
test nothing distinct. A single docker_ variant replaces them: an
environment with provider `docker` on buildpack-deps:noble, created on
an isolated server, skipping without the sandbox-driver-docker
executable or a daemon with the image unless
FABRO_REQUIRE_SANDBOX_PLUGINS is set.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:20:24 -04:00
Bryan Helmkamp
49a647e3a4
Install the sandbox-driver plugins in the Rust test jobs
Every Petri run takes its scope through a sandbox-driver plugin
executable that Petri finds on PATH, so the test jobs need
sandbox-driver-host and sandbox-driver-docker installed at the rev the
workspace pins. The three jobs share one from-source install through an
actions/cache entry keyed on the OS and the rev.

The Linux test job also pre-pulls Petri's default runner image, which
the suite's Docker scenarios leave to Petri: the plugin pulls it on
first use, but a 1 GiB pull inside a run's timeout is a flake.

The stdio plugin job was built for the deleted fabro-sandbox layer. It
becomes the Docker providers job: the `docker_` scenario variants and
the fabro-petri suite, with the fabro-sandbox and fabro-workflow steps
whose tests no longer exist removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:20:24 -04:00
Bryan Helmkamp
399aef8111
Keep the Daytona key rendering out of the test's assertion messages
CodeQL read the assertion messages as a log of the credentials'
Debug output. The test proves that output never holds the key, so the
messages added nothing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:05:01 -04:00
Bryan Helmkamp
ec5faab116
Merge remote-tracking branch 'origin/main' into petri-integration
# Conflicts:
#	lib/apps/fabro-cli/src/commands/run/runner.rs
2026-09-19 15:25:20 -04:00
Bryan Helmkamp
6082f82950
Fix the gate findings in the prune change
Clippy's absolute-paths lint on the rendered prune error, the sync
directory reads the prune test documents, and a redundant rustdoc link.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 15:02:20 -04:00
Bryan Helmkamp
dfd2458f76
Test an Ask Fabro turn against the container Petri created
A live-server scenario runs a one-stage workflow on Docker whose command
writes a file into the workspace, opens an Ask Fabro session on the
finished run, and sends one turn. The session attaches to the container
Petri created, stopped at the run's end, starts it again, and its tool
reads the file inside it; the turn succeeds, the tool's output and the
model's reply carry the file's content, and the twin's follow-up request
shows the model read it from the tool. Ask Fabro's tool policy is
read-only, so the shell tool is hidden from the model and refused; the
turn reads the file with the `read_file` tool, scripted on the twin,
instead of a shell `cat`.

The scenario skips, and says why, without the Docker plugin or a daemon,
as the other Docker scenarios do.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:58:09 -04:00
Bryan Helmkamp
c7aa50c943
Delete a run's sandboxes through Petri's lease ledger
Run deletion called the driver's `provider.delete(id)` under the run's
`petri.run` scope, a delete of Fabro's own over a sandbox whose lease
record Petri owns. It now goes the way `petri sandbox prune` goes:
`fabro_petri::prune` builds the run's Petri runtime over the server's
store (the run key, the run directory, the sandbox backend) and calls
Petri's prune, which opens the run for writing, checks each lease's
provider fingerprint, writes the delete intent and the tombstone beside
the run's other records, and lets each provider remove its managed
workspace, a host workspace included.

A run a live process holds answers 409 unless the delete is forced; a
lease Petri could not prune answers 409 with the problem text, or is
warned and skipped under force or a delete that already started. The
server drops the worker's handles before the prune, on the store
instance the prune opens, so the lease a stopped worker held is released
first. The projection reads only the coordinator and execution logs, so
the resource records change nothing it reports.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:54:03 -04:00
Bryan Helmkamp
809b3891b5
Forward every configured sandbox plugin to the Petri worker
The worker's environment carried only the host, Docker and Daytona plugin
variables from the server's own environment, so a run on a third-party
provider kind never learned where its plugin was, although the server
kept `[server.sandbox.providers.<kind>]` plugin settings for its own
attach. The launch spec now derives `PETRI_SANDBOX_<KIND>_PLUGIN` and
`PETRI_SANDBOX_<KIND>_SHA256` from every enabled kind's plugin settings,
and `PETRI_SANDBOX_PLUGIN_DEV=1` when any of them sets `dev`, set after
the allowlist so the settings win over an ambient variable of the same
name and the allowlist stays the fallback.

The server's default plugin binary name is `sandbox-driver-<kind>`, the
name Petri looks up, since one executable serves both sides.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:45:42 -04:00
Bryan Helmkamp
956feda0ca
Fix the gate findings in the sandbox tests
The grep test reads paths as the driver reports them for a resolved
absolute path; the local preflight check test gives the manifest a
source directory that exists; the Docker attach scenario accepts that
the in-process app has no daemon record for the run-tools client an Ask
Fabro turn builds after the sandbox attach, and asserts the turn got past
the sandbox. The inventory's lazy connection is boxed for clippy's
variant-size lint.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:07:52 -04:00
Bryan Helmkamp
2a4f2aa718
Test the server's attach to the container Petri created
A server scenario runs a command workflow on Docker through Petri's
plugin (skipped without the plugin or a daemon), then reaches the
container without Petri: the sandbox tab describes it under its
`petri.run` label, Run Files writes, lists and reads a file in its
workspace after starting the stopped container, a preview URL opens to a
port in it, and an Ask Fabro turn runs against it through the OpenAI
twin. A unit test attaches through the ownership seam with a scripted
provider: the run's own container attaches, another run's and one that
carries only Fabro's retired `sh.fabro.*` labels are refused as not
owned.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:00:56 -04:00
Bryan Helmkamp
f054082f86
Delete fabro-sandbox
Nothing imports it any more: the Pebble glue lives in
fabro-pebble-sandbox, the server reaches run sandboxes through
sandbox_access, and Petri creates every run sandbox. The crate, its
test-support, its integration tests and every dependency edge go with
it. The `[server.sandbox.providers.<kind>.plugin]` settings stay: the
server still launches a plugin executable through them to attach to a
sandbox of a non-bundled kind.

AGENTS.md names the new crate and the direct-access pattern in place of
`RunSandbox`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 13:57:30 -04:00
Bryan Helmkamp
23b8a6c449
Serve the sandbox tab, Run Files, and deletion on the driver handle
The sandbox handlers describe, list, download, upload, open a terminal
and build SSH and VNC access on the `Arc<dyn Sandbox>` the server attaches
to the run's record, with paths resolved against the recorded working
directory. Run Files holds the handle beside that directory and runs its
git through the driver's git facet and Fabro's exec policy;
fabro-workflow's sandbox git takes the same pair, and its
`GitCommandError` carries the driver's error. Run deletion deletes by id
through the provider scoped to the run's `petri.run` label, so a foreign
sandbox is refused and a designated host directory is left in place. Ask
Fabro wraps the attached, running handle. The legacy access shim is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 13:56:13 -04:00
Bryan Helmkamp
0cc645b2d1
Reach run sandboxes from the server through the sandbox driver
`fabro-server/src/sandbox_access.rs` is the server's own path to a run's
sandbox: it connects the record's provider (the driver's Host, Docker and
Daytona providers in process, a plugin executable for any other kind),
keys ownership on the `petri.run` label Petri stamps on every sandbox it
creates, attaches by the recorded id, and for a host record designates
the recorded directory again when the id lives only in the worker's
registry. The Docker client resolves its endpoint from the same variables
Petri forwards to its plugin, so both meet on one daemon.

The doctor's Docker check and the Daytona credential probe move here with
`DaytonaCredentials`, and the `/sandboxes` inventory is rebuilt over the
driver's `list`, narrowed to sandboxes that carry Petri's run label.
Preflight asks the provider for its health instead of creating and
deleting a throwaway sandbox in Fabro's own shape, which no run uses; the
git retry policy behind the repository probe moves into run_manifest.

The callers still on fabro-sandbox's reconnect read their access through
a `legacy_provider_access` shim until they move.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 13:47:02 -04:00
Bryan Helmkamp
c4ed995b44
Add fabro-pebble-sandbox: a driver handle as pebble's Environment
Petri creates and owns every run sandbox through the sandbox driver, so
what Fabro still needs around a driver handle is the Pebble glue: the
Environment pebble's coding agent runs its tools through, the exec policy
(stop grace, working directory, StripAll, the termination mapping, the
redacted output tail), pebble's port routes over the driver's preview
URLs, the secret redactor, the path helpers, and a log rendering that
appends a failed command's redacted tail. This crate holds that glue,
moved from fabro-sandbox, over `Arc<dyn Sandbox>` plus a working
directory instead of `RunSandbox`, with a `MockSandbox` double behind
`test-support`.

`fabro exec` creates its host sandbox directly on the driver's Host
provider and activates it; Ask Fabro wraps the attached handle. Both
keep the provider alive beside the sandbox where the session's processes
are the provider's process groups.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 13:37:38 -04:00
Bryan Helmkamp
467087998d
Read workflow graphs through Petri's DOT parser
Fabro's own DOT parser was left with one job after create-time compile
moved to Petri: walking a workflow's file references for the bundler and
the workflow-version store, and reading a name, a goal and two counts.
Petri's frontend parses the same language, so the parser goes and a small
crate reads the graph through Petri's.

`fabro-dot` is that crate: `WorkflowGraph::parse` over
`petri_frontend_attractor::dot` and its semantic model (defaults applied,
subgraphs flattened, chains expanded), `references(position)` as the one
walker over the static-reference vocabulary (each reference with its node
and position, file references checked to be template-free), and
`normalize_for_graphviz`, the re-emit of Fabro DOT with dotted attribute
keys quoted, which the SVG render needs. It sits beside `fabro-petri`
rather than inside it because `fabro-petri` depends on `fabro-workflow`,
which depends on `fabro-workflow-version`: the version store cannot reach
`fabro-petri` without a cycle, and the bundler should not pull the engine
in to read a graph.

Deleted: `fabro-graphviz`'s lexer, grammar, AST, semantic pass and
`parse_ast` (1,829 lines, plus the `nom` dependency); the DOT model in
`fabro-types::graph` (`Graph`, `Node`, `Edge`, `AttrValue`,
`shape_to_handler_type`), with only `ReferenceKind` kept, moved to
`fabro_types::reference`; `fabro-template`'s `visit_graph_references` and
the `GraphReference`/`GraphPosition` types, with the template-syntax rule
(`validate_static_reference`) kept there; the pull-request body's DOT
fallback summary, which was unreachable because the DOT source only
travels with the run spec whose display graph the summary already reads.
`fabro-graphviz` is now the render alone, over `fabro-dot`.

Parity: the old and new walkers were run over every `.fabro` and `.dot`
file in the repository (118) before the deletion. Every reference set is
identical. Five files differ in what Petri reads more correctly: a
backslash before a newline inside a quoted string is a line continuation
(four files, inline prompt text only), and a node named only by an edge
counts as a node (`test/edge_only_node.fabro`, 3 nodes rather than 2, so
the `fabro validate` snapshot moves). The checked-in bundles' shapes and
references are pinned by a snapshot in `fabro-dot`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 12:23:57 -04:00
Bryan Helmkamp
bd59f52e22
Build the run's display graph from Petri's admission
The rest of the change whose deletions the previous commit carries (its
`git add` stopped at an already-removed path): `fabro_types::RunGraph`
and the `fabro-petri` builder that reads it off the admitted graph, the
server's create, validate, preflight and render paths on Petri's check
alone, the consumers moved to the new shape, the OpenAPI `RunGraph`
schemas with their parity tests, the regenerated TS client, and the
docs naming Petri's diagnostic codes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:41:58 -04:00
Bryan Helmkamp
52aed8c642
Read the run's display graph off Petri's admitted graph
Every run is admitted by Petri, whose check lowers imports, file
references, templates and the model stylesheet, lints the workflow and
pins its models. Fabro then re-parsed the same workflow through its own
legacy pipeline (parse, transforms, structural validation) only to fill
`RunSpec.graph` for the read side. That second pass is gone: the run's
display graph is `fabro_types::RunGraph`, built once in `fabro-petri`
from the admitted graph's metadata (the workflow name and goal from the
graph params; each declared stage's label and handler kind; one edge per
routing arm as written, lowering artifacts left out), and stored on the
spec at create beside the DOT as `graph_source`.

Deleted: `fabro-workflow`'s `pipeline`, `transforms`, `file_resolver`,
`operations::{source, validate}`, `run_materialization` and the legacy
`compile_admitted_run`; the server's `compile_admitted`, the structural
manifest pass, `preflight_model` and the model probe `run_llm_check`
(Petri's admission raises `attractor.model.unknown`); `fabro-graphviz`'s
stylesheet parser; most of `fabro_types::graph` (the DOT model keeps
what the bundler, version registration and template walker read). The
DOT parser stays for the bundler and the SVG render.

`POST /validate`, `POST /preflight`, `POST /graph/render`, `fabro
validate` and `fabro preflight` run on Petri's check alone, so their
diagnostics carry Petri's codes (`attractor.unbound_input`,
`unsupported.template.unbound_input`) where Fabro's
`template_undefined_variable` and `goal_self_reference` were. A refused
workflow's summary still names the DOT as written. The OpenAPI `RunSpec`
schema gains `RunGraph`, `RunGraphNode` and `RunGraphEdge`, reused from
`fabro-types` with parity tests; the TS client is regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:41:38 -04:00
Bryan Helmkamp
eca2812602
Fix what the gates found after the removal sweep
The CLI artifact scenario seeded its run through the deleted upload
route; it now runs a real Petri workflow whose hooks collect the
artifacts, and the fabro artifact list and cp assertions read those.
A real command retry is not producible from a command node (a plain
failure or a timeout routes onward), so the retry dimension of the old
fixture goes; the stage, node, and retry filters, the tree copies, the
cross-stage ambiguity, and the filename collision stay covered. The
archive guard test drops its upload row (the blob write row covers an
octet-stream mutation). A dangling doc comment and two absolute paths
clippy flagged are fixed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 10:15:15 -04:00
Bryan Helmkamp
51138cda57
Drop the dependency edges with no production use
Each edge was checked with rg over the crate's sources outside its
test paths. fabro-cli keeps git2, regex, ulid, and shlex as
dev-dependencies for its integration tests. fabro-automation keeps
chrono, tokio, and tracing: its migrations compile into the crate
through #[path]. petri_testkit was already optional behind fabro-petri's
test-support feature and dual-listed as a dev-dependency. The workspace
loses the agent-client-protocol and AWS entries no crate references;
jsonschema stays for the fabro-api and fabro-tool tests. Cargo.lock
was refreshed by a plain build and only loses entries.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:33:31 -04:00
Bryan Helmkamp
05a14a6c9b
Delete the checkpoint endpoint, fabro parse, and the fabro-workflow shims
GET /runs/{id}/checkpoint duplicated what /state serves; the hidden
fabro parse command had no user; records, run_status, outcome, and
usage_rollup in fabro-workflow only re-exported fabro_types. The
importers now name fabro_types directly. format_cost keeps its two
callers (the pull request body and the CLI stage display) and moves to
fabro_types::usage; the usage rollup tests move beside the function in
fabro-types, with test_usage in its test support.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:31:09 -04:00
Bryan Helmkamp
0752d4c7c4
Delete the stage artifact upload endpoint
Artifacts reach the blob table through the hooks, so the POST on
/runs/{id}/stages/{stageId}/artifacts, its octet-stream and multipart
handlers, the RequireStageArtifact extractor, the client's upload
functions, and the generated TypeScript operation go. The spec loses
the operation, the multipart variant writeRunBlob never served, and
the batch manifest schemas; fabro-types loses ArtifactUpload, the
batch upload's only input type. Every list and download path stays,
and the server tests seed the artifact store directly to cover them.
fabro-server drops multer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:23:47 -04:00
Bryan Helmkamp
06f9cb8361
Delete fabro-sandbox's clone and push chain
The engine prepares every run's checkout, so fabro's clone
orchestration, the per-checkout GitHub credentials, the run-branch
setup, the push retries, and the push policies had no production
caller. RepoWorkspace::plan still validates the clone request and now
refuses one that asks for a clone; initialize creates an empty
workspace root. SandboxWorkspaceLayout and snapshot_info stay: the run
record projection in sandbox_spec.rs reads them. The run tool
regression keeps its assertion (a child targets the parent's pushed
run branch) over a plain git fixture instead of the deleted setup. The
Docker, Daytona, and Daytona-wire clone layout tests go: they proved
only the legacy clone. fabro-sandbox drops base64, uuid, fabro-proc,
serde, strum, and sandbox-driver-daytona-config; chrono is test-only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:18:29 -04:00
Bryan Helmkamp
01713c6aac
Hand Petri only the environment keys it reads
The settings layer carried every key of every catalog environment, so a
run in an environment with `lifecycle`, `labels`, `cwd`, `network` or a
Dockerfile warned `ignored.workflow_toml.environments.<id>.<key>` on
every admit. Those keys are the platform's and stay with the server's own
resolution; the layer now carries the provider, `image.docker` under
`docker` and `daytona`, `resources` under `daytona`, and `env`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:27:57 -04:00
Bryan Helmkamp
1166caed0b
Give up on an unreachable worker control stream and reap .ft- test daemons
A worker whose server is killed outright retried its control-stream
connect forever at a 5 s capped backoff, so it never finalized, never
stopped its sandbox, and lived until reboot. The worker now tracks the
start of each run of continuous connection failure and gives up after
60 s (10 s once its parent is pid 1), through the existing fatal
control-loss path that interrupts interviews and cancels the run. After
that fatal fires, the runner keeps driving the cancelled pipeline for a
bounded grace so `conclude` can stop the sandbox before the process
exits, instead of dropping the pipeline future mid-flight.

The test harness's stale-daemon reaper only matched `fabro server`
titles bound under `/tmp/.tmp*`, but `TestContext` roots are
`.ft-<label>-*` under `std::env::temp_dir()`, so daemons bound there
were never reaped. The regex now also matches sockets below a `.ft-`
directory component under any parent, and the unit test covers both
roots plus real-looking paths and `.ft-` outside a directory component.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:27:55 -04:00
Bryan Helmkamp
978a5b1b7e
Merge branch 'petri-integration' into petri-followup-envcat
# Conflicts:
#	Cargo.lock
#	Cargo.toml
2026-09-19 07:10:02 -04:00
Bryan Helmkamp
b1d95faa57
Hand Petri the server's environment and MCP catalogs
Petri's Fabro frontend refused a bundle naming an environment it did not
declare and every MCP catalog reference, so the fixtures declared
`[environments.local]` and the server's catalogs never reached Petri.

Pin Petri at c874b86, where the frontend reads `[environments.<id>]` and
`[run.environment]` from every settings layer (bundle over project over the
host's layer, key by key), takes the environment a launch selected over the
layers, and resolves `[run.agent.mcps.<name>] id = "..."` against a catalog
the host binds. The server hands Petri its environment catalog as
`[environments.<id>]` tables of the settings layer it already passes, the
intent's environment as the launch's selection (`Launch::environment`, as
the intent overrides the bundle in Fabro's own resolution), and its MCP
catalog as `RuntimeSpec::mcp_catalog_toml`, one inline entry per definition
keyed by id. Offline validation hands Petri the seeded catalog the same
way, so `fabro validate` accepts `[run.environment] id = "local"`.

The fixtures drop the `[environments.local]` tables they carried for this;
the secrets test keeps its own, on purpose. Scenario tests cover a bundle
naming a catalog environment (its image lowered, and run on Docker when the
plugin and a daemon are there), a bundle's own table winning key by key,
the server refusing an unknown environment before Petri, and a catalog MCP
reference whose tool the agent session lists (an echo server under
`test/mcp/`).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:08:08 -04:00
Bryan Helmkamp
776c50307f
Merge remote-tracking branch 'origin/main' into petri-integration
# Conflicts:
#	Cargo.lock
#	Cargo.toml
#	lib/components/fabro-validate/src/lib.rs
#	lib/components/fabro-workflow/src/handler/llm/fallback.rs
#	lib/components/fabro-workflow/src/handler/prompt.rs
#	lib/components/fabro-workflow/src/operations/start.rs
#	lib/components/fabro-workflow/src/pipeline/pull_request.rs
#	lib/components/fabro-workflow/src/transforms/model_resolution.rs
#	lib/components/fabro-workflow/tests/it/integration.rs
#	lib/components/fabro-workflow/tests/it/pebble_agent.rs
2026-09-19 06:37:02 -04:00
Bryan Helmkamp
cb26c5603d
Answer steer and interrupt with the worker's acknowledgement
The worker control bus was publish-only: the steer and interrupt
endpoints answered 202 once the control was forwarded, and a refusal
showed up only later as a `run.notice` on the run's stream.

A steer or an interrupt now carries a request id. The worker answers it
over the control stream it arrived on with `{request_id, outcome}`,
where the outcome is `delivered` (with the stage's label) or `refused`
(with the code and the reason). The server keeps the outstanding
requests in a registry and waits up to 5 s for the answer: the endpoint
answers 202 `{"outcome":"delivered","stage":…}`, 409 with the refusal's
code (`no_live_turn`, `no_such_stage`, `steer_refused`,
`interrupt_refused`) and message, or 202 `{"outcome":"pending"}` when
the worker gave no answer in time. The `run.notice` record on refusal
stays, under the same code, so a steer to a stage that is not running is
now `no_such_stage` there too. Pause and unpause are unchanged.

The in-process test path answers a steer or an interrupt from the run's
own controls at once. `FABRO_TEST_CONTROL_ACKS_MUTED=1` on the server
mutes the worker's answers, so a test can see the pending fallback.
`fabro steer` prints the worker's answer, and a refusal is its error.
The OpenAPI spec documents the 202 body and the 409 codes; the Rust and
TypeScript clients are regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 23:54:11 -04:00
Bryan Helmkamp
86ac13713b
Merge branch 'petri-followup-fork' into petri-integration 2026-09-18 23:18:14 -04:00
Bryan Helmkamp
22ebbf73e4
Cover fork, retry, rewind and the timeline with real-binary scenarios
Through a real server and its worker: a three-stage run forked at its
first stage continues with the other two on the first stage's restored
file and commits on the new run branch after the source's commits; a
retry of a run whose last stage failed transiently reruns that stage and
succeeds; a rewind archives and supersedes its source, and an archived
run is refused; the timeline lists every checkpoint record with its
commit; a fork at a checkpoint inside a parallel branch is refused and
creates nothing, while one after the join continues.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 23:17:25 -04:00
Bryan Helmkamp
8d4a1bf894
Add the fork, rewind, retry and timeline commands to the CLI
`fabro timeline <run>` prints the checkpoint table (or JSON); `fabro fork
<run> [target]` and `fabro rewind <run> <target>` make and start the new
run and name it, `--list` showing the timeline instead; `fabro retry
<run>` starts the retry and prints its id. The top-level help snapshot
and the generated CLI reference follow.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 23:17:25 -04:00
Bryan Helmkamp
e7d55d6ec7
Serve fork, rewind, retry and the timeline on the runs API
`GET /runs/{id}/timeline` lists the run's checkpoints with their Petri
positions, stages, commits and diff summaries, and its fork origin.
`POST /runs/{id}/fork` resolves a target on that timeline, creates the
new run, seeds it through `fabro_petri::fork` and queues it in resume
mode, so its worker restores the checkpoint into a fresh workspace and
continues from the position; a position inside a parallel branch is
refused with 400 before the run exists. `POST /runs/{id}/rewind` is that
fork of a terminal run followed by the source's archive and its
`run.superseded` record (207 when the archive fails); `POST
/runs/{id}/retry` forks a terminal run at its last checkpoint, rerunning
the stage that failed. The projection carries `forked_from`. The Rust
and TypeScript clients gain the four calls.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 23:17:25 -04:00
Bryan Helmkamp
a664b0ccce
Merge branch 'petri-followup-interrupt' into petri-integration 2026-09-18 22:45:31 -04:00
Bryan Helmkamp
696acc18d6
Merge branch 'petri-followup-interrupt' into petri-integration
# Conflicts:
#	lib/apps/fabro-cli/src/commands/run/petri_stream.rs
2026-09-18 22:45:06 -04:00