Commit graph

383 commits

Author SHA1 Message Date
Bryan Helmkamp
81b7427add
Merge pull request #888 from fabro-sh/settle-worker-run-on-terminal-record
Settle a worker's run at Petri's finish, not at the worker's exit
2026-09-21 08:08:31 -04:00
Bryan Helmkamp
2b30689e77
Settle a worker's run at Petri's finish, not at the worker's exit
A worker-backed run's in-memory managed run settled only when the
worker process exited. GET /runs/{id} reads the stored summary, which
the projector ends at Petri's own `run.finished` record, a moment
before the worker stores Fabro's terminal lifecycle record and exits.
The delete precheck prefers the managed run, so a delete issued in that
window was refused with 409 "cannot remove active run". Against a real
worker the window hit about six times in ten.

The server sees both records before they are stored: `run.finished` on
the coordinator log through the worker's records endpoint, and the
terminal lifecycle record through the platform-records endpoint. The
managed run now settles at either, ahead of the store, so the view
never reports the run ended while the managed run still says running.
The stream follower no longer reopens a settled run with the records
that precede its terminal one, and the worker's exit keeps the settled
status: it reaps the process, records a missing terminal record as it
did, and takes the store's status only when the store ended the run
differently. The mapping from Petri's finish to the run's status is the
projection's own, shared with its fold.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 03:49:44 -04:00
Bryan Helmkamp
1fda9633e4
Bind the run's resolved goal into Petri's check
An intent's goal override, and a `[run.goal]` layer, reached the run's
display graph and its settings but not Petri's check, so the agent stages
executed with the workflow's own goal while the run showed the override.
The launch now carries the run's resolved goal (the settings' inline
`run.goal`, layered as the create path layers it) as Petri's
`petri.launch_goal` compile variable, which Petri binds over the bundle's
`[run] goal` and the graph's own `goal`, so admission's frozen plan carries
the goal the run shows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 03:48:35 -04:00
Bryan Helmkamp
9a93bbcfbd
Wait for the managed run to settle before the prune scenario deletes it
The stored view reports a Petri run ended as soon as its own run.finished
record is folded, which is before the server stores the terminal
lifecycle record and settles the managed run in its map. The delete
precheck reads that map, so a delete sent as soon as the API reports the
run ended can be refused as active instead of by the held lease. The
scenario now waits for the managed run to settle, through a test-support
accessor for its status, before it deletes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 16:03:20 -04:00
Bryan Helmkamp
1f6e53c949
Share one poison-tolerant lock helper from fabro-util
Seven crates' files each carried the same four-line lock function that
recovers a poisoned mutex. fabro_util::sync::lock is that function, once;
the copies in fabro-petri and fabro-server are gone. fabro-template's
helper panics on poison instead, a different policy, and is left as is.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 14:47:46 -04:00
Bryan Helmkamp
c264a3567c
Settle the managed run as soon as its terminal record is stored
The in-process Petri path persisted the run's terminal lifecycle
record, settled the projector, aggregated usage, and only then settled
the managed run. GET /runs/{id} reads the stored summary, so it reported
the run as ended while the delete precheck, which prefers the managed
run, still saw it running and refused the delete as active. The prune
scenario hit that window about once in thirty runs. Settle the managed
run right after the record is stored, before the view catches up.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:56:48 -04:00
Bryan Helmkamp
399aef8111
Keep the Daytona key rendering out of the test's assertion messages
CodeQL read the assertion messages as a log of the credentials'
Debug output. The test proves that output never holds the key, so the
messages added nothing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:05:01 -04:00
Bryan Helmkamp
6082f82950
Fix the gate findings in the prune change
Clippy's absolute-paths lint on the rendered prune error, the sync
directory reads the prune test documents, and a redundant rustdoc link.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 15:02:20 -04:00
Bryan Helmkamp
c7aa50c943
Delete a run's sandboxes through Petri's lease ledger
Run deletion called the driver's `provider.delete(id)` under the run's
`petri.run` scope, a delete of Fabro's own over a sandbox whose lease
record Petri owns. It now goes the way `petri sandbox prune` goes:
`fabro_petri::prune` builds the run's Petri runtime over the server's
store (the run key, the run directory, the sandbox backend) and calls
Petri's prune, which opens the run for writing, checks each lease's
provider fingerprint, writes the delete intent and the tombstone beside
the run's other records, and lets each provider remove its managed
workspace, a host workspace included.

A run a live process holds answers 409 unless the delete is forced; a
lease Petri could not prune answers 409 with the problem text, or is
warned and skipped under force or a delete that already started. The
server drops the worker's handles before the prune, on the store
instance the prune opens, so the lease a stopped worker held is released
first. The projection reads only the coordinator and execution logs, so
the resource records change nothing it reports.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:54:03 -04:00
Bryan Helmkamp
809b3891b5
Forward every configured sandbox plugin to the Petri worker
The worker's environment carried only the host, Docker and Daytona plugin
variables from the server's own environment, so a run on a third-party
provider kind never learned where its plugin was, although the server
kept `[server.sandbox.providers.<kind>]` plugin settings for its own
attach. The launch spec now derives `PETRI_SANDBOX_<KIND>_PLUGIN` and
`PETRI_SANDBOX_<KIND>_SHA256` from every enabled kind's plugin settings,
and `PETRI_SANDBOX_PLUGIN_DEV=1` when any of them sets `dev`, set after
the allowlist so the settings win over an ambient variable of the same
name and the allowlist stays the fallback.

The server's default plugin binary name is `sandbox-driver-<kind>`, the
name Petri looks up, since one executable serves both sides.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:45:42 -04:00
Bryan Helmkamp
956feda0ca
Fix the gate findings in the sandbox tests
The grep test reads paths as the driver reports them for a resolved
absolute path; the local preflight check test gives the manifest a
source directory that exists; the Docker attach scenario accepts that
the in-process app has no daemon record for the run-tools client an Ask
Fabro turn builds after the sandbox attach, and asserts the turn got past
the sandbox. The inventory's lazy connection is boxed for clippy's
variant-size lint.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:07:52 -04:00
Bryan Helmkamp
2a4f2aa718
Test the server's attach to the container Petri created
A server scenario runs a command workflow on Docker through Petri's
plugin (skipped without the plugin or a daemon), then reaches the
container without Petri: the sandbox tab describes it under its
`petri.run` label, Run Files writes, lists and reads a file in its
workspace after starting the stopped container, a preview URL opens to a
port in it, and an Ask Fabro turn runs against it through the OpenAI
twin. A unit test attaches through the ownership seam with a scripted
provider: the run's own container attaches, another run's and one that
carries only Fabro's retired `sh.fabro.*` labels are refused as not
owned.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:00:56 -04:00
Bryan Helmkamp
23b8a6c449
Serve the sandbox tab, Run Files, and deletion on the driver handle
The sandbox handlers describe, list, download, upload, open a terminal
and build SSH and VNC access on the `Arc<dyn Sandbox>` the server attaches
to the run's record, with paths resolved against the recorded working
directory. Run Files holds the handle beside that directory and runs its
git through the driver's git facet and Fabro's exec policy;
fabro-workflow's sandbox git takes the same pair, and its
`GitCommandError` carries the driver's error. Run deletion deletes by id
through the provider scoped to the run's `petri.run` label, so a foreign
sandbox is refused and a designated host directory is left in place. Ask
Fabro wraps the attached, running handle. The legacy access shim is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 13:56:13 -04:00
Bryan Helmkamp
0cc645b2d1
Reach run sandboxes from the server through the sandbox driver
`fabro-server/src/sandbox_access.rs` is the server's own path to a run's
sandbox: it connects the record's provider (the driver's Host, Docker and
Daytona providers in process, a plugin executable for any other kind),
keys ownership on the `petri.run` label Petri stamps on every sandbox it
creates, attaches by the recorded id, and for a host record designates
the recorded directory again when the id lives only in the worker's
registry. The Docker client resolves its endpoint from the same variables
Petri forwards to its plugin, so both meet on one daemon.

The doctor's Docker check and the Daytona credential probe move here with
`DaytonaCredentials`, and the `/sandboxes` inventory is rebuilt over the
driver's `list`, narrowed to sandboxes that carry Petri's run label.
Preflight asks the provider for its health instead of creating and
deleting a throwaway sandbox in Fabro's own shape, which no run uses; the
git retry policy behind the repository probe moves into run_manifest.

The callers still on fabro-sandbox's reconnect read their access through
a `legacy_provider_access` shim until they move.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 13:47:02 -04:00
Bryan Helmkamp
c4ed995b44
Add fabro-pebble-sandbox: a driver handle as pebble's Environment
Petri creates and owns every run sandbox through the sandbox driver, so
what Fabro still needs around a driver handle is the Pebble glue: the
Environment pebble's coding agent runs its tools through, the exec policy
(stop grace, working directory, StripAll, the termination mapping, the
redacted output tail), pebble's port routes over the driver's preview
URLs, the secret redactor, the path helpers, and a log rendering that
appends a failed command's redacted tail. This crate holds that glue,
moved from fabro-sandbox, over `Arc<dyn Sandbox>` plus a working
directory instead of `RunSandbox`, with a `MockSandbox` double behind
`test-support`.

`fabro exec` creates its host sandbox directly on the driver's Host
provider and activates it; Ask Fabro wraps the attached handle. Both
keep the provider alive beside the sandbox where the session's processes
are the provider's process groups.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 13:37:38 -04:00
Bryan Helmkamp
467087998d
Read workflow graphs through Petri's DOT parser
Fabro's own DOT parser was left with one job after create-time compile
moved to Petri: walking a workflow's file references for the bundler and
the workflow-version store, and reading a name, a goal and two counts.
Petri's frontend parses the same language, so the parser goes and a small
crate reads the graph through Petri's.

`fabro-dot` is that crate: `WorkflowGraph::parse` over
`petri_frontend_attractor::dot` and its semantic model (defaults applied,
subgraphs flattened, chains expanded), `references(position)` as the one
walker over the static-reference vocabulary (each reference with its node
and position, file references checked to be template-free), and
`normalize_for_graphviz`, the re-emit of Fabro DOT with dotted attribute
keys quoted, which the SVG render needs. It sits beside `fabro-petri`
rather than inside it because `fabro-petri` depends on `fabro-workflow`,
which depends on `fabro-workflow-version`: the version store cannot reach
`fabro-petri` without a cycle, and the bundler should not pull the engine
in to read a graph.

Deleted: `fabro-graphviz`'s lexer, grammar, AST, semantic pass and
`parse_ast` (1,829 lines, plus the `nom` dependency); the DOT model in
`fabro-types::graph` (`Graph`, `Node`, `Edge`, `AttrValue`,
`shape_to_handler_type`), with only `ReferenceKind` kept, moved to
`fabro_types::reference`; `fabro-template`'s `visit_graph_references` and
the `GraphReference`/`GraphPosition` types, with the template-syntax rule
(`validate_static_reference`) kept there; the pull-request body's DOT
fallback summary, which was unreachable because the DOT source only
travels with the run spec whose display graph the summary already reads.
`fabro-graphviz` is now the render alone, over `fabro-dot`.

Parity: the old and new walkers were run over every `.fabro` and `.dot`
file in the repository (118) before the deletion. Every reference set is
identical. Five files differ in what Petri reads more correctly: a
backslash before a newline inside a quoted string is a line continuation
(four files, inline prompt text only), and a node named only by an edge
counts as a node (`test/edge_only_node.fabro`, 3 nodes rather than 2, so
the `fabro validate` snapshot moves). The checked-in bundles' shapes and
references are pinned by a snapshot in `fabro-dot`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 12:23:57 -04:00
Bryan Helmkamp
bd59f52e22
Build the run's display graph from Petri's admission
The rest of the change whose deletions the previous commit carries (its
`git add` stopped at an already-removed path): `fabro_types::RunGraph`
and the `fabro-petri` builder that reads it off the admitted graph, the
server's create, validate, preflight and render paths on Petri's check
alone, the consumers moved to the new shape, the OpenAPI `RunGraph`
schemas with their parity tests, the regenerated TS client, and the
docs naming Petri's diagnostic codes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:41:58 -04:00
Bryan Helmkamp
eca2812602
Fix what the gates found after the removal sweep
The CLI artifact scenario seeded its run through the deleted upload
route; it now runs a real Petri workflow whose hooks collect the
artifacts, and the fabro artifact list and cp assertions read those.
A real command retry is not producible from a command node (a plain
failure or a timeout routes onward), so the retry dimension of the old
fixture goes; the stage, node, and retry filters, the tree copies, the
cross-stage ambiguity, and the filename collision stay covered. The
archive guard test drops its upload row (the blob write row covers an
octet-stream mutation). A dangling doc comment and two absolute paths
clippy flagged are fixed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 10:15:15 -04:00
Bryan Helmkamp
05a14a6c9b
Delete the checkpoint endpoint, fabro parse, and the fabro-workflow shims
GET /runs/{id}/checkpoint duplicated what /state serves; the hidden
fabro parse command had no user; records, run_status, outcome, and
usage_rollup in fabro-workflow only re-exported fabro_types. The
importers now name fabro_types directly. format_cost keeps its two
callers (the pull request body and the CLI stage display) and moves to
fabro_types::usage; the usage rollup tests move beside the function in
fabro-types, with test_usage in its test support.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:31:09 -04:00
Bryan Helmkamp
0752d4c7c4
Delete the stage artifact upload endpoint
Artifacts reach the blob table through the hooks, so the POST on
/runs/{id}/stages/{stageId}/artifacts, its octet-stream and multipart
handlers, the RequireStageArtifact extractor, the client's upload
functions, and the generated TypeScript operation go. The spec loses
the operation, the multipart variant writeRunBlob never served, and
the batch manifest schemas; fabro-types loses ArtifactUpload, the
batch upload's only input type. Every list and download path stays,
and the server tests seed the artifact store directly to cover them.
fabro-server drops multer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:23:47 -04:00
Bryan Helmkamp
06f9cb8361
Delete fabro-sandbox's clone and push chain
The engine prepares every run's checkout, so fabro's clone
orchestration, the per-checkout GitHub credentials, the run-branch
setup, the push retries, and the push policies had no production
caller. RepoWorkspace::plan still validates the clone request and now
refuses one that asks for a clone; initialize creates an empty
workspace root. SandboxWorkspaceLayout and snapshot_info stay: the run
record projection in sandbox_spec.rs reads them. The run tool
regression keeps its assertion (a child targets the parent's pushed
run branch) over a plain git fixture instead of the deleted setup. The
Docker, Daytona, and Daytona-wire clone layout tests go: they proved
only the legacy clone. fabro-sandbox drops base64, uuid, fabro-proc,
serde, strum, and sandbox-driver-daytona-config; chrono is test-only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:18:29 -04:00
Bryan Helmkamp
01713c6aac
Hand Petri only the environment keys it reads
The settings layer carried every key of every catalog environment, so a
run in an environment with `lifecycle`, `labels`, `cwd`, `network` or a
Dockerfile warned `ignored.workflow_toml.environments.<id>.<key>` on
every admit. Those keys are the platform's and stay with the server's own
resolution; the layer now carries the provider, `image.docker` under
`docker` and `daytona`, `resources` under `daytona`, and `env`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:27:57 -04:00
Bryan Helmkamp
978a5b1b7e
Merge branch 'petri-integration' into petri-followup-envcat
# Conflicts:
#	Cargo.lock
#	Cargo.toml
2026-09-19 07:10:02 -04:00
Bryan Helmkamp
b1d95faa57
Hand Petri the server's environment and MCP catalogs
Petri's Fabro frontend refused a bundle naming an environment it did not
declare and every MCP catalog reference, so the fixtures declared
`[environments.local]` and the server's catalogs never reached Petri.

Pin Petri at c874b86, where the frontend reads `[environments.<id>]` and
`[run.environment]` from every settings layer (bundle over project over the
host's layer, key by key), takes the environment a launch selected over the
layers, and resolves `[run.agent.mcps.<name>] id = "..."` against a catalog
the host binds. The server hands Petri its environment catalog as
`[environments.<id>]` tables of the settings layer it already passes, the
intent's environment as the launch's selection (`Launch::environment`, as
the intent overrides the bundle in Fabro's own resolution), and its MCP
catalog as `RuntimeSpec::mcp_catalog_toml`, one inline entry per definition
keyed by id. Offline validation hands Petri the seeded catalog the same
way, so `fabro validate` accepts `[run.environment] id = "local"`.

The fixtures drop the `[environments.local]` tables they carried for this;
the secrets test keeps its own, on purpose. Scenario tests cover a bundle
naming a catalog environment (its image lowered, and run on Docker when the
plugin and a daemon are there), a bundle's own table winning key by key,
the server refusing an unknown environment before Petri, and a catalog MCP
reference whose tool the agent session lists (an echo server under
`test/mcp/`).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:08:08 -04:00
Bryan Helmkamp
776c50307f
Merge remote-tracking branch 'origin/main' into petri-integration
# Conflicts:
#	Cargo.lock
#	Cargo.toml
#	lib/components/fabro-validate/src/lib.rs
#	lib/components/fabro-workflow/src/handler/llm/fallback.rs
#	lib/components/fabro-workflow/src/handler/prompt.rs
#	lib/components/fabro-workflow/src/operations/start.rs
#	lib/components/fabro-workflow/src/pipeline/pull_request.rs
#	lib/components/fabro-workflow/src/transforms/model_resolution.rs
#	lib/components/fabro-workflow/tests/it/integration.rs
#	lib/components/fabro-workflow/tests/it/pebble_agent.rs
2026-09-19 06:37:02 -04:00
Bryan Helmkamp
cb26c5603d
Answer steer and interrupt with the worker's acknowledgement
The worker control bus was publish-only: the steer and interrupt
endpoints answered 202 once the control was forwarded, and a refusal
showed up only later as a `run.notice` on the run's stream.

A steer or an interrupt now carries a request id. The worker answers it
over the control stream it arrived on with `{request_id, outcome}`,
where the outcome is `delivered` (with the stage's label) or `refused`
(with the code and the reason). The server keeps the outstanding
requests in a registry and waits up to 5 s for the answer: the endpoint
answers 202 `{"outcome":"delivered","stage":…}`, 409 with the refusal's
code (`no_live_turn`, `no_such_stage`, `steer_refused`,
`interrupt_refused`) and message, or 202 `{"outcome":"pending"}` when
the worker gave no answer in time. The `run.notice` record on refusal
stays, under the same code, so a steer to a stage that is not running is
now `no_such_stage` there too. Pause and unpause are unchanged.

The in-process test path answers a steer or an interrupt from the run's
own controls at once. `FABRO_TEST_CONTROL_ACKS_MUTED=1` on the server
mutes the worker's answers, so a test can see the pending fallback.
`fabro steer` prints the worker's answer, and a refusal is its error.
The OpenAPI spec documents the 202 body and the 409 codes; the Rust and
TypeScript clients are regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 23:54:11 -04:00
Bryan Helmkamp
86ac13713b
Merge branch 'petri-followup-fork' into petri-integration 2026-09-18 23:18:14 -04:00
Bryan Helmkamp
e7d55d6ec7
Serve fork, rewind, retry and the timeline on the runs API
`GET /runs/{id}/timeline` lists the run's checkpoints with their Petri
positions, stages, commits and diff summaries, and its fork origin.
`POST /runs/{id}/fork` resolves a target on that timeline, creates the
new run, seeds it through `fabro_petri::fork` and queues it in resume
mode, so its worker restores the checkpoint into a fresh workspace and
continues from the position; a position inside a parallel branch is
refused with 400 before the run exists. `POST /runs/{id}/rewind` is that
fork of a terminal run followed by the source's archive and its
`run.superseded` record (207 when the archive fails); `POST
/runs/{id}/retry` forks a terminal run at its last checkpoint, rerunning
the stage that failed. The projection carries `forked_from`. The Rust
and TypeScript clients gain the four calls.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 23:17:25 -04:00
Bryan Helmkamp
1978257aa5
Interrupt a live agent stage's model turn over Petri
Petri 639ce3e added `ControlService::interrupt_firing` and the `LiveTurns`
capability a host installs beside the pause hooks. Fabro now drives it:
`RunControls::interrupt(stage, text)` resolves its stage the way a steer
does (a label, a node name, or the run's one live agent stage) and stops
that stage's current model turn, keeping the session; the text, when
given, is the stage's next input. `engine::run` installs the live-turn set
as a runtime capability, so without it no interrupt could ever land.

The worker maps `run.interrupt` and `run.interrupt_then_steer`, both of
which now carry an optional `stage`, to that call. The control bus is
one-way, so a refusal is recorded the way a refused steer is: a
`run.notice` on the run's stream whose code says why (`no_live_turn` when
Petri refuses a stage with no turn in flight, `no_such_stage`,
`interrupt_refused`).

The server's `POST /runs/{id}/interrupt` and `POST /runs/{id}/steer` with
`interrupt=true` forward the control and answer 202, replacing the 501
`interrupt_unsupported` stub. The interrupt endpoint takes an optional
body (`stage`, `text`), refuses a finished run with 409
`run_not_interruptible`, and forwards an interrupt of a blocked run, since
an agent stage may be running a turn beside the question and the worker
judges each stage itself. `fabro events --pretty` prints the delivered
interrupt and the stage's `attractor.turn.interrupted` report.

Verified on the twin: `fabro steer --interrupt` during a long tool call
ends the turn, the text is the agent's next request, and the stream
carries the `$interrupt` record and the interrupted-turn report; an
interrupt of a gate stage is refused with `no_live_turn` and the gate's
question is untouched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 22:38:17 -04:00
Bryan Helmkamp
3e065b806b
Bump lithos-llm to 43a42ac and migrate catalogs to the codecs schema
Move the lithos-llm pin from 55add459 to 43a42ac28e9d9bcf40a91abc02be4f12ca274ebb,
and the three Pebble pins from a39f43e to 67c9f48, Pebble `main`, which pins
that same lithos-llm revision so Cargo holds one lithos-llm crate. lithos-llm
`main` (ca19fac) is one commit further; that commit touches only its nightly
workflow, so this pin stays on the revision Pebble unifies with.

The `openai`, `anthropic`, `gemini`, and `openai-compatible` features are
gone upstream; each expanded to `runtime`, which `bedrock` implies, so the
four names leave the fabro-llm feature list. Every other manifest already
names `runtime`.

The catalog schema now names one adapter and many codecs per provider.
`adapter` defaults to `http`, `codecs = [...]` replaces `codec` and defaults
to `["openai-chat"]`, and the loader rejects the old `codec` key and the four
protocol-named adapter ids. Every inline catalog in tests and docs moves to
the new shape: the `openai-compatible` + `openai-chat` pair is dropped as the
default, `adapter = "openai"` + `codec = "openai-responses"` becomes
`codecs = ["openai-responses"]`, and the one test that swaps in a custom
adapter id now adds the line instead of replacing one. The settings
reference, the API schema's `Provider.adapter` description, and the SDK page
describe the new fields; the three `docs/superpowers/plans/` files that show
the old shape are dated, unchecked historical plans and are left as they are.

The implied agent profile for an operator provider that declares none used
to read the removed protocol adapter ids; it now reads the provider's first
codec (Anthropic Messages and Gemini map to their harnesses, the `bedrock`
adapter to Anthropic, everything else to OpenAI), with a test for the codec
path.

Absorbing the rest of the range: OpenRouter and Fireworks now ship enabled,
so the two fabro-llm tests that used OpenRouter as the disabled fixture use
`bedrock-openai`, and the docs and comments that said the two ship disabled
are corrected. The built-in catalog grew past 100 enabled model rows
(Vercel, TypeSafe, and the enabled OpenRouter and Fireworks rosters), so the
pagination shape test walks `page[offset]` to the last page instead of
assuming one page fits.

`cargo update -p` on the four crates also re-resolved a few already-locked
edges to match the lithos-llm lockfile: `windows-sys` 0.61.2/0.60.2 ->
0.59.0 under dirs-sys, errno, nu-ansi-term, quinn-udp, rustix,
rustls-platform-verifier, tempfile, terminal_size, and winapi-util;
`windows-core` 0.61.2 -> 0.62.2 under iana-time-zone; `errno` 0.2.8 ->
0.3.14 under signal-hook-registry; and `indexmap` 2.13.0 as a new public
dependency of lithos-llm. No package version was added or removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 22:07:00 -04:00
Bryan Helmkamp
a70750f3f6
Steer a Petri run by stage
`SteerRunRequest` takes an optional `stage`: the label the projection
shows (`node@visit`, or `node/e<execution>@visit` when two executions
share one) or the node's name. The server passes it on the worker control
message; the worker's `RunControls` resolves a label to the live agent
firing and steers that firing, and a node name through Petri's own
live-stage index. Unnamed, the one-live-agent rule stays, and the refusal
now names the live stages by their labels. `fabro steer --stage` sets it.
A controls scenario runs two agent stages side by side, sees the unnamed
steer refused with both named, and steers each apart, one over the API
and one through the flag.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 19:59:23 -04:00
Bryan Helmkamp
6aec33c4f5
Settle clippy and formatting after the sandbox merge
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 18:39:58 -04:00
Bryan Helmkamp
baa2fc8b5c
Keep every workspace under Petri's retention and say why
Petri's retention decides whether a released workspace is kept or removed.
Fabro's lifecycle settings decide whether a sandbox keeps running after
the run and whether a delete may remove it; none asks for removal at the
run's end, and the sandbox tab, `fabro cp`, the run's delete and the
sandbox scenarios read the container after the run. So the mapping is
`Retention::Always` for every setting, named once as `engine::RETENTION`
with the reasoning, instead of a per-setting function that released a
finished sandbox.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 18:39:58 -04:00
Bryan Helmkamp
56f7180dbf
Merge branch 'petri-integration' into petri-integration-gaps
# Conflicts:
#	lib/components/fabro-petri/README.md
#	lib/components/fabro-petri/src/projection.rs
2026-09-18 16:00:51 -04:00
Bryan Helmkamp
76461da780
Remove the [server.slatedb] settings and the SlateDB prefix probe
No store sits behind `[server.slatedb]` any more: the section leaves the
settings layer, the resolved server settings, the defaults, the API schema,
the TypeScript client, the install wizard and the docs, and `fabro install`
probes the bucket for the `artifacts/` prefix alone. A settings file that
still carries the section is rewritten once at startup by a temporary
migration that removes it with a backup beside the file.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 16:00:24 -04:00
Bryan Helmkamp
a5d4e96bbf
Serve collected artifacts from the projection and the blob table
The run and stage artifact listings, the download and the archive join the
artifacts the projection records, whose bytes are in the blob table, with
the ones uploaded to the artifact store, each once.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 16:00:11 -04:00
Bryan Helmkamp
67ba595b01
Record the run branch, identity, artifacts and diffs from the hooks
The commit that creates a workspace's run branch records `run.branch`
(the base commit, or the first checkpoint in a workspace with no history)
and `git.identity`. Every checkpoint record after the first carries the
stage's diff from its parent commit, with the patch as a text blob. After
the checkpoint record, the transition hook lists the stage's workspace
through the scope's environment, on the host and in a sandbox alike, and
collects every file under `[run.artifacts] include` into the blob table as
an `artifact.collected` record, skipping a file already collected under
the same path and digest. At the run's end the hooks diff the branch's
last checkpoint against its base in the snapshot repository and record
`run.diff`.

`engine::retention` maps the environment's lifecycle settings onto
Petri's workspace retention instead of always keeping every workspace:
`preserve`, `stop_on_terminal = false` and the local provider keep them,
anything else keeps a failed scope's only. The hooks docs no longer name a
redundant link target, so rustdoc passes with warnings denied.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 16:00:11 -04:00
Bryan Helmkamp
f0b23fe426
Project a Petri run's sandbox instance from its scope records
Petri a5906f6 records where each scope's sandbox ran (`scope.acquired`,
`scope.failed`) and how its lease was released (`scope.released`). The
projection folds the root invocation's records into `Run.sandbox`:
`initializing` from `run.started`, `ready` with the `RunSandboxInstance`
(the provider, Petri's `host` as Fabro's `local`, the provider's id, the
image and snapshot, the working directory) from `scope.acquired`, `failed`
from `scope.failed`; the retention outcome is kept in the fold state, since
the view has no field for it. Ask Fabro reconnect and `sandbox cp`,
`preview` and `ssh` reach the run's sandbox again.

A local reconnect designates the recorded working directory again when the
host provider does not know the id: the provider mints a registry-only id
for a workspace path too long for a path-derived one, and that registry
belongs to the run's worker. The stream listing redacts its items the way
the attached stream does, so a client that pages after a stream sees the
same items.

Pins move to Petri a5906f6 (run format 6, engine log v11, event contract
4). The attach stream snapshot is re-recorded with the new record and a
filter for the host provider's minted ids; `sandbox cp` reads an upload
back through the run's workspace, which is no longer the target folder.
Server scenario tests prove the projected instance on the host and Docker
providers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 15:57:12 -04:00
Bryan Helmkamp
60b503322c
Delete the legacy run event log, its reducer and its types
Step 4 of the legacy executor deletion, fourth commit: with no writer
and no reader left, the legacy event log goes.

- `fabro-types`: `run_event` (`EventBody`, `RunEvent` and every props
  struct), `EventEnvelope` and the `RunEventDetail*` types are deleted.
  What the projection and the API still use moves out of the event
  vocabulary: `AgentEventProps`, `AgentSessionActivatedProps`,
  `AgentToolsAvailableProps`, `StagePromptProps`, `SessionCapability`
  and the coding event names to `agent_props`; `RunNoticeLevel` and
  `RunNoticeCode` to `notice`; `InterviewOption` beside the question
  types; `RunRunnableSource` beside the run status. `Checkpoint` is
  what Fabro records for a Petri run: `timestamp`, `current_node`,
  `git_commit_sha`; the conclusion's stage summaries derive from the
  projection's stages instead of the checkpoint's node maps.
- `fabro-store`: the Slate bridge (`RunDatabase`, the Slate `Database`,
  `keys`, `record`, `EventPayload`) and the reducer (`run_state`) are
  deleted. `Database` is the blob table and the run summary store over
  one pool; the blob store is SQLite only; the run summary store keeps
  the `runs` row a projector writes and lists, and finds the pull
  request creation candidates over `platform_records`; `build_summary`
  and `projected_usage` live in `run_summary`. The SlateDB dependency
  is gone. Test fixtures build the store from its two SQLite stores.
- `fabro-workflow`: the `event` module (the `Event` enum, its
  conversion, sink, emitter, redaction, stored fields and names),
  `runtime_store`, `StageScope` and the legacy seeding test helpers are
  deleted; the tests that seeded legacy runs read platform records or
  a projection instead.
- `fabro-sandbox` owns `GitRetryReason`.
- The server builds the store without an object store; the legacy
  `POST /runs/{id}/events` tests go, an interrupt answers
  `interrupt_unsupported` in the tests as it does in the handler, and
  the tests that read a run back through the Slate handle read its
  projection or its platform records. The projection folds a block
  that lands while the run is paused as the pause's prior block, and a
  pause or unpause clears the pending control it answers; a control
  request's check-and-append holds a per-run lock so two concurrent
  cancels record one request.
- The CLI's final output is the response of the last stage that
  produced one; the workflow tests read completed nodes from the
  succeeded stages.
- The spec's `RunCheckpoint` carries the three fields the type keeps.

Still failing until the next commits: the CLI tests that seed runs
through `POST /runs/{id}/events` or wait for legacy event names, and
the two Ask Fabro resume tests (the sandbox instance gap).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 14:09:17 -04:00
Bryan Helmkamp
bc31b8d23a
Serve the run stream as the only run event API
Step 4 of the legacy executor deletion, third commit: the legacy event
API and every reader of it go, so that the next commits can delete the
event log, its reducer and the types beneath them.

The API:
- `GET /runs/{id}/events` pages the run stream only
  (`PaginatedRunStreamList` by `after`); the legacy `since_seq`,
  `before_seq` and `order` cursors, the `oneOf` envelope, the legacy
  `EventEnvelope`, `PaginatedEventList`, `RunEvent`, `EventSeq`,
  `AppendEventResponse` and `RunEventDetailResponse` schemas,
  `POST /runs/{id}/events`, `GET /runs/{id}/events/{seq}` and
  `GET /runs/{id}/stages/{stageId}/events` are deleted. `GET
  /runs/{id}/attach` and `GET /attach` frame `RunStreamItem`s only.
- The Rust and TypeScript clients regenerate; the removed models leave
  the TypeScript package.

The readers:
- `fabro-client` drops the legacy run event listing, tail and attach
  methods and `RunEventStream`; `list_run_stream_until` bounds a stream
  read.
- `fabro-tool`'s `fabro_run_events` lists, searches and details the run
  stream: `after` is the exclusive `stream_seq` cursor, `event_id` the
  item's id, filters match the item's name and `recorded_at`.
- `fabro-dump` writes the stream to `events.jsonl`; `fabro dump` reads
  it.
- The CLI's progress renderer keeps only what the run stream drives:
  the legacy event conversion, the sandbox and setup displays and their
  styles go. `fabro system events` prints stream items.
- The server's demo mode folds its agent fixture straight into the
  session projection and answers the attach stub with a stream item;
  the demo stage events endpoint is gone.
- The web app: every run is a Petri run. The legacy event hooks,
  renderer props, stage popover summary, run phases derivation and
  live-event payload handling are deleted or ported to `RunStreamItem`;
  toasts and board refreshes read the stream's platform records.
- Tests: the legacy API round trips and pagination tests are deleted;
  the CLI's MCP, attach and system event mocks serve stream pages; the
  CLI test helpers read stream items.

Still failing until the later commits: the CLI tests seeded through
`POST /runs/{id}/events`, the server tests over the legacy store, and
the legacy type tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 13:36:00 -04:00
Bryan Helmkamp
34996d630f
Give Ask Fabro sessions their own event log
Step 4 of the legacy executor deletion, second commit. Ask Fabro's
sessions were the last writer of `run_events`: a session's creation, its
turns and their messages, tool calls and endings went into the run's
legacy event log, keyed by the run's sequence. They now have a log of
their own.

- `run_session_events` (migration `2026091802`): one row per session
  event, numbered per session from 1, with the owning run, the turn, the
  event name and its properties. `RunSessionEventStore` appends under the
  write lock, lists a session from a sequence, names a session's owner
  from its creation event, deletes a run's sessions with the run, and
  publishes each committed event to its subscribers.
- `fabro_types::SessionEvent`: `seq`, `session_id`, `run_id`, `ts` and a
  flattened body (`event` naming the kind, `properties` its fields), with
  the same event names and property shapes the legacy events carried,
  so the web app and the CLI read the same JSON. The property structs
  move to `session_event`; `run_event::session` re-exports them under
  their old names until the legacy event log goes.
- The API: `GET /sessions/{id}/events` pages `PaginatedSessionEventList`
  by the session's own sequence, `GET /sessions/{id}/attach` replays and
  streams `SessionEvent` frames (subscribed before the replay, so no
  event falls between the two), the turn stream carries the same frames,
  and an interrupt answers with the recorded event. The session
  projection folds `SessionEvent`s; the legacy `find_session_owner` over
  `run_events` is gone.
- The CLI's `run ask` and the web app's session stream read
  `SessionEvent`; the web runtime no longer accepts the nested legacy
  envelope shape.

The two session resume tests in the server keep failing for a reason
this commit does not touch: Ask Fabro reconnects to the run's sandbox
from the projection's sandbox instance, which the Petri projection does
not carry yet (`VIEWS.md`, the `scope.acquired` gap).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 12:47:43 -04:00
Bryan Helmkamp
3e8b2ebfcc
Run the server's lifecycle over platform records instead of run_events
Step 4 of the legacy executor deletion, first commit of several: step 4
spans commits because the legacy event log and its consumers cannot go
in one compiling change. This commit moves every writer off `run_events`;
the reducer, `EventBody`, the Slate bridge and the API's event types
still exist for the readers the next commits port or delete.

Writers:
- The server records a run's lifecycle (submitted, runnable, starting,
  running, blocked, paused, control requests and effects, the terminal
  status), its title, parent link, archive state, notices and pull
  request state as platform records (`fabro_store::platform_records`),
  through the new `server::run_records` module. Every append wakes the
  projector and waits for its pass, so the read that follows a write
  holds the record.
- Pull request creation is recorded as `pull_request.requested`,
  `pull_request.created`, `pull_request.failed`, `pull_request.linked`
  and `pull_request.unlinked`; the projection folds them into the run's
  pull request and creation state.
- Answers to questions are recorded as `interview.answered` with the
  answering principal and the answer text; the interview adapter no
  longer posts legacy `interview.*` events (`QuestionSink` is now an
  optional observer).
- The worker (`fabro run __run-worker`) records its lifecycle, notices
  and pause state over `HttpPlatformRecords`; `HttpRunStore` for the
  legacy event log and the worker's `run_store` are gone.
- `persist_created_run` appends `run.created` and `run.submitted`.

Readers:
- A stream follower (`server::stream_follower`) follows each live run's
  stream (Petri events and platform records), folds lifecycle records
  into the in-memory run state, forwards items to the global attach
  broadcast, and syncs blocked and paused from the projection.
- Slack posts questions from the projection's pending interviews,
  finishes them on `interview.answered` or `question_expired`, and sends
  lifecycle notifications with `notification.sent` dedupe.
- `GET /runs/{id}/events` and the attach endpoints serve only the run
  stream; the per-event, per-stage and `POST /runs/{id}/events`
  endpoints and their tests are deleted.
- `Database::load_run_projection` reads the Petri projection only.

Deleted with the writers:
- The SQLite blob and run-history activation migrations and their
  legacy Slate imports (`legacy_blob_import`, `legacy_run_history_import`,
  the activation backup): a greenfield server has no Slate history to
  import, and the run-history verification refused to start a server
  whose runs have no legacy events.
- `fabro-workflow`'s `operations::archive` and `operations::run_store`.
- The server's legacy-event unit tests and the CLI's `HttpRunStore` tests.

The in-process answer transport is now set after the starting and
running records land, not gated on the live status still being
`Starting` (the records already moved it).

The manifest validation test for a `run.agent.mcps.<name>` catalog
reference now expects `unsupported.workflow_toml.run.agent.mcps.reference`:
Petri's Fabro frontend has no server catalog to resolve it against.

Legacy readers still fail their tests until the next commits: the
reducer and Slate tests in fabro-store, the fabro-workflow create tests
that read the run back through the legacy store, the CLI tests seeded
through `POST /runs/{id}/events`, the CLI's legacy attach and render
paths, the sessions API, the OpenAPI conformance test, and the web
fixtures.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 12:32:06 -04:00
Bryan Helmkamp
af38d68942
Delete fabro-validate and fabro-acp; validate on Petri's check
Petri judges a workflow at admission, so Fabro's lint rules go:
`fabro-validate` (its 35 rules and the `LintRule` trait) is deleted, and
with it `fabro-acp` (only a rule and two legacy executor tests used it),
the model-resolution transform, the legacy `create`, `compile_create_run`
and `materialize_create_run` stages, and `fabro-graphviz`'s `condition`
and `fidelity` modules. `Diagnostic`, `RelatedDiagnostic` and `Severity`
move to `fabro_types::diagnostic`, the one shape every diagnostic takes.

Validation is now the same question the create handler asks. A new
server module, `petri_check`, builds Petri's check request from a
workflow bundle and the run's settings (every workflow of the bundle at
its bundle-relative path, the inputs, the run variables, the launch),
runs the check, and maps the diagnostics; Fabro adds one rule of its
own, `fabro.model.no_ready_provider`, refusing a model node when no
provider is ready. Admission, the validate and preflight endpoints and
the offline `fabro validate` all go through it:

- `validate_prepared_manifest` runs Fabro's structural pass (parse and
  transform, whose diagnostics stay) and then Petri's check, on the
  blocking pool from the handlers;
- the offline `validate_manifest` checks with no model client and with
  an unbound input as a warning (`CheckRequest.unbound_is_warning`), so
  a workflow validates before its inputs exist; a collected workflow
  before upload checks with unbound inputs as errors, as before;
- preflight resolves each LLM node's selector against the ready
  providers and the catalog for its probe, as the deleted transform did,
  and no longer probes a model Petri refused;
- the graph render endpoint needs only the structural pass;
- a run manifest now carries its `[run.goal] file`, which Petri reads
  from the bundle as it does for a version.

The transforms keep the authored model selector (`sonnet` stays
`sonnet`): Petri pins the catalog model in the admitted graph, not in
the graph Fabro displays or in the settings snapshot. Tests assert that,
and the CLI's validate, preflight and graph snapshots carry Petri's
diagnostics (`attractor.no_start`, `attractor.undeclared_node`,
`attractor.bad_on_failure`, ...) in place of the lint rules' text.

Known gaps, Petri's side: a `workflow.toml` whose `[run.environment]`
names an environment the server catalog defines but the file does not
is refused (`unsupported.workflow_toml.run.environment`), as admission
already refused it; an unbound input inside an included template
partial is a render error rather than the unbound-input warning.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 11:32:04 -04:00
Bryan Helmkamp
e8e681adbf
Remove the retry, rewind, fork and timeline endpoints and commands
The legacy executor replayed a run from a checkpoint; Petri resumes a
run from its records instead, and the checkpoint timeline, rewind, fork
and retry were the operations that replay carried. The server dropped
their handlers with the executor; this removes the rest:

- the API spec's `/runs/{id}/retry`, `/rewind`, `/fork` and `/timeline`
  paths with the `ForkRequest`, `ForkResponse`, `RewindRequest`,
  `RewindResponse` and `TimelineEntryResponse` schemas, and the
  generated TypeScript models;
- `fabro rewind` and `fabro fork` (with the checkpoint timeline printer
  and the repo-origin check only they used), their reference pages and
  the checkpoints guide's rewind and fork sections;
- `fabro-client`'s `rewind_run`, `fork_run` and `run_timeline`;
- the web app's Retry action on the run list and the run page.

`Run.retried_from` stays on the run type: a run that was retried before
the cutover would still name its source.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 10:54:45 -04:00
Bryan Helmkamp
3e157d3356
Format the Petri test fixtures after the engine key removal
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 10:54:45 -04:00
Bryan Helmkamp
d90a5d9cbb
Delete fabro-core and the engine half of fabro-workflow
Every run executes on Petri, so the in-process legacy executor goes:
`fabro-core` and, in `fabro-workflow`, the handlers, lifecycle, pipeline
execution, routing, retry, conditions, node handlers, steering, agent
memory, artifacts, checkpoints, command log, and the `start`, `resume`,
`retry`, `fork`, `rewind` and `timeline` operations. The two are deleted
together because the engine half of `fabro-workflow` was the only user of
`fabro-core` and `fabro-core` the only runtime of that half; neither
compiles without the other.

Kept in `fabro-workflow`, narrowed: the parse/transform/validate/persist
pipeline and `create`, `archive`, `validate` (workflow definitions still
come from DOT and settings); the run tools (`run_tools`, moved from
`handler/llm/fabro_tools.rs`) for Ask Fabro, `fabro exec` and Petri's
host tools; the pull request pipeline (`pull_request`, moved from
`pipeline/`, for the step 0 port); Run Files' diff helpers in
`sandbox_git`; `git_identity`, `usage_rollup`, `run_status`,
`run_materialization`, `web_search` and `workflow_bundle`.

Server: `RegistryFactoryOverride` becomes `execute_in_process`;
`RunAnswerTransport::InProcess` carries only the interviewer; the
interrupt endpoint answers 501 `interrupt_unsupported` and every pair
endpoint 501 `pair_unsupported` (status lists none); rewind, fork, retry
and timeline handlers and routes are removed; the command log is served
from the stage output blob; usage rollups accumulate from the settled
projection after an in-process run as after a worker exit.

Ported while here:

- `materialize_admitted_run` materializes the goal and drops a disabled
  pull request block, as the legacy materializer did.
- A run whose admitted graph has an agent or prompt node is refused at
  create when no LLM provider is ready (`fabro.model.no_ready_provider`);
  a workflow of commands and gates needs no model and is admitted.
- The projection's question type falls back on the options, as the
  interview adapter does, so a gate with edge-label options answers as
  multiple choice.

Tests: the server scenarios (lifecycle, run completion, SSE, helpers)
run in process on Petri and assert Petri's stage labels and stream
names; the reconcile tests assert Petri's relaunch semantics; legacy
unit tests of the deleted executor are removed; three server unit tests
the removal took with it are restored; the pair fixtures go with the
pair feature. Petri test fixtures no longer name `[workflow] engine`.

Still red after this commit, all legacy consumers the next steps
delete or port: fabro-store's Slate/reducer fixtures and fabro-types
legacy JSON tests (step 4); server unit tests over legacy run events
(retry endpoints, list_run_events, artifacts, per-event pause/unpause,
run history activation, legacy sandbox fixtures) (steps 3-4); CLI tests
that parse legacy event envelopes, the legacy `events`/`attach`/`diff`/
`dump`/`inspect` snapshots, `run rewind`/`run fork`, the ACP and
git-identity workflow tests, and the runner tests that drive the legacy
worker by hand (steps 3-4); the web app's Petri fixtures still carry
`engine` (regenerate with `FABRO_CAPTURE_PETRI_FIXTURES` in step 4).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 10:44:41 -04:00
Bryan Helmkamp
a36bea15d2
Remove the engine flag: every run is a Petri run
Delete `Engine`, `RunEngine`, `[workflow] engine`, `[server.execution]
engine`, `FABRO_SERVER_ENGINE` and `fabro server start --engine`. The run
spec records what Petri admitted as `admission: PetriAdmission`; the
create handler always admits through `Runtime::check`; `execute_run`
always launches the Petri worker (or executes in process under the test
override); the CLI runner takes only the Petri worker path, and its
legacy control arm, artifact uploader, signal pause handlers and
credential helpers go with it. The CLI's `attach` and `events` read the
run stream only.

Two gaps this surfaced are closed here: the check adapter binds the
server's run variables as Petri compile variables (`{{ vars.* }}` in a
prompt no longer fails admission), and deleting a run removes its Petri
records, lease, platform records, projection and stream.

Tests: the config engine tests are replaced (an engine key is unknown),
the API round-trip test covers `PetriAdmission`, the server and CLI
Petri scenarios drop their engine settings, and the API tests that read
legacy event names now read the run stream or the session events. The
remaining red tests are fixtures and scenarios of the legacy executor
and the legacy event store (`fabro-store` `slate` and `run_state`,
`fabro-types` legacy `run.created` JSON, the server's handler-registry
scenarios, the CLI dry-run snapshots), which the next steps of the F4.3
series delete or port.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 09:58:21 -04:00
Bryan Helmkamp
b5fdc015f0
Merge branch 'petri-integration-sandbox' into petri-integration
# Conflicts:
#	lib/apps/fabro-cli/tests/it/scenario/mod.rs
2026-09-18 09:11:19 -04:00
Bryan Helmkamp
97f74c3ab0
Forward the Docker daemon selection and Daytona credentials to Petri workers
A Petri run's worker launches the sandbox-driver plugins itself, and the
Docker plugin forwards `DOCKER_HOST`, `DOCKER_TLS_VERIFY`,
`DOCKER_CERT_PATH`, `DOCKER_API_VERSION`, `DOCKER_CONFIG` and
`DOCKER_CONTEXT` from the process that launches it. They now cross the
worker's environment allowlist, so the worker's sandboxes go to the daemon
the server uses. The concern that kept them out, the legacy worker's own
Docker client picking up a daemon it was not meant to, is moot: the
legacy executor is being deleted. The same variables pass through the
test harness's isolation, so a developer's daemon selection reaches the
servers tests start.

Daytona's non-secret selectors, `DAYTONA_API_URL` and
`DAYTONA_ORGANIZATION_ID`, cross the allowlist too. The API key stays the
vault's: a Daytona run's worker command carries it the way the GitHub app
key travels, and the Daytona plugin reads it from the worker's process.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 09:10:29 -04:00
Bryan Helmkamp
5093265efb
Wire pause, unpause and steer into the Petri worker
A Petri run answered only cancel and answers; pause, unpause and steer
were ignored with a warning. `fabro_petri::controls::RunControls` now
wraps Petri's `ControlService` per run: `engine::run` installs its pause
gate over the run's hooks, observes the run through it and wires it to
the coordinator, on a start and a resume alike, so a run paused when its
worker died resumes paused.

The worker's control channel takes a `WorkerControls` enum: the legacy
hub and pause flag, or the Petri run's controls. On Petri, `run.pause`
holds admission, `run.unpause` releases it once the record is durable,
and `run.steer` goes to the one live agent stage (Fabro's steer names no
stage); with none or several it is refused with a `run.notice` record.
The paused state is mirrored to Fabro's lifecycle as `run.paused` and
`run.unpaused` events, so the server's live status and the projection
follow Petri's own records.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 07:49:54 -04:00