Commit graph

205 commits

Author SHA1 Message Date
Scott Werner
04309977e5 Commit checkpoint failures through required finalization
A failed checkpoint cancels the run, so Petri recorded it as cancelled
and the worker overrode its own outcome in memory. Runs that checkpoint
now declare required finalization, and finalize_run rejects with
checkpoint_failed before publishing. The projection and the engine
outcome report a cancelled finish carrying that failure as a workflow
failure with the checkpoint's message, and the in-memory override is
gone. A run whose checkpoint failed is never published.

Retry a host failure that storage rejects with a short backoff, so a
brief storage fault does not leave the run active and holding its
scheduler slot. A failure that never commits still leaves the run's
status alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 11:00:01 -04:00
Scott Werner
3443a923c0 Carry required publication into fork declarations
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 12:45:25 -04:00
Scott Werner
d3d8893bf9 Bind Petri scenario servers to OS-assigned ports 2026-10-08 10:33:38 -04:00
Scott Werner
8b773cdd00
Merge pull request #788 from chunga-ict/fix/schema-mismatch-error-message
Explain schema mismatches as a version skew
2026-10-07 16:33:03 -04:00
Scott Werner
ee86a497d5 Update runner test for schema mismatch diagnostics 2026-10-07 16:21:51 -04:00
Scott Werner
e889f9978d Simplify the store-fault handling
- stop_lock_holder returns the stopped pid as Option<u32> instead of a
  LockHolder enum that only wrapped it
- stop_previous_worker uses with_context, and one run_scratch helper
  replaces three spellings of the run's scratch path
- relaunch builds its runnable record through run_records::runnable,
  moved out of the lifecycle handler so both callers share it
- WorkerExit derives success from its exit code instead of storing both
- the scenario tests share one worker-pid lookup and wait loop
- small readability fixes in the engine's resume arm and the resume test

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 16:07:36 -04:00
Bryan Helmkamp
cb54a0df90 Stop a surviving worker before ending its lease from outside
A worker leads a process group of its own, so it outlives a server
crash. The restarted server released the run's lease from outside and
launched a resume while that worker could still be running: a worker
whose lease is released keeps acting until its next write, beside its
successor. A delete likewise dropped the lease without knowing the
worker was gone.

Now each worker holds a lock on `worker.lock` in its run's scratch
directory for its whole life, taken before anything else. It is a POSIX
record lock: the kernel frees it only when the worker exits, and names
the process that holds it. Before the server ends a lease from outside
(the relaunch after a restart or a store interruption, and a delete), it
kills whatever process still holds the lock, with its process group, and
waits until the lock is free. A worker that finds the lock held does not
start.

- fabro-proc: ProcessLock::try_hold and stop_lock_holder, tested with
  this test binary as the holding process.
- fabro-config: RunScratch::worker_lock_path.
- New scenario: a worker that outlives the server is gone before the
  resume's worker launches, and the run succeeds once. It fails without
  the server-side stop.

A host stage process runs in a process group of its own and still
outlives its killed worker, as it did at a worker crash.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 12:29:34 -04:00
Scott Werner
9e4af50863 Move stage credentials to Petri's spawn env and let the managed token win
Pin Petri at e46845b, the merge of its executor layer change. The stage
credential layer is now a Petri `SpawnEnv` applied with
`EnvHandle::with_spawn_env`, so Petri forwards every other environment
method and applies the layer to one-shot containers as well as processes.

A container gets the managed GITHUB_TOKEN, but no credential store refresh
or Git helper configuration: the store lives in the scope's sandbox, which
the container does not share.

The managed GITHUB_TOKEN now replaces one set in the workflow environment,
an ACP agent's environment or the sandbox's own, as Fabro's stage
environment did before Petri. The token carries exactly the access the run
declares; a stage that needs other access changes its declaration.

The new pin also keeps a timeout as the failure a partial success came
from: `PartialSuccess.underlying` is now an `UnderlyingFailure`, so the
projection reports "the step timed out" for a partial success converted
from a timeout. The run format moves from 7 to 8, which the attach JSON
snapshot records; runs stored before this pin are refused, as with earlier
format changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 19:26:08 -04:00
Scott Werner
4e71554773 Apply CLI model and provider overrides to agent execution 2026-10-02 11:28:16 -04:00
Scott Werner
3c942d8a79 Fix attach test deadlines around workflow completion 2026-10-01 16:33:44 -04:00
Scott Werner
5f22936437 Expect checkpoint commits from a retried local run
Retry now starts the workflow over, and a local-folder run commits its
checkpoints, so the retry scenario checks that every stage commits again
under the retry's run id. The scenario for retrying without Git
checkpoints goes: local runs have them now, and the retry scenario
covers starting over.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:10:13 -04:00
Scott Werner
6e0ef2c402 Keep Git checkpoints for runs on host workspaces
Checkpoints were only enabled for runs with a GitHub source, so local
folder runs, empty Local runs and dry runs stopped committing. `fabro
diff` then failed for them, and their checkpoint, run branch and diff
records disappeared from the event stream.

A run whose workspace is on the host now commits checkpoints there
again, without pushing, as on main. Docker and Daytona runs with no
GitHub source still record execution checkpoints without Git commits,
so a sandbox image without `git` cannot fail the run.

The scenario tests for crash recovery go back to asserting commits. A
workspace deleted while the run is down now fails the resumed run,
since the server keeps no copy to restore it from.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:01:53 -04:00
Scott Werner
aad954693e Keep checkpoints and run branch pushes in the sandbox 2026-10-01 13:01:13 -04:00
Scott Werner
3c62d851fe Retry a finished run from the start
Retry forked the source run at its last checkpoint and reran the failed
stage. It now creates a new run from the source's saved spec and starts
the workflow from the beginning in a fresh workspace, as retry did
before the Petri cutover. The new run records `retried_from` and no
`fork_source_ref`.

Retry no longer needs a checkpoint, a retained workspace, or a published
run branch, so it works for any terminal run that is not archived. To
continue from where a run stopped, fork it at a checkpoint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 11:14:45 -04:00
Scott Werner
0537fbab8f Emit pending questions before JSON attach exits 2026-09-29 13:12:43 -04:00
Scott Werner
d26b2324f0 Merge remote-tracking branch 'origin/main' into codex/restore-configured-artifact-storage
# Conflicts:
#	lib/apps/fabro-cli/tests/it/scenario/artifacts.rs
2026-09-28 16:27:30 -04:00
Scott Werner
ead2ca53b8 Keep local artifacts under the storage directory
Servers have always moved a local artifact store under the storage
directory at startup, whatever `local.root` said. Browser-wizard installs
write `local.root = "<storage>/objects"`, so honoring that root would move
their store and hide every artifact already written, with nothing to
migrate it. Restore the storage-directory override for local roots and
leave honoring custom roots to a change that migrates existing objects.

Installer metadata still goes through the override, so it lands where the
server reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 14:30:36 -04:00
Scott Werner
a1a98c69d0 fix: address review of the in-process sandbox providers
- Keep plugin-era Daytona lease fingerprints: read only DAYTONA_API_URL and
  DAYTONA_ORGANIZATION_ID (no URL alias, no placement target), and stop
  forwarding DAYTONA_SERVER_URL and DAYTONA_TARGET to the worker.
- Take the Docker fingerprint and network from this process's DOCKER_HOST,
  the endpoint the Docker client actually connects to; make the provider
  configuration's fields private.
- Return an error instead of panicking when Petri supplies no Host registry.
- Run deletion reads the Daytona key only for a Daytona run, and a forced
  or restarted delete goes on when the secret store fails, as it does for
  every other prune failure.
- Stop putting DAYTONA_API_KEY in the worker's environment; the worker reads
  it from the vault. Give the worker's Daytona client the shared HTTP client.
- Fork, rewind and retry no longer read the vault: a fork acquires no sandbox.
- Remove the dead worker plugin forwarding and document that runs execute
  only on the built-in providers.
- Build every Petri runtime through providers::standard_runtime or
  bare_runtime, with a Clippy lint against Runtime::standard/bare.
- Share the Docker require-or-skip policy in fabro-test, tighten the Host
  scope assertion.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 12:31:37 -04:00
Scott Werner
71b08b61a1 Preserve custom local artifact roots when overriding storage 2026-09-25 11:56:59 -04:00
Scott Werner
8a09e8fa08 refactor: tidy the in-process sandbox provider wiring
Load the Daytona key for fork and prune through one AppState method
instead of two copied vault reads, and pass the sandbox configuration
into runtime_spec rather than building it and overwriting it. The
worker reuses the CLI's process_env_var lookup.

Share one Docker availability check and the backend-requirement
variable through fabro-test, drop the built-in plugin path and pin
constants nothing reads any more, and let enabled_plugins() exclude the
bundled kinds itself. Refresh the comments and the spawn_env test that
still described built-in plugins.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 11:29:29 -04:00
Scott Werner
6ac6d5495f fix: run built-in sandbox providers in process
Register lazy Host, Docker, and Daytona factories for Petri execution,
fork, and prune. Share server provider configuration, preserve lease
fingerprints, and source Daytona credentials from the vault.

Remove built-in plugin setup and skip gates; add a release-mode worker
and prune regression to catch the failure that blocked nightly builds.

Co-Authored-By: Codex <noreply@openai.com>
2026-09-24 17:14:47 -04:00
Scott Werner
a1926c8632 Restore configured artifact storage for workflow captures 2026-09-24 13:53:46 -04:00
Scott Werner
f75dc8828c docs: shorten the Lithos git dependency convention comments
Reduce the workspace Cargo.toml convention block to three lines: Lithos
libraries track `main`, Cargo.lock picks the commits (move one with
`cargo update -p <crate>`), and unmerged library work is tried with an
uncommitted `[patch]`. Drop the instruction to hand-review lockfile diffs
and trim the restatements in the pebble and petri comments, the
fabro-petri README and module doc, AGENTS.md, and the docker test doc.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 12:09:47 -04:00
Scott Werner
00984ce241 build: track Lithos git dependencies on branch main
Every lithoscomputer git dependency (sandbox-driver, pebble, petri,
lithos-llm, twins) now uses `branch = "main"` instead of an exact rev,
matching the libraries, so the workspace resolves one Cargo source per
repository. Cargo.lock is the single place the commits are chosen; move
one with `cargo update -p <crate>`.

The lockfile keeps every commit except sandbox-driver, which moves from
583a164 to b30203c: Daytona removed its paginated sandbox listing, and
b30203c lists through cursors instead (it also moves the driver's
daytona-sdk-rust dependency to 0e69058). The Daytona auth-probe test
mocks now serve the cursor endpoint the driver calls.

CI reads the sandbox-driver commit for the plugin install from the
lockfile through cargo metadata instead of from Cargo.toml.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 12:40:52 -04:00
Bryan Helmkamp
2b30689e77
Settle a worker's run at Petri's finish, not at the worker's exit
A worker-backed run's in-memory managed run settled only when the
worker process exited. GET /runs/{id} reads the stored summary, which
the projector ends at Petri's own `run.finished` record, a moment
before the worker stores Fabro's terminal lifecycle record and exits.
The delete precheck prefers the managed run, so a delete issued in that
window was refused with 409 "cannot remove active run". Against a real
worker the window hit about six times in ten.

The server sees both records before they are stored: `run.finished` on
the coordinator log through the worker's records endpoint, and the
terminal lifecycle record through the platform-records endpoint. The
managed run now settles at either, ahead of the store, so the view
never reports the run ended while the managed run still says running.
The stream follower no longer reopens a settled run with the records
that precede its terminal one, and the worker's exit keeps the settled
status: it reaps the process, records a missing terminal record as it
did, and takes the store's status only when the store ended the run
differently. The mapping from Petri's finish to the run's status is the
projection's own, shared with its fold.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 03:49:44 -04:00
Bryan Helmkamp
e010563ad5
Read the resume action off the unified restore log line
The Docker recovery scenarios scraped the worker log for "sandbox
workspace brought to its durable snapshot"; since the host and sandbox
restores share one path, the line reads "workspace brought to its
durable snapshot" with the site as a field.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 15:31:39 -04:00
Bryan Helmkamp
634891ded2
Filter the working-directory depth out of the partial include snapshot
The include error names the partial relative to the run's working
directory, so its `../` run is as long as that directory is deep: eight
on this machine's temp dir, three on the CI runner's. Collapse the run
to a token before the snapshot compares.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:39:18 -04:00
Bryan Helmkamp
fdf5141917
Wait for the terminal lifecycle record before reading the cancelled run
The detached cancel test waited for the run's status to read `failed`
and then asserted on the stored `run.lifecycle` record. The projection
concludes the run from Petri's `run.finished` coordinator record, and
the worker stores the platform's terminal lifecycle record a moment
later, so the read raced the write and the assertion failed about once
in thirty runs. Wait for the record itself, and print the stored events
and the run state when the assertion fails.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:31:49 -04:00
Bryan Helmkamp
beac547a1f
Run the workflow scenarios on the Docker provider instead of the stdio plugins
The host_plugin_ and docker_plugin_ variants ran each scenario under
Fabro's old plugin transport with the provider kinds `host` and
`docker-plugin`, which Petri's Fabro frontend rejects. Under Petri every
provider is already served by a sandbox-driver plugin, so those variants
test nothing distinct. A single docker_ variant replaces them: an
environment with provider `docker` on buildpack-deps:noble, created on
an isolated server, skipping without the sandbox-driver-docker
executable or a daemon with the image unless
FABRO_REQUIRE_SANDBOX_PLUGINS is set.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:20:24 -04:00
Bryan Helmkamp
49a647e3a4
Install the sandbox-driver plugins in the Rust test jobs
Every Petri run takes its scope through a sandbox-driver plugin
executable that Petri finds on PATH, so the test jobs need
sandbox-driver-host and sandbox-driver-docker installed at the rev the
workspace pins. The three jobs share one from-source install through an
actions/cache entry keyed on the OS and the rev.

The Linux test job also pre-pulls Petri's default runner image, which
the suite's Docker scenarios leave to Petri: the plugin pulls it on
first use, but a 1 GiB pull inside a run's timeout is a flake.

The stdio plugin job was built for the deleted fabro-sandbox layer. It
becomes the Docker providers job: the `docker_` scenario variants and
the fabro-petri suite, with the fabro-sandbox and fabro-workflow steps
whose tests no longer exist removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 18:20:24 -04:00
Bryan Helmkamp
dfd2458f76
Test an Ask Fabro turn against the container Petri created
A live-server scenario runs a one-stage workflow on Docker whose command
writes a file into the workspace, opens an Ask Fabro session on the
finished run, and sends one turn. The session attaches to the container
Petri created, stopped at the run's end, starts it again, and its tool
reads the file inside it; the turn succeeds, the tool's output and the
model's reply carry the file's content, and the twin's follow-up request
shows the model read it from the tool. Ask Fabro's tool policy is
read-only, so the shell tool is hidden from the model and refused; the
turn reads the file with the `read_file` tool, scripted on the twin,
instead of a shell `cat`.

The scenario skips, and says why, without the Docker plugin or a daemon,
as the other Docker scenarios do.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:58:09 -04:00
Bryan Helmkamp
f054082f86
Delete fabro-sandbox
Nothing imports it any more: the Pebble glue lives in
fabro-pebble-sandbox, the server reaches run sandboxes through
sandbox_access, and Petri creates every run sandbox. The crate, its
test-support, its integration tests and every dependency edge go with
it. The `[server.sandbox.providers.<kind>.plugin]` settings stay: the
server still launches a plugin executable through them to attach to a
sandbox of a non-bundled kind.

AGENTS.md names the new crate and the direct-access pattern in place of
`RunSandbox`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 13:57:30 -04:00
Bryan Helmkamp
467087998d
Read workflow graphs through Petri's DOT parser
Fabro's own DOT parser was left with one job after create-time compile
moved to Petri: walking a workflow's file references for the bundler and
the workflow-version store, and reading a name, a goal and two counts.
Petri's frontend parses the same language, so the parser goes and a small
crate reads the graph through Petri's.

`fabro-dot` is that crate: `WorkflowGraph::parse` over
`petri_frontend_attractor::dot` and its semantic model (defaults applied,
subgraphs flattened, chains expanded), `references(position)` as the one
walker over the static-reference vocabulary (each reference with its node
and position, file references checked to be template-free), and
`normalize_for_graphviz`, the re-emit of Fabro DOT with dotted attribute
keys quoted, which the SVG render needs. It sits beside `fabro-petri`
rather than inside it because `fabro-petri` depends on `fabro-workflow`,
which depends on `fabro-workflow-version`: the version store cannot reach
`fabro-petri` without a cycle, and the bundler should not pull the engine
in to read a graph.

Deleted: `fabro-graphviz`'s lexer, grammar, AST, semantic pass and
`parse_ast` (1,829 lines, plus the `nom` dependency); the DOT model in
`fabro-types::graph` (`Graph`, `Node`, `Edge`, `AttrValue`,
`shape_to_handler_type`), with only `ReferenceKind` kept, moved to
`fabro_types::reference`; `fabro-template`'s `visit_graph_references` and
the `GraphReference`/`GraphPosition` types, with the template-syntax rule
(`validate_static_reference`) kept there; the pull-request body's DOT
fallback summary, which was unreachable because the DOT source only
travels with the run spec whose display graph the summary already reads.
`fabro-graphviz` is now the render alone, over `fabro-dot`.

Parity: the old and new walkers were run over every `.fabro` and `.dot`
file in the repository (118) before the deletion. Every reference set is
identical. Five files differ in what Petri reads more correctly: a
backslash before a newline inside a quoted string is a line continuation
(four files, inline prompt text only), and a node named only by an edge
counts as a node (`test/edge_only_node.fabro`, 3 nodes rather than 2, so
the `fabro validate` snapshot moves). The checked-in bundles' shapes and
references are pinned by a snapshot in `fabro-dot`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 12:23:57 -04:00
Bryan Helmkamp
bd59f52e22
Build the run's display graph from Petri's admission
The rest of the change whose deletions the previous commit carries (its
`git add` stopped at an already-removed path): `fabro_types::RunGraph`
and the `fabro-petri` builder that reads it off the admitted graph, the
server's create, validate, preflight and render paths on Petri's check
alone, the consumers moved to the new shape, the OpenAPI `RunGraph`
schemas with their parity tests, the regenerated TS client, and the
docs naming Petri's diagnostic codes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:41:58 -04:00
Bryan Helmkamp
52aed8c642
Read the run's display graph off Petri's admitted graph
Every run is admitted by Petri, whose check lowers imports, file
references, templates and the model stylesheet, lints the workflow and
pins its models. Fabro then re-parsed the same workflow through its own
legacy pipeline (parse, transforms, structural validation) only to fill
`RunSpec.graph` for the read side. That second pass is gone: the run's
display graph is `fabro_types::RunGraph`, built once in `fabro-petri`
from the admitted graph's metadata (the workflow name and goal from the
graph params; each declared stage's label and handler kind; one edge per
routing arm as written, lowering artifacts left out), and stored on the
spec at create beside the DOT as `graph_source`.

Deleted: `fabro-workflow`'s `pipeline`, `transforms`, `file_resolver`,
`operations::{source, validate}`, `run_materialization` and the legacy
`compile_admitted_run`; the server's `compile_admitted`, the structural
manifest pass, `preflight_model` and the model probe `run_llm_check`
(Petri's admission raises `attractor.model.unknown`); `fabro-graphviz`'s
stylesheet parser; most of `fabro_types::graph` (the DOT model keeps
what the bundler, version registration and template walker read). The
DOT parser stays for the bundler and the SVG render.

`POST /validate`, `POST /preflight`, `POST /graph/render`, `fabro
validate` and `fabro preflight` run on Petri's check alone, so their
diagnostics carry Petri's codes (`attractor.unbound_input`,
`unsupported.template.unbound_input`) where Fabro's
`template_undefined_variable` and `goal_self_reference` were. A refused
workflow's summary still names the DOT as written. The OpenAPI `RunSpec`
schema gains `RunGraph`, `RunGraphNode` and `RunGraphEdge`, reused from
`fabro-types` with parity tests; the TS client is regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:41:38 -04:00
Bryan Helmkamp
eca2812602
Fix what the gates found after the removal sweep
The CLI artifact scenario seeded its run through the deleted upload
route; it now runs a real Petri workflow whose hooks collect the
artifacts, and the fabro artifact list and cp assertions read those.
A real command retry is not producible from a command node (a plain
failure or a timeout routes onward), so the retry dimension of the old
fixture goes; the stage, node, and retry filters, the tree copies, the
cross-stage ambiguity, and the filename collision stay covered. The
archive guard test drops its upload row (the blob write row covers an
octet-stream mutation). A dangling doc comment and two absolute paths
clippy flagged are fixed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 10:15:15 -04:00
Bryan Helmkamp
05a14a6c9b
Delete the checkpoint endpoint, fabro parse, and the fabro-workflow shims
GET /runs/{id}/checkpoint duplicated what /state serves; the hidden
fabro parse command had no user; records, run_status, outcome, and
usage_rollup in fabro-workflow only re-exported fabro_types. The
importers now name fabro_types directly. format_cost keeps its two
callers (the pull request body and the CLI stage display) and moves to
fabro_types::usage; the usage rollup tests move beside the function in
fabro-types, with test_usage in its test support.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:31:09 -04:00
Bryan Helmkamp
978a5b1b7e
Merge branch 'petri-integration' into petri-followup-envcat
# Conflicts:
#	Cargo.lock
#	Cargo.toml
2026-09-19 07:10:02 -04:00
Bryan Helmkamp
b1d95faa57
Hand Petri the server's environment and MCP catalogs
Petri's Fabro frontend refused a bundle naming an environment it did not
declare and every MCP catalog reference, so the fixtures declared
`[environments.local]` and the server's catalogs never reached Petri.

Pin Petri at c874b86, where the frontend reads `[environments.<id>]` and
`[run.environment]` from every settings layer (bundle over project over the
host's layer, key by key), takes the environment a launch selected over the
layers, and resolves `[run.agent.mcps.<name>] id = "..."` against a catalog
the host binds. The server hands Petri its environment catalog as
`[environments.<id>]` tables of the settings layer it already passes, the
intent's environment as the launch's selection (`Launch::environment`, as
the intent overrides the bundle in Fabro's own resolution), and its MCP
catalog as `RuntimeSpec::mcp_catalog_toml`, one inline entry per definition
keyed by id. Offline validation hands Petri the seeded catalog the same
way, so `fabro validate` accepts `[run.environment] id = "local"`.

The fixtures drop the `[environments.local]` tables they carried for this;
the secrets test keeps its own, on purpose. Scenario tests cover a bundle
naming a catalog environment (its image lowered, and run on Docker when the
plugin and a daemon are there), a bundle's own table winning key by key,
the server refusing an unknown environment before Petri, and a catalog MCP
reference whose tool the agent session lists (an echo server under
`test/mcp/`).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:08:08 -04:00
Bryan Helmkamp
776c50307f
Merge remote-tracking branch 'origin/main' into petri-integration
# Conflicts:
#	Cargo.lock
#	Cargo.toml
#	lib/components/fabro-validate/src/lib.rs
#	lib/components/fabro-workflow/src/handler/llm/fallback.rs
#	lib/components/fabro-workflow/src/handler/prompt.rs
#	lib/components/fabro-workflow/src/operations/start.rs
#	lib/components/fabro-workflow/src/pipeline/pull_request.rs
#	lib/components/fabro-workflow/src/transforms/model_resolution.rs
#	lib/components/fabro-workflow/tests/it/integration.rs
#	lib/components/fabro-workflow/tests/it/pebble_agent.rs
2026-09-19 06:37:02 -04:00
Bryan Helmkamp
cb26c5603d
Answer steer and interrupt with the worker's acknowledgement
The worker control bus was publish-only: the steer and interrupt
endpoints answered 202 once the control was forwarded, and a refusal
showed up only later as a `run.notice` on the run's stream.

A steer or an interrupt now carries a request id. The worker answers it
over the control stream it arrived on with `{request_id, outcome}`,
where the outcome is `delivered` (with the stage's label) or `refused`
(with the code and the reason). The server keeps the outstanding
requests in a registry and waits up to 5 s for the answer: the endpoint
answers 202 `{"outcome":"delivered","stage":…}`, 409 with the refusal's
code (`no_live_turn`, `no_such_stage`, `steer_refused`,
`interrupt_refused`) and message, or 202 `{"outcome":"pending"}` when
the worker gave no answer in time. The `run.notice` record on refusal
stays, under the same code, so a steer to a stage that is not running is
now `no_such_stage` there too. Pause and unpause are unchanged.

The in-process test path answers a steer or an interrupt from the run's
own controls at once. `FABRO_TEST_CONTROL_ACKS_MUTED=1` on the server
mutes the worker's answers, so a test can see the pending fallback.
`fabro steer` prints the worker's answer, and a refusal is its error.
The OpenAPI spec documents the 202 body and the 409 codes; the Rust and
TypeScript clients are regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 23:54:11 -04:00
Bryan Helmkamp
86ac13713b
Merge branch 'petri-followup-fork' into petri-integration 2026-09-18 23:18:14 -04:00
Bryan Helmkamp
22ebbf73e4
Cover fork, retry, rewind and the timeline with real-binary scenarios
Through a real server and its worker: a three-stage run forked at its
first stage continues with the other two on the first stage's restored
file and commits on the new run branch after the source's commits; a
retry of a run whose last stage failed transiently reruns that stage and
succeeds; a rewind archives and supersedes its source, and an archived
run is refused; the timeline lists every checkpoint record with its
commit; a fork at a checkpoint inside a parallel branch is refused and
creates nothing, while one after the join continues.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 23:17:25 -04:00
Bryan Helmkamp
8d4a1bf894
Add the fork, rewind, retry and timeline commands to the CLI
`fabro timeline <run>` prints the checkpoint table (or JSON); `fabro fork
<run> [target]` and `fabro rewind <run> <target>` make and start the new
run and name it, `--list` showing the timeline instead; `fabro retry
<run>` starts the retry and prints its id. The top-level help snapshot
and the generated CLI reference follow.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 23:17:25 -04:00
Bryan Helmkamp
a664b0ccce
Merge branch 'petri-followup-interrupt' into petri-integration 2026-09-18 22:45:31 -04:00
Bryan Helmkamp
696acc18d6
Merge branch 'petri-followup-interrupt' into petri-integration
# Conflicts:
#	lib/apps/fabro-cli/src/commands/run/petri_stream.rs
2026-09-18 22:45:06 -04:00
Bryan Helmkamp
ea85359a12
Name the stage and Petri's reason in a refused interrupt's notice
The worker's `run.notice` for an interrupt it could not deliver now reads
"Interrupt of stage `gate` refused: the stage has no model turn to
interrupt" (or "Interrupt refused: …" when the control named no stage),
so the web and the CLI can show which stage and why, with the reason as
Petri's `ControlError` spells it. The gate scenario asserts both messages.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 22:40:51 -04:00
Bryan Helmkamp
1978257aa5
Interrupt a live agent stage's model turn over Petri
Petri 639ce3e added `ControlService::interrupt_firing` and the `LiveTurns`
capability a host installs beside the pause hooks. Fabro now drives it:
`RunControls::interrupt(stage, text)` resolves its stage the way a steer
does (a label, a node name, or the run's one live agent stage) and stops
that stage's current model turn, keeping the session; the text, when
given, is the stage's next input. `engine::run` installs the live-turn set
as a runtime capability, so without it no interrupt could ever land.

The worker maps `run.interrupt` and `run.interrupt_then_steer`, both of
which now carry an optional `stage`, to that call. The control bus is
one-way, so a refusal is recorded the way a refused steer is: a
`run.notice` on the run's stream whose code says why (`no_live_turn` when
Petri refuses a stage with no turn in flight, `no_such_stage`,
`interrupt_refused`).

The server's `POST /runs/{id}/interrupt` and `POST /runs/{id}/steer` with
`interrupt=true` forward the control and answer 202, replacing the 501
`interrupt_unsupported` stub. The interrupt endpoint takes an optional
body (`stage`, `text`), refuses a finished run with 409
`run_not_interruptible`, and forwards an interrupt of a blocked run, since
an agent stage may be running a turn beside the question and the worker
judges each stage itself. `fabro events --pretty` prints the delivered
interrupt and the stage's `attractor.turn.interrupted` report.

Verified on the twin: `fabro steer --interrupt` during a long tool call
ends the turn, the text is the agent's next request, and the stream
carries the `$interrupt` record and the interrupted-turn report; an
interrupt of a gate stage is refused with `no_live_turn` and the gate's
question is untouched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 22:38:17 -04:00
Bryan Helmkamp
bfe94038fd
Record the two dry-run snapshots the earlier gate timed out on
`dry_run_parallel` and `dry_run_styled` timed out under load in the run
that found the other eight, so they were not recorded with them. Alone they
show the same one-line change: the Start stage's completion now precedes the
run branch line, the order the positioned records give every run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 22:30:57 -04:00
Bryan Helmkamp
08b7a4fdd9
Record the snapshots the Petri pin and the record positions changed
The pin to Petri 639ce3e moves the run's format version from 6 to 7, which
the attach JSON snapshot records. The other eight snapshots had recorded the
run branch and Git identity lines before the Start stage's completion: the
order the clock gave them before "Place the run branch and git identity
records with their checkpoint" positioned the two records after the firing's
finish. That commit refreshed only two files, and these eight already
differed the same way at the commit before the pin bump; they now record the
one order every run produces.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 22:17:15 -04:00