Commit graph

323 commits

Author SHA1 Message Date
Bryan Helmkamp
467087998d
Read workflow graphs through Petri's DOT parser
Fabro's own DOT parser was left with one job after create-time compile
moved to Petri: walking a workflow's file references for the bundler and
the workflow-version store, and reading a name, a goal and two counts.
Petri's frontend parses the same language, so the parser goes and a small
crate reads the graph through Petri's.

`fabro-dot` is that crate: `WorkflowGraph::parse` over
`petri_frontend_attractor::dot` and its semantic model (defaults applied,
subgraphs flattened, chains expanded), `references(position)` as the one
walker over the static-reference vocabulary (each reference with its node
and position, file references checked to be template-free), and
`normalize_for_graphviz`, the re-emit of Fabro DOT with dotted attribute
keys quoted, which the SVG render needs. It sits beside `fabro-petri`
rather than inside it because `fabro-petri` depends on `fabro-workflow`,
which depends on `fabro-workflow-version`: the version store cannot reach
`fabro-petri` without a cycle, and the bundler should not pull the engine
in to read a graph.

Deleted: `fabro-graphviz`'s lexer, grammar, AST, semantic pass and
`parse_ast` (1,829 lines, plus the `nom` dependency); the DOT model in
`fabro-types::graph` (`Graph`, `Node`, `Edge`, `AttrValue`,
`shape_to_handler_type`), with only `ReferenceKind` kept, moved to
`fabro_types::reference`; `fabro-template`'s `visit_graph_references` and
the `GraphReference`/`GraphPosition` types, with the template-syntax rule
(`validate_static_reference`) kept there; the pull-request body's DOT
fallback summary, which was unreachable because the DOT source only
travels with the run spec whose display graph the summary already reads.
`fabro-graphviz` is now the render alone, over `fabro-dot`.

Parity: the old and new walkers were run over every `.fabro` and `.dot`
file in the repository (118) before the deletion. Every reference set is
identical. Five files differ in what Petri reads more correctly: a
backslash before a newline inside a quoted string is a line continuation
(four files, inline prompt text only), and a node named only by an edge
counts as a node (`test/edge_only_node.fabro`, 3 nodes rather than 2, so
the `fabro validate` snapshot moves). The checked-in bundles' shapes and
references are pinned by a snapshot in `fabro-dot`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 12:23:57 -04:00
Bryan Helmkamp
bd59f52e22
Build the run's display graph from Petri's admission
The rest of the change whose deletions the previous commit carries (its
`git add` stopped at an already-removed path): `fabro_types::RunGraph`
and the `fabro-petri` builder that reads it off the admitted graph, the
server's create, validate, preflight and render paths on Petri's check
alone, the consumers moved to the new shape, the OpenAPI `RunGraph`
schemas with their parity tests, the regenerated TS client, and the
docs naming Petri's diagnostic codes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:41:58 -04:00
Bryan Helmkamp
52aed8c642
Read the run's display graph off Petri's admitted graph
Every run is admitted by Petri, whose check lowers imports, file
references, templates and the model stylesheet, lints the workflow and
pins its models. Fabro then re-parsed the same workflow through its own
legacy pipeline (parse, transforms, structural validation) only to fill
`RunSpec.graph` for the read side. That second pass is gone: the run's
display graph is `fabro_types::RunGraph`, built once in `fabro-petri`
from the admitted graph's metadata (the workflow name and goal from the
graph params; each declared stage's label and handler kind; one edge per
routing arm as written, lowering artifacts left out), and stored on the
spec at create beside the DOT as `graph_source`.

Deleted: `fabro-workflow`'s `pipeline`, `transforms`, `file_resolver`,
`operations::{source, validate}`, `run_materialization` and the legacy
`compile_admitted_run`; the server's `compile_admitted`, the structural
manifest pass, `preflight_model` and the model probe `run_llm_check`
(Petri's admission raises `attractor.model.unknown`); `fabro-graphviz`'s
stylesheet parser; most of `fabro_types::graph` (the DOT model keeps
what the bundler, version registration and template walker read). The
DOT parser stays for the bundler and the SVG render.

`POST /validate`, `POST /preflight`, `POST /graph/render`, `fabro
validate` and `fabro preflight` run on Petri's check alone, so their
diagnostics carry Petri's codes (`attractor.unbound_input`,
`unsupported.template.unbound_input`) where Fabro's
`template_undefined_variable` and `goal_self_reference` were. A refused
workflow's summary still names the DOT as written. The OpenAPI `RunSpec`
schema gains `RunGraph`, `RunGraphNode` and `RunGraphEdge`, reused from
`fabro-types` with parity tests; the TS client is regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:41:38 -04:00
Bryan Helmkamp
51138cda57
Drop the dependency edges with no production use
Each edge was checked with rg over the crate's sources outside its
test paths. fabro-cli keeps git2, regex, ulid, and shlex as
dev-dependencies for its integration tests. fabro-automation keeps
chrono, tokio, and tracing: its migrations compile into the crate
through #[path]. petri_testkit was already optional behind fabro-petri's
test-support feature and dual-listed as a dev-dependency. The workspace
loses the agent-client-protocol and AWS entries no crate references;
jsonschema stays for the fabro-api and fabro-tool tests. Cargo.lock
was refreshed by a plain build and only loses entries.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:33:31 -04:00
Bryan Helmkamp
05a14a6c9b
Delete the checkpoint endpoint, fabro parse, and the fabro-workflow shims
GET /runs/{id}/checkpoint duplicated what /state serves; the hidden
fabro parse command had no user; records, run_status, outcome, and
usage_rollup in fabro-workflow only re-exported fabro_types. The
importers now name fabro_types directly. format_cost keeps its two
callers (the pull request body and the CLI stage display) and moves to
fabro_types::usage; the usage rollup tests move beside the function in
fabro-types, with test_usage in its test support.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:31:09 -04:00
Bryan Helmkamp
efb43b45aa
Delete the executor-era error and Git helpers in fabro-workflow
The failure classifiers, the handler and publish error builders, the
FailureDetail projections, and the LLM error conversions served the
deleted executor; Petri classifies failures now. Error keeps the
variants the create and read side construct, and the Engine variant
replaces the three-stage Stage shape. git_identity goes: the hooks
record git.identity through fabro_checkpoint. git.rs keeps the
observe, head, non-interactive push, and sync helpers the server and
fabro-manifest call, and loses the push half. fabro-llm loses the
failure signature hint whose only reader was the deleted classifier.
fabro-workflow drops regex, strum, and fabro-checkpoint.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:10:34 -04:00
Bryan Helmkamp
c0fb71a467
Delete the interviewers the Petri cutover left unused
ConsoleInterviewer, RecordingInterviewer, ReplayInterviewer,
QueueInterviewer, CallbackInterviewer, and ask_with_timeout had no
production caller once every run executes on Petri. review_target_line
moves to lib.rs for the CLI's attach prompt. fabro-interview drops
dialoguer and fabro-util.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 09:05:52 -04:00
Bryan Helmkamp
776c50307f
Merge remote-tracking branch 'origin/main' into petri-integration
# Conflicts:
#	Cargo.lock
#	Cargo.toml
#	lib/components/fabro-validate/src/lib.rs
#	lib/components/fabro-workflow/src/handler/llm/fallback.rs
#	lib/components/fabro-workflow/src/handler/prompt.rs
#	lib/components/fabro-workflow/src/operations/start.rs
#	lib/components/fabro-workflow/src/pipeline/pull_request.rs
#	lib/components/fabro-workflow/src/transforms/model_resolution.rs
#	lib/components/fabro-workflow/tests/it/integration.rs
#	lib/components/fabro-workflow/tests/it/pebble_agent.rs
2026-09-19 06:37:02 -04:00
Bryan Helmkamp
f18f02206b
Restore the timeline, fork, rewind and retry operations over checkpoints
The operations the cutover removed come back in their Petri shape, with
no engine of their own: the timeline is the run's checkpoint records
labelled by stage (`RunTimeline::build`), and a target (`@ordinal`, a
node, or `node@visit`) resolves to one of them. A fork's run row is the
source's spec under a new id with `fork_source_ref` naming the source and
the checkpoint's commit (`persist_forked_run`), and `retried_from` when
it is a retry. A rewind needs a terminal source and records
`run.superseded` on it; a retry needs a terminal source and reruns the
last checkpointed stage when the run failed on it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 23:17:25 -04:00
Bryan Helmkamp
3e065b806b
Bump lithos-llm to 43a42ac and migrate catalogs to the codecs schema
Move the lithos-llm pin from 55add459 to 43a42ac28e9d9bcf40a91abc02be4f12ca274ebb,
and the three Pebble pins from a39f43e to 67c9f48, Pebble `main`, which pins
that same lithos-llm revision so Cargo holds one lithos-llm crate. lithos-llm
`main` (ca19fac) is one commit further; that commit touches only its nightly
workflow, so this pin stays on the revision Pebble unifies with.

The `openai`, `anthropic`, `gemini`, and `openai-compatible` features are
gone upstream; each expanded to `runtime`, which `bedrock` implies, so the
four names leave the fabro-llm feature list. Every other manifest already
names `runtime`.

The catalog schema now names one adapter and many codecs per provider.
`adapter` defaults to `http`, `codecs = [...]` replaces `codec` and defaults
to `["openai-chat"]`, and the loader rejects the old `codec` key and the four
protocol-named adapter ids. Every inline catalog in tests and docs moves to
the new shape: the `openai-compatible` + `openai-chat` pair is dropped as the
default, `adapter = "openai"` + `codec = "openai-responses"` becomes
`codecs = ["openai-responses"]`, and the one test that swaps in a custom
adapter id now adds the line instead of replacing one. The settings
reference, the API schema's `Provider.adapter` description, and the SDK page
describe the new fields; the three `docs/superpowers/plans/` files that show
the old shape are dated, unchecked historical plans and are left as they are.

The implied agent profile for an operator provider that declares none used
to read the removed protocol adapter ids; it now reads the provider's first
codec (Anthropic Messages and Gemini map to their harnesses, the `bedrock`
adapter to Anthropic, everything else to OpenAI), with a test for the codec
path.

Absorbing the rest of the range: OpenRouter and Fireworks now ship enabled,
so the two fabro-llm tests that used OpenRouter as the disabled fixture use
`bedrock-openai`, and the docs and comments that said the two ship disabled
are corrected. The built-in catalog grew past 100 enabled model rows
(Vercel, TypeSafe, and the enabled OpenRouter and Fireworks rosters), so the
pagination shape test walks `page[offset]` to the last page instead of
assuming one page fits.

`cargo update -p` on the four crates also re-resolved a few already-locked
edges to match the lithos-llm lockfile: `windows-sys` 0.61.2/0.60.2 ->
0.59.0 under dirs-sys, errno, nu-ansi-term, quinn-udp, rustix,
rustls-platform-verifier, tempfile, terminal_size, and winapi-util;
`windows-core` 0.61.2 -> 0.62.2 under iana-time-zone; `errno` 0.2.8 ->
0.3.14 under signal-hook-registry; and `indexmap` 2.13.0 as a new public
dependency of lithos-llm. No package version was added or removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 22:07:00 -04:00
Bryan Helmkamp
0fe066d420
Fix the rustdoc link warnings and gate rustdoc in CI
Every intra-doc link `cargo doc --workspace --no-deps` warned on now
resolves or is plain code: the private constant and helper, the removed
`InterpString::resolve`, the lithos `Message`, the sandbox-driver facets,
the `RunOptions::git_author` the cutover removed, and the stale
`platform_record_for` paragraph on the platform records. The `[@REF]`
segment of `fabro run`'s help text is allowed as help, not a link. A
`Rustdoc` job runs the same command with `-D warnings`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 19:24:34 -04:00
Bryan Helmkamp
60b503322c
Delete the legacy run event log, its reducer and its types
Step 4 of the legacy executor deletion, fourth commit: with no writer
and no reader left, the legacy event log goes.

- `fabro-types`: `run_event` (`EventBody`, `RunEvent` and every props
  struct), `EventEnvelope` and the `RunEventDetail*` types are deleted.
  What the projection and the API still use moves out of the event
  vocabulary: `AgentEventProps`, `AgentSessionActivatedProps`,
  `AgentToolsAvailableProps`, `StagePromptProps`, `SessionCapability`
  and the coding event names to `agent_props`; `RunNoticeLevel` and
  `RunNoticeCode` to `notice`; `InterviewOption` beside the question
  types; `RunRunnableSource` beside the run status. `Checkpoint` is
  what Fabro records for a Petri run: `timestamp`, `current_node`,
  `git_commit_sha`; the conclusion's stage summaries derive from the
  projection's stages instead of the checkpoint's node maps.
- `fabro-store`: the Slate bridge (`RunDatabase`, the Slate `Database`,
  `keys`, `record`, `EventPayload`) and the reducer (`run_state`) are
  deleted. `Database` is the blob table and the run summary store over
  one pool; the blob store is SQLite only; the run summary store keeps
  the `runs` row a projector writes and lists, and finds the pull
  request creation candidates over `platform_records`; `build_summary`
  and `projected_usage` live in `run_summary`. The SlateDB dependency
  is gone. Test fixtures build the store from its two SQLite stores.
- `fabro-workflow`: the `event` module (the `Event` enum, its
  conversion, sink, emitter, redaction, stored fields and names),
  `runtime_store`, `StageScope` and the legacy seeding test helpers are
  deleted; the tests that seeded legacy runs read platform records or
  a projection instead.
- `fabro-sandbox` owns `GitRetryReason`.
- The server builds the store without an object store; the legacy
  `POST /runs/{id}/events` tests go, an interrupt answers
  `interrupt_unsupported` in the tests as it does in the handler, and
  the tests that read a run back through the Slate handle read its
  projection or its platform records. The projection folds a block
  that lands while the run is paused as the pause's prior block, and a
  pause or unpause clears the pending control it answers; a control
  request's check-and-append holds a per-run lock so two concurrent
  cancels record one request.
- The CLI's final output is the response of the last stage that
  produced one; the workflow tests read completed nodes from the
  succeeded stages.
- The spec's `RunCheckpoint` carries the three fields the type keeps.

Still failing until the next commits: the CLI tests that seed runs
through `POST /runs/{id}/events` or wait for legacy event names, and
the two Ask Fabro resume tests (the sandbox instance gap).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 14:09:17 -04:00
Bryan Helmkamp
3e8b2ebfcc
Run the server's lifecycle over platform records instead of run_events
Step 4 of the legacy executor deletion, first commit of several: step 4
spans commits because the legacy event log and its consumers cannot go
in one compiling change. This commit moves every writer off `run_events`;
the reducer, `EventBody`, the Slate bridge and the API's event types
still exist for the readers the next commits port or delete.

Writers:
- The server records a run's lifecycle (submitted, runnable, starting,
  running, blocked, paused, control requests and effects, the terminal
  status), its title, parent link, archive state, notices and pull
  request state as platform records (`fabro_store::platform_records`),
  through the new `server::run_records` module. Every append wakes the
  projector and waits for its pass, so the read that follows a write
  holds the record.
- Pull request creation is recorded as `pull_request.requested`,
  `pull_request.created`, `pull_request.failed`, `pull_request.linked`
  and `pull_request.unlinked`; the projection folds them into the run's
  pull request and creation state.
- Answers to questions are recorded as `interview.answered` with the
  answering principal and the answer text; the interview adapter no
  longer posts legacy `interview.*` events (`QuestionSink` is now an
  optional observer).
- The worker (`fabro run __run-worker`) records its lifecycle, notices
  and pause state over `HttpPlatformRecords`; `HttpRunStore` for the
  legacy event log and the worker's `run_store` are gone.
- `persist_created_run` appends `run.created` and `run.submitted`.

Readers:
- A stream follower (`server::stream_follower`) follows each live run's
  stream (Petri events and platform records), folds lifecycle records
  into the in-memory run state, forwards items to the global attach
  broadcast, and syncs blocked and paused from the projection.
- Slack posts questions from the projection's pending interviews,
  finishes them on `interview.answered` or `question_expired`, and sends
  lifecycle notifications with `notification.sent` dedupe.
- `GET /runs/{id}/events` and the attach endpoints serve only the run
  stream; the per-event, per-stage and `POST /runs/{id}/events`
  endpoints and their tests are deleted.
- `Database::load_run_projection` reads the Petri projection only.

Deleted with the writers:
- The SQLite blob and run-history activation migrations and their
  legacy Slate imports (`legacy_blob_import`, `legacy_run_history_import`,
  the activation backup): a greenfield server has no Slate history to
  import, and the run-history verification refused to start a server
  whose runs have no legacy events.
- `fabro-workflow`'s `operations::archive` and `operations::run_store`.
- The server's legacy-event unit tests and the CLI's `HttpRunStore` tests.

The in-process answer transport is now set after the starting and
running records land, not gated on the live status still being
`Starting` (the records already moved it).

The manifest validation test for a `run.agent.mcps.<name>` catalog
reference now expects `unsupported.workflow_toml.run.agent.mcps.reference`:
Petri's Fabro frontend has no server catalog to resolve it against.

Legacy readers still fail their tests until the next commits: the
reducer and Slate tests in fabro-store, the fabro-workflow create tests
that read the run back through the legacy store, the CLI tests seeded
through `POST /runs/{id}/events`, the CLI's legacy attach and render
paths, the sessions API, the OpenAPI conformance test, and the web
fixtures.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 12:32:06 -04:00
Bryan Helmkamp
af38d68942
Delete fabro-validate and fabro-acp; validate on Petri's check
Petri judges a workflow at admission, so Fabro's lint rules go:
`fabro-validate` (its 35 rules and the `LintRule` trait) is deleted, and
with it `fabro-acp` (only a rule and two legacy executor tests used it),
the model-resolution transform, the legacy `create`, `compile_create_run`
and `materialize_create_run` stages, and `fabro-graphviz`'s `condition`
and `fidelity` modules. `Diagnostic`, `RelatedDiagnostic` and `Severity`
move to `fabro_types::diagnostic`, the one shape every diagnostic takes.

Validation is now the same question the create handler asks. A new
server module, `petri_check`, builds Petri's check request from a
workflow bundle and the run's settings (every workflow of the bundle at
its bundle-relative path, the inputs, the run variables, the launch),
runs the check, and maps the diagnostics; Fabro adds one rule of its
own, `fabro.model.no_ready_provider`, refusing a model node when no
provider is ready. Admission, the validate and preflight endpoints and
the offline `fabro validate` all go through it:

- `validate_prepared_manifest` runs Fabro's structural pass (parse and
  transform, whose diagnostics stay) and then Petri's check, on the
  blocking pool from the handlers;
- the offline `validate_manifest` checks with no model client and with
  an unbound input as a warning (`CheckRequest.unbound_is_warning`), so
  a workflow validates before its inputs exist; a collected workflow
  before upload checks with unbound inputs as errors, as before;
- preflight resolves each LLM node's selector against the ready
  providers and the catalog for its probe, as the deleted transform did,
  and no longer probes a model Petri refused;
- the graph render endpoint needs only the structural pass;
- a run manifest now carries its `[run.goal] file`, which Petri reads
  from the bundle as it does for a version.

The transforms keep the authored model selector (`sonnet` stays
`sonnet`): Petri pins the catalog model in the admitted graph, not in
the graph Fabro displays or in the settings snapshot. Tests assert that,
and the CLI's validate, preflight and graph snapshots carry Petri's
diagnostics (`attractor.no_start`, `attractor.undeclared_node`,
`attractor.bad_on_failure`, ...) in place of the lint rules' text.

Known gaps, Petri's side: a `workflow.toml` whose `[run.environment]`
names an environment the server catalog defines but the file does not
is refused (`unsupported.workflow_toml.run.environment`), as admission
already refused it; an unbound input inside an included template
partial is a render error rather than the unbound-input warning.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 11:32:04 -04:00
Bryan Helmkamp
d90a5d9cbb
Delete fabro-core and the engine half of fabro-workflow
Every run executes on Petri, so the in-process legacy executor goes:
`fabro-core` and, in `fabro-workflow`, the handlers, lifecycle, pipeline
execution, routing, retry, conditions, node handlers, steering, agent
memory, artifacts, checkpoints, command log, and the `start`, `resume`,
`retry`, `fork`, `rewind` and `timeline` operations. The two are deleted
together because the engine half of `fabro-workflow` was the only user of
`fabro-core` and `fabro-core` the only runtime of that half; neither
compiles without the other.

Kept in `fabro-workflow`, narrowed: the parse/transform/validate/persist
pipeline and `create`, `archive`, `validate` (workflow definitions still
come from DOT and settings); the run tools (`run_tools`, moved from
`handler/llm/fabro_tools.rs`) for Ask Fabro, `fabro exec` and Petri's
host tools; the pull request pipeline (`pull_request`, moved from
`pipeline/`, for the step 0 port); Run Files' diff helpers in
`sandbox_git`; `git_identity`, `usage_rollup`, `run_status`,
`run_materialization`, `web_search` and `workflow_bundle`.

Server: `RegistryFactoryOverride` becomes `execute_in_process`;
`RunAnswerTransport::InProcess` carries only the interviewer; the
interrupt endpoint answers 501 `interrupt_unsupported` and every pair
endpoint 501 `pair_unsupported` (status lists none); rewind, fork, retry
and timeline handlers and routes are removed; the command log is served
from the stage output blob; usage rollups accumulate from the settled
projection after an in-process run as after a worker exit.

Ported while here:

- `materialize_admitted_run` materializes the goal and drops a disabled
  pull request block, as the legacy materializer did.
- A run whose admitted graph has an agent or prompt node is refused at
  create when no LLM provider is ready (`fabro.model.no_ready_provider`);
  a workflow of commands and gates needs no model and is admitted.
- The projection's question type falls back on the options, as the
  interview adapter does, so a gate with edge-label options answers as
  multiple choice.

Tests: the server scenarios (lifecycle, run completion, SSE, helpers)
run in process on Petri and assert Petri's stage labels and stream
names; the reconcile tests assert Petri's relaunch semantics; legacy
unit tests of the deleted executor are removed; three server unit tests
the removal took with it are restored; the pair fixtures go with the
pair feature. Petri test fixtures no longer name `[workflow] engine`.

Still red after this commit, all legacy consumers the next steps
delete or port: fabro-store's Slate/reducer fixtures and fabro-types
legacy JSON tests (step 4); server unit tests over legacy run events
(retry endpoints, list_run_events, artifacts, per-event pause/unpause,
run history activation, legacy sandbox fixtures) (steps 3-4); CLI tests
that parse legacy event envelopes, the legacy `events`/`attach`/`diff`/
`dump`/`inspect` snapshots, `run rewind`/`run fork`, the ACP and
git-identity workflow tests, and the runner tests that drive the legacy
worker by hand (steps 3-4); the web app's Petri fixtures still carry
`engine` (regenerate with `FABRO_CAPTURE_PETRI_FIXTURES` in step 4).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 10:44:41 -04:00
Bryan Helmkamp
a36bea15d2
Remove the engine flag: every run is a Petri run
Delete `Engine`, `RunEngine`, `[workflow] engine`, `[server.execution]
engine`, `FABRO_SERVER_ENGINE` and `fabro server start --engine`. The run
spec records what Petri admitted as `admission: PetriAdmission`; the
create handler always admits through `Runtime::check`; `execute_run`
always launches the Petri worker (or executes in process under the test
override); the CLI runner takes only the Petri worker path, and its
legacy control arm, artifact uploader, signal pause handlers and
credential helpers go with it. The CLI's `attach` and `events` read the
run stream only.

Two gaps this surfaced are closed here: the check adapter binds the
server's run variables as Petri compile variables (`{{ vars.* }}` in a
prompt no longer fails admission), and deleting a run removes its Petri
records, lease, platform records, projection and stream.

Tests: the config engine tests are replaced (an engine key is unknown),
the API round-trip test covers `PetriAdmission`, the server and CLI
Petri scenarios drop their engine settings, and the API tests that read
legacy event names now read the run stream or the session events. The
remaining red tests are fixtures and scenarios of the legacy executor
and the legacy event store (`fabro-store` `slate` and `run_state`,
`fabro-types` legacy `run.created` JSON, the server's handler-registry
scenarios, the CLI dry-run snapshots), which the next steps of the F4.3
series delete or port.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 09:58:21 -04:00
Bryan Helmkamp
d2f0e70dca
Run a workflow through Petri in the server, behind the engine flag
When a run's engine is Petri, the create handler hands the bundle, inputs
and launch to Petri's check instead of the legacy compile, lint and model
pinning, refuses the run with the validation error the legacy validator
uses (Petri's codes as the rules, listed in the API detail), and records
the admission on the run spec. The Fabro graph the read side displays is
parsed without validation. The scheduler executes a Petri run in the server
process through fabro_petri::engine, appending only the run lifecycle
events the read side needs (run.starting, run.running, run.completed or
run.failed); no stage or agent event is projected yet.

Scenario tests run the hello bundle on the OpenAI twin under the version
flag and a command-only bundle under the server setting, check Petri's
record agrees, and cover the refusals for an unknown attribute, an
undeclared node and an unknown model.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:27:44 -04:00
Bryan Helmkamp
c383b6a70b
Add the engine flag and record the engine on the run spec
A workflow version names its engine with `engine = "petri"` in the
`[workflow]` table of `workflow.toml`, and `[server.execution] engine`
(`FABRO_SERVER_ENGINE`, `--engine`) defaults it for every version that
names none. The choice, with what Petri admitted (the lowered root graph
and its children by blob and digest), is recorded on the run spec as
`RunEngine`, carried on `run.created`, and replayed into the projection.
A legacy run's spec omits the field, so existing specs decode unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:02:47 -04:00
Bryan Helmkamp
27f16f89c4
Delete fabro's own pricing: lithos-llm prices every response once
lithos-llm attaches a cost to every response at the client: the codec
keeps a provider-reported cost when the provider supplies one, and the
resolver fills the catalog's price for the route when it does not.
Pebble records that priced usage on every assistant turn and sums it,
so each AssistantMessage on the stream, and the store fold's live stage
usage, already carries the cost. Fabro's catalog re-pricing of the same
tokens was redundant, and is gone.

model_usage_from_llm, with_reported_cost, and every estimate_cost call
in fabro are deleted. The pebble handler's stage_usage groups pebble's
accounts by route and sums them with Usage::saturating_add, keeping the
cost and source pebble carried, so the terminal stage.completed usage is
the live fold's sum; it no longer fails when the catalog does not know a
provider. A one-shot prompt stage records the response's own usage and
cost as lithos-llm returned it. The per-model price cards in fabro-llm's
API module stay.

Tests: the pebble handler sums Catalog and Provider costs per row and
leaves a row's and the total's cost unknown once an answer was unpriced;
the store fold shows the same tokens and cost live and at completion,
and None at both for an unpriced answer; the agent integration test
compares the whole completed Usage with the live fold, cost included;
a one-shot prompt stage on a mocked OpenAI-compatible provider reports a
Catalog cost that lithos-llm's resolver attached, with no fabro pricing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:46:44 -06:00
Bryan Helmkamp
e9ee0aaeaa
Merge remote-tracking branch 'origin/main' into one-usage-type
# Conflicts:
#	Cargo.lock
#	Cargo.toml
2026-09-14 12:40:49 -06:00
Bryan Helmkamp
ef86ad278a
Carry usage as lithos-llm's Usage and rename billing to usage
Re-pin lithos-llm to 55add4596b861a0623d00c3a54aa5c147c8d504b and
pebble to c91810fe51aece80359b9cd8efea971af0c46925, where token usage
and cost travel together as Usage { tokens: TokenCounts, cost:
Option<Cost> }. Fabro now carries that one type everywhere it used to
carry BilledTokenCounts, BilledModelUsage, UsdMicros, or a token count
beside a cost_usd_micros.

fabro-types: billing.rs is usage.rs with ModelRef, ModelUsage { model,
usage }, sum_usage, and usage_is_empty; billing_rollup.rs is
usage_rollup.rs with ProjectionUsageStage, ProjectionUsageByModel,
ProjectionUsageRollup, and usage_rollup_from_projection. Every usage
field is named usage: StageProjection.usage and usage_by_model,
Outcome<Option<ModelUsage>>, stage.completed and stage.failed usage and
usage_by_model, prompt.completed usage, run.completed and run.failed
usage (total_usd_micros is gone), Conclusion.usage, StageSummary.usage,
Run.usage. RunSize buckets by Cost.

fabro-workflow: model_usage_from_llm prices tokens from the catalog with
a Catalog cost source, with_reported_cost keeps a provider cost, and the
pebble handler's stage_usage groups pebble's accounts by model and sums
rows with Usage::saturating_add, so a total has a cost only when every
priced part was priced. The store fold's live usage is the agent's
usage plus its descendants'.

API: the OpenAPI spec deletes BilledTokenCounts, BilledModelUsage,
CompletionUsage, CompletionCost, TokenUsage, and RunBillingSummary,
adds TokenCounts, Cost, Usage, and ModelUsage, and renames every
billing schema, property, tag, path, and operation to usage. fabro-api
reuses lithos-llm's and fabro-types' types through with_replacement,
with a round-trip test per replacement.

Old stored runs get no migration: their pebble events in the old shape
read back with zero usage, and their rebuilt projections lose agent
usage.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:31:34 -06:00
Bryan Helmkamp
3b8d712edb
Delete a live Daytona test's sandbox even when the test panics
The live Daytona tests create a provider sandbox and delete it on their
last line, so any panic or failed assertion before that line leaks a
running, billed sandbox. Two leaked that way on 2026-09-14 when a
sandbox-driver decoder flake panicked daytona_playwright_mcp_sandbox_transport.

Add fabro_sandbox::test_support::DeletedOnDrop, a guard that owns the
RunSandbox (Deref keeps the tests reading unchanged), offers an explicit
delete(self) for the happy path, and deletes from Drop otherwise. The
drop-time delete runs on its own thread and runtime because the test's
runtime may be unwinding. It reconnects the provider through the
ProviderAccess the test built the sandbox with, because the sandbox's
own handle pools HTTP connections whose tasks live on the test's runtime;
a live check of that path timed out after 10s.

Every Daytona test that creates a sandbox now holds it through the guard.
Unit tests over the scripted double prove delete-on-drop runs once, an
explicit delete runs once, and a panic inside catch_unwind still deletes
with and without a runtime.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 11:01:24 -06:00
Scott Werner
012b556367 Remove tests and fixtures tied to retired manifest fields 2026-09-13 08:55:55 -06:00
Scott Werner
ea11538038 Retire legacy manifest run creation 2026-09-13 08:49:50 -06:00
Bryan Helmkamp
f1118eead2 Delete the agent mirrors and move failover to prompt stages
Pebble's stream is the agent event contract. The run's own agent.mcp.ready,
agent.mcp.failed, and agent.mcp.disconnected events, which mirrored pebble's
McpServer* events, are gone with their props, the sink arms that emitted
them, and their conversion and naming entries; pebble's stored
agent.mcp.server.* events are the only record and feed the stage's fold.

The sink no longer mirrors RouteFailover onto agent.failover either: an
agent stage's moves are pebble's agent.route.failover. The event is now
prompt.failover, emitted only by a one-shot prompt stage that walks its
fallback plan itself, and its props are trimmed to the two routes, the
attempt, and the error; nothing read the rest.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:36:24 -06:00
Bryan Helmkamp
0c0e589a78 Bill a failed agent stage what it spent
An agent stage that failed billed nothing: the backend returned a bare
error and the outcome built from it carried no usage. A terminal failure
now becomes the stage's failed outcome from the same fold that bills a
completed stage, with the tree's usage, the rows by model, the files it
wrote, and its active time; stage.failed carries billing and
billing_by_model and the store keeps both. Cancellation and retryable
failures still go up as the error.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:47 -06:00
Bryan Helmkamp
df8762663b Bill an agent stage's whole session tree from one fold
One usage rule: a stage's usage is its session tree's, the root and every
subagent, live and at completion. The worker's event sink folds pebble's
SessionProjection over the events it records and the stage's billing and
files come from that fold at stage end, so the completed values are what
the run showed live. The store's live usage is the fold's tree usage, and
completion brings the catalog's price for the same tokens instead of
resetting them to the root's.

Fabro keeps catalog pricing: the root at its route, each descendant at its
own route where the catalog knows it and at the root's otherwise, a
provider-reported cost standing in where pebble has one. The rows travel
as billing_by_model on stage.completed and the stage projection, and the
billing rollup splits by_model by them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:47 -06:00
Bryan Helmkamp
ecafe1e807 Store ProcessingEnd and the mirrored pebble events in the run log
Pebble's SessionProjection reads ProcessingEnd to complete a prompt and mark
the session idle, so a projection rebuilt from the run's log needs it: one
small event per prompt. The four pebble events the sink mirrored onto
fabro's own agent.failover and agent.mcp.* are now stored verbatim as well,
so the fold sees the route moves and the MCP outcomes; the mirrors stay
until every reader is on the projection.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:07 -06:00
Bryan Helmkamp
697d8622e9
Merge pull request #843 from fabro-sh/remove/run-metadata-branches
Some checks are pending
TypeScript / Build (push) Waiting to run
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox plugins (stdio) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
TypeScript / Typecheck (push) Waiting to run
TypeScript / Test (push) Waiting to run
Remove Git run metadata branches
2026-09-12 16:41:24 -06:00
Bryan Helmkamp
e5d5c534ab
Merge remote-tracking branch 'origin/main' into remove/run-metadata-branches
Resolve conflicts between the metadata-branch removal and the
sandbox-driver adoption on main:

- fabro-sandbox docker.rs, sandbox.rs, daytona/mod.rs: take main's driver
  rewrite. The Sandbox trait is gone, so the PR's push_token_source
  removal now applies to RunSandbox instead; drop that accessor and the
  RepoCredentials::source helper that only served it.
- run_metadata.rs: keep deleted. Main's edits there were adaptations to
  the driver API and the run git identity field.
- lifecycle/git.rs, finalize.rs: keep the PR's removal of metadata
  snapshots and write_finalize_commit; carry main's RunSandbox,
  GitRetryPolicy, git_identity, local_sandbox, and test catalog changes.
- sandbox_git.rs: take main's version and drop the shadow_sha parameter
  and Fabro-Checkpoint trailer.
- git_integration.rs: remove meta_branch from the new git identity test.
- Cargo.toml: main's dependency set with fabro-dump kept as a
  dev-dependency.
- checkpoints.mdx: keep both the git identity paragraph and the durable
  execution state section.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:28:01 -06:00
Bryan Helmkamp
9550ff0806
Cover the failover continuation and the stopped failover
The existing failover test asserts the backup continued the turn after
the primary committed a tool result. A new test exhausts a two-route
chain and checks that one agent.failover and one
agent.route.failover.stopped are stored on the work stage, the stop
after the error it reports, with the exhausted reason and the failing
route. The events catalog documents both.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:26:32 -06:00
Bryan Helmkamp
57a746235e
Pin pebble to main after lithoscomputer/pebble#12
Pebble's RouteFailover event now describes the route that failed and how
the new route carried the prompt on. Record the continuation on fabro's
agent.failover event as an optional string (replay_prompt or
continue_turn); events written before it existed, and one-shot prompt
stages that walk the plan themselves, read as absent. The failed route's
usage, cost, and timing are not mirrored: the stage's totals already
include them through the prompt report, and no fabro run event carries
per-route usage yet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:26:32 -06:00
Bryan Helmkamp
7879e6d223
Merge pull request #856 from fabro-sh/brynary/run-git-identity
Resolve one Git identity per run and inject it into every workflow command
2026-09-12 14:23:30 -06:00
Bryan Helmkamp
ff7821dcd2
Adopt pebble's MCP disconnect event and startup timing
Pebble reports an MCP server whose connection closed mid-session once,
as McpServerDisconnected, and carries startup_ms on McpServerReady and
McpServerFailed. The workflow event sink mirrors the disconnect onto a
new agent.mcp.disconnected run event shaped like agent.mcp.failed, and
passes startup_ms through on agent.mcp.ready and agent.mcp.failed. The
raw pebble event is not stored for these, so the timing would otherwise
be dropped at the boundary.

The stage projection's McpServerStatus gains a `disconnected` kind next
to `ready` and `failed`. The fold keeps the server's tool count and
sticky invoked flag and only moves the status. The OpenAPI
McpServerStatus oneOf gains McpServerStatusDisconnected, and the
fabro-api round-trip test covers its JSON shape. A new
session_projection_parity test folds the same MCP events through
pebble's SessionProjection and fabro's stage projection and compares
them, including pebble's `disconnected`.

ToolErrorKind::Timeout needs no fabro change: the kind is stored as
pebble serializes it and never matched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 12:39:31 -06:00
Bryan Helmkamp
1aee59585b
Pin pebble to mcp-embedder-events and accept startup_ms on MCP outcomes
Pebble's McpServerReady and McpServerFailed events now carry startup_ms.
The workflow event sink destructured both variants by name, so it stops
listing every field. The pin is temporary: it moves to pebble main once
lithoscomputer/pebble merges the branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 12:29:36 -06:00
Bryan Helmkamp
ee6576cff7
Resolve one Git identity per run and inject it everywhere
A run now resolves a single author and committer identity once, after its
GitHub credentials are selected and before anything can commit, and uses it
for every commit it creates. Resolution order: a complete explicit
`run.git.author`; the run's GitHub App bot account
(`<slug>[bot] <id+slug[bot]@users.noreply.github.com>`); the authenticated
user of the run's PAT; the generic `Fabro <noreply@fabro.sh>`. A partial
explicit author overlays the fields it supplies. Only the selected
credential is consulted; a failed lookup is a setup error. A standalone
installation token falls back to the generic identity with a warning.

The resolved identity is carried on `RunOptions` and `EngineServices`,
recorded as a `git.identity.resolved` event and `RunProjection.git_identity`
so resume reuses it, and exposed through the run state API. Engine
checkpoints and metadata commits read it through `RunOptions::git_author`.
Every workflow execution path receives it as `GIT_AUTHOR_NAME`,
`GIT_AUTHOR_EMAIL`, `GIT_COMMITTER_NAME`, and `GIT_COMMITTER_EMAIL`, applied
last so it wins over inherited host variables and `[run.environment]`
entries: prepare steps, command stages, native agent shell tools, and ACP
launches. The identity is injected even without a Git origin, and the old
local `git config user.*` write is removed.

fabro-github gains `GET /user` and `/users/{slug}[bot]` lookups with mocked
tests for success, unauthorized, malformed, and transient cases. Real-Git
integration tests commit in the primary checkout, a clone, and a fresh
repository under conflicting local config, `[run.environment]`, and host
variables, and prove concurrent runs do not leak identities. CLI workflow
tests cover host script stages and ACP launch env through `fabro run`.

Docs and generated option metadata now describe the credential-derived
defaults instead of the stale `fabro`/`fabro@local` values.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 11:35:46 -06:00
Scott Werner
21a5e5b86f Simplify run creation to registered workflow versions 2026-09-12 11:04:15 -06:00
Scott Werner
c67c60eeba Create run tools from immutable workflow versions 2026-09-12 10:57:22 -06:00
Scott Werner
d611ef2bb1 Port workflow version creation to Pebble native tools 2026-09-12 10:23:34 -06:00
Bryan Helmkamp
957fc97c5c
Merge origin/main into pebble-agent-loop
Main merged the sandbox-driver adoption (#849) in a later form than this
branch was stacked on: the driver's own exec types replace fabro-sandbox's,
shell quoting moved to fabro-util, the sandbox lifecycle collapsed, and the
driver's events are stored as run events. This branch had deleted
`fabro-agent` and put the coding agent, the environment adapter, and the
steering hub on pebble.

The resolution takes main's sandbox API and re-applies pebble on top: the
`RunSandbox` `Environment` adapter moves to `pebble_environment.rs` (main's
`environment.rs` is the sandbox spec) and runs commands through `ExecSpec`
and `ExecControls`, feeding pebble's output sink from the driver's; the
driver-era `sandbox.*` names leave the known-event list, as on main, so a
stored event with that name and no driver shape is `Unknown` rather than an
error; `program_exit_code` matches pebble's non-exhaustive termination; the
Docker and Daytona smokes use main's constructor and credentials; the
remaining `fabro_agent` paths point at fabro-sandbox.

Pebble's `mcp` feature pins sandbox-driver, and the preview-url trait
objects only cross when both sides name one revision, so pebble moved to
main's `a92c0db6` (lithoscomputer/pebble#10) and fabro pins that pebble
revision until it lands on pebble main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 09:40:00 -06:00
Bryan Helmkamp
a33c9069cf
Run steering over pebble's bus and keep only fabro's attribution
Pebble's `SteeringBus` now owns the map of live sessions, the buffer for
steers that arrive between sessions, fan-out of steers and interrupts, the
close-the-door detach, and the hold that keeps a paired session open. The
hub keeps what only fabro knows: pair records, principals, stage ids, and
the run events that put bus activity on the run's stream in the order its
consumers expect. `PebbleControlHandle` is gone, since the coding agent's
control handle is a bus session natively; the ACP session joins the bus
through a small adapter and carries pebble's steering message end to end,
so a human steer keeps its author on the ACP `agent.steering.injected`
event. A pair message that evicted an older steer is now accepted and the
eviction recorded, where before it was queued and reported as refused.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 20:35:31 -06:00
Bryan Helmkamp
6592505d0e
Take environment helpers, compaction accounting, and search providers from pebble
Steps 4, 6, and 7 of .ai/plans/pebble-absorbs-embedder-concerns.md,
pinning pebble fc907a1 with its `search-providers` feature.

The sandbox's `Environment` adapter uses pebble's `environment::support`
for the glob grammar check, the tree-order sort of a listing, and the
capture accounting, in place of its own copies; the adapter itself stays.
The stage reads the compactions a prompt performed from the report, as a
breakdown of the usage it already billed. `web_search.rs` keeps the
vault-backed credentials and the Brave-over-Venice preference, and hands
pebble's `Brave` or `Venice` provider a fabro HTTP client; the providers
and their tests are pebble's now.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 19:51:12 -06:00
Bryan Helmkamp
d4e483925c
Let pebble discover memory and skills, and name the reuse rules
Steps 5 and 8 of .ai/plans/pebble-absorbs-embedder-concerns.md, pinning
pebble 49da137.

Agent stages and `fabro exec` ask pebble for the profile's instruction
files from the repository root down to the working directory
(`MemoryDiscovery::from_git_root`), which fabro lacked: it read the
working directory alone. Skill directories are pebble's to resolve too:
the user's skills directory, then `.fabro/skills` and `skills` under the
repository root. Prompt stages keep reading the working directory alone,
through the same discovery and loader, so `agent_memory.rs` keeps only
that call; the filename table is pebble's now.

A retained thread's export comes from `export_for_reuse`, which closes
the session and hands back an export whose cursor is already past the
close, in place of export, shutdown, and a cursor advance by hand. Ask
Fabro resumes a stored record with `resume_after`, the rule it applied
under the older name.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 19:40:36 -06:00
Bryan Helmkamp
64c578d044
Take MCP servers from pebble
Step 3 of .ai/plans/pebble-absorbs-embedder-concerns.md, pinning pebble
6cdb30a with its `mcp` feature.

Pebble starts the stage's MCP servers while the agent is built, registers
their tools under `mcp__{server}__{tool}` with `ToolSource::Mcp`, and
closes them with the agent, for all three placements: a child of the run
worker over stdio, a server over HTTP (streamable or SSE), and a server
launched in the run sandbox and reached through the sandbox's preview
URL. `fabro_mcp::pebble::pebble_server` maps `McpServerSettings` onto
pebble's `McpServer`, keeping fabro's `/sse` path for sandbox-hosted SSE
servers; `RunSandbox::port_routes` hands pebble the driver's `PreviewUrls`
facet as the route to a sandbox port. The stage's event sink mirrors
`McpServerReady` and `McpServerFailed` onto the run's `agent.mcp.ready`
and `agent.mcp.failed` events, as it mirrors `RouteFailover` onto
`agent.failover`, and stores no second copy of a mirrored fact.

Deleted: `sandbox_mcp.rs`, the MCP branches of `pebble.rs` and `fabro
exec`, and fabro-mcp's client, connection manager, HTTP helpers, and SSE
transport, whose tests moved to pebble. fabro-mcp keeps the settings
re-export, the mapping, and a stdio client behind `test-support` for the
tests of fabro's own MCP server. `fabro exec` reports each server's
outcome from the agent's snapshot. The Daytona Playwright live test now
drives the sandbox-hosted server through an agent.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 19:26:41 -06:00
Bryan Helmkamp
b5fd907c3d
Take files touched and route failover from pebble
Steps 1 and 2 of .ai/plans/pebble-absorbs-embedder-concerns.md, pinning
pebble 1a5abe4.

Files touched come from `PromptReport`: pebble computes them from every
successful write, edit, and patch across the prompt, subagents included,
so the event-fed `FileTracking` and the tracking half of
`WorkflowEventSink` go. The stage unions the reports of its prompts.

Route failover is pebble's. The stage resolves its plan from the catalog
as before and hands pebble the remaining routes through
`fallback_routes`, each with its controls and the stage's output limit.
Pebble keeps the conversation, moves it to the next route, requeues
pending steering, and continues the prompt; the stage's plan follows the
route the report says the prompt ended on, re-activates the session
there, and mirrors pebble's `RouteFailover` as the run's `agent.failover`
event with the same payload as before. `prompt_with_failover`,
`resume_agent_on_route`, and the route bookkeeping in `LiveAgent` go.
One-shot prompt stages still walk the plan themselves.

`agent.route.failover` and `agent.tool.rounds.exhausted` join the derived
event names; both variants were falling back to `agent.event`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 18:47:12 -06:00
Bryan Helmkamp
8d32b472db
Pin the merged pebble main and adopt the pebble features petri already uses
Step 0 of .ai/plans/pebble-absorbs-embedder-concerns.md. Pebble's
fabro-exec-sink-max-turns branch is merged into its main (pebble 4db661c,
on the embedder-concerns branch until it lands), so fabro and petri pin
one line. Both prompt budgets survive the merge because they count
different things.

- Agent hooks use `with_max_tool_rounds`, which is what the
  `max_tool_rounds` setting names and what petri does: `max_tool_rounds`
  model turns may run, and the hook proceeds on a turn that still asks for
  tools, so pebble gets `max_tool_rounds - 1` rounds and `ToolRoundsExhausted`
  fails open. Zero rounds proceeds without an agent, as the old loop did.
  Agent stages set no turn budget; the stage timeout and stall watchdog
  bound them.
- Prompt stages load project memory through pebble's `ProjectMemory`, the
  loader agent stages already run over `with_memory_files`, instead of a
  hand-rolled copy of its budget, deduplication, and truncation.
- `fabro_sandbox::SecretRedactor` puts fabro's secret scanner on pebble's
  text seams (process output tails, failed tool messages) for agent
  stages, Ask Fabro, hook evaluators, and `fabro exec`. The final redaction
  pass over every stored `RunEvent` stays; this does not replace it.
- Agent stages state their compaction policy explicitly: the 80 percent
  threshold and six preserved turns fabro's own agent loop applied.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 18:16:03 -06:00
Bryan Helmkamp
6599f7cdc0
Cover the pebble agent loop with workflow-level tests
Each behaviour the pebble backend owes a run now has one test that drives
it through the workflow engine against a scripted OpenAI-compatible model:

- an agent stage under every harness profile (openai, anthropic, claude-5,
  gemini, kimi) writing a file with that profile's own tool spelling, with
  the event sequence, files touched, response, usage, and cost checked; the
  codex vocabulary applies a patch through the OpenAI twin's custom tool
  call, which `TwinToolCall::custom` now scripts
- steering delivered mid-stage, an interrupt with a steer, run cancellation,
  and the executor-enforced stage timeout
- a question answered through the interviewer, a subagent whose events carry
  its parent's session id, and an MCP tool served by a stdio server
- failover to a second provider after a tool ran, continuing the recorded
  conversation without running the tool again
- a failing event sink ending the stage with the sink's error
- Ask Fabro resuming a stored record across turns, with the cursor moved
  past the run's event log when the record's own cursor fell behind
- Docker and Daytona smokes running an agent stage through the provider
  sandboxes

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 16:32:52 -06:00
Bryan Helmkamp
350abac014
Build the local sandbox through the one provider path
SandboxSpec had a Local variant beside the provider spec, and a local
sandbox was created by hand over a bare Host provider: no workspace, no
provider connection, its own reconnect, and its own push rule for the
designated directory. The local kind is now one more SandboxSpec:
SandboxSpec::local names the directory on a HostDirectory spec with a
skip clone, and provider_sandbox builds it like a plugin kind, creating
the directory when missing since the Host provider requires it to exist.
Every RunSandbox carries a workspace; a handle wrapped as is gets the
workspace of its own working directory.

The push rule is one rule for every checkout: a checkout fabro cloned
pushes with the credentials it was cloned with, and any other checkout
pushes when it has an origin, with whatever credentials it carries. A
local run therefore pushes the same way before and after a resume;
before, a reconnected local sandbox carried an attached workspace that
never pushed while a fresh one did.

Reconnect uses the recorded id for every kind. The recompute of a local
id from its directory, kept for records written before directories had
ids, is gone, and test fixtures that wrote made-up local ids derive them
through test_support::local_sandbox_id instead. A local run's record now
carries its workspace layout like every provider-chosen directory, and
the sandbox.initializing event precedes the driver's create events for
local as for every other kind.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 16:28:17 -06:00
Bryan Helmkamp
981253f990
Collapse the run sandbox's lifecycle surface
RunSandbox grew its lifecycle methods one adopter at a time and ended up
with several names for each step. Reconnecting from a run record had
four entry points (reconnect, reconnect_for_run,
reconnect_for_run_with_events, reconnect_driver_for_run) that all
forwarded to the last one. Bringing a sandbox back had two (start and
activate) over the same make_ready, and releasing it had two (delete and
cleanup) over the same release. Two more methods had no callers at all:
set_autostop_interval, which nothing set after the driver took over
lifecycle timers, and resume_setup_commands, which resume stopped using
when checkout moved to the git facet.

There is now one of each. reconnect_for_run takes the record, the
provider access, an optional run id, and an optional event context;
callers that need none pass None. activate is the single "make usable"
step: a running sandbox only learns its platform when it has not yet, a
stopped or paused one is started and its Bash verified, and resume calls
it like every access-time caller. delete is the single release; for a
designated host directory it frees the handle and leaves the directory in
place, as cleanup did. The tests and server call sites follow the
renames; behavior is unchanged except that resuming an already running
sandbox no longer re-runs the Bash probe.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 15:05:06 -06:00
Bryan Helmkamp
8771ef6d8e
Move the Read tool's line numbering and the shell quoting wrapper out of fabro-sandbox
format_lines_numbered renders a file the way the agent's Read tool shows
it: every line prefixed with its number, from an offset for a limit. That
is the agent's presentation, not a sandbox concern, and it only lived in
fabro-sandbox so RunSandbox::read_file(path, offset, limit) could call
it. The function now lives in fabro-agent next to the Read tool, with its
tests, and the Read, ReadManyFiles, and Kimi ReadFile tools number the
text they get from read_file_text themselves. RunSandbox::read_file goes
away; the sandbox returns bytes or text and nothing else.

fabro-sandbox also re-exported shell_quote through a one-line wrapper so
callers could reach it from the sandbox crate or from fabro-agent. The
audited implementation is fabro_util:🐚:shell_quote; the six
importers now use it directly and the wrapper and both re-exports are
gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 15:00:17 -06:00