`dry_run_parallel` and `dry_run_styled` timed out under load in the run
that found the other eight, so they were not recorded with them. Alone they
show the same one-line change: the Start stage's completion now precedes the
run branch line, the order the positioned records give every run.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The pin to Petri 639ce3e moves the run's format version from 6 to 7, which
the attach JSON snapshot records. The other eight snapshots had recorded the
run branch and Git identity lines before the Start stage's completion: the
order the clock gave them before "Place the run branch and git identity
records with their checkpoint" positioned the two records after the firing's
finish. That commit refreshed only two files, and these eight already
differed the same way at the commit before the pin bump; they now record the
one order every run produces.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`run events --pretty` reads the flattened `git.identity` fields, and
shows a `run.diff` record as its summary and an `artifact.collected`
record as its path and size. The snapshot filters redact the base commit a
`Branch:` line names. The CLI snapshots now carry the `Base:` line, the
`run.branch` and `git.identity` stream items, the dry run's simulated
response and the two response files a dry run dumps.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The TypeScript client's `RunCheckpoint` still carried the legacy
executor's resume fields; regenerating it from the spec gives it the
slim shape (`timestamp`, `current_node`, `git_commit_sha`). The CLI
test helpers that read a run's stream from its directory or the API
are named `run_stream_items`, so nothing but the dropped table's
migration history is still called `run_events`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The CLI's integration tests seeded runs by appending legacy run events
and waited on legacy event names. Now every seeded run is a real dry
run: the fixtures start the run through the CLI, read the run id from
its output and wait for the stream's terminal lifecycle record. Waits,
assertions and snapshots read `RunStreamItem`s (`run.finished`, the
platform `run.lifecycle` record, `derived.parsed.kind == "question"`).
Test changes:
- support.rs: `run_completed_dry_run`, `wait_for_run_finished`,
`wait_for_lifecycle`, `wait_for_stream_item`; the `append_seeded_*`
writers, `wait_for_event_names` and the git-backed seeded fixtures
are gone (the checkpoint patch is not in the projection yet).
- diff.rs keeps only the help test; inspect.rs drops the git-backed
checkpoint test; events.rs, dump.rs, create.rs, attach.rs and
dry_run_examples.rs snapshots are re-recorded over Petri's rendering
with redactions for epoch millis, digests and commit shas.
- run.rs: the remote foreground mock serves stream pages and a run
state with a conclusion and a `report` stage response; the event
history test checks `run.finished` and the terminal lifecycle item.
- runner.rs / attach.rs: question ids containing `#` are percent-encoded
in answer URLs.
Production fixes the ports surfaced:
- petri_worker.rs: a cancelled run exits without reporting a failure.
- runner.rs: resuming a run that already finished fails its precondition
instead of starting a worker.
Left failing on purpose, each bound to a Petri-side gap reported to the
lead rather than to the port: sandbox_cp (4), sandbox_preview and
sandbox_ssh (the projection carries no sandbox instance), the artifact
collection tests in workflow::artifacts and run.rs (no artifact
collection for Petri runs yet), and the two dump blob-ref tests (blob
refs are not visible in the inspect output).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Step 4 of the legacy executor deletion, fourth commit: with no writer
and no reader left, the legacy event log goes.
- `fabro-types`: `run_event` (`EventBody`, `RunEvent` and every props
struct), `EventEnvelope` and the `RunEventDetail*` types are deleted.
What the projection and the API still use moves out of the event
vocabulary: `AgentEventProps`, `AgentSessionActivatedProps`,
`AgentToolsAvailableProps`, `StagePromptProps`, `SessionCapability`
and the coding event names to `agent_props`; `RunNoticeLevel` and
`RunNoticeCode` to `notice`; `InterviewOption` beside the question
types; `RunRunnableSource` beside the run status. `Checkpoint` is
what Fabro records for a Petri run: `timestamp`, `current_node`,
`git_commit_sha`; the conclusion's stage summaries derive from the
projection's stages instead of the checkpoint's node maps.
- `fabro-store`: the Slate bridge (`RunDatabase`, the Slate `Database`,
`keys`, `record`, `EventPayload`) and the reducer (`run_state`) are
deleted. `Database` is the blob table and the run summary store over
one pool; the blob store is SQLite only; the run summary store keeps
the `runs` row a projector writes and lists, and finds the pull
request creation candidates over `platform_records`; `build_summary`
and `projected_usage` live in `run_summary`. The SlateDB dependency
is gone. Test fixtures build the store from its two SQLite stores.
- `fabro-workflow`: the `event` module (the `Event` enum, its
conversion, sink, emitter, redaction, stored fields and names),
`runtime_store`, `StageScope` and the legacy seeding test helpers are
deleted; the tests that seeded legacy runs read platform records or
a projection instead.
- `fabro-sandbox` owns `GitRetryReason`.
- The server builds the store without an object store; the legacy
`POST /runs/{id}/events` tests go, an interrupt answers
`interrupt_unsupported` in the tests as it does in the handler, and
the tests that read a run back through the Slate handle read its
projection or its platform records. The projection folds a block
that lands while the run is paused as the pause's prior block, and a
pause or unpause clears the pending control it answers; a control
request's check-and-append holds a per-run lock so two concurrent
cancels record one request.
- The CLI's final output is the response of the last stage that
produced one; the workflow tests read completed nodes from the
succeeded stages.
- The spec's `RunCheckpoint` carries the three fields the type keeps.
Still failing until the next commits: the CLI tests that seed runs
through `POST /runs/{id}/events` or wait for legacy event names, and
the two Ask Fabro resume tests (the sandbox instance gap).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Step 4 of the legacy executor deletion, third commit: the legacy event
API and every reader of it go, so that the next commits can delete the
event log, its reducer and the types beneath them.
The API:
- `GET /runs/{id}/events` pages the run stream only
(`PaginatedRunStreamList` by `after`); the legacy `since_seq`,
`before_seq` and `order` cursors, the `oneOf` envelope, the legacy
`EventEnvelope`, `PaginatedEventList`, `RunEvent`, `EventSeq`,
`AppendEventResponse` and `RunEventDetailResponse` schemas,
`POST /runs/{id}/events`, `GET /runs/{id}/events/{seq}` and
`GET /runs/{id}/stages/{stageId}/events` are deleted. `GET
/runs/{id}/attach` and `GET /attach` frame `RunStreamItem`s only.
- The Rust and TypeScript clients regenerate; the removed models leave
the TypeScript package.
The readers:
- `fabro-client` drops the legacy run event listing, tail and attach
methods and `RunEventStream`; `list_run_stream_until` bounds a stream
read.
- `fabro-tool`'s `fabro_run_events` lists, searches and details the run
stream: `after` is the exclusive `stream_seq` cursor, `event_id` the
item's id, filters match the item's name and `recorded_at`.
- `fabro-dump` writes the stream to `events.jsonl`; `fabro dump` reads
it.
- The CLI's progress renderer keeps only what the run stream drives:
the legacy event conversion, the sandbox and setup displays and their
styles go. `fabro system events` prints stream items.
- The server's demo mode folds its agent fixture straight into the
session projection and answers the attach stub with a stream item;
the demo stage events endpoint is gone.
- The web app: every run is a Petri run. The legacy event hooks,
renderer props, stage popover summary, run phases derivation and
live-event payload handling are deleted or ported to `RunStreamItem`;
toasts and board refreshes read the stream's platform records.
- Tests: the legacy API round trips and pagination tests are deleted;
the CLI's MCP, attach and system event mocks serve stream pages; the
CLI test helpers read stream items.
Still failing until the later commits: the CLI tests seeded through
`POST /runs/{id}/events`, the server tests over the legacy store, and
the legacy type tests.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri judges a workflow at admission, so Fabro's lint rules go:
`fabro-validate` (its 35 rules and the `LintRule` trait) is deleted, and
with it `fabro-acp` (only a rule and two legacy executor tests used it),
the model-resolution transform, the legacy `create`, `compile_create_run`
and `materialize_create_run` stages, and `fabro-graphviz`'s `condition`
and `fidelity` modules. `Diagnostic`, `RelatedDiagnostic` and `Severity`
move to `fabro_types::diagnostic`, the one shape every diagnostic takes.
Validation is now the same question the create handler asks. A new
server module, `petri_check`, builds Petri's check request from a
workflow bundle and the run's settings (every workflow of the bundle at
its bundle-relative path, the inputs, the run variables, the launch),
runs the check, and maps the diagnostics; Fabro adds one rule of its
own, `fabro.model.no_ready_provider`, refusing a model node when no
provider is ready. Admission, the validate and preflight endpoints and
the offline `fabro validate` all go through it:
- `validate_prepared_manifest` runs Fabro's structural pass (parse and
transform, whose diagnostics stay) and then Petri's check, on the
blocking pool from the handlers;
- the offline `validate_manifest` checks with no model client and with
an unbound input as a warning (`CheckRequest.unbound_is_warning`), so
a workflow validates before its inputs exist; a collected workflow
before upload checks with unbound inputs as errors, as before;
- preflight resolves each LLM node's selector against the ready
providers and the catalog for its probe, as the deleted transform did,
and no longer probes a model Petri refused;
- the graph render endpoint needs only the structural pass;
- a run manifest now carries its `[run.goal] file`, which Petri reads
from the bundle as it does for a version.
The transforms keep the authored model selector (`sonnet` stays
`sonnet`): Petri pins the catalog model in the admitted graph, not in
the graph Fabro displays or in the settings snapshot. Tests assert that,
and the CLI's validate, preflight and graph snapshots carry Petri's
diagnostics (`attractor.no_start`, `attractor.undeclared_node`,
`attractor.bad_on_failure`, ...) in place of the lint rules' text.
Known gaps, Petri's side: a `workflow.toml` whose `[run.environment]`
names an environment the server catalog defines but the file does not
is refused (`unsupported.workflow_toml.run.environment`), as admission
already refused it; an unbound input inside an included template
partial is a render error rather than the unbound-input warning.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The four hook tests and arc_e2e_with_real_llm in workflow/hooks.rs failed
in twin mode for three reasons, all in the test fixtures.
The hooked workflows were written as `<name>.toml` beside `<name>.fabro`.
Version packaging accepts a config only as `workflow.toml` beside its
graph (44dccfa3d), so `fabro run` failed at collection. Each hooked
workflow now lives in its own `<name>/` directory as `workflow.toml`.
The isolated server never learned the twin's base URL. The CLI command
carried `OPENAI_BASE_URL`, but the run executes in the server, which does
not see the test process environment, so it called the real OpenAI API
with the namespace as its key. The twin-mode server settings now repoint
the `openai` provider at the twin through the operator `[llm]` overlay,
the same way `run_uses_vault_credentials_for_worker_execution` does.
With the server reaching the twin, the hook scenarios were consumed by
the wrong request: the server asks the model for a run title in the same
namespace before the hook fires, and the scenarios had no matcher. The
block test then saw the twin's default response and the hook failed open,
so the run succeeded. Hook scenarios now match on the `Hook prompt:`
prefix of the evaluator's user message.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Add a `sandbox_tests!` scenario that initializes a repository inside the
sandbox from a script stage, commits, and prints the author and committer
the commit object carries. It runs on the local host and, when the plugin
executables are on PATH, on the host and Docker sandbox plugins, with
conflicting `GIT_*` variables inherited from the launching shell.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A run now resolves a single author and committer identity once, after its
GitHub credentials are selected and before anything can commit, and uses it
for every commit it creates. Resolution order: a complete explicit
`run.git.author`; the run's GitHub App bot account
(`<slug>[bot] <id+slug[bot]@users.noreply.github.com>`); the authenticated
user of the run's PAT; the generic `Fabro <noreply@fabro.sh>`. A partial
explicit author overlays the fields it supplies. Only the selected
credential is consulted; a failed lookup is a setup error. A standalone
installation token falls back to the generic identity with a warning.
The resolved identity is carried on `RunOptions` and `EngineServices`,
recorded as a `git.identity.resolved` event and `RunProjection.git_identity`
so resume reuses it, and exposed through the run state API. Engine
checkpoints and metadata commits read it through `RunOptions::git_author`.
Every workflow execution path receives it as `GIT_AUTHOR_NAME`,
`GIT_AUTHOR_EMAIL`, `GIT_COMMITTER_NAME`, and `GIT_COMMITTER_EMAIL`, applied
last so it wins over inherited host variables and `[run.environment]`
entries: prepare steps, command stages, native agent shell tools, and ACP
launches. The identity is injected even without a Git origin, and the old
local `git config user.*` write is removed.
fabro-github gains `GET /user` and `/users/{slug}[bot]` lookups with mocked
tests for success, unauthorized, malformed, and transient cases. Real-Git
integration tests commit in the primary checkout, a clone, and a fresh
repository under conflicting local config, `[run.environment]`, and host
variables, and prove concurrent runs do not leak identities. CLI workflow
tests cover host script stages and ACP launch env through `fabro run`.
Docs and generated option metadata now describe the credential-derived
defaults instead of the stale `fabro`/`fabro@local` values.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The hook tests wrote their `[[run.hooks]]` entries into the user's
settings file. `fabro run` no longer transmits `run` settings from
there — it warns and points at `workflow.toml` — so no hook ran in any of
these tests. The three that expect the run to proceed kept passing for
the wrong reason.
Each hooked test now writes a workflow config that names its graph and
carries the hooks, and runs that config. The twin-mode server settings
stay in the settings file, which is where they belong.
The hooks reach the run now, but the twin-mode tests still cannot pass
on this branch: the run executes in the isolated server, which never
learns the twin's base URL and so calls the real OpenAI API with the
namespace as a key. That plumbing belongs with the lithos credential
resolution on main, not here.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The plugin scenarios launched two executables fabro built itself,
`fabro-sandbox-host` and `fabro-sandbox-docker`, that only wrapped the
driver's providers in a stdio server the driver already ships as
`sandbox-driver-host` and `sandbox-driver-docker`. Fabro now finds the
driver's executables on PATH: CI installs them at the rev the workspace
pins, read from Cargo.toml so the plugins and the in-process providers
are one build, and a developer installs them the same way. The plugin
proof in fabro-sandbox skips without the executable unless the CI
environment forbids skipping; it was also never running in CI, which
ran it under `--run-ignored only` although it is not ignored, so the
job now runs it on its own.
The pin moves to the head of the sandbox-driver PR stack #9 through
#15: tag pins, classified git failures, the stop grace ladder, snapshot
ensure, the ownership scope, and the testing doubles, which the next
commits adopt. The `sandbox-driver-testing` crate joins the workspace
dependencies for them.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
sandbox-driver PR #9 removes the check that a plugin's declared kind match
the configured one: an operator who configures a path and pins its
checksum has already chosen the executable, so the configured kind is
fabro's name for whatever it serves. With that in the driver, fabro no
longer needs plugin settings on a bundled kind to reach Docker over the
wire. Bundled kinds reject plugin keys again, `connect_provider` links a
bundled kind in-process and launches everything else, and the CLI
scenarios run the Docker executable under the non-bundled `docker-plugin`
kind. The driver pin moves to the PR head until it merges.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Closes the sandbox-driver adoption: any provider a sandbox-driver plugin
executable serves can now host a fabro run, and fabro's own bundled
providers can be served the same way.
- `SandboxSpec::Plugin` builds a normalized driver spec from the
environment (image or Dockerfile source, or a provider-managed
directory; resources; network policy; labels; env) and lays fabro's
repository checkout out inside the provider's working directory. The
layout is recorded on the run through the new `workspace_layout` trait
method.
- Plugin settings on a bundled kind (`[server.sandbox.providers.docker]
path = ...`) serve that kind out of process through the driver's
executable; the config layer no longer rejects them.
- `ProviderAccess` carries the server's provider settings and the vault's
Daytona credentials to every reconnect: run resume, sandbox details,
terminals, previews, and the worker's start path. The worker receives
the settings through `StartServices`. No "plugin not wired" errors
remain.
- The CLI worker requires GitHub credentials only when a repository will
be cloned; a `none` target on a clone-based provider creates an empty
workspace and needs none.
- fabro-db tracks its migrations directory so a new migration file
recompiles the crate; the environment provider migration had been
silently missing from stale builds. Environment store 500s now log
their cause.
- The CLI workflow scenarios run against `host-plugin` (the driver's
Host executable under the non-bundled `host` kind) and `docker-plugin`
(the bundled `docker` kind served over stdio), each on an isolated
server, printing the server log on failure. A live Daytona gate runs
the native git clone over the JSON-RPC wire. A new CI job runs the
plugin scenarios and the driver-backed Docker integration tests with
the plugin executables built.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A node with no `shape` defaulted to `box`, which resolves to the agent
handler. That made a shapeless `script` node run as an LLM call prompted
with its own label, while the `script` was reported as inert — wrong
behavior behind a warning.
`script` is read by the command handler and by nothing else, so a
shapeless node that sets it is unambiguously a command node. `shape()`
now infers `parallelogram` in that case. An explicit `shape` still wins.
Two rules keep the inference honest:
- `script_prompt_conflict` — setting both `script` and `prompt` is an
error. No handler reads both. It fires regardless of shape so that
adding one cannot downgrade the error to a warning.
- `command_requires_script` — a command node without a script is an
error. Without this the original trap just moves: a node meant as a
command that omits its script silently becomes an agent again.
Also drops the `tool_command` alias in favor of `script` alone, routing
the six read sites through a new `Node::script()` accessor.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Diagnostics have carried a `fix` field all along, but the CLI renderer
never printed it — the suggestion was only reachable through --json. The
actionable half of every validation failure was invisible to the person
running the command.
print_diagnostics now emits the fix as a dim-labelled continuation line
under any diagnostic that has one, at both error and warning severity.
Gating it behind --verbose would defeat the point, and printing it only
for errors would read as "this warning has no fix" — the warning
suggestions are useful on their own. Diagnostics that set no fix simply
omit the line.
The severity match moved into print_diagnostic so the fix line is
appended once in the loop rather than copied into all five arms; the
rest of the diff is reindentation.
print_diagnostics is shared by validate, preflight, graph, exec, and
dry-run, so this covers all five. Eleven inline snapshots across four
files gain a fix line; every change is additive.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>