In-process tests over the memory store: every finish is committed on the
run branch with its identity trailers and recorded with its commit, the
run-end hooks reach Petri's local service through Fabro's wrapper, a
stage that fails on its own terms is committed and its failure route runs
on the committed files, a failed checkpoint records `checkpoint_failed`
with no route taken and a restart reports the run failed, and a
`[[run.hooks]]` hook blocks an agent's tool call through the forwarded
service, with the model told why.
Real-binary scenarios crash the server and its worker with SIGKILL: after
a durable finish the stage's commit is not repeated and the interrupted
stage reruns on its snapshot; a crash held before the commit reruns the
stage once; a crash held after the commit but before its record
reconciles the record from the snapshot repository; a deleted workspace
is restored; a failure route sees the same committed files after a
crash; a failed checkpoint fails the run and a restart leaves it failed.
Recovery selects the executions `inspect_run` reports incomplete, and a
run whose coordinator log is still empty is left to the worker's resume.
The worker's platform record endpoints get an API test and the generated
TypeScript client.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The interview adapter derived its own question id from Petri's identity
and posted it on `interview.started`, while the projection over Petri's
records serves the pending question under Petri's `Question.id` with the
firing's stage label. The answer endpoint validates against the
projection, so an answer under the projection's id never reached the
adapter's wait.
The adapter now waits under Petri's id and labels the question's stage
through the projection's own rule: `stage_label`, `is_shown` and
`visit_of` move out of `start_visit` into shared functions, and the
adapter's observer derives each firing's `visit.started` through Petri's
`Projection`, as the projector does, so the label matches by
construction. The full Petri identity stays on `AskedQuestion`.
The legacy `interview.*` events are still posted, under Petri's id, for
the readers that follow the event stream rather than the projection: the
Slack service, `run attach`, the web app's Q&A renderer and the server's
answer claim. The store already derives the `interview.answered` platform
record from `interview.completed` for a Petri run, so who answered is
recorded under Petri's id with the answering principal.
The gate scenarios assert the new identity and encode the id as one path
segment, as the generated clients do. Projection tests cover an expired
question and an auto-approved answer.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fabro's hooks on a Petri run wrap the hooks the runtime installed for
`[[run.hooks]]` and forward every point. In `prepare_result`, before the
finish is recorded, they commit the stage's files on the run branch of
its host workspace with Fabro's author identity and the run, execution,
firing and attempt as trailers, and publish the commit to a snapshot
repository beside the run's workspaces under a ref per checkpoint. A
stage that failed on its own terms is committed like a successful one; a
commit that fails is fatal: the outcome becomes a `checkpoint_failed`
failure, the run is cancelled through the coordinator handle, and the
transition refuses the firing's routes. In `transition` they write the
platform checkpoint record, keyed on the Petri position and the
checkpoint's operation identity, and a failed write is a recorded
problem.
On restart the server runs the recovery protocol before it relaunches a
worker: a run with a failed checkpoint is reported failed; otherwise
every live execution's last durable finish names the snapshot its
workspace is verified against, reset to, or restored from, with a lost
record reconciled from the snapshot repository, and a finish with no
snapshot fails the run rather than resume it on stale files.
The worker reaches the platform records over two new worker-scoped
endpoints; the server reaches the table directly. A test gate directory
lets the CLI scenarios hold a checkpoint at a named point.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The startup pass called the pass directly while a signalled pass could
run for the same run, so both read one committed stream sequence and
the second insert into the stream failed on its primary key, which
stopped the restarted server. Passes now take a per-run lock, and a run
whose startup pass fails is logged and left for its next signal instead
of stopping the server. A test races four passes, a signal and the
startup pass over one run and checks the stream stays contiguous.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The projector takes two pools: the one Petri's records live in and the
one the view tables live in. In the server both are the one database;
a test fixture keeps the runs row, the platform records and the
projection tables in the run summary store's own pool, which the
projector was not reading, so a run projected in a test server folded
its Petri events before its run.created record. The startup run-history
verification checks only a Petri run's identity and legacy guard, since
its row is the projector's. An agent stage's response is the
response.<node> its outcome wrote into the run context, as the prompt
step writes it. The scenario tests assert each branch's own index.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The adapter tests run whole workflows on the host sandbox through the
plugin, and one calls the twin; under a full parallel run one of them was
killed at the default 3 s.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The server holds one projector over its database and signals it after
each committed worker append, after each committed platform record
(through the run summary store's hook), at worker exit, and over every
Petri run at startup after the restart reconcile. A run executing in
the server process under the test override appends through the
projector's observing store, so it is signalled the same way. The
scenario tests read GET /runs/{id}/state after the view settles: the
hello prompt stage with its response, the command stage with its
output, and a two-branch parallel bundle whose branches are grouped
under the fork with the fork's results.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The projection folds Petri's public events (replay_since over the run's
stored records) and Fabro's platform records into the RunProjection the
API serves, row by row as VIEWS.md maps them. The stage key is the
execution and firing; the StageId label is node@visit, made unique with
the execution when two child invocations would share one. A stage's
first_event_seq is the milliseconds from the run's creation to its
visit.started, so the view built live equals the view rebuilt from the
records whatever order two logs' records were committed in.
The projector is the view pass and its wake-up. Records first: an append
returns before any view work; a pass reads what is committed, folds the
items past the committed positions, and writes the projection document,
the ordered stream (one stream_seq per Petri event or platform record,
with the item's own identity beside it) and the narrowed runs row in one
later transaction. Signals coalesce per run, a lost signal costs only
latency, the startup pass folds every run the view trails, and a pass
that races a platform record leaves the view alone and runs again. A
torn tail holds the view where it stands and reports the run incomplete
with the replay's error; inspect_run decides completeness once the run
recorded its finish.
The tests build the view live for the hello bundle, a command workflow
and a two-branch parallel workflow and compare it with the rebuild; drop
every wake-up and catch up by a signal and by the startup pass; crash
between the record commit and the view transaction and apply only the
suffix; restart the projector over child executions; and hold at a torn
tail.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The run's creation commits its first event on the create path, not the
append path, so the platform record hook never fired for run.created.
The store now notifies after that commit as well, and a test proves a
Petri run's lifecycle events leave platform records beside them with one
wake-up per record while a legacy run leaves none.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The server scenario answers a gate in the in-process run through the
questions API and checks the branch it routed and the cleared pending
question. The CLI scenarios drive the real worker: a gate answered through
the API over the worker's control channel, two parallel gates each bound
to their own answer, and an unanswered gate that expires with its default
and records `interview.timeout`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Integration tests in `fabro-petri` run workflows through `engine::run` on
the host sandbox: a gate answered under the posted question id, two
parallel gates each bound to their own answer, an expired question
completed as a timeout with the gate's default, an auto-approved run, a
cancelled run; a secret resolved from a vault into a command and masked in
every `petri_records` row; a command's large output round-tripped through
the `blobs` table under `blob://sha256/<hex>`; and the `hello` bundle on
the OpenAI twin with a model client over a vault that holds the key,
whose skills step searched the configured home.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A Petri run in the worker, and in the server under its test override, now
gets Fabro's platform adapters instead of the standalone defaults:
- `fabro_petri::interview`: Petri's `Interviewer` over the questions API
and the worker's control channel. A human gate's question is posted as
the `interview.started` event a legacy stage emits, keyed by an id
derived from Petri's identity (node, execution, firing, occurrence,
ask), so the API, the web app and Slack list it; the answer posted to
the questions endpoint reaches the control interviewer the adapter waits
on and is mapped onto Petri's answer. An expiry the gate reports is
completed as `interview.timeout`, a cancel as `interview.interrupted`,
and an auto-approved run answers itself. The hook points the read side
takes over are marked.
- `fabro_petri::secrets`: Petri's `SecretProvider` over the vault's token
entries, so `{{ secrets.NAME }}` resolves at spawn and is masked in every
record; a sensitive answer registers as a dynamic secret.
- `fabro_petri::blobs`: Petri's `OutputStore` over Fabro's `blobs` table,
through the server's blob store or the worker's client.
- The Fabro home the server resolved travels to the worker as
`--fabro-home`, so the skills step reads it whatever the worker's
environment says.
`engine::RunRequest` takes the interviewer, its observers, the secret
provider and the blob table from the caller; `interviewer::Unattended` is
gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A Petri run's own Fabro facts (its lifecycle before and after the
engine, a checkpoint commit, a pull request, a notification, a pairing)
are platform records in a table beside Petri's records, one typed enum
of kinds tagged on the wire, each keyed to a Petri stage where it
belongs to one and carrying the operation identity of the effect it
records. The run summary store derives the lifecycle kinds from the
legacy run events a Petri run still appends, in the event's
transaction, and calls a hook after the commit so the run's projector
can wake up.
Two more tables serve the projection that follows: the per-run
projection document with its committed positions, and the ordered
stream of everything the view consumed. The run summary store reads the
Petri projection back for the API and writes the narrowed runs row from
it without touching the legacy concurrency guard.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`fabro_petri::check` materialized the version's bundle into a temporary
directory because `Runtime::check` read the workflow and its settings
files from disk. Petri now has `Runtime::check_source`, which takes the
workflow's repository-relative path, its text and a `FileSource`, so the
bundle goes into a `frontend::MapFiles` map instead: every file at its
bundle-relative path, `workflow.toml` beside the workflow, and
`.fabro/project.toml` at the root when the caller has one. Nothing is
written to disk, and the diagnostics name the bundle-relative paths
directly, with no root to strip.
The compile inputs are unchanged: the intent's inputs, the launch model
and provider, and `petri.repository` bound by Fabro itself (the launch's
path, or `null`). An entrypoint that is not one of the bundle's files is
now `CheckError::MissingEntrypoint`; `CheckError::Materialize` goes away.
`tempfile` becomes a dev-dependency, as only the tests use it.
New tests: a version whose `workflow.toml` names `engine = "petri"` is
admitted (the new Petri pin knows the key), an unknown `[workflow]` key
is refused with `unsupported.workflow_toml.key` and named in
`workflow.toml`, the project settings are read from the map, and a
missing entrypoint is an error.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri's `fabro-integration-p1` branch at 83345a8 adds
`Runtime::check_source`, the in-memory check entry point over a
`frontend::FileSource`, and makes `[workflow] engine` a known
`workflow.toml` key in its Fabro frontend, refusing any other unknown
`[workflow]` key with `unsupported.workflow_toml.key`.
Every `petri_*` entry moves from a0d2ceb to 83345a8. The lockfile changes
only the source line of the fifteen Petri packages; no other crate moves.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
When `fabro run __run-worker` finds its run's stored spec names Petri, the
new `petri_worker` module executes it through `fabro_petri::engine` over
`HttpRunStore`, leased for a launch id the worker mints and logs at start.
`--mode start` loads the admitted graphs through the client's blob read;
`--mode resume` continues the run from its records. The worker's existing
services carry over: the control channel's cancel and SIGTERM/SIGINT cancel
Petri's root invocation politely, a lost control channel cancels the run
and is reported once it settles, and pause, unpause and steer are received
and ignored with a warning until their adapters land. The model client
comes from the worker's catalog and vault snapshot for the providers whose
credentials resolve, and the lifecycle events (`run.starting`,
`run.running`, then `run.completed` or `run.failed`) go through the client
as the legacy worker's do.
Scenario tests against the real binary: a command-only Petri run executes
in the worker a foreground server launched, its records reach
`petri_records` over the HTTP store and its lease ends with the worker; and
a run whose server and worker are both killed mid-stage resumes in a new
worker after the server restarts, with one `run.completed`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`execute_run` no longer runs a Petri run in the server process by default:
it takes the subprocess path a legacy run takes, and `worker_exited` still
releases the worker's lease when the process ends. The in-process path
stays under the handler-registry test override, so the scenario tests need
no worker binary; it now honours the managed run's execution mode.
At startup, `reconcile_incomplete_runs_on_startup` hands a Petri run the
previous server left in flight (runnable, starting, running, blocked or
paused, with no cancel pending) back to a worker instead of failing it:
`PetriRuns::release_for_restart` ends the dead worker's lease from outside,
which fences it should it still be alive, the run is asked to start again
as a resume (`run.start_requested` with `resume`, then `run.runnable`, the
pair the API's resume appends), and the managed run is registered in
resume mode when Petri's store holds the run, else in start mode. Full
workspace recovery is the plan's F3.5 and is noted in the module docs.
Tests: the restart reconcile releases the lease, rewrites the history, and
launches the worker with `--mode resume`; a worker's HTTP store leases for
its launch id over the loopback server.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A Petri run executes in the worker process, which resolves the
sandbox-driver plugins itself. The `PETRI_SANDBOX_*` variables (plugin
paths, checksum overrides, dev mode, the Docker host address and the action
host image) now have `EnvVars` names, cross the worker's environment
allowlist with `PATH`, and pass through the test harness's isolation so a
developer's plugin override reaches the servers tests start and the workers
those servers launch.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`fabro_petri::engine` is now the one assembly the worker process and the
server share: `RunRequest` takes the run's store as `Arc<dyn RunStore>` and
an `Execution`, either `Start` with the admitted graphs or `Resume` from the
run's records through `host::resume_configured`, with the same interview
observer a start installs. A resume whose record has no root invocation is
refused with a named error instead of a panic in the host. The outcome is
mapped to a `Conclusion` (succeeded, or failed with Fabro's reason and a
message) so both callers record the same terminal event.
`admission::load_with` loads the admitted graphs through any blob read, so
a worker loads them through its client; `admission::load` over the server's
`BlobStore` delegates to it.
`HttpRunStore::for_worker` takes every lease for the worker's launch id,
whatever owner Petri minted for the run runtime, and the open logs the
owner. The module docs state the rule.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The create error now names each validation diagnostic as `rule: message`
after "Validation failed", and `fabro server start --help` lists
`--engine`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
When a run's engine is Petri, the create handler hands the bundle, inputs
and launch to Petri's check instead of the legacy compile, lint and model
pinning, refuses the run with the validation error the legacy validator
uses (Petri's codes as the rules, listed in the API detail), and records
the admission on the run spec. The Fabro graph the read side displays is
parsed without validation. The scheduler executes a Petri run in the server
process through fabro_petri::engine, appending only the run lifecycle
events the read side needs (run.starting, run.running, run.completed or
run.failed); no stage or agent event is projected yet.
Scenario tests run the hello bundle on the OpenAI twin under the version
flag and a command-only bundle under the server setting, check Petri's
record agrees, and cover the refusals for an unknown attribute, an
undeclared node and an unknown model.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`check` materializes a workflow version's bundle into a temporary directory
(`Runtime::check` reads files from disk), lowers it with the run's inputs
and launch, and returns the admitted graphs or Petri's diagnostics in a
shape the server maps onto Fabro's. `admission` keeps the admitted graphs
in the blob store, named on the run spec and verified by digest on load.
`runtime` assembles the same Petri runtime at create and at execution: the
Fabro frontend with the server's settings layer, the Attractor step kinds,
the model client as the PebbleClient capability so admission pins every
model. `engine` runs the admitted graph in the server process over
SqliteRunStore under the Fabro run id, with the standalone defaults, an
interviewer that fails any question, and cancel on a token, and derives the
outcome from inspect_run over the run's record.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri's store conformance suite runs over `HttpRunStore` talking to an
axum listener on a loopback port. The suite opens runs under keys of
its own, while a key over the API is a Fabro run id the worker's token
names, so an adapter gives each suite key a fresh run with a token
minted for that run alone: the least a worker holds.
Three more tests cover what the suite cannot: the operator release
through the server's store turns the worker's handle stale; a
middleware swallows the reply of one committed append and the store's
resend leaves each record once; and two workers with owners of their
own never hold one run's lease at the same time.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The server answers the `/api/v1/runs/{id}/petri/*` endpoints from one
`SqliteRunStore` over its pool. `PetriRuns` in `AppState` keeps the
writer handle each worker opened, keyed by the run and the worker's
owner id, so the lease semantics stay the store's: the handle drops on
the worker's `release`, and every handle of a run drops when the server
observes the run's worker exit, in the subprocess wait path. Never by
timeout. A write from an owner with no held handle reopens only when
the lease row still names that owner, so a server restart or a lost
open reply recovers, and an owner the lease moved away from gets
`petri_stale_owner`.
Every endpoint is worker-scoped through the existing worker auth; a
new `RequireWorkerRunSegment` extractor covers the two-segment routes.
Store errors answer with a machine-readable code, the leased owner and
the conflict position under `meta`, and a backend failure's cause goes
to the server log rather than the worker.
A test drives a held worker through the scheduler, opens the run over
the API with its token, ends the worker, and sees the lease end.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`HttpRunStore` is Petri's `RunStore` and `RunLogs` as a worker process
reaches them: over `fabro_client::Client` with the worker's token,
against the server's SQLite store. A run key is a Fabro run id, the
`{id}` of every request, which is the plan's rule that Petri's run key
is Fabro's run id.
The lease rules are the store's. A same-owner reopen shares the live
handle in the process, and the server makes a same-owner reopen after
a lost reply the same lease. Dropping the last handle of an owner sends
`release` on the current Tokio runtime, and the store awaits every such
release before its next open, so a drop followed by an open observes
it. The server's worker-exit release is the backstop.
A reply that never arrives, a transport error or the client's request
timeout, is retried by resending the same request up to three times.
Every request is idempotent on the server, so that is safe; a reply
that did arrive is never retried. Each server error code maps back to
its `StoreError` variant, with the leased owner and the conflict
position read from `meta`.
`fabro_petri::petri` re-exports the store vocabulary for the server,
and the `test-support` feature re-exports Petri's test kit so the
server's tests can run the conformance suite over the wire. Both keep
this crate the one place that names a Petri package.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The worker shape of the integration plan (F1.3) needs a run's worker to
reach the run's Petri records over the server's API. This adds the
contract: six worker-scoped endpoints under `/api/v1/runs/{id}/petri/`
(open, release, list and append records of one log, write and read a
blob), their request and response schemas, and the generated Rust and
TypeScript clients.
A store error needs more than a code: `petri_run_leased` names the
holding owner and `petri_record_conflict` names the refused position.
`ErrorResponseEntry` gains an optional `meta` object for such
code-specific members, `ApiError` can carry it, and the client's
`ApiFailure` parses it beside the code so a caller can act on it.
Records travel as `{seq, recorded_at, record}`, the store's own unit,
with `seq` and `recorded_at` as `uint64`. The log path segment is the
log id's text (`coordinator`, `resources`, `execution <n>`), which the
generated client percent-encodes. The blob write reuses
`WriteBlobResponse`, since Petri's digest is Fabro's blob hash.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A workflow version names its engine with `engine = "petri"` in the
`[workflow]` table of `workflow.toml`, and `[server.execution] engine`
(`FABRO_SERVER_ENGINE`, `--engine`) defaults it for every version that
names none. The choice, with what Petri admitted (the lowered root graph
and its children by blob and digest), is recorded on the run spec as
`RunEngine`, carried on `run.created`, and replayed into the projection.
A legacy run's spec omits the field, so existing specs decode unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Plan item F2.1: every Fabro view of a run, the Petri event or platform
record that supplies each fact, and the identity it is keyed on. Ends
with the two completeness checks (every EVENTS.md family, every Fabro
view) and the gaps table.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri is a private repository, so Cargo's fetch of its pinned revision
needs the user's git credentials. The git CLI reads them; libgit2 does not.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`SqliteRunStore` implements Petri's `RunStore` and `RunLogs` on the pool
Fabro's other stores share. A run's existence and writer lease live in
`petri_runs`; every record of every log lives in `petri_records`, keyed by
(run, log, seq) with the record stored as JSON and read back unchanged;
blobs share the `blobs` table with `BlobStore`. The lease is taken
idempotently per owner, ends when the last handle drops or when an
operator releases it, and never by timeout; every write checks it inside
its own transaction. An append is one `BEGIN IMMEDIATE` transaction per
batch: a repeated record is accepted, a different record at a taken seq or
a seq past the head is a conflict that stores nothing.
Petri's store conformance suite passes against it, with the operator
release, lease exclusivity, a crash between appends, and blob
interoperation checked beside it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fabro runs its workflows on Petri. The six Petri packages and the testkit
are pinned by revision in the workspace manifest under `petri_*` keys, and
`fabro-petri` is the one crate that depends on them. The crate's tests run
the `hello` bundle in memory on the stub registry and a command-only
workflow on the host sandbox; both skip without the sandbox-driver host
plugin, and the sandbox-plugins CI job requires it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
lithos-llm attaches a cost to every response at the client: the codec
keeps a provider-reported cost when the provider supplies one, and the
resolver fills the catalog's price for the route when it does not.
Pebble records that priced usage on every assistant turn and sums it,
so each AssistantMessage on the stream, and the store fold's live stage
usage, already carries the cost. Fabro's catalog re-pricing of the same
tokens was redundant, and is gone.
model_usage_from_llm, with_reported_cost, and every estimate_cost call
in fabro are deleted. The pebble handler's stage_usage groups pebble's
accounts by route and sums them with Usage::saturating_add, keeping the
cost and source pebble carried, so the terminal stage.completed usage is
the live fold's sum; it no longer fails when the catalog does not know a
provider. A one-shot prompt stage records the response's own usage and
cost as lithos-llm returned it. The per-model price cards in fabro-llm's
API module stay.
Tests: the pebble handler sums Catalog and Provider costs per row and
leaves a row's and the total's cost unknown once an answer was unpriced;
the store fold shows the same tokens and cost live and at completion,
and None at both for an unpriced answer; the agent integration test
compares the whole completed Usage with the live fold, cost included;
a one-shot prompt stage on a mocked OpenAI-compatible provider reports a
Catalog cost that lithos-llm's resolver attached, with no fabro pricing.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A row with no model usage keeps its dash; a usage with tokens but no
cost, such as a model the catalog cannot price or a total with an
unpriced part, reads "unknown" rather than looking like zero.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Pebble main a39f43e26effdf99635eaf343f095c17157c9c93 (pebble #22) carries
an assistant turn's usage as Usage in the session record and moves the
record format to version 5. CodingRuntime::from_record refuses a record
in another format with UnsupportedRecord { version, supported } before
it reads the route. Fabro persists those records in SQLite for Ask Fabro
resume, and old runs get no migration, so a record written by an older
build is read back as stored and refused on the next turn.
Two tests pin that down. The store reads pebble's own version 4 fixture
back through get without a parse error and reports it unsupported. A
resumed Ask Fabro session whose stored record declares the previous
format fails its next turn with the agent_error code and the message
"session record format version 4 is not supported (this build requires
5)", runs no turn, and leaves the stored record in place.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
run.completed, run.failed, stage.completed, stage.failed, and
agent.message carry lithos-llm's Usage; the API reference navigation
names the run usage endpoint.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The run detail tab, route, query hook, and query key say usage. Every
read of input_tokens, output_tokens, total_tokens, reasoning_tokens,
cache_read_tokens, cache_write_tokens, and total_usd_micros moves to
usage.tokens and usage.cost, with lib/usage.ts replacing lib/billing.ts:
totalTokens sums the five buckets, and costSourceTag names a cost that
the provider reported or that was summed from differently sourced parts.
The Usage tab and the stage popover show that tag next to such a cost.
Test fixtures build a Usage through makeUsage; the failing set of the
web tests is unchanged from main (the same 13 environment failures).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The generated models follow the spec: TokenCounts, Cost, Usage,
ModelUsage, UsageModelRef, UsageStageRef, Speed, RunUsage,
RunUsageStage, RunUsageTotals, UsageByModel, AggregateUsage, and
AggregateUsageTotals replace the billing models, and UsageApi replaces
BillingApi. The stale billing model files are deleted.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Re-pin lithos-llm to 55add4596b861a0623d00c3a54aa5c147c8d504b and
pebble to c91810fe51aece80359b9cd8efea971af0c46925, where token usage
and cost travel together as Usage { tokens: TokenCounts, cost:
Option<Cost> }. Fabro now carries that one type everywhere it used to
carry BilledTokenCounts, BilledModelUsage, UsdMicros, or a token count
beside a cost_usd_micros.
fabro-types: billing.rs is usage.rs with ModelRef, ModelUsage { model,
usage }, sum_usage, and usage_is_empty; billing_rollup.rs is
usage_rollup.rs with ProjectionUsageStage, ProjectionUsageByModel,
ProjectionUsageRollup, and usage_rollup_from_projection. Every usage
field is named usage: StageProjection.usage and usage_by_model,
Outcome<Option<ModelUsage>>, stage.completed and stage.failed usage and
usage_by_model, prompt.completed usage, run.completed and run.failed
usage (total_usd_micros is gone), Conclusion.usage, StageSummary.usage,
Run.usage. RunSize buckets by Cost.
fabro-workflow: model_usage_from_llm prices tokens from the catalog with
a Catalog cost source, with_reported_cost keeps a provider cost, and the
pebble handler's stage_usage groups pebble's accounts by model and sums
rows with Usage::saturating_add, so a total has a cost only when every
priced part was priced. The store fold's live usage is the agent's
usage plus its descendants'.
API: the OpenAPI spec deletes BilledTokenCounts, BilledModelUsage,
CompletionUsage, CompletionCost, TokenUsage, and RunBillingSummary,
adds TokenCounts, Cost, Usage, and ModelUsage, and renames every
billing schema, property, tag, path, and operation to usage. fabro-api
reuses lithos-llm's and fabro-types' types through with_replacement,
with a round-trip test per replacement.
Old stored runs get no migration: their pebble events in the old shape
read back with zero usage, and their rebuilt projections lose agent
usage.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
When the driver reports an output loss on a streaming exec, the pebble
Environment adapter now appends one line to the stderr it hands back:
[sandbox] N output frame(s), M bytes dropped by the provider
Pebble renders the result's stderr into the tool output, so the model
and the run log both see that the command's output is incomplete rather
than reading a silently shortened stream. The adapter also logs one
`warn!` with `dropped_frames`, `dropped_bytes`, and the command's first
word, bounded, so the operator can find the event without the log
carrying the command itself.
The loss is not folded into pebble's per-stream capture stats: the
dropped frames' stream is unknown and the counts are of encoded bytes,
so attributing them to stdout or stderr would be a guess. The driver's
`truncated` flags on both captures already say the counts undercount.
`ExecOutputTail` is pebble-owned and mirrored in the OpenAPI spec, so
the run events are left alone.
Tests cover the appended line with and without existing stderr, a
lossless command over the scripted double staying unchanged, the
bounded program name, and a lossy command end to end through
`Environment::exec` over a scripted sandbox whose exec facet reports a
loss (the driver's scripted double has no knob for it).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Moves the eight sandbox-driver crates from ddb32e1 to 64c14b8, which
brings sandbox-driver PR #22 (Daytona output resync): the Daytona
plugin's encoded-exec decoder no longer fails a command on a torn frame.
It discards through the next newline, counts the loss in the new
`sandbox_driver::OutputLoss { dropped_frames, dropped_bytes }`, and
carries it as `ExecStreamingResult::output_loss` (serde default) across
the plugin wire. A loss also sets `truncated` on both captures, so
`into_complete()` refuses while `run_streaming` completes. PR #21
(supervisor generations) was already on main under the previous pin.
Nothing in fabro needed a source change for the bump. The lockfile
changes only the eight driver source lines; the unrelated windows-sys,
windows-core, and errno flips `cargo update` proposed are not taken, and
`cargo metadata --locked` accepts the result.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>