Commit graph

17 commits

Author SHA1 Message Date
Bryan Helmkamp
b1d95faa57
Hand Petri the server's environment and MCP catalogs
Petri's Fabro frontend refused a bundle naming an environment it did not
declare and every MCP catalog reference, so the fixtures declared
`[environments.local]` and the server's catalogs never reached Petri.

Pin Petri at c874b86, where the frontend reads `[environments.<id>]` and
`[run.environment]` from every settings layer (bundle over project over the
host's layer, key by key), takes the environment a launch selected over the
layers, and resolves `[run.agent.mcps.<name>] id = "..."` against a catalog
the host binds. The server hands Petri its environment catalog as
`[environments.<id>]` tables of the settings layer it already passes, the
intent's environment as the launch's selection (`Launch::environment`, as
the intent overrides the bundle in Fabro's own resolution), and its MCP
catalog as `RuntimeSpec::mcp_catalog_toml`, one inline entry per definition
keyed by id. Offline validation hands Petri the seeded catalog the same
way, so `fabro validate` accepts `[run.environment] id = "local"`.

The fixtures drop the `[environments.local]` tables they carried for this;
the secrets test keeps its own, on purpose. Scenario tests cover a bundle
naming a catalog environment (its image lowered, and run on Docker when the
plugin and a daemon are there), a bundle's own table winning key by key,
the server refusing an unknown environment before Petri, and a catalog MCP
reference whose tool the agent session lists (an echo server under
`test/mcp/`).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:08:08 -04:00
Bryan Helmkamp
08b7a4fdd9
Record the snapshots the Petri pin and the record positions changed
The pin to Petri 639ce3e moves the run's format version from 6 to 7, which
the attach JSON snapshot records. The other eight snapshots had recorded the
run branch and Git identity lines before the Start stage's completion: the
order the clock gave them before "Place the run branch and git identity
records with their checkpoint" positioned the two records after the firing's
finish. That commit refreshed only two files, and these eight already
differed the same way at the commit before the pin bump; they now record the
one order every run produces.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 22:17:15 -04:00
Bryan Helmkamp
b354a2f948
Render the branch, identity, diff and artifact records and re-record the run snapshots
`run events --pretty` reads the flattened `git.identity` fields, and
shows a `run.diff` record as its summary and an `artifact.collected`
record as its path and size. The snapshot filters redact the base commit a
`Branch:` line names. The CLI snapshots now carry the `Base:` line, the
`run.branch` and `git.identity` stream items, the dry run's simulated
response and the two response files a dry run dumps.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 18:39:58 -04:00
Bryan Helmkamp
0d74fdf01d
Port the CLI tests to Petri runs and the run stream
The CLI's integration tests seeded runs by appending legacy run events
and waited on legacy event names. Now every seeded run is a real dry
run: the fixtures start the run through the CLI, read the run id from
its output and wait for the stream's terminal lifecycle record. Waits,
assertions and snapshots read `RunStreamItem`s (`run.finished`, the
platform `run.lifecycle` record, `derived.parsed.kind == "question"`).

Test changes:
- support.rs: `run_completed_dry_run`, `wait_for_run_finished`,
  `wait_for_lifecycle`, `wait_for_stream_item`; the `append_seeded_*`
  writers, `wait_for_event_names` and the git-backed seeded fixtures
  are gone (the checkpoint patch is not in the projection yet).
- diff.rs keeps only the help test; inspect.rs drops the git-backed
  checkpoint test; events.rs, dump.rs, create.rs, attach.rs and
  dry_run_examples.rs snapshots are re-recorded over Petri's rendering
  with redactions for epoch millis, digests and commit shas.
- run.rs: the remote foreground mock serves stream pages and a run
  state with a conclusion and a `report` stage response; the event
  history test checks `run.finished` and the terminal lifecycle item.
- runner.rs / attach.rs: question ids containing `#` are percent-encoded
  in answer URLs.

Production fixes the ports surfaced:
- petri_worker.rs: a cancelled run exits without reporting a failure.
- runner.rs: resuming a run that already finished fails its precondition
  instead of starting a worker.

Left failing on purpose, each bound to a Petri-side gap reported to the
lead rather than to the port: sandbox_cp (4), sandbox_preview and
sandbox_ssh (the projection carries no sandbox instance), the artifact
collection tests in workflow::artifacts and run.rs (no artifact
collection for Petri runs yet), and the two dump blob-ref tests (blob
refs are not visible in the inspect output).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 14:30:18 -04:00
Bryan Helmkamp
ef86ad278a
Carry usage as lithos-llm's Usage and rename billing to usage
Re-pin lithos-llm to 55add4596b861a0623d00c3a54aa5c147c8d504b and
pebble to c91810fe51aece80359b9cd8efea971af0c46925, where token usage
and cost travel together as Usage { tokens: TokenCounts, cost:
Option<Cost> }. Fabro now carries that one type everywhere it used to
carry BilledTokenCounts, BilledModelUsage, UsdMicros, or a token count
beside a cost_usd_micros.

fabro-types: billing.rs is usage.rs with ModelRef, ModelUsage { model,
usage }, sum_usage, and usage_is_empty; billing_rollup.rs is
usage_rollup.rs with ProjectionUsageStage, ProjectionUsageByModel,
ProjectionUsageRollup, and usage_rollup_from_projection. Every usage
field is named usage: StageProjection.usage and usage_by_model,
Outcome<Option<ModelUsage>>, stage.completed and stage.failed usage and
usage_by_model, prompt.completed usage, run.completed and run.failed
usage (total_usd_micros is gone), Conclusion.usage, StageSummary.usage,
Run.usage. RunSize buckets by Cost.

fabro-workflow: model_usage_from_llm prices tokens from the catalog with
a Catalog cost source, with_reported_cost keeps a provider cost, and the
pebble handler's stage_usage groups pebble's accounts by model and sums
rows with Usage::saturating_add, so a total has a cost only when every
priced part was priced. The store fold's live usage is the agent's
usage plus its descendants'.

API: the OpenAPI spec deletes BilledTokenCounts, BilledModelUsage,
CompletionUsage, CompletionCost, TokenUsage, and RunBillingSummary,
adds TokenCounts, Cost, Usage, and ModelUsage, and renames every
billing schema, property, tag, path, and operation to usage. fabro-api
reuses lithos-llm's and fabro-types' types through with_replacement,
with a round-trip test per replacement.

Old stored runs get no migration: their pebble events in the old shape
read back with zero usage, and their rebuilt projections lose agent
usage.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:31:34 -06:00
Scott Werner
86bcc8f128 Simplify CLI repository selectors and add workflow shorthand 2026-09-12 14:18:34 -06:00
Scott Werner
7c97ba1b70 Move the run-driven remote workflow test to cmd/run.rs
remote_workflow_run_starts_once_create_leaves_submitted_and_failures_do_not_refetch
lived in cmd/create.rs but drove fabro run in four of its five
iterations and asserted the start call, which is fabro run's contract.
Split it: cmd/create.rs keeps the single create invocation that must
leave the run submitted without starting it, cmd/run.rs owns the
run-driven success and failure iterations, and the workflow and remote
repository fixtures move to the shared command test support module.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 11:48:21 -06:00
Scott Werner
0e0d986634 Select CLI workflow sources and run targets independently 2026-09-12 11:48:21 -06:00
Bryan Helmkamp
705f411c56
Expect the driver's stop event in the stored event history snapshot
A dry run's stored history ends with the sandbox stop, which now carries
the driver's event instead of fabro's provider and duration fields.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 13:58:09 -06:00
Scott Werner
ccfd23104d Harden CLI run intent creation 2026-09-01 17:08:34 -04:00
Scott Werner
694d1981ff Simplify the CLI run-intent create path
Quality pass over the intent-producer changes, no behavior changes
intended:

- Move the TOML->JSON scalar conversion into fabro-types as
  toml_scalar_to_json_value, next to its inverse, with typed errors and
  round-trip tests; the CLI now calls the shared helper.
- Reuse goal_layer_from_args for --goal/--goal-file resolution instead
  of a second copy of the exclusivity check and cwd anchoring.
- Delete the dead run_manifest_args helper and the test that kept it
  compiling; preflight_manifest_args is the remaining real builder.
- Make run_target_for_environment a pure (provider, cwd) -> target
  mapping using is_clone_based(), warning at the call site, and default
  the environment id from DEFAULT_ENVIRONMENT_ID instead of a literal.
- Resolve the parent run and retrieve the environment concurrently.
- Drop the ResolvedCommandSettings pass-through struct and the
  duplicated parse-error mapping in the project settings presence read.
- Share the environment/workflow-version/git test mocks from the cmd
  test support module instead of three per-file copies.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 16:14:34 -04:00
Scott Werner
afe1133878 Fix CLI RunIntent producer CI failures 2026-09-01 16:14:34 -04:00
Scott Werner
2507a0075f Create CLI runs from immutable workflow intents 2026-09-01 16:14:34 -04:00
Scott Werner
fa796f7f24 Fix labeled run lookup in CLI test 2026-08-01 13:55:58 -04:00
Scott Werner
44eadca29c Keep removed flag coverage server-free 2026-08-01 11:47:11 -04:00
Scott Werner
3a558b225e Remove client-selected run IDs from the CLI 2026-08-01 11:47:11 -04:00
Scott Werner
47bc772f7b refactor: organize crates into three layers 2026-07-23 17:59:34 -04:00
Renamed from lib/crates/fabro-cli/tests/it/cmd/run.rs (Browse further)