Commit graph

14 commits

Author SHA1 Message Date
Bryan Helmkamp
0d74fdf01d
Port the CLI tests to Petri runs and the run stream
The CLI's integration tests seeded runs by appending legacy run events
and waited on legacy event names. Now every seeded run is a real dry
run: the fixtures start the run through the CLI, read the run id from
its output and wait for the stream's terminal lifecycle record. Waits,
assertions and snapshots read `RunStreamItem`s (`run.finished`, the
platform `run.lifecycle` record, `derived.parsed.kind == "question"`).

Test changes:
- support.rs: `run_completed_dry_run`, `wait_for_run_finished`,
  `wait_for_lifecycle`, `wait_for_stream_item`; the `append_seeded_*`
  writers, `wait_for_event_names` and the git-backed seeded fixtures
  are gone (the checkpoint patch is not in the projection yet).
- diff.rs keeps only the help test; inspect.rs drops the git-backed
  checkpoint test; events.rs, dump.rs, create.rs, attach.rs and
  dry_run_examples.rs snapshots are re-recorded over Petri's rendering
  with redactions for epoch millis, digests and commit shas.
- run.rs: the remote foreground mock serves stream pages and a run
  state with a conclusion and a `report` stage response; the event
  history test checks `run.finished` and the terminal lifecycle item.
- runner.rs / attach.rs: question ids containing `#` are percent-encoded
  in answer URLs.

Production fixes the ports surfaced:
- petri_worker.rs: a cancelled run exits without reporting a failure.
- runner.rs: resuming a run that already finished fails its precondition
  instead of starting a worker.

Left failing on purpose, each bound to a Petri-side gap reported to the
lead rather than to the port: sandbox_cp (4), sandbox_preview and
sandbox_ssh (the projection carries no sandbox instance), the artifact
collection tests in workflow::artifacts and run.rs (no artifact
collection for Petri runs yet), and the two dump blob-ref tests (blob
refs are not visible in the inspect output).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 14:30:18 -04:00
Bryan Helmkamp
ef86ad278a
Carry usage as lithos-llm's Usage and rename billing to usage
Re-pin lithos-llm to 55add4596b861a0623d00c3a54aa5c147c8d504b and
pebble to c91810fe51aece80359b9cd8efea971af0c46925, where token usage
and cost travel together as Usage { tokens: TokenCounts, cost:
Option<Cost> }. Fabro now carries that one type everywhere it used to
carry BilledTokenCounts, BilledModelUsage, UsdMicros, or a token count
beside a cost_usd_micros.

fabro-types: billing.rs is usage.rs with ModelRef, ModelUsage { model,
usage }, sum_usage, and usage_is_empty; billing_rollup.rs is
usage_rollup.rs with ProjectionUsageStage, ProjectionUsageByModel,
ProjectionUsageRollup, and usage_rollup_from_projection. Every usage
field is named usage: StageProjection.usage and usage_by_model,
Outcome<Option<ModelUsage>>, stage.completed and stage.failed usage and
usage_by_model, prompt.completed usage, run.completed and run.failed
usage (total_usd_micros is gone), Conclusion.usage, StageSummary.usage,
Run.usage. RunSize buckets by Cost.

fabro-workflow: model_usage_from_llm prices tokens from the catalog with
a Catalog cost source, with_reported_cost keeps a provider cost, and the
pebble handler's stage_usage groups pebble's accounts by model and sums
rows with Usage::saturating_add, so a total has a cost only when every
priced part was priced. The store fold's live usage is the agent's
usage plus its descendants'.

API: the OpenAPI spec deletes BilledTokenCounts, BilledModelUsage,
CompletionUsage, CompletionCost, TokenUsage, and RunBillingSummary,
adds TokenCounts, Cost, Usage, and ModelUsage, and renames every
billing schema, property, tag, path, and operation to usage. fabro-api
reuses lithos-llm's and fabro-types' types through with_replacement,
with a round-trip test per replacement.

Old stored runs get no migration: their pebble events in the old shape
read back with zero usage, and their rebuilt projections lose agent
usage.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:31:34 -06:00
Scott Werner
86bcc8f128 Simplify CLI repository selectors and add workflow shorthand 2026-09-12 14:18:34 -06:00
Scott Werner
7c97ba1b70 Move the run-driven remote workflow test to cmd/run.rs
remote_workflow_run_starts_once_create_leaves_submitted_and_failures_do_not_refetch
lived in cmd/create.rs but drove fabro run in four of its five
iterations and asserted the start call, which is fabro run's contract.
Split it: cmd/create.rs keeps the single create invocation that must
leave the run submitted without starting it, cmd/run.rs owns the
run-driven success and failure iterations, and the workflow and remote
repository fixtures move to the shared command test support module.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 11:48:21 -06:00
Scott Werner
0e0d986634 Select CLI workflow sources and run targets independently 2026-09-12 11:48:21 -06:00
Bryan Helmkamp
705f411c56
Expect the driver's stop event in the stored event history snapshot
A dry run's stored history ends with the sandbox stop, which now carries
the driver's event instead of fabro's provider and duration fields.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 13:58:09 -06:00
Scott Werner
ccfd23104d Harden CLI run intent creation 2026-09-01 17:08:34 -04:00
Scott Werner
694d1981ff Simplify the CLI run-intent create path
Quality pass over the intent-producer changes, no behavior changes
intended:

- Move the TOML->JSON scalar conversion into fabro-types as
  toml_scalar_to_json_value, next to its inverse, with typed errors and
  round-trip tests; the CLI now calls the shared helper.
- Reuse goal_layer_from_args for --goal/--goal-file resolution instead
  of a second copy of the exclusivity check and cwd anchoring.
- Delete the dead run_manifest_args helper and the test that kept it
  compiling; preflight_manifest_args is the remaining real builder.
- Make run_target_for_environment a pure (provider, cwd) -> target
  mapping using is_clone_based(), warning at the call site, and default
  the environment id from DEFAULT_ENVIRONMENT_ID instead of a literal.
- Resolve the parent run and retrieve the environment concurrently.
- Drop the ResolvedCommandSettings pass-through struct and the
  duplicated parse-error mapping in the project settings presence read.
- Share the environment/workflow-version/git test mocks from the cmd
  test support module instead of three per-file copies.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 16:14:34 -04:00
Scott Werner
afe1133878 Fix CLI RunIntent producer CI failures 2026-09-01 16:14:34 -04:00
Scott Werner
2507a0075f Create CLI runs from immutable workflow intents 2026-09-01 16:14:34 -04:00
Scott Werner
fa796f7f24 Fix labeled run lookup in CLI test 2026-08-01 13:55:58 -04:00
Scott Werner
44eadca29c Keep removed flag coverage server-free 2026-08-01 11:47:11 -04:00
Scott Werner
3a558b225e Remove client-selected run IDs from the CLI 2026-08-01 11:47:11 -04:00
Scott Werner
47bc772f7b refactor: organize crates into three layers 2026-07-23 17:59:34 -04:00
Renamed from lib/crates/fabro-cli/tests/it/cmd/run.rs (Browse further)