Commit graph

313 commits

Author SHA1 Message Date
Bryan Helmkamp
0fe066d420
Fix the rustdoc link warnings and gate rustdoc in CI
Every intra-doc link `cargo doc --workspace --no-deps` warned on now
resolves or is plain code: the private constant and helper, the removed
`InterpString::resolve`, the lithos `Message`, the sandbox-driver facets,
the `RunOptions::git_author` the cutover removed, and the stale
`platform_record_for` paragraph on the platform records. The `[@REF]`
segment of `fabro run`'s help text is allowed as help, not a link. A
`Rustdoc` job runs the same command with `-D warnings`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 19:24:34 -04:00
Bryan Helmkamp
60b503322c
Delete the legacy run event log, its reducer and its types
Step 4 of the legacy executor deletion, fourth commit: with no writer
and no reader left, the legacy event log goes.

- `fabro-types`: `run_event` (`EventBody`, `RunEvent` and every props
  struct), `EventEnvelope` and the `RunEventDetail*` types are deleted.
  What the projection and the API still use moves out of the event
  vocabulary: `AgentEventProps`, `AgentSessionActivatedProps`,
  `AgentToolsAvailableProps`, `StagePromptProps`, `SessionCapability`
  and the coding event names to `agent_props`; `RunNoticeLevel` and
  `RunNoticeCode` to `notice`; `InterviewOption` beside the question
  types; `RunRunnableSource` beside the run status. `Checkpoint` is
  what Fabro records for a Petri run: `timestamp`, `current_node`,
  `git_commit_sha`; the conclusion's stage summaries derive from the
  projection's stages instead of the checkpoint's node maps.
- `fabro-store`: the Slate bridge (`RunDatabase`, the Slate `Database`,
  `keys`, `record`, `EventPayload`) and the reducer (`run_state`) are
  deleted. `Database` is the blob table and the run summary store over
  one pool; the blob store is SQLite only; the run summary store keeps
  the `runs` row a projector writes and lists, and finds the pull
  request creation candidates over `platform_records`; `build_summary`
  and `projected_usage` live in `run_summary`. The SlateDB dependency
  is gone. Test fixtures build the store from its two SQLite stores.
- `fabro-workflow`: the `event` module (the `Event` enum, its
  conversion, sink, emitter, redaction, stored fields and names),
  `runtime_store`, `StageScope` and the legacy seeding test helpers are
  deleted; the tests that seeded legacy runs read platform records or
  a projection instead.
- `fabro-sandbox` owns `GitRetryReason`.
- The server builds the store without an object store; the legacy
  `POST /runs/{id}/events` tests go, an interrupt answers
  `interrupt_unsupported` in the tests as it does in the handler, and
  the tests that read a run back through the Slate handle read its
  projection or its platform records. The projection folds a block
  that lands while the run is paused as the pause's prior block, and a
  pause or unpause clears the pending control it answers; a control
  request's check-and-append holds a per-run lock so two concurrent
  cancels record one request.
- The CLI's final output is the response of the last stage that
  produced one; the workflow tests read completed nodes from the
  succeeded stages.
- The spec's `RunCheckpoint` carries the three fields the type keeps.

Still failing until the next commits: the CLI tests that seed runs
through `POST /runs/{id}/events` or wait for legacy event names, and
the two Ask Fabro resume tests (the sandbox instance gap).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 14:09:17 -04:00
Bryan Helmkamp
3e8b2ebfcc
Run the server's lifecycle over platform records instead of run_events
Step 4 of the legacy executor deletion, first commit of several: step 4
spans commits because the legacy event log and its consumers cannot go
in one compiling change. This commit moves every writer off `run_events`;
the reducer, `EventBody`, the Slate bridge and the API's event types
still exist for the readers the next commits port or delete.

Writers:
- The server records a run's lifecycle (submitted, runnable, starting,
  running, blocked, paused, control requests and effects, the terminal
  status), its title, parent link, archive state, notices and pull
  request state as platform records (`fabro_store::platform_records`),
  through the new `server::run_records` module. Every append wakes the
  projector and waits for its pass, so the read that follows a write
  holds the record.
- Pull request creation is recorded as `pull_request.requested`,
  `pull_request.created`, `pull_request.failed`, `pull_request.linked`
  and `pull_request.unlinked`; the projection folds them into the run's
  pull request and creation state.
- Answers to questions are recorded as `interview.answered` with the
  answering principal and the answer text; the interview adapter no
  longer posts legacy `interview.*` events (`QuestionSink` is now an
  optional observer).
- The worker (`fabro run __run-worker`) records its lifecycle, notices
  and pause state over `HttpPlatformRecords`; `HttpRunStore` for the
  legacy event log and the worker's `run_store` are gone.
- `persist_created_run` appends `run.created` and `run.submitted`.

Readers:
- A stream follower (`server::stream_follower`) follows each live run's
  stream (Petri events and platform records), folds lifecycle records
  into the in-memory run state, forwards items to the global attach
  broadcast, and syncs blocked and paused from the projection.
- Slack posts questions from the projection's pending interviews,
  finishes them on `interview.answered` or `question_expired`, and sends
  lifecycle notifications with `notification.sent` dedupe.
- `GET /runs/{id}/events` and the attach endpoints serve only the run
  stream; the per-event, per-stage and `POST /runs/{id}/events`
  endpoints and their tests are deleted.
- `Database::load_run_projection` reads the Petri projection only.

Deleted with the writers:
- The SQLite blob and run-history activation migrations and their
  legacy Slate imports (`legacy_blob_import`, `legacy_run_history_import`,
  the activation backup): a greenfield server has no Slate history to
  import, and the run-history verification refused to start a server
  whose runs have no legacy events.
- `fabro-workflow`'s `operations::archive` and `operations::run_store`.
- The server's legacy-event unit tests and the CLI's `HttpRunStore` tests.

The in-process answer transport is now set after the starting and
running records land, not gated on the live status still being
`Starting` (the records already moved it).

The manifest validation test for a `run.agent.mcps.<name>` catalog
reference now expects `unsupported.workflow_toml.run.agent.mcps.reference`:
Petri's Fabro frontend has no server catalog to resolve it against.

Legacy readers still fail their tests until the next commits: the
reducer and Slate tests in fabro-store, the fabro-workflow create tests
that read the run back through the legacy store, the CLI tests seeded
through `POST /runs/{id}/events`, the CLI's legacy attach and render
paths, the sessions API, the OpenAPI conformance test, and the web
fixtures.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 12:32:06 -04:00
Bryan Helmkamp
af38d68942
Delete fabro-validate and fabro-acp; validate on Petri's check
Petri judges a workflow at admission, so Fabro's lint rules go:
`fabro-validate` (its 35 rules and the `LintRule` trait) is deleted, and
with it `fabro-acp` (only a rule and two legacy executor tests used it),
the model-resolution transform, the legacy `create`, `compile_create_run`
and `materialize_create_run` stages, and `fabro-graphviz`'s `condition`
and `fidelity` modules. `Diagnostic`, `RelatedDiagnostic` and `Severity`
move to `fabro_types::diagnostic`, the one shape every diagnostic takes.

Validation is now the same question the create handler asks. A new
server module, `petri_check`, builds Petri's check request from a
workflow bundle and the run's settings (every workflow of the bundle at
its bundle-relative path, the inputs, the run variables, the launch),
runs the check, and maps the diagnostics; Fabro adds one rule of its
own, `fabro.model.no_ready_provider`, refusing a model node when no
provider is ready. Admission, the validate and preflight endpoints and
the offline `fabro validate` all go through it:

- `validate_prepared_manifest` runs Fabro's structural pass (parse and
  transform, whose diagnostics stay) and then Petri's check, on the
  blocking pool from the handlers;
- the offline `validate_manifest` checks with no model client and with
  an unbound input as a warning (`CheckRequest.unbound_is_warning`), so
  a workflow validates before its inputs exist; a collected workflow
  before upload checks with unbound inputs as errors, as before;
- preflight resolves each LLM node's selector against the ready
  providers and the catalog for its probe, as the deleted transform did,
  and no longer probes a model Petri refused;
- the graph render endpoint needs only the structural pass;
- a run manifest now carries its `[run.goal] file`, which Petri reads
  from the bundle as it does for a version.

The transforms keep the authored model selector (`sonnet` stays
`sonnet`): Petri pins the catalog model in the admitted graph, not in
the graph Fabro displays or in the settings snapshot. Tests assert that,
and the CLI's validate, preflight and graph snapshots carry Petri's
diagnostics (`attractor.no_start`, `attractor.undeclared_node`,
`attractor.bad_on_failure`, ...) in place of the lint rules' text.

Known gaps, Petri's side: a `workflow.toml` whose `[run.environment]`
names an environment the server catalog defines but the file does not
is refused (`unsupported.workflow_toml.run.environment`), as admission
already refused it; an unbound input inside an included template
partial is a render error rather than the unbound-input warning.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 11:32:04 -04:00
Bryan Helmkamp
d90a5d9cbb
Delete fabro-core and the engine half of fabro-workflow
Every run executes on Petri, so the in-process legacy executor goes:
`fabro-core` and, in `fabro-workflow`, the handlers, lifecycle, pipeline
execution, routing, retry, conditions, node handlers, steering, agent
memory, artifacts, checkpoints, command log, and the `start`, `resume`,
`retry`, `fork`, `rewind` and `timeline` operations. The two are deleted
together because the engine half of `fabro-workflow` was the only user of
`fabro-core` and `fabro-core` the only runtime of that half; neither
compiles without the other.

Kept in `fabro-workflow`, narrowed: the parse/transform/validate/persist
pipeline and `create`, `archive`, `validate` (workflow definitions still
come from DOT and settings); the run tools (`run_tools`, moved from
`handler/llm/fabro_tools.rs`) for Ask Fabro, `fabro exec` and Petri's
host tools; the pull request pipeline (`pull_request`, moved from
`pipeline/`, for the step 0 port); Run Files' diff helpers in
`sandbox_git`; `git_identity`, `usage_rollup`, `run_status`,
`run_materialization`, `web_search` and `workflow_bundle`.

Server: `RegistryFactoryOverride` becomes `execute_in_process`;
`RunAnswerTransport::InProcess` carries only the interviewer; the
interrupt endpoint answers 501 `interrupt_unsupported` and every pair
endpoint 501 `pair_unsupported` (status lists none); rewind, fork, retry
and timeline handlers and routes are removed; the command log is served
from the stage output blob; usage rollups accumulate from the settled
projection after an in-process run as after a worker exit.

Ported while here:

- `materialize_admitted_run` materializes the goal and drops a disabled
  pull request block, as the legacy materializer did.
- A run whose admitted graph has an agent or prompt node is refused at
  create when no LLM provider is ready (`fabro.model.no_ready_provider`);
  a workflow of commands and gates needs no model and is admitted.
- The projection's question type falls back on the options, as the
  interview adapter does, so a gate with edge-label options answers as
  multiple choice.

Tests: the server scenarios (lifecycle, run completion, SSE, helpers)
run in process on Petri and assert Petri's stage labels and stream
names; the reconcile tests assert Petri's relaunch semantics; legacy
unit tests of the deleted executor are removed; three server unit tests
the removal took with it are restored; the pair fixtures go with the
pair feature. Petri test fixtures no longer name `[workflow] engine`.

Still red after this commit, all legacy consumers the next steps
delete or port: fabro-store's Slate/reducer fixtures and fabro-types
legacy JSON tests (step 4); server unit tests over legacy run events
(retry endpoints, list_run_events, artifacts, per-event pause/unpause,
run history activation, legacy sandbox fixtures) (steps 3-4); CLI tests
that parse legacy event envelopes, the legacy `events`/`attach`/`diff`/
`dump`/`inspect` snapshots, `run rewind`/`run fork`, the ACP and
git-identity workflow tests, and the runner tests that drive the legacy
worker by hand (steps 3-4); the web app's Petri fixtures still carry
`engine` (regenerate with `FABRO_CAPTURE_PETRI_FIXTURES` in step 4).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 10:44:41 -04:00
Bryan Helmkamp
a36bea15d2
Remove the engine flag: every run is a Petri run
Delete `Engine`, `RunEngine`, `[workflow] engine`, `[server.execution]
engine`, `FABRO_SERVER_ENGINE` and `fabro server start --engine`. The run
spec records what Petri admitted as `admission: PetriAdmission`; the
create handler always admits through `Runtime::check`; `execute_run`
always launches the Petri worker (or executes in process under the test
override); the CLI runner takes only the Petri worker path, and its
legacy control arm, artifact uploader, signal pause handlers and
credential helpers go with it. The CLI's `attach` and `events` read the
run stream only.

Two gaps this surfaced are closed here: the check adapter binds the
server's run variables as Petri compile variables (`{{ vars.* }}` in a
prompt no longer fails admission), and deleting a run removes its Petri
records, lease, platform records, projection and stream.

Tests: the config engine tests are replaced (an engine key is unknown),
the API round-trip test covers `PetriAdmission`, the server and CLI
Petri scenarios drop their engine settings, and the API tests that read
legacy event names now read the run stream or the session events. The
remaining red tests are fixtures and scenarios of the legacy executor
and the legacy event store (`fabro-store` `slate` and `run_state`,
`fabro-types` legacy `run.created` JSON, the server's handler-registry
scenarios, the CLI dry-run snapshots), which the next steps of the F4.3
series delete or port.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 09:58:21 -04:00
Bryan Helmkamp
d2f0e70dca
Run a workflow through Petri in the server, behind the engine flag
When a run's engine is Petri, the create handler hands the bundle, inputs
and launch to Petri's check instead of the legacy compile, lint and model
pinning, refuses the run with the validation error the legacy validator
uses (Petri's codes as the rules, listed in the API detail), and records
the admission on the run spec. The Fabro graph the read side displays is
parsed without validation. The scheduler executes a Petri run in the server
process through fabro_petri::engine, appending only the run lifecycle
events the read side needs (run.starting, run.running, run.completed or
run.failed); no stage or agent event is projected yet.

Scenario tests run the hello bundle on the OpenAI twin under the version
flag and a command-only bundle under the server setting, check Petri's
record agrees, and cover the refusals for an unknown attribute, an
undeclared node and an unknown model.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:27:44 -04:00
Bryan Helmkamp
c383b6a70b
Add the engine flag and record the engine on the run spec
A workflow version names its engine with `engine = "petri"` in the
`[workflow]` table of `workflow.toml`, and `[server.execution] engine`
(`FABRO_SERVER_ENGINE`, `--engine`) defaults it for every version that
names none. The choice, with what Petri admitted (the lowered root graph
and its children by blob and digest), is recorded on the run spec as
`RunEngine`, carried on `run.created`, and replayed into the projection.
A legacy run's spec omits the field, so existing specs decode unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:02:47 -04:00
Bryan Helmkamp
27f16f89c4
Delete fabro's own pricing: lithos-llm prices every response once
lithos-llm attaches a cost to every response at the client: the codec
keeps a provider-reported cost when the provider supplies one, and the
resolver fills the catalog's price for the route when it does not.
Pebble records that priced usage on every assistant turn and sums it,
so each AssistantMessage on the stream, and the store fold's live stage
usage, already carries the cost. Fabro's catalog re-pricing of the same
tokens was redundant, and is gone.

model_usage_from_llm, with_reported_cost, and every estimate_cost call
in fabro are deleted. The pebble handler's stage_usage groups pebble's
accounts by route and sums them with Usage::saturating_add, keeping the
cost and source pebble carried, so the terminal stage.completed usage is
the live fold's sum; it no longer fails when the catalog does not know a
provider. A one-shot prompt stage records the response's own usage and
cost as lithos-llm returned it. The per-model price cards in fabro-llm's
API module stay.

Tests: the pebble handler sums Catalog and Provider costs per row and
leaves a row's and the total's cost unknown once an answer was unpriced;
the store fold shows the same tokens and cost live and at completion,
and None at both for an unpriced answer; the agent integration test
compares the whole completed Usage with the live fold, cost included;
a one-shot prompt stage on a mocked OpenAI-compatible provider reports a
Catalog cost that lithos-llm's resolver attached, with no fabro pricing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:46:44 -06:00
Bryan Helmkamp
e9ee0aaeaa
Merge remote-tracking branch 'origin/main' into one-usage-type
# Conflicts:
#	Cargo.lock
#	Cargo.toml
2026-09-14 12:40:49 -06:00
Bryan Helmkamp
ef86ad278a
Carry usage as lithos-llm's Usage and rename billing to usage
Re-pin lithos-llm to 55add4596b861a0623d00c3a54aa5c147c8d504b and
pebble to c91810fe51aece80359b9cd8efea971af0c46925, where token usage
and cost travel together as Usage { tokens: TokenCounts, cost:
Option<Cost> }. Fabro now carries that one type everywhere it used to
carry BilledTokenCounts, BilledModelUsage, UsdMicros, or a token count
beside a cost_usd_micros.

fabro-types: billing.rs is usage.rs with ModelRef, ModelUsage { model,
usage }, sum_usage, and usage_is_empty; billing_rollup.rs is
usage_rollup.rs with ProjectionUsageStage, ProjectionUsageByModel,
ProjectionUsageRollup, and usage_rollup_from_projection. Every usage
field is named usage: StageProjection.usage and usage_by_model,
Outcome<Option<ModelUsage>>, stage.completed and stage.failed usage and
usage_by_model, prompt.completed usage, run.completed and run.failed
usage (total_usd_micros is gone), Conclusion.usage, StageSummary.usage,
Run.usage. RunSize buckets by Cost.

fabro-workflow: model_usage_from_llm prices tokens from the catalog with
a Catalog cost source, with_reported_cost keeps a provider cost, and the
pebble handler's stage_usage groups pebble's accounts by model and sums
rows with Usage::saturating_add, so a total has a cost only when every
priced part was priced. The store fold's live usage is the agent's
usage plus its descendants'.

API: the OpenAPI spec deletes BilledTokenCounts, BilledModelUsage,
CompletionUsage, CompletionCost, TokenUsage, and RunBillingSummary,
adds TokenCounts, Cost, Usage, and ModelUsage, and renames every
billing schema, property, tag, path, and operation to usage. fabro-api
reuses lithos-llm's and fabro-types' types through with_replacement,
with a round-trip test per replacement.

Old stored runs get no migration: their pebble events in the old shape
read back with zero usage, and their rebuilt projections lose agent
usage.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:31:34 -06:00
Bryan Helmkamp
3b8d712edb
Delete a live Daytona test's sandbox even when the test panics
The live Daytona tests create a provider sandbox and delete it on their
last line, so any panic or failed assertion before that line leaks a
running, billed sandbox. Two leaked that way on 2026-09-14 when a
sandbox-driver decoder flake panicked daytona_playwright_mcp_sandbox_transport.

Add fabro_sandbox::test_support::DeletedOnDrop, a guard that owns the
RunSandbox (Deref keeps the tests reading unchanged), offers an explicit
delete(self) for the happy path, and deletes from Drop otherwise. The
drop-time delete runs on its own thread and runtime because the test's
runtime may be unwinding. It reconnects the provider through the
ProviderAccess the test built the sandbox with, because the sandbox's
own handle pools HTTP connections whose tasks live on the test's runtime;
a live check of that path timed out after 10s.

Every Daytona test that creates a sandbox now holds it through the guard.
Unit tests over the scripted double prove delete-on-drop runs once, an
explicit delete runs once, and a panic inside catch_unwind still deletes
with and without a runtime.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 11:01:24 -06:00
Scott Werner
012b556367 Remove tests and fixtures tied to retired manifest fields 2026-09-13 08:55:55 -06:00
Scott Werner
ea11538038 Retire legacy manifest run creation 2026-09-13 08:49:50 -06:00
Bryan Helmkamp
f1118eead2 Delete the agent mirrors and move failover to prompt stages
Pebble's stream is the agent event contract. The run's own agent.mcp.ready,
agent.mcp.failed, and agent.mcp.disconnected events, which mirrored pebble's
McpServer* events, are gone with their props, the sink arms that emitted
them, and their conversion and naming entries; pebble's stored
agent.mcp.server.* events are the only record and feed the stage's fold.

The sink no longer mirrors RouteFailover onto agent.failover either: an
agent stage's moves are pebble's agent.route.failover. The event is now
prompt.failover, emitted only by a one-shot prompt stage that walks its
fallback plan itself, and its props are trimmed to the two routes, the
attempt, and the error; nothing read the rest.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:36:24 -06:00
Bryan Helmkamp
0c0e589a78 Bill a failed agent stage what it spent
An agent stage that failed billed nothing: the backend returned a bare
error and the outcome built from it carried no usage. A terminal failure
now becomes the stage's failed outcome from the same fold that bills a
completed stage, with the tree's usage, the rows by model, the files it
wrote, and its active time; stage.failed carries billing and
billing_by_model and the store keeps both. Cancellation and retryable
failures still go up as the error.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:47 -06:00
Bryan Helmkamp
df8762663b Bill an agent stage's whole session tree from one fold
One usage rule: a stage's usage is its session tree's, the root and every
subagent, live and at completion. The worker's event sink folds pebble's
SessionProjection over the events it records and the stage's billing and
files come from that fold at stage end, so the completed values are what
the run showed live. The store's live usage is the fold's tree usage, and
completion brings the catalog's price for the same tokens instead of
resetting them to the root's.

Fabro keeps catalog pricing: the root at its route, each descendant at its
own route where the catalog knows it and at the root's otherwise, a
provider-reported cost standing in where pebble has one. The rows travel
as billing_by_model on stage.completed and the stage projection, and the
billing rollup splits by_model by them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:47 -06:00
Bryan Helmkamp
ecafe1e807 Store ProcessingEnd and the mirrored pebble events in the run log
Pebble's SessionProjection reads ProcessingEnd to complete a prompt and mark
the session idle, so a projection rebuilt from the run's log needs it: one
small event per prompt. The four pebble events the sink mirrored onto
fabro's own agent.failover and agent.mcp.* are now stored verbatim as well,
so the fold sees the route moves and the MCP outcomes; the mirrors stay
until every reader is on the projection.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:07 -06:00
Bryan Helmkamp
697d8622e9
Merge pull request #843 from fabro-sh/remove/run-metadata-branches
Some checks are pending
TypeScript / Build (push) Waiting to run
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox plugins (stdio) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
TypeScript / Typecheck (push) Waiting to run
TypeScript / Test (push) Waiting to run
Remove Git run metadata branches
2026-09-12 16:41:24 -06:00
Bryan Helmkamp
e5d5c534ab
Merge remote-tracking branch 'origin/main' into remove/run-metadata-branches
Resolve conflicts between the metadata-branch removal and the
sandbox-driver adoption on main:

- fabro-sandbox docker.rs, sandbox.rs, daytona/mod.rs: take main's driver
  rewrite. The Sandbox trait is gone, so the PR's push_token_source
  removal now applies to RunSandbox instead; drop that accessor and the
  RepoCredentials::source helper that only served it.
- run_metadata.rs: keep deleted. Main's edits there were adaptations to
  the driver API and the run git identity field.
- lifecycle/git.rs, finalize.rs: keep the PR's removal of metadata
  snapshots and write_finalize_commit; carry main's RunSandbox,
  GitRetryPolicy, git_identity, local_sandbox, and test catalog changes.
- sandbox_git.rs: take main's version and drop the shadow_sha parameter
  and Fabro-Checkpoint trailer.
- git_integration.rs: remove meta_branch from the new git identity test.
- Cargo.toml: main's dependency set with fabro-dump kept as a
  dev-dependency.
- checkpoints.mdx: keep both the git identity paragraph and the durable
  execution state section.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:28:01 -06:00
Bryan Helmkamp
9550ff0806
Cover the failover continuation and the stopped failover
The existing failover test asserts the backup continued the turn after
the primary committed a tool result. A new test exhausts a two-route
chain and checks that one agent.failover and one
agent.route.failover.stopped are stored on the work stage, the stop
after the error it reports, with the exhausted reason and the failing
route. The events catalog documents both.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:26:32 -06:00
Bryan Helmkamp
57a746235e
Pin pebble to main after lithoscomputer/pebble#12
Pebble's RouteFailover event now describes the route that failed and how
the new route carried the prompt on. Record the continuation on fabro's
agent.failover event as an optional string (replay_prompt or
continue_turn); events written before it existed, and one-shot prompt
stages that walk the plan themselves, read as absent. The failed route's
usage, cost, and timing are not mirrored: the stage's totals already
include them through the prompt report, and no fabro run event carries
per-route usage yet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:26:32 -06:00
Bryan Helmkamp
7879e6d223
Merge pull request #856 from fabro-sh/brynary/run-git-identity
Resolve one Git identity per run and inject it into every workflow command
2026-09-12 14:23:30 -06:00
Bryan Helmkamp
ff7821dcd2
Adopt pebble's MCP disconnect event and startup timing
Pebble reports an MCP server whose connection closed mid-session once,
as McpServerDisconnected, and carries startup_ms on McpServerReady and
McpServerFailed. The workflow event sink mirrors the disconnect onto a
new agent.mcp.disconnected run event shaped like agent.mcp.failed, and
passes startup_ms through on agent.mcp.ready and agent.mcp.failed. The
raw pebble event is not stored for these, so the timing would otherwise
be dropped at the boundary.

The stage projection's McpServerStatus gains a `disconnected` kind next
to `ready` and `failed`. The fold keeps the server's tool count and
sticky invoked flag and only moves the status. The OpenAPI
McpServerStatus oneOf gains McpServerStatusDisconnected, and the
fabro-api round-trip test covers its JSON shape. A new
session_projection_parity test folds the same MCP events through
pebble's SessionProjection and fabro's stage projection and compares
them, including pebble's `disconnected`.

ToolErrorKind::Timeout needs no fabro change: the kind is stored as
pebble serializes it and never matched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 12:39:31 -06:00
Bryan Helmkamp
1aee59585b
Pin pebble to mcp-embedder-events and accept startup_ms on MCP outcomes
Pebble's McpServerReady and McpServerFailed events now carry startup_ms.
The workflow event sink destructured both variants by name, so it stops
listing every field. The pin is temporary: it moves to pebble main once
lithoscomputer/pebble merges the branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 12:29:36 -06:00
Bryan Helmkamp
ee6576cff7
Resolve one Git identity per run and inject it everywhere
A run now resolves a single author and committer identity once, after its
GitHub credentials are selected and before anything can commit, and uses it
for every commit it creates. Resolution order: a complete explicit
`run.git.author`; the run's GitHub App bot account
(`<slug>[bot] <id+slug[bot]@users.noreply.github.com>`); the authenticated
user of the run's PAT; the generic `Fabro <noreply@fabro.sh>`. A partial
explicit author overlays the fields it supplies. Only the selected
credential is consulted; a failed lookup is a setup error. A standalone
installation token falls back to the generic identity with a warning.

The resolved identity is carried on `RunOptions` and `EngineServices`,
recorded as a `git.identity.resolved` event and `RunProjection.git_identity`
so resume reuses it, and exposed through the run state API. Engine
checkpoints and metadata commits read it through `RunOptions::git_author`.
Every workflow execution path receives it as `GIT_AUTHOR_NAME`,
`GIT_AUTHOR_EMAIL`, `GIT_COMMITTER_NAME`, and `GIT_COMMITTER_EMAIL`, applied
last so it wins over inherited host variables and `[run.environment]`
entries: prepare steps, command stages, native agent shell tools, and ACP
launches. The identity is injected even without a Git origin, and the old
local `git config user.*` write is removed.

fabro-github gains `GET /user` and `/users/{slug}[bot]` lookups with mocked
tests for success, unauthorized, malformed, and transient cases. Real-Git
integration tests commit in the primary checkout, a clone, and a fresh
repository under conflicting local config, `[run.environment]`, and host
variables, and prove concurrent runs do not leak identities. CLI workflow
tests cover host script stages and ACP launch env through `fabro run`.

Docs and generated option metadata now describe the credential-derived
defaults instead of the stale `fabro`/`fabro@local` values.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 11:35:46 -06:00
Scott Werner
21a5e5b86f Simplify run creation to registered workflow versions 2026-09-12 11:04:15 -06:00
Scott Werner
c67c60eeba Create run tools from immutable workflow versions 2026-09-12 10:57:22 -06:00
Scott Werner
d611ef2bb1 Port workflow version creation to Pebble native tools 2026-09-12 10:23:34 -06:00
Bryan Helmkamp
957fc97c5c
Merge origin/main into pebble-agent-loop
Main merged the sandbox-driver adoption (#849) in a later form than this
branch was stacked on: the driver's own exec types replace fabro-sandbox's,
shell quoting moved to fabro-util, the sandbox lifecycle collapsed, and the
driver's events are stored as run events. This branch had deleted
`fabro-agent` and put the coding agent, the environment adapter, and the
steering hub on pebble.

The resolution takes main's sandbox API and re-applies pebble on top: the
`RunSandbox` `Environment` adapter moves to `pebble_environment.rs` (main's
`environment.rs` is the sandbox spec) and runs commands through `ExecSpec`
and `ExecControls`, feeding pebble's output sink from the driver's; the
driver-era `sandbox.*` names leave the known-event list, as on main, so a
stored event with that name and no driver shape is `Unknown` rather than an
error; `program_exit_code` matches pebble's non-exhaustive termination; the
Docker and Daytona smokes use main's constructor and credentials; the
remaining `fabro_agent` paths point at fabro-sandbox.

Pebble's `mcp` feature pins sandbox-driver, and the preview-url trait
objects only cross when both sides name one revision, so pebble moved to
main's `a92c0db6` (lithoscomputer/pebble#10) and fabro pins that pebble
revision until it lands on pebble main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 09:40:00 -06:00
Bryan Helmkamp
a33c9069cf
Run steering over pebble's bus and keep only fabro's attribution
Pebble's `SteeringBus` now owns the map of live sessions, the buffer for
steers that arrive between sessions, fan-out of steers and interrupts, the
close-the-door detach, and the hold that keeps a paired session open. The
hub keeps what only fabro knows: pair records, principals, stage ids, and
the run events that put bus activity on the run's stream in the order its
consumers expect. `PebbleControlHandle` is gone, since the coding agent's
control handle is a bus session natively; the ACP session joins the bus
through a small adapter and carries pebble's steering message end to end,
so a human steer keeps its author on the ACP `agent.steering.injected`
event. A pair message that evicted an older steer is now accepted and the
eviction recorded, where before it was queued and reported as refused.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 20:35:31 -06:00
Bryan Helmkamp
6592505d0e
Take environment helpers, compaction accounting, and search providers from pebble
Steps 4, 6, and 7 of .ai/plans/pebble-absorbs-embedder-concerns.md,
pinning pebble fc907a1 with its `search-providers` feature.

The sandbox's `Environment` adapter uses pebble's `environment::support`
for the glob grammar check, the tree-order sort of a listing, and the
capture accounting, in place of its own copies; the adapter itself stays.
The stage reads the compactions a prompt performed from the report, as a
breakdown of the usage it already billed. `web_search.rs` keeps the
vault-backed credentials and the Brave-over-Venice preference, and hands
pebble's `Brave` or `Venice` provider a fabro HTTP client; the providers
and their tests are pebble's now.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 19:51:12 -06:00
Bryan Helmkamp
d4e483925c
Let pebble discover memory and skills, and name the reuse rules
Steps 5 and 8 of .ai/plans/pebble-absorbs-embedder-concerns.md, pinning
pebble 49da137.

Agent stages and `fabro exec` ask pebble for the profile's instruction
files from the repository root down to the working directory
(`MemoryDiscovery::from_git_root`), which fabro lacked: it read the
working directory alone. Skill directories are pebble's to resolve too:
the user's skills directory, then `.fabro/skills` and `skills` under the
repository root. Prompt stages keep reading the working directory alone,
through the same discovery and loader, so `agent_memory.rs` keeps only
that call; the filename table is pebble's now.

A retained thread's export comes from `export_for_reuse`, which closes
the session and hands back an export whose cursor is already past the
close, in place of export, shutdown, and a cursor advance by hand. Ask
Fabro resumes a stored record with `resume_after`, the rule it applied
under the older name.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 19:40:36 -06:00
Bryan Helmkamp
64c578d044
Take MCP servers from pebble
Step 3 of .ai/plans/pebble-absorbs-embedder-concerns.md, pinning pebble
6cdb30a with its `mcp` feature.

Pebble starts the stage's MCP servers while the agent is built, registers
their tools under `mcp__{server}__{tool}` with `ToolSource::Mcp`, and
closes them with the agent, for all three placements: a child of the run
worker over stdio, a server over HTTP (streamable or SSE), and a server
launched in the run sandbox and reached through the sandbox's preview
URL. `fabro_mcp::pebble::pebble_server` maps `McpServerSettings` onto
pebble's `McpServer`, keeping fabro's `/sse` path for sandbox-hosted SSE
servers; `RunSandbox::port_routes` hands pebble the driver's `PreviewUrls`
facet as the route to a sandbox port. The stage's event sink mirrors
`McpServerReady` and `McpServerFailed` onto the run's `agent.mcp.ready`
and `agent.mcp.failed` events, as it mirrors `RouteFailover` onto
`agent.failover`, and stores no second copy of a mirrored fact.

Deleted: `sandbox_mcp.rs`, the MCP branches of `pebble.rs` and `fabro
exec`, and fabro-mcp's client, connection manager, HTTP helpers, and SSE
transport, whose tests moved to pebble. fabro-mcp keeps the settings
re-export, the mapping, and a stdio client behind `test-support` for the
tests of fabro's own MCP server. `fabro exec` reports each server's
outcome from the agent's snapshot. The Daytona Playwright live test now
drives the sandbox-hosted server through an agent.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 19:26:41 -06:00
Bryan Helmkamp
b5fd907c3d
Take files touched and route failover from pebble
Steps 1 and 2 of .ai/plans/pebble-absorbs-embedder-concerns.md, pinning
pebble 1a5abe4.

Files touched come from `PromptReport`: pebble computes them from every
successful write, edit, and patch across the prompt, subagents included,
so the event-fed `FileTracking` and the tracking half of
`WorkflowEventSink` go. The stage unions the reports of its prompts.

Route failover is pebble's. The stage resolves its plan from the catalog
as before and hands pebble the remaining routes through
`fallback_routes`, each with its controls and the stage's output limit.
Pebble keeps the conversation, moves it to the next route, requeues
pending steering, and continues the prompt; the stage's plan follows the
route the report says the prompt ended on, re-activates the session
there, and mirrors pebble's `RouteFailover` as the run's `agent.failover`
event with the same payload as before. `prompt_with_failover`,
`resume_agent_on_route`, and the route bookkeeping in `LiveAgent` go.
One-shot prompt stages still walk the plan themselves.

`agent.route.failover` and `agent.tool.rounds.exhausted` join the derived
event names; both variants were falling back to `agent.event`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 18:47:12 -06:00
Bryan Helmkamp
8d32b472db
Pin the merged pebble main and adopt the pebble features petri already uses
Step 0 of .ai/plans/pebble-absorbs-embedder-concerns.md. Pebble's
fabro-exec-sink-max-turns branch is merged into its main (pebble 4db661c,
on the embedder-concerns branch until it lands), so fabro and petri pin
one line. Both prompt budgets survive the merge because they count
different things.

- Agent hooks use `with_max_tool_rounds`, which is what the
  `max_tool_rounds` setting names and what petri does: `max_tool_rounds`
  model turns may run, and the hook proceeds on a turn that still asks for
  tools, so pebble gets `max_tool_rounds - 1` rounds and `ToolRoundsExhausted`
  fails open. Zero rounds proceeds without an agent, as the old loop did.
  Agent stages set no turn budget; the stage timeout and stall watchdog
  bound them.
- Prompt stages load project memory through pebble's `ProjectMemory`, the
  loader agent stages already run over `with_memory_files`, instead of a
  hand-rolled copy of its budget, deduplication, and truncation.
- `fabro_sandbox::SecretRedactor` puts fabro's secret scanner on pebble's
  text seams (process output tails, failed tool messages) for agent
  stages, Ask Fabro, hook evaluators, and `fabro exec`. The final redaction
  pass over every stored `RunEvent` stays; this does not replace it.
- Agent stages state their compaction policy explicitly: the 80 percent
  threshold and six preserved turns fabro's own agent loop applied.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 18:16:03 -06:00
Bryan Helmkamp
6599f7cdc0
Cover the pebble agent loop with workflow-level tests
Each behaviour the pebble backend owes a run now has one test that drives
it through the workflow engine against a scripted OpenAI-compatible model:

- an agent stage under every harness profile (openai, anthropic, claude-5,
  gemini, kimi) writing a file with that profile's own tool spelling, with
  the event sequence, files touched, response, usage, and cost checked; the
  codex vocabulary applies a patch through the OpenAI twin's custom tool
  call, which `TwinToolCall::custom` now scripts
- steering delivered mid-stage, an interrupt with a steer, run cancellation,
  and the executor-enforced stage timeout
- a question answered through the interviewer, a subagent whose events carry
  its parent's session id, and an MCP tool served by a stdio server
- failover to a second provider after a tool ran, continuing the recorded
  conversation without running the tool again
- a failing event sink ending the stage with the sink's error
- Ask Fabro resuming a stored record across turns, with the cursor moved
  past the run's event log when the record's own cursor fell behind
- Docker and Daytona smokes running an agent stage through the provider
  sandboxes

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 16:32:52 -06:00
Bryan Helmkamp
350abac014
Build the local sandbox through the one provider path
SandboxSpec had a Local variant beside the provider spec, and a local
sandbox was created by hand over a bare Host provider: no workspace, no
provider connection, its own reconnect, and its own push rule for the
designated directory. The local kind is now one more SandboxSpec:
SandboxSpec::local names the directory on a HostDirectory spec with a
skip clone, and provider_sandbox builds it like a plugin kind, creating
the directory when missing since the Host provider requires it to exist.
Every RunSandbox carries a workspace; a handle wrapped as is gets the
workspace of its own working directory.

The push rule is one rule for every checkout: a checkout fabro cloned
pushes with the credentials it was cloned with, and any other checkout
pushes when it has an origin, with whatever credentials it carries. A
local run therefore pushes the same way before and after a resume;
before, a reconnected local sandbox carried an attached workspace that
never pushed while a fresh one did.

Reconnect uses the recorded id for every kind. The recompute of a local
id from its directory, kept for records written before directories had
ids, is gone, and test fixtures that wrote made-up local ids derive them
through test_support::local_sandbox_id instead. A local run's record now
carries its workspace layout like every provider-chosen directory, and
the sandbox.initializing event precedes the driver's create events for
local as for every other kind.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 16:28:17 -06:00
Bryan Helmkamp
981253f990
Collapse the run sandbox's lifecycle surface
RunSandbox grew its lifecycle methods one adopter at a time and ended up
with several names for each step. Reconnecting from a run record had
four entry points (reconnect, reconnect_for_run,
reconnect_for_run_with_events, reconnect_driver_for_run) that all
forwarded to the last one. Bringing a sandbox back had two (start and
activate) over the same make_ready, and releasing it had two (delete and
cleanup) over the same release. Two more methods had no callers at all:
set_autostop_interval, which nothing set after the driver took over
lifecycle timers, and resume_setup_commands, which resume stopped using
when checkout moved to the git facet.

There is now one of each. reconnect_for_run takes the record, the
provider access, an optional run id, and an optional event context;
callers that need none pass None. activate is the single "make usable"
step: a running sandbox only learns its platform when it has not yet, a
stopped or paused one is started and its Bash verified, and resume calls
it like every access-time caller. delete is the single release; for a
designated host directory it frees the handle and leaves the directory in
place, as cleanup did. The tests and server call sites follow the
renames; behavior is unchanged except that resuming an already running
sandbox no longer re-runs the Bash probe.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 15:05:06 -06:00
Bryan Helmkamp
8771ef6d8e
Move the Read tool's line numbering and the shell quoting wrapper out of fabro-sandbox
format_lines_numbered renders a file the way the agent's Read tool shows
it: every line prefixed with its number, from an offset for a limit. That
is the agent's presentation, not a sandbox concern, and it only lived in
fabro-sandbox so RunSandbox::read_file(path, offset, limit) could call
it. The function now lives in fabro-agent next to the Read tool, with its
tests, and the Read, ReadManyFiles, and Kimi ReadFile tools number the
text they get from read_file_text themselves. RunSandbox::read_file goes
away; the sandbox returns bytes or text and nothing else.

fabro-sandbox also re-exported shell_quote through a one-line wrapper so
callers could reach it from the sandbox crate or from fabro-agent. The
audited implementation is fabro_util:🐚:shell_quote; the six
importers now use it directly and the wrapper and both re-exports are
gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 15:00:17 -06:00
Bryan Helmkamp
18a3c4741e
Run agent stages, Ask Fabro, and fabro exec on pebble's CodingAgent
Replace fabro's hand-written agent loop with pebble's `CodingAgent` and
delete the `fabro-agent` crate.

Workflow: `PebbleBackend` builds one agent per stage over `RunSandbox`,
binds the stage's hooks as tool middleware, the interviewer as the
human-input provider, and a durable `EventSink` that writes every agent
event through the run event log before the agent goes on. Full-fidelity
threads continue across stages through `export`/`resume_from_export`.
Model failover takes the session record after the failed prompt and
continues it on the next route with `ResumeMode::UseModel`, so no tool
effect repeats. The steering hub targets pebble's control handle, with
a steering lease holding completion open while a human is paired.

Events: `EventBody::Agent` carries pebble's `CodingAgentEvent` envelope;
the per-variant bodies, the transcript projection, and the fabro-only
context-window, tool-summary, and skill types are gone in favor of
pebble's. The OpenAPI schemas, generated Rust and TypeScript clients,
and web readers follow.

Ask Fabro: the session runs a `CodingAgent` under a read-only permission
policy and a system prompt transform. Its conversation lives in a new
`run_session_records` table and resumes on the recorded model with the
event cursor advanced past the run log.

`fabro exec` builds the same agent over a local sandbox with pebble's
permission middleware and an interactive approval service.

The catalog fills in `metadata.agent.profile` for operator providers
that declare none, so pebble's lookup is the one resolution path.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 14:19:15 -06:00
Bryan Helmkamp
895735a829
Redact the driver's event ids in CLI snapshots and leave a local sandbox unnamed
The attach snapshots now carry the driver's create events, whose event
source and operation ids are minted per process and whose durations run
to the nanosecond, and the local sandbox's id is derived from a temporary
directory; the shared snapshot filters cover all three. A local sandbox's
ready event no longer names that id: the record already holds the
directory, and the id is nothing a person reads.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 13:55:45 -06:00
Bryan Helmkamp
d986fc9751
Keep the sandbox driver's events whole as run events
A bridge translated the driver's events into thirteen lifecycle variants
of fabro's own (start, stop, and delete phases, image pulls, snapshot
builds) and dropped everything else the driver reported, pairing an image
pull's first progress report with the create's completion to invent a
duration. The driver's event is now stored as the run event itself, under
a name derived from it: subject, action, and phase (sandbox.stop.completed,
sandbox.create.progress for an image pull, snapshot.create.started), or
<subject>.state and <subject>.notice. Every operation the driver performs
on the run's sandbox lands on the run, including creates and state
observations the bridge skipped. The CLI reads image pulls and snapshot
builds from the driver's event for its setup progress and pretty output,
the thirteen variants and their props go, and a run stored under the old
names still reads as an unknown body. Checkpoint file numbers in a dump
shift because the run records more events before each checkpoint.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 13:32:54 -06:00
Bryan Helmkamp
7d9d2bf5f3
Run fabro's git plumbing through the driver's verbs
Fabro assembled its own hardened git command lines (maintenance, hooks,
fsmonitor, path quoting, signing, the file transport, external diff
drivers) in three crates and parsed raw diff, numstat, cat-file, and log
output itself. The driver's git facet now carries fetch, rev-parse,
ancestry, diff entries, numstat, patch, log, blob sizes and contents,
config, untracked files, and stage-all, hardened by default and typed, so
the checkpoint commit, the run diffs, the Run Files listing and blob
reads, the commit log, the fork fetch, the agent's changed-files
detection, and the git identity setup go through it. The parsers and the
command prefixes go; the per-run capability probe keeps its own plumbing
script. Checkpoint commits never run repository hooks now, so
skip_git_hooks and commit_timeout are accepted for compatibility only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 12:55:39 -06:00
Bryan Helmkamp
22cd5d3382
Let the driver retry git operations and decide from the credential's age
Fabro carried its own retry loop, its own reading of what a git failure
class means for the credentials in hand, and a credential context derived
from the token snapshot. The driver now owns the loop and the decision:
a rejected credential retries only while its mint time is within the
replication horizon, a remote that could not be reached retries on its
own, a static credential fails fast, and an operation whose outcome is
unknown is never replayed. Fabro keeps its budgets as retry policies
(clone, repository probe, checkpoint push, publish push), hands the mint
time along with the token, and records the driver's attempt history as
the push attempts the events carry. Host-side git (the repository probe
and the metadata push classification) goes through the same decision
from its rendered message.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 12:44:58 -06:00
Bryan Helmkamp
a39b97c940
Reconnect a local sandbox by attaching to its directory
A local sandbox was rebuilt by creating a fresh Host sandbox over the
recorded working directory, so reconnect, sandbox details, the console
URL, the recorded id, and the terminal each carried a local branch. The
Host provider now derives a designated directory's id from its path and
attaches to it from any provider instance, so reconnect goes through the
one attach path: the record carries that id, a record written before
directories had ids recomputes it from the directory, and describe works
for local like every other kind. The local provider skips the ownership
scope because a designated directory carries no labels and nothing else
shares the host's directories with fabro.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 12:36:29 -06:00
Bryan Helmkamp
5e04495740
Make RunSandbox pebble's Environment and pin pebble
Add pebble-agent and pebble-coding-agent as git dependencies pinned to
the pebble branch that carries the live exec output sink and max_turns,
and align the shared crate versions with lithos-llm's lockfile.

RunSandbox implements pebble_coding_agent::environment::Environment
directly: rename_file over the driver's rename with a pre-created
destination parent, grep rendered as path:line:text, pebble's glob
grammar enforced before the driver sees a pattern, directory listings in
tree order, and exec over the streaming path with the output sink mapped
onto the driver's OutputSink. Pebble's EnvironmentContract runs against
the Host provider in the unit tests and against Docker in the live suite.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 10:16:24 -06:00
Bryan Helmkamp
370a6c96d5
Give the checkout its GitHub credentials through the driver's store
The GitHub App token reached the agent's git commands through the origin
URL: after the clone fabro ran `git remote set-url origin` with the
token embedded, then tracked which generation the URL carried, held an
embed lease across every push so a refresh could not rewrite the URL
mid-operation, re-embedded on the first auth-shaped push failure in
case the agent had rewritten origin, and redacted the URL out of every
log line and output tail. The token showed in `git remote -v` and
`.git/config`.

The driver now installs ambient credentials for a checkout: one
credential-store line beside the checkout and a `credential.helper`
entry pointing at it, with the remote URL untouched. Fabro's part is
`credentials.rs`: the token source, one mint for the clone, one resolve
per push operation, and the facet call. The clone carries the token per
call and installs it afterwards; the ACP refresh tick rewrites the store
instead of the URL; fabro's own pushes pin one resolved token for the
whole operation and pass it per call, so nothing is ever re-embedded and
a retry after replication lag presents the same token by construction.

Gone with the URL: `push_credentials.rs`, `redact.rs`, the lease and
drift repair in `git_push`, `RefreshOutcome`, and the `credential_action`
and `refresh_error` fields on push attempt events. Stored events that
carry those keys still read. A failed store install after the clone now
fails setup, where a failed `set-url` used to be logged and repaired by
the first push. The one remaining caller of the URL redactor, the
server's repository probe, uses `DisplaySafeUrl::redact_in`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 10:02:01 -06:00
Bryan Helmkamp
20a5500510
Read a mock sandbox's recordings from the driver double
MockSandbox forwarded a dozen read-backs to the driver's scripted doubles
one line each: the commands run, the term stops, the stdin fed, the files
written and deleted, the lifecycle counts, whether a walk ran. Tests now
ask the double through MockSandbox::driver. The accessors that convert a
recorded spec into the shape a test asserts on stay: the last command, the
timeouts in milliseconds, the caller's environment without the exec
policy's BASH_ENV blank, and written files as text.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-10 16:55:05 -06:00
Bryan Helmkamp
054a37f830
Hold Daytona credentials as the SDK's configuration
DaytonaCredentials mirrored DaytonaConfig field for field and was copied
into one at connect time. It is now a newtype over the SDK configuration
with the API key always present and a Debug that never prints it; the
driver's Daytona provider connects with the configuration as it is.
Callers build it from an API key, a settings lookup, and the optional
control-plane URL, organization, and HTTP client.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-10 16:53:20 -06:00