Commit graph

5146 commits

Author SHA1 Message Date
Bryan Helmkamp
c98785d84a
Pin pebble a39f43e and refuse an older stored session record clearly
Pebble main a39f43e26effdf99635eaf343f095c17157c9c93 (pebble #22) carries
an assistant turn's usage as Usage in the session record and moves the
record format to version 5. CodingRuntime::from_record refuses a record
in another format with UnsupportedRecord { version, supported } before
it reads the route. Fabro persists those records in SQLite for Ask Fabro
resume, and old runs get no migration, so a record written by an older
build is read back as stored and refused on the next turn.

Two tests pin that down. The store reads pebble's own version 4 fixture
back through get without a parse error and reports it unsupported. A
resumed Ask Fabro session whose stored record declares the previous
format fails its next turn with the agent_error code and the message
"session record format version 4 is not supported (this build requires
5)", runs no turn, and leaves the stored record in place.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:58:31 -06:00
Bryan Helmkamp
e9ee0aaeaa
Merge remote-tracking branch 'origin/main' into one-usage-type
# Conflicts:
#	Cargo.lock
#	Cargo.toml
2026-09-14 12:40:49 -06:00
Bryan Helmkamp
9d15e96fd3
Document the Usage shape on run and stage events
run.completed, run.failed, stage.completed, stage.failed, and
agent.message carry lithos-llm's Usage; the API reference navigation
names the run usage endpoint.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:40:04 -06:00
Bryan Helmkamp
1cdd78926e
Read Usage in the web app and rename the Billing tab to Usage
The run detail tab, route, query hook, and query key say usage. Every
read of input_tokens, output_tokens, total_tokens, reasoning_tokens,
cache_read_tokens, cache_write_tokens, and total_usd_micros moves to
usage.tokens and usage.cost, with lib/usage.ts replacing lib/billing.ts:
totalTokens sums the five buckets, and costSourceTag names a cost that
the provider reported or that was summed from differently sourced parts.
The Usage tab and the stage popover show that tag next to such a cost.

Test fixtures build a Usage through makeUsage; the failing set of the
web tests is unchanged from main (the same 13 environment failures).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:40:04 -06:00
Bryan Helmkamp
aca9bbbf97
Regenerate the TypeScript API client for the usage schemas
The generated models follow the spec: TokenCounts, Cost, Usage,
ModelUsage, UsageModelRef, UsageStageRef, Speed, RunUsage,
RunUsageStage, RunUsageTotals, UsageByModel, AggregateUsage, and
AggregateUsageTotals replace the billing models, and UsageApi replaces
BillingApi. The stale billing model files are deleted.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:31:34 -06:00
Bryan Helmkamp
ef86ad278a
Carry usage as lithos-llm's Usage and rename billing to usage
Re-pin lithos-llm to 55add4596b861a0623d00c3a54aa5c147c8d504b and
pebble to c91810fe51aece80359b9cd8efea971af0c46925, where token usage
and cost travel together as Usage { tokens: TokenCounts, cost:
Option<Cost> }. Fabro now carries that one type everywhere it used to
carry BilledTokenCounts, BilledModelUsage, UsdMicros, or a token count
beside a cost_usd_micros.

fabro-types: billing.rs is usage.rs with ModelRef, ModelUsage { model,
usage }, sum_usage, and usage_is_empty; billing_rollup.rs is
usage_rollup.rs with ProjectionUsageStage, ProjectionUsageByModel,
ProjectionUsageRollup, and usage_rollup_from_projection. Every usage
field is named usage: StageProjection.usage and usage_by_model,
Outcome<Option<ModelUsage>>, stage.completed and stage.failed usage and
usage_by_model, prompt.completed usage, run.completed and run.failed
usage (total_usd_micros is gone), Conclusion.usage, StageSummary.usage,
Run.usage. RunSize buckets by Cost.

fabro-workflow: model_usage_from_llm prices tokens from the catalog with
a Catalog cost source, with_reported_cost keeps a provider cost, and the
pebble handler's stage_usage groups pebble's accounts by model and sums
rows with Usage::saturating_add, so a total has a cost only when every
priced part was priced. The store fold's live usage is the agent's
usage plus its descendants'.

API: the OpenAPI spec deletes BilledTokenCounts, BilledModelUsage,
CompletionUsage, CompletionCost, TokenUsage, and RunBillingSummary,
adds TokenCounts, Cost, Usage, and ModelUsage, and renames every
billing schema, property, tag, path, and operation to usage. fabro-api
reuses lithos-llm's and fabro-types' types through with_replacement,
with a round-trip test per replacement.

Old stored runs get no migration: their pebble events in the old shape
read back with zero usage, and their rebuilt projections lose agent
usage.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:31:34 -06:00
Bryan Helmkamp
5264227ca8
Merge pull request #872 from fabro-sh/exec-verbose-rendering
Pin pebble main 6d802a9 and print tool calls and the transcript under fabro exec --verbose
2026-09-14 13:07:37 -04:00
Bryan Helmkamp
86219355b6
Merge pull request #871 from fabro-sh/daytona-sandbox-guard
Delete a live Daytona test's sandbox even when the test panics
2026-09-14 13:07:24 -04:00
Bryan Helmkamp
3b8d712edb
Delete a live Daytona test's sandbox even when the test panics
The live Daytona tests create a provider sandbox and delete it on their
last line, so any panic or failed assertion before that line leaks a
running, billed sandbox. Two leaked that way on 2026-09-14 when a
sandbox-driver decoder flake panicked daytona_playwright_mcp_sandbox_transport.

Add fabro_sandbox::test_support::DeletedOnDrop, a guard that owns the
RunSandbox (Deref keeps the tests reading unchanged), offers an explicit
delete(self) for the happy path, and deletes from Drop otherwise. The
drop-time delete runs on its own thread and runtime because the test's
runtime may be unwinding. It reconnects the provider through the
ProviderAccess the test built the sandbox with, because the sandbox's
own handle pools HTTP connections whose tasks live on the test's runtime;
a live check of that path timed out after 10s.

Every Daytona test that creates a sandbox now holds it through the guard.
Unit tests over the scripted double prove delete-on-drop runs once, an
explicit delete runs once, and a panic inside catch_unwind still deletes
with and without a runtime.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 11:01:24 -06:00
Bryan Helmkamp
b46c293ceb
Print tool calls and the transcript under fabro exec --verbose
Since #852, `fabro exec --verbose` only turned on the request/response
middleware on the LLM client and no longer printed tool calls, tool
results, or the transcript. Pebble #18 gives pebble-cli-core rendering
options, so `--verbose` now runs the prompt through
`run_prompt_with(..., RenderOptions::verbose())`: each tool call's
arguments and result in full under its `[tool]` and `[result]` lines,
plus the transcript. The middleware is enabled as before. Without the
flag the renderer gets the default options, so the output is unchanged.

The twin shell test now scripts the tool call and the final answer as
two turns, so the answer on stdout is the scripted one rather than the
twin's fallback echo, and it asserts that stderr carries no result,
reasoning, or verbose blocks. A new twin test runs the same prompt with
`--verbose` and asserts the tool and result blocks and the request dump.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 10:56:53 -06:00
Bryan Helmkamp
2c2746c762
Pin pebble main 6d802a9
Pebble #18 adds tool-result and transcript rendering options to
pebble-cli-core (RenderOptions, Renderer::options, and
session::run_prompt_with). Pebble #19 keeps a paired human's hold across
a route failover; no embedder change is needed for it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 10:46:26 -06:00
Bryan Helmkamp
5599ec9e95
Merge pull request #869 from fabro-sh/fix/daytona-e2e-slow-timeout
Give Daytona live tests a slow-timeout that fits a sandbox and a Playwright install
2026-09-14 12:39:03 -04:00
Bryan Helmkamp
899b5f42e2
Merge pull request #870 from fabro-sh/fix/twin-mode-failures
Fix the four twin-mode test failures on main
2026-09-14 12:38:54 -04:00
Bryan Helmkamp
e3ea3ff2e6
Assert the graph name field in local_run_lifecycle
`ps --json` reports the digraph name as `workflow_graph_name` and reserves
`workflow_name` for an explicit `[workflow] name` (6a86ced77). That change
updated the ps tests but not this ignored e2e test, which still expected
the graph name under `workflow_name`. The test now asserts the contract
the ps tests assert: `workflow_name` is null for a bare graph file and
`workflow_graph_name` is the digraph name.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 09:41:05 -06:00
Bryan Helmkamp
b46ed0a429
Point twin_doctor's isolated server at the twin
The provider probe runs inside the isolated server, which never sees the
test process environment. The test used to store `OPENAI_BASE_URL` in the
vault, and cd74013d0 dropped that entry without replacing it, so the
server probed the real OpenAI API with the namespace as its key and the
doctor reported the provider as failed. The server settings now repoint
the `openai` provider at the twin through the operator `[llm]` overlay.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 09:41:05 -06:00
Bryan Helmkamp
325fe0c4bb
Fix the twin-mode hook and arc e2e tests
The four hook tests and arc_e2e_with_real_llm in workflow/hooks.rs failed
in twin mode for three reasons, all in the test fixtures.

The hooked workflows were written as `<name>.toml` beside `<name>.fabro`.
Version packaging accepts a config only as `workflow.toml` beside its
graph (44dccfa3d), so `fabro run` failed at collection. Each hooked
workflow now lives in its own `<name>/` directory as `workflow.toml`.

The isolated server never learned the twin's base URL. The CLI command
carried `OPENAI_BASE_URL`, but the run executes in the server, which does
not see the test process environment, so it called the real OpenAI API
with the namespace as its key. The twin-mode server settings now repoint
the `openai` provider at the twin through the operator `[llm]` overlay,
the same way `run_uses_vault_credentials_for_worker_execution` does.

With the server reaching the twin, the hook scenarios were consumed by
the wrong request: the server asks the model for a run title in the same
namespace before the hook fires, and the scenarios had no matcher. The
block test then saw the twin's default response and the hook failed open,
so the run succeeded. Hook scenarios now match on the `Hook prompt:`
prefix of the evaluator's user message.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 09:41:05 -06:00
Bryan Helmkamp
87fd6d5cd4 Give Daytona live tests a slow-timeout that fits a sandbox and a Playwright install
The e2e nextest profile flags a test as slow after 10s and kills it after
three periods, 30s in all. A Daytona live test creates a remote sandbox,
installs tools in it, and waits for the provider; the Playwright MCP test
also fetches a 114 MiB browser and completes an MCP handshake through the
preview URL. A measured live run took 40.6s, so the documented
`--profile e2e --run-ignored only` command killed it before it could
report its own result.

Add an e2e override for every `daytona_` test that raises the slow period
to 60s and the kill to 20 periods, so a slow provider has room while a hung
test still ends.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 09:40:28 -06:00
fabro-releases[bot]
15198ca23a Bump version to 0.356.0-nightly.0 2026-09-14 09:42:30 +00:00
Bryan Helmkamp
8e67377a9b
Merge pull request #868 from fabro-sh/pins/sandbox-driver-ddb32e19
Some checks are pending
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
TypeScript / Typecheck (push) Waiting to run
TypeScript / Test (push) Waiting to run
TypeScript / Build (push) Waiting to run
Rust / Sandbox plugins (stdio) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
Pin sandbox-driver main ddb32e19
2026-09-13 13:39:19 -06:00
Scott Werner
4c10253b2b
Merge pull request #859 from fabro-sh/codex/retire-manifest-create
Retire legacy manifest run creation
2026-09-13 11:36:29 -04:00
Scott Werner
0dcec5af77 Drop redundant create tests and reuse the production client in fixtures 2026-09-13 09:16:31 -06:00
Bryan Helmkamp
49987c69cb Pin sandbox-driver main ddb32e19
The pin moves from the head of the section-4-driver-items branch (a92c0db6,
since merged as #20) to main, which adds #21: the protocol crate's
PluginSupervisor gains numbered generations, a health probe before a
generation serves, and a refusal of a replacement that reports another
resource namespace. Petri pins the same revision, so the two runners share
one sandbox-driver.

cargo build --workspace and cargo nextest run -p fabro-sandbox -p
fabro-workflow pass (1502 tests, 48 skipped).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DvCtm3CHs5TFbQqUWtkoMX
2026-09-13 09:16:24 -06:00
Scott Werner
012b556367 Remove tests and fixtures tied to retired manifest fields 2026-09-13 08:55:55 -06:00
Scott Werner
ea11538038 Retire legacy manifest run creation 2026-09-13 08:49:50 -06:00
Bryan Helmkamp
05ebd0fd1b
Merge pull request #867 from fabro-sh/drop-parity-tests
Drop the projection parity tests and fold agent_control into agent.activity
2026-09-13 08:42:18 -06:00
Bryan Helmkamp
0046d82e30
Merge origin/main into drop-parity-tests
The stack this branch grew from was rebase-merged, so main carries the same
content under new commits; every conflict resolves to this branch's side.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:40:19 -06:00
Bryan Helmkamp
a7f543dd9f
Drop agent_control from the OpenAPI spec
StageProjection.agent_control and the AgentControlState schema go with the
field; the agent's activity on AgentSessionProjection carries the fact.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:39:56 -06:00
Bryan Helmkamp
a2ace0888b
Drop the projection parity tests and fold agent_control into agent.activity
The session_projection_parity module pinned fabro's stage fold to pebble's
SessionProjection while both existed; the stage view reads the fold now,
and the one usage rule has its own tests. StageProjection.agent_control
and AgentControlState go too: pebble's fold carries the interrupted and
steered facts as agent.activity, and the stage's state says whether the
stage still runs, which is what the reset on fabro's own stage events was
for. The run-detail banner reads activity plus state.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:38:52 -06:00
Bryan Helmkamp
c9820cf0ad
Merge pull request #865 from fabro-sh/agent-events-docs
State the agent event contract and show the agent sidebar in demo mode
2026-09-13 10:36:58 -04:00
Bryan Helmkamp
68283e6413 Show the agent sidebar's sections in demo mode
The demo agent stage's stored events now read as one pebble session: MCP
servers up and failed, skills, a subagent, a failover, a compaction, and a
written file, ending with ProcessingEnd. Demo mode serves the run state it
answered not_implemented to, with the agent stage carrying the coding
agent's fold of those events, so the stage sidebar renders them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:36:37 -06:00
Bryan Helmkamp
4645181dbf State the agent event contract in the docs
Pebble's CodingAgentEvent stream is the agent event contract: every event
except streaming deltas is stored verbatim under its derived name and
folded into StageProjection.agent with pebble's SessionProjection. The
events doc, the events strategy, and the v2 shape doc say so, list the
agent events fabro still emits for facts pebble cannot know, and tell
consumers to read the fold rather than fold the events again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:36:37 -06:00
Bryan Helmkamp
866cf08d39
Merge pull request #864 from fabro-sh/delete-agent-mirrors
Delete the agent mirrors and move failover to prompt stages
2026-09-13 10:36:34 -04:00
Bryan Helmkamp
f1118eead2 Delete the agent mirrors and move failover to prompt stages
Pebble's stream is the agent event contract. The run's own agent.mcp.ready,
agent.mcp.failed, and agent.mcp.disconnected events, which mirrored pebble's
McpServer* events, are gone with their props, the sink arms that emitted
them, and their conversion and naming entries; pebble's stored
agent.mcp.server.* events are the only record and feed the stage's fold.

The sink no longer mirrors RouteFailover onto agent.failover either: an
agent stage's moves are pebble's agent.route.failover. The event is now
prompt.failover, emitted only by a one-shot prompt stage that walks its
fallback plan itself, and its props are trimmed to the two routes, the
attempt, and the error; nothing read the rest.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:36:24 -06:00
Bryan Helmkamp
f146312cda
Merge pull request #863 from fabro-sh/stage-view-from-agent
Read the stage view from the agent's fold
2026-09-13 10:36:21 -04:00
Bryan Helmkamp
393424b05e Show the agent sidebar from stage.agent
The stage insights sidebar reads every session fact from the coding
agent's fold: the root agent's todo list with subagent lists counted
apart, MCP status derived from disconnected, error, and tools, skills, and
new Files and Subagents sections, a failover badge naming the route the
session moved to and why it stopped, and a compactions row under the
context window. The run state refreshes on the agent's own events those
sections read.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:36:10 -06:00
Bryan Helmkamp
318fdf7206 Read the stage view from the agent's fold
StageProjection loses todos, subagents, skills, mcp_servers, and
context_window, the types behind them, their fold arms and helpers, and
their OpenAPI schemas: every one of those facts is pebble's fold in
StageProjection.agent now. The context-window endpoint reads the fold's
snapshot, whose event_seq is the agent's own sequence. The parity module
keeps its assertions on the surviving own fields, usage and model, and
checks that what the stage view reads from agent is the whole-session
fold's for the stage's events. The TypeScript client is regenerated and
its stale models removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:36:10 -06:00
Bryan Helmkamp
d5009d976b
Merge pull request #862 from fabro-sh/one-usage-rule
One usage rule: bill an agent stage's session tree from one fold
2026-09-13 10:36:07 -04:00
Bryan Helmkamp
0c0e589a78 Bill a failed agent stage what it spent
An agent stage that failed billed nothing: the backend returned a bare
error and the outcome built from it carried no usage. A terminal failure
now becomes the stage's failed outcome from the same fold that bills a
completed stage, with the tree's usage, the rows by model, the files it
wrote, and its active time; stage.failed carries billing and
billing_by_model and the store keeps both. Cancellation and retryable
failures still go up as the error.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:47 -06:00
Bryan Helmkamp
9101a90471 Describe billing_by_model on the API and the Billing tab
StageProjection.billing_by_model and a BilledModelUsage schema that reuses
fabro's type; the TypeScript client regenerated; the Billing tab's token
tooltip says subagent tokens are included and priced at each subagent's
model; the stage.completed docs describe the rows and the one usage rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:47 -06:00
Bryan Helmkamp
df8762663b Bill an agent stage's whole session tree from one fold
One usage rule: a stage's usage is its session tree's, the root and every
subagent, live and at completion. The worker's event sink folds pebble's
SessionProjection over the events it records and the stage's billing and
files come from that fold at stage end, so the completed values are what
the run showed live. The store's live usage is the fold's tree usage, and
completion brings the catalog's price for the same tokens instead of
resetting them to the root's.

Fabro keeps catalog pricing: the root at its route, each descendant at its
own route where the catalog knows it and at the root's otherwise, a
provider-reported cost standing in where pebble has one. The rows travel
as billing_by_model on stage.completed and the stage projection, and the
billing rollup splits by_model by them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:47 -06:00
Bryan Helmkamp
d3a2f2b0c4
Merge pull request #861 from fabro-sh/stage-agent-projection
Embed pebble's SessionProjection in StageProjection
2026-09-13 10:35:42 -04:00
Bryan Helmkamp
5d2b7cc4aa Describe StageProjection.agent on the API as pebble's own types
AgentSessionProjection and the schemas nested in it reuse pebble's types
through with_replacement; the AgentSession prefix marks the projection's
own types where fabro already has a schema of that name, and pebble's
event-level types keep their names. The round-trip test builds a
projection over a scripted stream, validates it against the spec with the
spec as the root document, checks every serialized key is declared, and
validates every enum variant this build knows. The TypeScript client is
regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:07 -06:00
Bryan Helmkamp
245d2c411d Embed pebble's SessionProjection in StageProjection
StageProjection.agent is pebble's fold of the stage's agent events, fed
every stored agent event before the fabro-only arms run. Every existing
field and arm stays for now. The parity tests prove each old field is
derivable from the embedded fold: the tree's usage, the route as the model,
the context window without fabro's stamped seq, the root's todo list, the
subagent rows, the skills, and the MCP servers under the disconnected,
error, ready rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:07 -06:00
Bryan Helmkamp
ecafe1e807 Store ProcessingEnd and the mirrored pebble events in the run log
Pebble's SessionProjection reads ProcessingEnd to complete a prompt and mark
the session idle, so a projection rebuilt from the run's log needs it: one
small event per prompt. The four pebble events the sink mirrored onto
fabro's own agent.failover and agent.mcp.* are now stored verbatim as well,
so the fold sees the route moves and the MCP outcomes; the mirrors stay
until every reader is on the projection.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:35:07 -06:00
Bryan Helmkamp
6cd4bb417b
Merge pull request #860 from fabro-sh/pebble-port-routes
Pin pebble main 6996942 and implement its PortRoutes trait
2026-09-13 10:35:03 -04:00
Bryan Helmkamp
32dca54cd7
Show the agent sidebar's sections in demo mode
The demo agent stage's stored events now read as one pebble session: MCP
servers up and failed, skills, a subagent, a failover, a compaction, and a
written file, ending with ProcessingEnd. Demo mode serves the run state it
answered not_implemented to, with the agent stage carrying the coding
agent's fold of those events, so the stage sidebar renders them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:28:32 -06:00
Bryan Helmkamp
2c3c2d8b7f
State the agent event contract in the docs
Pebble's CodingAgentEvent stream is the agent event contract: every event
except streaming deltas is stored verbatim under its derived name and
folded into StageProjection.agent with pebble's SessionProjection. The
events doc, the events strategy, and the v2 shape doc say so, list the
agent events fabro still emits for facts pebble cannot know, and tell
consumers to read the fold rather than fold the events again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:28:32 -06:00
Bryan Helmkamp
7a6eea0399
Delete the agent mirrors and move failover to prompt stages
Pebble's stream is the agent event contract. The run's own agent.mcp.ready,
agent.mcp.failed, and agent.mcp.disconnected events, which mirrored pebble's
McpServer* events, are gone with their props, the sink arms that emitted
them, and their conversion and naming entries; pebble's stored
agent.mcp.server.* events are the only record and feed the stage's fold.

The sink no longer mirrors RouteFailover onto agent.failover either: an
agent stage's moves are pebble's agent.route.failover. The event is now
prompt.failover, emitted only by a one-shot prompt stage that walks its
fallback plan itself, and its props are trimmed to the two routes, the
attempt, and the error; nothing read the rest.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:21:41 -06:00
Bryan Helmkamp
232d347d39
Show the agent sidebar from stage.agent
The stage insights sidebar reads every session fact from the coding
agent's fold: the root agent's todo list with subagent lists counted
apart, MCP status derived from disconnected, error, and tools, skills, and
new Files and Subagents sections, a failover badge naming the route the
session moved to and why it stopped, and a compactions row under the
context window. The run state refreshes on the agent's own events those
sections read.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:13:10 -06:00
Bryan Helmkamp
b5ee16b37a
Read the stage view from the agent's fold
StageProjection loses todos, subagents, skills, mcp_servers, and
context_window, the types behind them, their fold arms and helpers, and
their OpenAPI schemas: every one of those facts is pebble's fold in
StageProjection.agent now. The context-window endpoint reads the fold's
snapshot, whose event_seq is the agent's own sequence. The parity module
keeps its assertions on the surviving own fields, usage and model, and
checks that what the stage view reads from agent is the whole-session
fold's for the stage's events. The TypeScript client is regenerated and
its stale models removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 08:13:10 -06:00