During a long LLM turn the durable event stream was silent: between `agent.tool.completed` and the next `agent.message` nothing was emitted, so "the model is generating" and "the worker is wedged" were indistinguishable from the run store, SSE, or the UI. The signal already existed. `AssistantTextStart` fired at exactly the right point — after `build_request()`, after compaction, immediately before the stream opens — then was classified as streaming noise and thrown away. This promotes it rather than inventing a new one. Two events, each asserting only what is provable when it is emitted: - `agent.llm.started` carries the *requested* provider/model. No usage, no cost, no context window: none of it exists yet, and failover can re-target, so `agent.message` stays authoritative for what answered. - `agent.llm.first_output` is edge-triggered on the first output of an attempt and names what arrived. `ToolCall` is required, not optional: a turn that opens with a tool call produces no text or reasoning delta, so a latch keyed on those two would stay silent for exactly the tool-heavy rounds where liveness matters most. `agent.llm.retry` now also fires on the one previously invisible mid-turn path — a stream that ends without a finish event, which replays the turn and discards its output with nothing to show for it. Its `attempt` field was already fed by two independent counters, so an optional `phase` (open | consume) names which loop it counts. `StageProjection.inference` projects the open bracket. `Some` means "the event log contains an unclosed inference bracket", not "the model is computing now" — a SIGKILLed worker leaves it open, which is the truthful statement of what we know, and `watchdog.timeout` remains the authority on actually-stuck. The close is the subtle part. Terminal cancel and wall-clock timeout tear the session down through `discard_session` without emitting a message, error, or interrupt, so a session-lifecycle backstop is required. It has to be `agent.session.ended`, not `agent.session.deactivated`: deactivation is emitted by `lease.release()` *before* the forwarder drains queued agent events, so a queued `agent.llm.started` can arrive after it and re-open the bracket. But `agent.session.ended` carries no stage identity, so the close takes ordering from the event and identity from the projection, scanning for brackets the ending session opened. A normal stage lookup there finds no target and silently no-ops. Presentation states what the log proves and nothing more: no progress bar or ETA (no completion estimate exists), "reasoning" only when the provider sent reasoning output, elapsed counted since the request opened, and no live animation once the run is terminal. Scope is session-backed agent stages. One-shot completions call `client.complete` directly and never build a session; covering them means moving the emit point into `fabro-llm`, filed as a follow-up. `agent.output.start` was never persisted — it existed in a name map, an `unreachable!` arm, and docs — so the rename carries no migration risk. Corrects `events.md`, which documented it as a real emitted event, and the v2 proposal, which mapped it to `message.part.started` despite it firing before the request opens. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
14 KiB
Fabro Event Schema V2 Proposal
Date: 2026-04-08
Status: proposal
Assumptions:
- greenfield redesign
- no production deployments
- no backward-compatibility constraints
- optimize for the best long-term public event contract
This proposal turns the earlier ideation into concrete schema changes.
Design Goal
Fabro should expose:
- a durable, append-only event log for audit, storage, replay, and projections
- a separate live stream for UI-oriented snapshots, deltas, and fast progress
They should share IDs and correlation fields, but they should not be the same contract.
Top 10 Concrete Improvements
1. Split the single event story into two concrete public APIs
Proposal
Introduce two top-level event contracts:
DurableEventLiveEvent
Endpoints:
GET /runs/:run_id/events- append-only durable events
- replayable
- no keep-alive payload events
GET /runs/:run_id/live- live UI stream
- snapshots + deltas + keep-alives
- resumable with cursor
Durable event shape
{
"kind": "durable",
"id": "evt_01960d0c...",
"seq": 182,
"ts": "2026-04-08T15:01:02.123Z",
"run_id": "run_01JQ...",
"event": "agent.tool.started",
"session_id": "ses_123",
"node_id": "code",
"properties": {
"tool_call_id": "tool_abc",
"tool_name": "read_file",
"arguments": { "path": "src/main.rs" }
}
}
Live event shape
{
"kind": "live",
"id": "levt_01960d0d...",
"seq": 991,
"ts": "2026-04-08T15:01:03.000Z",
"run_id": "run_01JQ...",
"event": "message.delta",
"session_id": "ses_123",
"message_id": "msg_456",
"part_id": "part_1",
"properties": {
"block_type": "text",
"delta": "Let me check that file..."
}
}
Why this is better
- Durable events stay stable and analyzable.
- Live events can be noisy and UI-oriented without polluting projections.
- Keeps Fabro from repeating the Claude Code / OpenCode problem of mixing control, transport, and product semantics.
2. Add explicit stream ordering, replay, and recovery semantics
Proposal
Every durable and live stream event gets:
seq: u64- SSE
id:=seq - replay semantics based on
Last-Event-ID
Server rules:
- if
Last-Event-IDis present and still buffered, replayseq > cursor - if cursor is too old, return a structured reset event in live streams and
409 replay_reset_requiredin durable streams - durable streams never emit synthetic snapshots
- live streams may start with a
*.snapshotevent after reconnect
New live-only events
stream.heartbeatrun.snapshotsession.snapshotnode.snapshotstream.reset_required
Example stream.reset_required
{
"kind": "live",
"id": "levt_01960d0e...",
"seq": 1200,
"ts": "2026-04-08T15:02:00.000Z",
"run_id": "run_01JQ...",
"event": "stream.reset_required",
"properties": {
"reason": "cursor_too_old",
"expected_from_seq": 1170
}
}
Why this is better
- Reattach behavior becomes deterministic.
- Clients no longer guess whether they missed data.
- Replay is part of the contract, not an implementation detail.
3. Expand the envelope into a first-class correlation model
Proposal
Extend the shared envelope with these optional fields:
workflow_idstage_idbranch_idcheckpoint_idsession_idparent_session_idturn_idmessage_idpart_idtool_call_idrequest_idcausation_idcorrelation_id
Rules:
idis the event's own identitycausation_idpoints to the immediate triggering event, if anycorrelation_idgroups a whole logical operation, for example one user request or one retry attempt treerequest_idis transport/API request scoped, not workflow scoped
Concrete change
Move these IDs out of ad hoc properties payloads when they are structural identifiers.
Good:
{
"event": "agent.tool.completed",
"tool_call_id": "tool_abc",
"message_id": "msg_456",
"properties": {
"tool_name": "read_file",
"is_error": false
}
}
Bad:
{
"event": "agent.tool.completed",
"properties": {
"tool_call_id": "tool_abc",
"message_id": "msg_456",
"tool_name": "read_file"
}
}
Why this is better
- Correlation becomes universal instead of event-family-specific.
- UI and analytics consumers can join without parsing
properties. - Parent/child agent and retry trees become much easier to reason about.
4. Replace stringly state with concrete tagged unions
Proposal
Define explicit union types for stateful fields.
Examples:
type StopReason =
| { type: "completed" }
| { type: "requires_input"; request_id: string }
| { type: "interrupted"; interrupt_reason: InterruptReason }
| { type: "failed"; error_kind: ErrorKind }
| { type: "retries_exhausted"; attempts: number };
type RetryStatus =
| { type: "not_retrying" }
| { type: "retry_scheduled"; attempt: number; next_retry_at: string }
| { type: "retrying"; attempt: number }
| { type: "retries_exhausted"; attempts: number };
type ApprovalStatus =
| { type: "not_required" }
| { type: "requested"; approval_id: string }
| { type: "approved"; approval_id: string; actor: string }
| { type: "denied"; approval_id: string; actor: string; reason?: string };
Concrete fields to replace
statusreasonfailure_classinterrupt_reasonstop_reasonapproval_status
Why this is better
- Eliminates string drift.
- Makes reducers and policy engines much safer.
- Makes test fixtures much more stable.
5. Standardize event family grammar across the entire product
Proposal
Use one lifecycle vocabulary:
.created.started.snapshot.delta.updated.completed.failed.cancelled.interrupted.deleted
Apply it consistently to the same kinds of things:
run.*stage.*session.*turn.*message.*message.part.*tool.*command.*checkpoint.*parallel.branch.*retro.*
Concrete renames
Current style is already decent, but V2 should be stricter.
Examples:
agent.text.delta->message.part.deltaagent.tool.output.delta->tool.output.deltaagent.processing.end->turn.completedorsession.idle, depending on actual semantics
Why this is better
- Consumers can infer behavior from naming alone.
- Reduces one-off event families that encode bespoke lifecycle semantics.
6. Introduce typed content blocks and block-level deltas
Proposal
Represent streamable content as typed message parts.
Base union:
type MessagePart =
| { type: "text"; part_id: string; text: string }
| { type: "reasoning"; part_id: string; text: string }
| { type: "tool_call"; part_id: string; tool_call_id: string; tool_name: string; input: unknown }
| { type: "tool_result"; part_id: string; tool_call_id: string; output: unknown; is_error: boolean }
| { type: "patch"; part_id: string; patch_ref: string }
| { type: "file_ref"; part_id: string; file_id: string; path: string }
| { type: "artifact_ref"; part_id: string; artifact_id: string; label: string }
| { type: "plan"; part_id: string; items: PlanItem[] }
| { type: "todo"; part_id: string; items: TodoItem[] }
| { type: "command_output"; part_id: string; command_id: string; stream: "stdout" | "stderr"; text: string };
Live delta event:
{
"event": "message.part.delta",
"message_id": "msg_456",
"part_id": "part_1",
"properties": {
"part_type": "text",
"delta": "checking src/main.rs"
}
}
Durable completion event:
{
"event": "message.completed",
"message_id": "msg_456",
"properties": {
"parts": [
{ "type": "text", "part_id": "part_1", "text": "checking src/main.rs" }
]
}
}
Why this is better
- Supports rich UI without reparsing free-form text.
- Supports structured summarization, compaction, and retro generation.
- Aligns Fabro with the best parts of Claude Sessions and pi-mono.
7. Make approvals, questions, and operator interventions first-class durable events
Proposal
Add explicit event families:
approval.requestedapproval.respondedquestion.askedquestion.answeredinterrupt.requestedinterrupt.appliedresume.requiredresume.applied
Example approval.requested
{
"kind": "durable",
"id": "evt_01960d0f...",
"seq": 201,
"ts": "2026-04-08T15:03:00.000Z",
"run_id": "run_01JQ...",
"session_id": "ses_123",
"tool_call_id": "tool_abc",
"event": "approval.requested",
"properties": {
"approval_id": "apr_1",
"scope": "tool_call",
"tool_name": "exec_command",
"request": {
"cmd": "git push origin branch"
}
}
}
Example approval.responded
{
"kind": "durable",
"id": "evt_01960d10...",
"seq": 202,
"ts": "2026-04-08T15:03:10.000Z",
"run_id": "run_01JQ...",
"event": "approval.responded",
"properties": {
"approval_id": "apr_1",
"result": {
"type": "approved",
"actor": "user"
}
}
}
Why this is better
- Human-in-loop behavior becomes queryable and replayable.
- Workflow interruption is no longer hidden in transport or UI state.
8. Add real snapshot events instead of relying on ad hoc reconstruction
Proposal
Define explicit snapshot events for live attach and projection recovery:
run.snapshotsession.snapshotnode.snapshotcheckpoint.saved
Example session.snapshot
{
"kind": "live",
"id": "levt_01960d11...",
"seq": 1500,
"ts": "2026-04-08T15:04:00.000Z",
"run_id": "run_01JQ...",
"session_id": "ses_123",
"event": "session.snapshot",
"properties": {
"state": { "type": "running" },
"turn_id": "turn_9",
"messages": [
{
"message_id": "msg_456",
"role": "assistant",
"parts": [
{ "type": "text", "part_id": "part_1", "text": "checking src/main.rs" }
]
}
],
"active_tool_calls": [
{
"tool_call_id": "tool_abc",
"tool_name": "read_file",
"status": "running"
}
]
}
}
Concrete rule
- snapshots are authoritative replacement state for live consumers
- snapshots are optional in durable streams
- checkpoints are durable domain snapshots, not just UI snapshots
Why this is better
- Fast attach becomes trivial.
- Projections can self-heal from snapshots.
- Checkpoint semantics become explicit rather than emergent.
9. Make model, tool, command, and MCP work first-class span families
Proposal
Create event families with shared semantics:
model.request.startedmodel.request.completedmodel.request.failedtool.startedtool.output.deltatool.completedtool.failedcommand.startedcommand.output.deltacommand.completedcommand.failedmcp.call.startedmcp.call.progressmcp.call.completedmcp.call.failed
Example model.request.completed
{
"kind": "durable",
"id": "evt_01960d12...",
"seq": 220,
"ts": "2026-04-08T15:05:00.000Z",
"run_id": "run_01JQ...",
"session_id": "ses_123",
"turn_id": "turn_9",
"request_id": "req_llm_1",
"event": "model.request.completed",
"properties": {
"provider": "anthropic",
"model": "claude-sonnet-4",
"latency_ms": 1834,
"usage": {
"input_tokens": 1400,
"output_tokens": 380,
"reasoning_tokens": 120,
"cache_read_tokens": 900,
"cache_write_tokens": 0
},
"retry_status": { "type": "not_retrying" }
}
}
Why this is better
- Cost and latency analysis become first-class.
- Policy engines can reason about real operations, not just stage summaries.
- Cross-provider comparison gets much easier.
10. Generate and enforce the public schema, docs, and examples from one registry
Proposal
Build a single event_schema_registry source that defines:
- envelope fields
- event families
- payload types
- union types
- versioning
- example payloads
Artifacts generated from it:
- Rust types
- TypeScript types
- JSON Schema
- OpenAPI / SSE docs
- sample event fixtures
- validation tests
Concrete rules
- every public event must have:
- one schema definition
- one example payload
- one validation test
- no endpoint may inject extra consumer-visible fields outside the schema
- keep-alive frames are documented separately from payload events
Why this is better
- Prevents the OpenCode and Goose class of drift.
- Makes Fabro's event API publishable and stable from day one.
Recommended V2 Event Families
If Fabro were starting from scratch, I would structure the public families like this:
run.*stage.*checkpoint.*parallel.branch.*session.*turn.*message.*message.part.*model.request.*tool.*command.*mcp.call.*approval.*question.*interrupt.*resume.*compaction.*retro.*artifact.*stream.*(live only)
Recommended Field Placement Rules
Top-level envelope:
- identity and correlation
- ordering
- timestamps
- scope
properties:
- event-family-specific payload
- business data
- structured state payloads
Never in properties if they are structural:
run_idseqeventsession_idmessage_idtool_call_idrequest_idcausation_idcorrelation_id
Bottom Line
The best greenfield version of Fabro is not "the current schema plus more events."
It is:
- separate durable and live contracts
- replayable ordered streams
- a richer envelope
- typed state unions
- typed content blocks
- explicit snapshots
- first-class HITL events
- first-class span families
- generated schema/docs/tests from one registry
That would give Fabro a better event platform than any of the compared systems.