- Add AppState::cached_run with the standard 500/404 mapping and use it
everywhere handlers read the shared run-projection cache. This also
normalizes two inconsistencies: graph-source cache errors now map to
500 (was 502), and a missing projection in PR create/unlink now maps
to the canonical 404 (was a bespoke 500).
- Extract an EventScan cursor shared by the four run-event scan loops,
delegate list_events_from to the paginated variant, and stop the
stage-event scan once its page is full instead of walking the rest of
the log.
- Hold Arc<RunProjection> in the local projection cache so opening a run
no longer deep-copies the projection (copy-on-write via Arc::make_mut),
and drop the now-unreachable shared-cache branch in last_event_seq.
- Trim hot-path clones: run_files serves the projection Arc directly,
run-state serializes by reference, artifacts only checks existence, and
the command-log handler opens a reader only for the CAS-blob branch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A node cancelled (or lost to a crash) mid-flight and then resumed now
starts a new stage execution with the next StageId ordinal (work@2)
instead of reusing and clearing the cancelled execution's projection.
The old execution stays immutable with its own events, session, output,
timing, billing, and termination state.
Engine:
- Add a run-scoped StageExecutionTracker on RunServices with per-node
high-water marks. Ordinals are reserved after the StageStart hook
passes on the first attempt (retries reuse the reservation), ensured
at the composite checkpoint pre-step for hook-skips, and reserved in
on_terminal_reached for terminal nodes' synthetic events.
- Keep three concepts distinct: graph visit (max_visits/checkpoints,
unchanged), stage execution ordinal (the @N in StageId), and handler
attempt. The tracker is not checkpointed; the append-only stage event
history is its durable source of truth.
- resume() seeds the allocator from the run projection and computes a
node -> StageId provenance map of executions observed after the
selected checkpoint, threaded through execute_persisted_run,
RunSession, and InitOptions.
Events and projections:
- stage.started, parallel.branch.started, and checkpoint.completed
carry optional graph_visit and resumed_from_stage_id; StageProjection
stores both. Old events deserialize with None and legacy duplicate
stage.started replays keep last-attempt behavior.
- The CheckpointCompleted reducer is envelope-first: diffs and
skipped-stage synthesis attach to the exact execution StageId, an
existing Retrying projection finalizes as Skipped without losing
identity, and historical node_outcomes no longer create or collide
with newer ordinals (node_visits remains a legacy fallback).
Handlers:
- Parallel fan-out reserves child ordinals through the shared tracker,
derives worktree pass{N} from the parent's execution ordinal, and
seeds branch contexts with explicit child stage scopes so branch
lifecycle and nested handler events agree.
- Artifact capture and manager-loop child logs follow the ordinal.
API and UI:
- RunStage documents visit as the execution ordinal and adds optional
graph_visit and resumed_from_stage_id; Rust and TypeScript clients
regenerated.
- The web sidebar lists both executions chronologically; resumed stages
show a "Resumed from" link in the stage detail header and hover
popover, with the graph visit surfaced when it diverges.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Extract a shared fetch_run_events_page helper so the three client
paging loops (full list, until, tail) no longer repeat the request/
convert/has_more skeleton; fold the tail loop's two descending-order
checks into one and drop its redundant had_events flag.
- Skip the latest-seq lookup in list_events_before_with_limit when the
caller supplies a before_seq cursor, so a cold projection cache costs
at most one full history scan per pagination session instead of one
per page.
- Remove the dead before_seq max(1) clamp and the passthrough order()
accessor from EventListParams.
- Document the CLI --tail 0 --follow seeding trick and the reader
event_seq placeholder invariant.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Unify list_events_from with list_events_from_with_limit so projection
replay shares the seek path instead of duplicating the decode loop
- Bound the event scan with keys::run_events_range instead of an
unbounded range plus a manual prefix break, so slatedb never touches
SSTs belonging to other runs or namespaces
- Store reader event_seq as None instead of a valid-looking sentinel of
1, so appends through a reader-built inner fail as ReadOnly rather
than writing duplicate sequences
- Borrow keys during scans instead of allocating a String per entry,
drop a dead branch in cached_events_from, collapse recover_next_seq's
single-caller parameters, and document the zero-padded key ordering
invariant the seek depends on
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>