Commit graph

377 commits

Author SHA1 Message Date
Scott Werner
7edef76d77 Revoke replayed auth sessions transactionally
Delete the owning auth session inside the refresh-token rotation transaction when a spent token is replayed. Return the replay outcome only after the revocation commits, and propagate database failures without claiming the chain was revoked.
2026-08-21 13:55:41 -04:00
Scott Werner
a08dff1ce8 Merge main into feat/refresh-tokens-sqlite 2026-08-21 13:16:14 -04:00
Scott Werner
9d3aa7a4d4 Consolidate workflow-version lineage test coverage
The lineage field's `skip_serializing_if` behavior was asserted five times
across three crates. Keep the two assertions in fabro-types, which owns the
attribute, and drop the duplicates:

- Delete `run_created_omits_absent_workflow_version_id` from event/convert.rs,
  a copy of the test above it that re-checked another crate's serde attribute.
  convert.rs's own responsibility is covered by the existing field assertion.
- Delete `legacy_create_input_persists_without_workflow_version_id`, which ran
  the full create() pipeline to prove a hardcoded `None` literal is `None`.
  `CreateRunInput` has no such field, so no input could change the result.
- Fold `run_spec_omits_absent_workflow_version_id` into the adjacent legacy-spec
  test, which already holds an all-`None` record.
- Drop the off-topic spec re-serialization from run_state.rs's retried_from test.

Add `test_support::test_workflow_version_id()` alongside `test_run_provenance()`
and use it everywhere, replacing eight copies of the same magic seed across five
crates plus two assertion sites that recomputed the hash inline. This also
subsumes retry.rs's private helper of the same shape.

Revert the `run_spec_json` parameterization in the projection round-trip test:
`RunProjection` is a `with_replacement` alias for the canonical type, so the
`Some` and `None` call sites exercise identical code.

Have the two run.created literals that mirror a `RunSpec` read the spec's
lineage field instead of hardcoding `None`, so the mirrors stay accurate once a
producer populates it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 12:39:02 -04:00
Scott Werner
27fd48c603 Persist workflow version lineage on runs 2026-08-21 12:39:02 -04:00
Scott Werner
75fa8eca8b Merge remote-tracking branch 'origin/main' into codex/exact-target-checkout
# Conflicts:
#	lib/components/fabro-sandbox/src/clone_retry.rs
#	lib/components/fabro-sandbox/src/daytona/mod.rs
#	lib/components/fabro-sandbox/src/docker.rs
#	lib/components/fabro-sandbox/src/provider/docker.rs
2026-08-21 12:15:59 -04:00
Scott Werner
5104a787ce Bound Daytona post-clone setup 2026-08-21 12:08:23 -04:00
Bryan Helmkamp
db1faf02ec
Fix catalog dispatch invariant for shared models 2026-08-21 11:19:28 -04:00
Bryan Helmkamp
ded92a215d
Update Venice model catalog 2026-08-21 10:37:04 -04:00
Bryan Helmkamp
7d771e9b96
Merge remote-tracking branch 'origin/main' into pr-764
# Conflicts:
#	lib/components/fabro-sandbox/src/daytona/mod.rs
#	lib/components/fabro-sandbox/src/push_credentials.rs
#	lib/components/fabro-sandbox/src/sandbox.rs
2026-08-20 22:07:20 -04:00
Bryan Helmkamp
78cb0d1348
Clean up git push retry handling 2026-08-20 21:59:47 -04:00
Bryan Helmkamp
66dc30424a
Merge pull request #729 from fabro-sh/fix/doctor-timeout-budget
Bound doctor diagnostics within client timeout
2026-08-20 21:58:22 -04:00
Bryan Helmkamp
c021d37155
Merge pull request #740 from fabro-sh/fix/redaction-corrupts-executable-spec
Keep the executable run spec out of reach of event redaction
2026-08-20 21:58:09 -04:00
Release Repro
0355a13db8
Merge origin/main into fix/doctor-timeout-budget 2026-08-20 20:51:17 -04:00
Bryan Helmkamp
ca6d9a46da
Merge remote-tracking branch 'origin/main' into fix/redaction-corrupts-executable-spec
# Conflicts:
#	lib/components/fabro-dump/src/lib.rs
#	lib/components/fabro-store/src/run_state.rs
#	lib/components/fabro-store/tests/serializable_projection.rs
#	lib/components/fabro-workflow/src/billing_rollup.rs
#	lib/components/fabro-workflow/src/run_lookup.rs
#	lib/components/fabro-workflow/src/runtime_store.rs
#	lib/foundation/fabro-api/tests/run_projection_round_trip.rs
#	lib/foundation/fabro-test/src/lib.rs
#	lib/foundation/fabro-types/src/run.rs
#	lib/foundation/fabro-types/src/run_event/run.rs
#	lib/foundation/fabro-types/src/run_projection.rs
#	lib/foundation/fabro-types/tests/run_spec_methods.rs
2026-08-20 20:42:20 -04:00
Bryan Helmkamp
2b095612c8
Address run spec persistence review findings 2026-08-20 20:35:52 -04:00
Bryan Helmkamp
f8dd7b1daf
Merge origin/main into daytona-activate-state-transitions 2026-08-20 20:35:00 -04:00
Release Repro
b7e3b660ff
Simplify diagnostics timeout handling 2026-08-20 20:34:38 -04:00
Bryan Helmkamp
82218a228a
Merge pull request #763 from fabro-sh/feat/github-token-source
Add a cached GitHub installation-token source for push credentials
2026-08-20 20:33:04 -04:00
Bryan Helmkamp
e5c0301ccb
Simplify Daytona lifecycle retries 2026-08-20 19:57:19 -04:00
Bryan Helmkamp
f8a82d6865
fix: harden GitHub token refresh handling 2026-08-20 19:56:06 -04:00
Bryan Helmkamp
2456356a9f
Label sandbox git execs with git_op tracing spans
Sandbox exec logs previously required command_len fingerprinting to tell a
push from a credential refresh or a checkpoint commit. The shared git
helpers now instrument their futures with a git_op span, so Daytona's and
Docker's `exec_command: entered` lines inherit the operation label and the
log renders as `git_op{op=push}: exec_command: entered timeout_ms=...`.

Ops: push (git_push_via_exec), refresh-credentials (both providers'
refresh_push_credentials), checkpoint-commit (checked_git_checkpoint),
fetch (fetch_source_run_ref), and metadata-push (the run-metadata snapshot
write). Spans are attached with #[tracing::instrument] — attached to the
future, never an entered() guard held across an await — so they follow the
task across worker threads. No trait or signature changes.

Plan: .ai/plans/git-push-token-resilience.md (PR 3: item 10).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 19:37:44 -04:00
Bryan Helmkamp
234cac93ef
Merge pull request #728 from fabro-sh/codex/skip-daytona-edit-folder-post
Skip Daytona folder creation for file edits
2026-08-20 18:30:00 -04:00
Bryan Helmkamp
e688ee59a7
Merge main into feat/live-run-billing-totals 2026-08-20 18:20:10 -04:00
Bryan Helmkamp
1dc31771c6
Merge pull request #757 from fabro-sh/fireworks-412-failover
Classify provider 412s as failover-eligible account lockouts
2026-08-20 18:17:10 -04:00
Bryan Helmkamp
6885ed40cb
Merge pull request #768 from fabro-sh/sandbox-errors-transient-infra
Classify sandbox state-change rejections as transient infra
2026-08-20 18:16:47 -04:00
Bryan Helmkamp
f88df59163
Classify sandbox state-change rejections as transient infra
A Daytona "Sandbox state change in progress" rejection surfacing
through the pipeline lifecycle path ("Pipeline lifecycle operation
failed") matched no transient-infra hint, so the run failure was
categorized deterministic. The condition is a provider lifecycle
transition that finishes on its own — the definition of transient
infrastructure — and the deterministic label misinforms retry
machinery and anyone reading the failure.

Add two transient-infra hints: the provider rejection ("state change
in progress") and the bounded-wait timeout an activation reports when
a stop transition outlives its budget ("sandbox stop still in
progress").

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 17:39:19 -04:00
Bryan Helmkamp
0eedb1798c
Treat Daytona state transitions as wait-and-retry in activate/start/stop
Daytona rejects start/stop with HTTP 400 "State change in progress"
while a lifecycle transition is in flight, and activate() only handled
the Started and Starting states: any other state fell through to
start(), which surfaced the rejection as a hard failure. A run died
exactly this way when an inactivity auto-stop began seconds before the
stage finished — activate() saw the sandbox mid-stop and failed the
whole run 35ms later. The cleanup stop() then failed on the same
rejection.

Transitions finish on their own within seconds, so treat them as
wait-and-retry conditions:

- activate() now waits out a Stopping sandbox and dispatches on
  whatever state the transition lands on.
- start() and stop() retry the rejected call within a bounded budget,
  re-inspecting state between attempts: a transition that lands on
  Started needs no further start, and one that lands on Stopped or
  Destroyed needs no further stop.

All call sites go through these three provider methods, so no
lifecycle-layer changes are needed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 17:35:11 -04:00
Bryan Helmkamp
0845c331cb
Default Daytona auto-stop to 120 minutes
Omitting autoStopInterval from the create-sandbox request inherits
Daytona's server-side default of 15 idle minutes. Daytona counts
inactivity from the last sandbox interaction, and LLM inference never
touches the sandbox, so a single long inference call is enough for the
sandbox to auto-stop mid-run: a workflow failed exactly this way, with
the sandbox entering its stop transition 15 minutes after the last
command while the agent was still thinking.

Send an explicit 120-minute default when lifecycle.auto_stop is unset.
That clears any realistic inference call while still reclaiming
sandboxes leaked by a dead worker. An explicit auto_stop = "0s" still
disables auto-stop entirely.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 17:26:13 -04:00
Scott Werner
ebf6f92724 Harden exact-commit checkout in clone-based sandboxes
Run the local git steps of the Docker exact checkout under the shared
clone deadline instead of a fixed 10s timeout, so materializing a large
working tree cannot time out and abandon a running checkout in the
container.

Check the admitted commit out onto the admitted branch rather than
detaching. A detached HEAD makes `rev-parse --abbrev-ref HEAD` return
"HEAD", which the git setup helper maps to no base branch, silently
dropping it for callers that rely on it. Daytona does the same after its
native clone and now verifies the resulting HEAD the way Docker does.

Fetch the exact commit at the same depth a branch clone uses, so both
paths can reach the same number of parent commits, and stop suggesting
GitHub App credentials when a purely local git step fails.

Document that reachability of the commit from the branch is an
admission-time invariant that the sandbox layer does not re-verify.
2026-08-20 13:41:54 -04:00
Scott Werner
e7a32d12d5 Use native Daytona exact commit checkout 2026-08-20 12:16:12 -04:00
Bryan Helmkamp
1688cd5b91
Retry git pushes with a pinned token and record attempt history
Run 01M0DH033P2XSTHAGVBHG6922F completed 2.8 hours of work, then failed
terminally because four consecutive publish pushes hit GitHub's
token-replication lag (404 "Repository not found") — the push path had no
retry, the failure was misclassified as deterministic, and the same
fresh-mint-then-push pattern silently disabled metadata snapshots. This
generalizes the clone retry machinery to pushes and makes attempt detail
durable.

- clone_retry -> git_retry: the classifier's boolean becomes a
  CredentialContext derived from the token snapshot (fresh App tokens retry
  404s as replication lag, mature ones as transient infra, static
  credentials fail fast), and the attempt/backoff limits become a RetryPlan
  with layered optional bounds. Clone behavior is preserved: Docker keeps
  its absolute five-minute deadline, Daytona keeps no deadline.
- Pushes take a scoped CredentialLease before the first attempt: it owns
  the embed mutex for the whole operation, pins the single successful
  resolve, retries only failed resolves, falls back to the last embedded
  token when a mint fails, and force-re-embeds the pinned token once after
  the first auth-shaped failure (drift repair). The margin invariant
  (REFRESH_MARGIN > every push plan's max_elapsed) guarantees the pinned
  token outlives the operation; a unit test asserts it.
- Sandbox::git_push_ref now takes a RetryPlan and returns PushReport /
  PushError with per-attempt records (classification, redacted output tail,
  token generation/provenance/age, credential action, refresh errors).
  Checkpoint pushes use a 90-second budget; the terminal publish push gets
  5 attempts over at most 4 minutes.
- The single durable git.push event per push gains a nested attempts array
  (GitPushAttemptProps, token snapshot flattened to flat fields); stored
  events without it still deserialize. Publish push failures now carry an
  explicit failure category — exhausted transient retries stay
  transient_infra instead of deterministic — plus one bounded cause line
  per attempt and the last successful push time in the message.
- Metadata snapshot degradation records why it degraded: push failures with
  retryable classifications leave the writer eligible to re-probe at each
  later checkpoint, and a successful snapshot clears the degraded state and
  re-arms the warning. Permanent failures keep today's latch.

Plan: .ai/plans/git-push-token-resilience.md (PR 2: items 1, 2, 4, 7 and
the metadata re-probe).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 10:16:10 -04:00
Bryan Helmkamp
579f3db26f
Add a cached GitHub installation-token source for push credentials
Every push previously re-minted a fresh GitHub App installation token and
embedded it in the origin URL, so pushes routinely landed inside GitHub's
token-replication lag window (run 01M0DH033P2XSTHAGVBHG6922F failed
terminally on four consecutive fresh-token 404s). Reusing mature tokens
removes the failure trigger and saves two GitHub API calls plus one sandbox
exec per push.

- New fabro_github::token_source::InstallationTokenSource: one cached,
  single-flight source per origin repo. Static credentials pass through
  (generation 0); App credentials mint through the cache and reuse tokens
  until REFRESH_MARGIN (10 min) before expiry. Every resolve returns a
  non-secret TokenSnapshot (generation + Minted/Reused/Static provenance),
  and the source logs mints at INFO and reuses at DEBUG.
- Docker and Daytona share the source through PushCredentialState: an embed
  mutex serializes compare -> set-url -> record, a matching generation skips
  the set-url exec, and the generation is recorded only after a successful
  exec. The clone still mints its own token, but now seeds the source cache
  (generation 1) and the last-embedded state, so a refresh mint failure
  falls back to the known embedded token instead of believing nothing was
  ever embedded.
- RefreshOutcome now reports the remote action (embedded/unchanged/none)
  separately from the token snapshot; git_push_via_exec logs token age and
  provenance with each push, and refresh failures log the last embedded
  generation.
- The run-metadata writer resolves through the sandbox's shared source
  instead of minting per snapshot (with its own cached source on resume).
- The ACP refresh-ahead loop reschedules from the embedded token's
  expires_at minus the margin instead of a fixed 45-minute interval, which
  a cached source would have broken for long turns; static credentials stop
  the loop.

Plan: .ai/plans/git-push-token-resilience.md (PR 1: items 3 and 6).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 09:33:43 -04:00
Scott Werner
2168d902f0
Merge pull request #762 from fabro-sh/refactor/shared-run-spec-test-fixture
Some checks are pending
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
Add a shared RunSpec test fixture so additive fields stop churning tests
2026-08-19 17:55:59 -04:00
Scott Werner
eea868647f
Merge pull request #760 from fabro-sh/codex/blob-roundtrip-tests
Test blob offloads through production hydration
2026-08-19 17:44:19 -04:00
Scott Werner
19aa5940ea Add a shared RunSpec test fixture and adopt it
`RunSpec` has 13 fields and no `Default`, so every test that needed one
spelled out all 13 even when it cared about one or two. That put 64
hand-rolled `RunSpec { .. }` literals in `lib/`, and made a single
additive field cost a mechanical edit at roughly 30 sites.

Add `test_run_spec()` to `fabro-types`' feature-gated `test_support`
module: fixed `fixtures::RUN_1`, default settings, a minimal `test`
graph, `test_run_provenance()`, and every optional field unset. Tests
now spread it and only spell out what they assert on.

Adopt it at the 13 literals where the spread removes real duplication,
including the crate-local `test_run_spec` helpers in `fabro-store` and
`fabro-workflow`, which are now defined in terms of the shared fixture.
Tests that populate every field on purpose — the exhaustive `RunSpec`
serde round-trip in particular — keep spelling it out.

No production code and no behavior changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 16:50:37 -04:00
Scott Werner
65616b4557 Simplify exact-commit checkout across sandbox providers
- Fold Docker's exact-checkout path into clone_github_repo so the auth,
  retry, symlink, and bookkeeping skeleton is shared with branch clones
- Skip the Daytona SDK clone for exact checkouts: init and shallow-fetch
  the admitted commit directly instead of cloning the default branch and
  discarding it
- Combine the detach checkout and HEAD verification into one shell
  command, saving an exec round trip per init
- Drop the spec-level decide_clone pre-checks that duplicated the
  constructors' fail-fast validation
- Share a CloneAttemptFailure struct in clone_retry and the GIT command
  prefix constant across git command builders

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 15:10:44 -04:00
Scott Werner
facc6a02f2 Test blob offloads through production hydration 2026-08-19 14:10:02 -04:00
Scott Werner
7b47ef2d05 Cover missing SQLite blob reads 2026-08-18 17:43:22 -04:00
Scott Werner
01efe7c883 Document BlobBackend as a transitional enum
Mark the Slate arm as temporary and record that the SQLite arm's
verified-read and hash-conflict semantics are the intended end state,
so the dual-backend enum reads as a rollout vehicle rather than a
permanent abstraction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 17:41:29 -04:00
Scott Werner
cc16362528 Add SQLite blob store foundation 2026-08-18 17:41:29 -04:00
Scott Werner
2575ab85fc
Merge pull request #747 from fabro-sh/codex/blob-hash-vocabulary
Some checks failed
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
TypeScript / Typecheck (push) Has been cancelled
TypeScript / Test (push) Has been cancelled
TypeScript / Build (push) Has been cancelled
Unify blob hash vocabulary
2026-08-18 15:55:43 -04:00
Scott Werner
3e6b23ce76 Keep loaded workflow-version closures out of implicit copies
LoadedWorkflowVersionClosure owns every file of every version in the
dependency graph, so an advertised Clone invites accidental deep copies
of the whole set. Drop the derive until a consumer needs owned copies.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 13:02:54 -04:00
Scott Werner
1e29347227 Cover rejection of broken transitive includes under file run goals
The positive run-goal tests only asserted fixture shape, so a regression
that stopped pushing the file-goal template root would keep them green
while broken nested includes were silently accepted. Pin the rejection
path directly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 13:02:54 -04:00
Scott Werner
2f2097be54 Anchor run-goal template validation at the version entrypoint
Create-time validation of workflow.toml run goals anchored includes at
workflow.toml for inline goals and at the goal file's directory for file
goals, while the run engine inlines the effective goal into the
entrypoint graph and renders it under the entrypoint's template source.
That divergence rejected layouts `fabro run` executes fine and accepted
layouts that fail at render time. Anchor both goal forms at the
entrypoint so validation matches the runtime, and pin the anchor with a
nested-entrypoint test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 13:02:54 -04:00
Scott Werner
5a5cfbdaa0 Parse template dependencies whose paths collide with discovery roots
Batched dependency discovery pre-seeded roots into the path-keyed result
map and reused that map as the traversal-dedup set, so a loaded include
target whose path matched a root (e.g. a goal template including the
graph file that anchors an inline prompt) was recorded but never parsed,
silently accepting invalid template content that per-root discovery used
to reject. Dedup traversal on the full (path, root, content) occurrence
instead, which also stops re-parsing identical duplicate roots.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 13:02:25 -04:00
Scott Werner
408cd2f745 Simplify workflow-version closure validation and loading
- Replace the discarded dependency-closure map in put/get with a
  visitor-based walk so only get_closure retains loaded versions
- Hold the closure root structurally in LoadedWorkflowVersionClosure
  instead of asserting its presence in the map with expect()
- Drop the visited-set parameter that guarded against impossible
  content-address cycles
- Move template-discovery error source-name extraction into
  TemplateDiscoveryError::source_name() where the variants are owned
- Collapse repeated TemplateSource construction into a TemplateRoots
  collector and share the config file-reference validation pipeline
  between dockerfile and run-goal references
- Deduplicate test helpers (version_id, version_with_goal_file,
  impl Into<String> config fixtures)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 13:02:11 -04:00
Scott Werner
63025bb748 Close workflow goals over version dependencies 2026-08-18 13:01:03 -04:00
Scott Werner
a0845d8346
Merge pull request #756 from fabro-sh/refactor/graph-reference-kind
Classify graph attributes with a graph-only reference kind
2026-08-18 12:57:32 -04:00
Bryan Helmkamp
1226ed7377
Cite Fireworks' documentation for the 412 mapping
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 12:10:50 -04:00
Bryan Helmkamp
18a71ac310
Classify provider 412s as failover-eligible account lockouts
Fireworks reports an account suspension (spending cap reached or unpaid
invoices) as HTTP 412 with code PRECONDITION_FAILED. The status had no
explicit mapping, and the openai_compatible dialect extracts error.type
("error") as the code, so the suspension fell through to InvalidRequest
-- a deterministic request defect -- which suppressed both retry and the
configured model fallback chain. A live run then died mid-stage with
five healthy fallback candidates configured.

No LLM request carries conditional-request preconditions, so a 412 is
never about the request. Map it to AccessDenied, the same family as the
account_deactivated error code: non-retryable on the same provider,
eligible for failover to a provider with independent billing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 12:08:36 -04:00