Add ANSI color to `arc models` and `arc models test` output when stdout
is a TTY: bold model IDs, dim provider/aliases, cyan speed, green/red
test results. Add `Styles::detect_stdout()` to arc-util.
Also fix pre-existing clippy warnings: derive Default instead of manual
impls for enums in server_config, remove unused FailureDetail imports
in arc-workflows error tests, inline print literal in test_models header.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Extract parse_sandbox_provider, resolve_sandbox_provider, and
resolve_daytona_config helpers to eliminate 4 near-identical sandbox
parsing chains. Compute sandbox_provider once with proper error handling
instead of twice (preview silently swallowed errors). Fix bug where
run_preflight did not fall back to run_defaults for daytona config.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Rename AppConfig → ServerConfig with flattened RunDefaults so users can
set default llm, sandbox, setup, directory and vars in ~/.arc/arc.toml.
Precedence: CLI flags > workflow TOML > server config defaults > DOT
graph attrs > hardcoded defaults. Vars merge (defaults first, task
config overwrites collisions).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Better reflects domain semantics: the config describes a workflow run,
and the free-text field is the run's goal, not a generic "task".
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Remove RunMode enum (only Preflight was checked; DryRun/Normal unused)
— use args.preflight directly
- Switch default_model_for_provider() to take Provider enum instead of
Option<&str> for exhaustive matching; adds missing Inception default
- Extract sandbox init/cleanup into creation + single check pattern,
eliminating 3x copy-paste of identical init/cleanup boilerplate
- Replace setup_commands Vec clone with direct count (only .len() used)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Verifies sandbox boot (local/docker/daytona), LLM provider availability,
and model/provider resolution chain, then prints a structured report.
Extracts model/provider resolution into reusable helpers
(default_model_for_provider, resolve_model_provider) with tests.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Parallel branches now include the node visit count (pass1, pass2, etc.)
in their ref names, preventing silent overwrite when a parallel node is
re-executed via retry or loop_restart.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Rename node visit suffix from `-attempt_{V}` to `-visit_{V}` and move
asset collection under `artifacts/assets/` alongside artifact values in
`artifacts/values/`. Retry directories use `retry_{N}` instead of the
ambiguous `attempt_{N}`.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Automatically discovers and collects well-known output files (test reports,
screenshots, trace files) from each node's execution into its log directory.
Works across all three sandbox types (local, Docker, Daytona) via a new
`download_file_to_local` trait method that handles binary files correctly.
- New `asset_snapshot` module with pure functions for find command generation,
output parsing, candidate matching, and budget enforcement
- `Sandbox::download_file_to_local` implemented for Local (fs::copy),
Docker (host bind-mount resolution), and Daytona (SDK download)
- `AssetsCaptured` event variant with DEBUG-level tracing
- Engine integration: baseline snapshot before handler, collection after
(both success and error paths), non-fatal on errors
- E2e tests for local sandbox (2 tests), Docker (#[ignore]), Daytona (#[ignore])
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add node_id to Stage* enum variants so both the programmatic ID and
display label are available. Add rename_fields() post-processing in
flatten_event() to give flattened JSONL fields self-describing names:
- timestamp → ts (save space)
- name → node_label (Stage*), workflow_name, snapshot_name
- index → stage_index, branch_index, command_index
- stage → node_id (Agent.*, Interview*, Prompt)
- branch → node_id (ParallelBranch*)
- node → node_id (StallWatchdogTimeout)
- from_node/to_node → from_node_id/to_node_id
- start_node → start_node_id
- provider → sandbox_provider (Sandbox.*)
- text → prompt_text (Prompt)
- Insert node_label defaulting to node_id where only an id exists
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Instead of flattening the inner event's fields into the top level
(which causes collisions when SubAgentEvent wraps SubAgentEvent —
losing inner agent_id, depth, and nested event body), keep the inner
event as a `nested_event` JSON value. The dot-notation event name
still reflects the inner type for easy filtering.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
WorkflowRunStarted and GitCheckpoint have a `run_id` field that
collides with the envelope's top-level `run_id`. Filter out
`timestamp`, `run_id`, and `event` from event fields to avoid
duplicate keys in the JSONL output.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Event data was nested inside a tagged enum (`"event": {"StageStarted": {fields}}`).
Now `event` is a string name and fields merge into the top-level object
(`"event": "StageStarted", "name": "plan", ...`).
Nested events use dot notation:
- Agent wrapper: "Agent.ToolCallStarted" with stage at top level
- Sandbox wrapper: "Sandbox.Initializing" with fields at top level
- SubAgentEvent: "Agent.SubAgentEvent.ToolCallStarted" flattened one level
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The `-vv` (full detail) mode was not useful in practice. This collapses
the two-tier `-v`/`-vv` into a single `--verbose` boolean and removes
the now-dead `format_event_detail` function and its tests.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adopt cleaner terminology: Sandbox (resource providing disk + execution),
SandboxProvider (Local/Docker/Daytona), SandboxEvent (lifecycle events),
and Snapshot (pre-built environment images). Flatten DaytonaSandboxConfig
into DaytonaConfig, rename CLI flag to --sandbox, and update TOML config
sections from [execution] to [sandbox].
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
No DOT pipelines reference the sub_pipeline handler type. The manager
loop handler now covers the child pipeline spawning use case.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The manager loop (house node) now parses a child DOT pipeline, spawns a
real PipelineEngine, and monitors it via a tokio::select poll loop. Context
is cloned into the child and diffed on completion to propagate updates back
to the parent.
- Add PipelineEngine::from_services() to share parent's Arc services
- Add PipelineEngine::run_with_context() returning (Outcome, Context)
- Rewrite ManagerLoopHandler as unit struct with real child engine spawning
- Support both inline DOT (stack.child_dot_source) and file (stack.child_dotfile)
- Delete ChildObserver trait entirely
- Update all existing tests and add new e2e tests for context flow and dotfile
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace hardcoded /home/daytona/workspace/.arc-parallel path with
{working_directory}/.arc/logs/{run_id}/parallel/{node_id}/{branch_key},
matching the host worktree layout and scoping worktrees per-run.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add `labels: HashMap<String, String>` to RunConfig, written to manifest.json
- Add `--label KEY=VALUE` flag to `arc run` (repeatable)
- New `arc runs` command: list pipeline runs with table or --json output
- New `arc runs prune` command: delete old runs with --before, --pipeline,
--label, --orphans filters (dry-run by default, --yes to confirm)
- 13 new tests covering scan_runs, filter_runs, prune, and manifest labels
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
A background tokio task polls EventEmitter.last_event_at(). When idle
time exceeds the graph-level stall_timeout (default 600s), it cancels
a CancellationToken that races against execute_with_retry via
tokio::select!, dropping the hung handler future.
- EventEmitter: AtomicI64 last_event_at field, touch() to seed, emit()
auto-updates
- Graph::stall_timeout(): reads Duration attr, defaults 600s, None for 0
- StallWatchdogTimeout event variant with warn-level trace()
- Engine: watchdog spawn before main loop, select! at handler call,
shutdown after loop
- CLI: format arms for summary and detail views
- Unit tests for emitter, graph accessor, event serialization, and
engine watchdog behavior (hung, keepalive, disabled)
- E2e integration tests: DOT-parsed pipelines with 100-200ms stall
timeouts verifying trigger, keepalive, disable, and timing
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Every emitted event now produces a structured tracing log line so
developers can debug after the fact via ~/.arc/logs/. Each event
variant gets an appropriate log level (info/debug/warn/error) with
structured fields. Streaming noise variants (TextDelta,
ToolCallOutputDelta) are no-ops, and wrapper variants (Agent,
ExecutionEnv on PipelineEvent) delegate to the inner event's trace.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
SdkError now produces hand-crafted signature hints (e.g.
"api_deterministic|openai|authentication") that are identical regardless
of error message wording, replacing fragile regex-normalized signatures
for API errors. ArcError.to_fail_outcome() centralizes fail outcome
construction with failure_class and failure_signature context_updates.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Guard loop_restart edges to only allow transient_infra failures, matching
Kilroy's behavior. Non-transient failures (deterministic, structural,
budget_exhausted, canceled, compilation_loop) are now blocked immediately
instead of getting restart attempts before the circuit breaker fires.
Fix normalize_failure_reason truncation to use floor_char_boundary(240)
instead of a raw byte slice, preventing panics on multi-byte UTF-8 chars.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
max_node_visits used > while signature circuit breakers used >=, creating
an inconsistency where "limit" meant different things depending on the
mechanism. Kilroy uses >= for all three checks. This aligns Arc to match.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add wall_clock_timeout field to SessionConfig that spawns a tokio timer
to cancel the session after a duration, reusing existing Aborted path
- Set 120s wall-clock timeout on retro agent to prevent unbounded runs
- Replace silent `let _ =` with logged warnings for checkpoint load and
retro save failures
- Reuse LLM client from initial from_env() call instead of creating a
second one for the retro agent
- Extract retro generation into generate_retro() helper and call it from
both run_command and run_from_branch (resume path)
- Tolerate mutex poisoning in retro agent with unwrap_or_else
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add git_cmd() helper disabling maintenance.auto and gc.auto on all
host-side git commands; add GIT_REMOTE constant for remote commands
- Use --force on branch creation for idempotent retry/resume
- Add replace_worktree() that does best-effort remove before add
- Add reset_hard() after parallel worktree setup for deterministic state
- Add sanitize_ref_component() to clean node IDs in branch names
- Add git_replace_worktree_remote() for remote sandbox environments
- Add tests for sanitize_ref_component, replace_worktree, reset_hard
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Instead of skipping the retro agent entirely in dry-run, use a
placeholder narrative so derive → apply_narrative → save all run.
This catches bugs in the merge/persistence path without LLM calls.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
After each pipeline run, auto-derive stats from the checkpoint (stages,
retries, cost, files touched) then run an Opus agent session that
explores progress.ndjson to produce qualitative analysis: smoothness
rating, intent, outcome, learnings, friction points, and open items.
Backend:
- retro.rs: data model, save/load, derive_retro(), extract_stage_durations()
- retro_agent.rs: post-pipeline agent session with submit_retro tool
- cli/run.rs: hook retro generation after final.json, before engine_result?
- server.rs: GET /pipelines/{id}/retro endpoint, auto-derive on completion
Frontend:
- data/retros.ts: TS types + mock data + smoothness color config
- routes/retros.tsx: list page with smoothness badges
- routes/run-retro.tsx: detail view (stats, intent, stages, learnings)
- routes.ts + run-detail.tsx: wire up retro route and tab
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Parallel branches now get isolated git worktrees so concurrent file
writes don't collide. Works across Local, Docker (bind-mount), and
Daytona (remote exec_command) environments.
Key changes:
- git.rs: add create_branch_at() and merge_ff_only() helpers
- engine.rs: add GitState struct, remote worktree helpers
(git_create_branch_at_remote, git_add_worktree_remote, etc.)
- handler/mod.rs: add git_state field to EngineServices (RwLock)
- handler/parallel.rs: WorktreeEnv wrapper, per-branch worktree
setup/teardown, checkpoint commits per branch, ff-merge winner
before returning to engine
- handler/fan_in.rs: ff-merge to winner's HEAD, set best_head_sha
- E2E tests for Host (local) and Daytona (remote) modes
When git_state is None, behavior is unchanged.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
LLM-authored output can set failure_class to non-canonical strings like
"retryable", "transient", or "permanent". Expand FromStr to accept 30+
aliases with case-insensitive trimmed matching, matching Kilroy's
normalizedFailureClass(). Unknown values fail-closed to Deterministic
instead of returning Err.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Absolute `cd` paths in shell commands (script/tool_command attributes) silently
override the engine's worktree CWD, breaking portability across machines,
containers, and worktrees. Ported from kilroy (danshapiro/kilroy d9c1fec).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Fix 4 test failures: add unconditional fallback edges to branching.dot
and conditions.dot to satisfy all_conditional_edges validation rule
- Fix clippy await_holding_lock: scope MutexGuard before await in
daytona_integration.rs
- Fix clippy unnecessary_get_then_check: use contains_key in script.rs
- Fix clippy expect_fun_call: use unwrap_or_else in integration.rs
- Run cargo fmt across entire workspace
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add 8 missing transient_infra patterns (crates.io registry, toolchain,
cross-device link errors), 2 structural hints (write_scope_violation
variants), and reorder heuristic priority to check transient_infra
before budget_exhausted to match Kilroy's classification behavior.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Raise the Anthropic adapter fallback from 4096 to 16384 to prevent
truncation of large tool call JSON when the model isn't in the catalog.
Add max_tokens as a configurable DOT node attribute that flows through
SessionConfig to LLM requests, following the same pattern as
reasoning_effort. Priority: node attribute > catalog > provider default.
Ported from kilroy (danshapiro/kilroy) commits 99a5cd7 and 78fadad.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Extract hint patterns into const arrays (TRANSIENT_INFRA_HINTS,
BUDGET_EXHAUSTED_HINTS, STRUCTURAL_HINTS) and add 28 new patterns
from Kilroy to prevent transient/budget failures from misclassifying
as deterministic. Add comprehensive regression test per pattern plus
count-guard tests to catch accidental additions/removals.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>