Previously sandbox env vars only reached CLI backend agents but not
API backend tool calls. Thread tool_env through ToolContext so shell
and web_fetch tools pass env vars to exec_command for all backends.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Hoist OneShotCapturingBackend to module scope and remove two identical
CapturingBackend definitions that duplicated its functionality.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Prompt nodes now discover project docs (AGENTS.md, CLAUDE.md, etc.)
and pass them as a system prompt to one_shot LLM calls. The
project_memory attribute defaults to true and can be set to false
to disable this behavior.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Allow passing environment variables into sandbox command execution via
`[sandbox.env]` in TOML configs. Supports literal values and host env
passthrough via `${env.VARNAME}` syntax (whole-value only, missing vars
are hard errors). Env vars are injected into command nodes via
`cmd.envs()` and into CLI backend agents via the sandbox env file.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Extract read_existing_id() helper to deduplicate file-read-trim-check
pattern, and extract private Telemetry::new() to consolidate for_server/for_cli.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Server is long-lived on a fixed host, so a persisted UUID at ~/.arc/.id is
appropriate. CLI runs ephemerally, so an MD5 of the MAC address avoids file
I/O and is stable per-machine. CLI falls back to ~/.arc/.id if it exists
for migration.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds a telemetry library for product analytics with Segment integration.
Includes Track/User wire types, persistent anonymous ID (~/.arc/anonymous_id),
OS/arch/locale context, fire-and-forget sender via tokio::spawn, and
ARC_TELEMETRY env var control (off/errors/all). No CLI or server integration yet.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Cloud sandboxes clone the repo, so untracked local files referenced via
prompt="@path/to/file.md" won't exist. This inlines those file contents
at prepare time while leaving git-tracked @references for the agent to
read from the sandbox.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Extract ContentPart::is_opaque_openai() to deduplicate matches! patterns
- Use std::mem::take to avoid cloning reasoning/message items
- Replace hand-rolled DockerfileSource serde with derive + untagged enum
- Add --fail-with-body to imagegen curl for better error reporting
- Add tmp to .gitignore
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Allows overriding the workflow goal from the command line, which is
exposed as $goal in node prompts via VariableExpansionTransform.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add DockerfileSource enum (Inline/Path) with custom serde to allow
referencing a Dockerfile by path instead of embedding content inline.
Paths are resolved relative to the config file directory during loading.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace 16 raw string literal usages of "openai_reasoning" and
"openai_message" across openai.rs and history.rs with constants
defined on ContentPart, eliminating typo risk and centralizing
the kind identifiers.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
When an assistant turn has reasoning + text + tool calls, the message
output item was reconstructed without its original `id` and `status`
fields. The Responses API requires reasoning items to be followed by
a valid output item identified by `id`, so the reconstructed message
was not recognized, causing "Item 'rs_...' was provided without its
required following item" errors.
Preserve the full message output item as an opaque `openai_message`
provider part (like we already do for `openai_reasoning`), and use it
in translate_input instead of constructing a new message from text.
Also strip `openai_message` items during compaction alongside reasoning.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Implement the server-side session handlers (create, retrieve, send message,
stream events, list) with in-memory storage and LLM generation, wire them
into the router replacing not_implemented stubs, add run_chat_via_server
CLI function with SSE streaming, and add mode dispatch so `arc llm chat
--mode server` delegates to the API server.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds a completions API endpoint that supports both streaming (SSE) and
non-streaming (JSON) modes, with structured output via JSON Schema.
Wires up the CLI `arc llm prompt` command to use the server when
`--mode server` is specified.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
After compaction replaces old turns with a summary, preserved Assistant
turns may contain opaque openai_reasoning items that reference the now-
removed context. These orphaned items violate the OpenAI Responses API
constraint that reasoning items must be followed by their paired output,
causing "Item 'rs_...' was provided without its required following item"
errors. Strip them in a new strip_opaque_reasoning() method called at
the end of compact(). Anthropic thinking blocks are left untouched.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Implements the Sandbox trait backed by the `sprite` CLI binary.
Includes 29 unit tests with mock runner and an e2e integration test
against the live Sprites service.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
GPT-5.4 frequently emits bare @@ instead of @@ context @@, causing all
Update File patches to fail. The parser now accepts bare @@ and locates
the hunk position from the first remove/context change line. Also
improves the system prompt with an explicit @@ context @@ example.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Consolidates the two separate String parameters (git_author_name,
git_author_email) threaded through ~10 function signatures into a
single GitAuthor struct with Default providing "arc"/"arc@local".
Also quotes git config values in parallel.rs shell commands.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The OpenAI Responses API sends reasoning_summary_text.delta and
reasoning_text.delta SSE events during extended reasoning, but the
OpenAI provider silently swallowed them. This caused the tight-loop
in process_next_sse_events to consume events without yielding any
StreamEvent, so emitter.touch() was never called and the 600s stall
watchdog fired during long reasoning phases (e.g. GPT-5.4-pro).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Users can now configure the git author name/email used for checkpoint
commits via [git.author] in server.toml (default) and cli.toml (override).
Defaults to "arc" / "arc@local" preserving current behavior.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Move emitter.touch() before the event filter in spawn_event_forwarder
so streaming events (TextDelta, AssistantTextStart, etc.) also reset
the watchdog timer. Previously these were filtered out, causing the
600s stall watchdog to fire during long generation turns with no
tool calls.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
After each checkpoint commit, the metadata branch (checkpoint.json, manifest,
graph DOT, artifacts) is now pushed from the host process to the GitHub remote
using a GitHub App installation token. The local custom ref (refs/arc/{run_id})
is mapped to refs/heads/arc/meta/{run_id} on the remote since GitHub rejects
branch names starting with "refs/".
Changes:
- Move ssh_url_to_https to github_app.rs as pub fn for reuse
- Add push_ref() to git.rs for pushing a ref to an explicit URL
- Add github_app field to RunConfig to thread credentials into the engine
- Add git_push_meta_host() async wrapper in engine.rs
- Call git_push_meta_host after each remote checkpoint
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Introduces a [checkpoint] config table with exclude_globs in both run.toml
(per-run) and server.toml (defaults). Globs are merged (union + dedup) when
both are present. Non-empty excludes use git pathspec :(glob,exclude) syntax
to prevent staging matching files during checkpoint commits.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
`arc serve` writes to `serve-YYYY-MM-DD.log` and all other commands
write to `cli-YYYY-MM-DD.log` so the two are easy to tail independently.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Enables `arc models test` to work in server mode by adding an API
endpoint that sends "Say OK" (max_tokens=16, 30s timeout) to a model
and reports pass/fail. Dry-run mode returns synthetic "ok" status.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Commands can now delegate to a running Arc API server instead of
executing in-process. Adds ExecutionMode, ServerDefaults, and
ClientTlsConfig to cli.toml parsing with CLI flag > config > default
precedence. The models list command fetches from GET /models when in
server mode, with mTLS client certificate auth when [server.tls] is
configured.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replaces flat ModelInfo fields with nested sub-structs (ModelLimits,
ModelFeatures, ModelCosts) and adds family, training, and
cache_input_cost_per_mtok fields to enrich the model catalog.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Filter engine-internal keys (internal.*, graph.*, thread.*, current*)
from the context diff returned by SubWorkflowHandler, preventing child
run state from overwriting parent values. Pass the parent's preamble
into the child context so child workflows have awareness of what the
parent already accomplished.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
ModelInfo already derives Serialize with matching field names, so the
manual json!({...}) mapping was redundant. Matches the pattern used by
every other paginated handler in demo/mod.rs.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Exposes the embedded model catalog (same data as `arc models list`) via
a new authenticated API endpoint so the web UI can display available models.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Restructure the verification API from a flat `/verifications` namespace to
`/verification/criteria` and `/verification/controls` as distinct resources.
Singularize the run sub-resource path to `/runs/{id}/verification`.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The extract_status_fields function recognized "outcome" as a field for
detecting status JSON objects but never read its value — LLM responses
like {"outcome": "fail", "failure_reason": "tests failed"} were silently
ignored and the outcome was always Success.
Now extract_status_fields reads the outcome field to set the node status
and failure_reason to populate the failure detail. Also adds a fallback:
if no routing directives are found in the response text, the handler
reads status.json from the sandbox CWD (written by agents that prefer
file output over inline JSON). Response text always takes priority.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
New `arc-exe` crate that runs agent tool operations inside ephemeral
exe.dev VMs via SSH. Uses two SSH connections: a management plane
(`ssh exe.dev`) for VM lifecycle and a data plane (`ssh vmname.exe.xyz`)
for command execution and file I/O.
Includes SshRunner trait with MockSshRunner for unit tests and
OpensshRunner for real SSH, with raw_mode for the exe.dev management
plane's custom command handler.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
OpenAI's Codex Spark is a smaller, faster codex derivative on Cerebras
hardware (1000 tok/s, 128K context, text-only, no pricing yet).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The active_stages map was only populated inside a TTY renderer guard,
so Plain mode never tracked counts. Move counters to a separate
stage_counts map that is always populated regardless of renderer.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
resolve_clone_credentials was short-circuiting for public repos,
returning no token. This broke git push from the sandbox since push
requires authentication regardless of repo visibility. Always generate
an installation access token.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Replace apt-get/NodeSource install (requires root) with direct Node.js
binary download to ~/.local (works as non-root daytona user)
- Run node install + npm install in single shell so PATH persists
- Add ~/.local/bin to PATH in env file and version check
- Fall back to stdout for error details when stderr is empty (Daytona
always returns empty stderr)
- Add e2e assertion that cli_stdout.log is written during poll
- Verified on Daytona with haiku: ensure_cli installs in 2s, full
workflow succeeds
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
AgentCliBackend now detects missing CLIs at runtime and installs them
on-demand (including Node.js via NodeSource if needed), removing the
need for custom Dockerfiles that pre-install CLI tools.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Remove hardcoded default_model_for_provider() from arc-workflows and
default_model() from arc-agent, delegating both to the catalog via
arc_llm::catalog::default_model_for_provider(). Add claude-sonnet-4-6
to catalog and move "sonnet"/"claude-sonnet" aliases to it.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Sync cli_stdout.log and cli_stderr.log to stage_dir each poll iteration
so there is visibility into what the CLI agent is doing before it finishes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Daytona's POST /process/execute blocks until all descendant processes
exit. The backgrounded claude process kept the API hanging, causing a
60-second HTTP timeout. Using setsid creates a new session so the child
is fully detached — the API now returns in ~200ms.
Also touch the event emitter during the poll loop to prevent the stall
watchdog from killing the stage while waiting for claude to finish.
Falls back gracefully on macOS where setsid isn't available (not needed
since the local exec implementation doesn't wait for grandchildren).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Daytona's POST /process/execute is synchronous and blocks until the
command finishes. Long-running CLI agent sessions (claude, codex, gemini)
cause HTTP proxy timeouts. Replace the single blocking exec_command with
a background launch + poll pattern:
- Generate UUID-based temp file paths to avoid collisions between
concurrent CLI nodes
- Disable sandbox auto-stop before launching (new Sandbox trait method)
- Launch command in background, capture PID
- Poll every 5s for exit code file
- Read stdout/stderr from temp files after completion
- Cleanup temp files
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>