Each SSE frame now includes a sequential numeric id: field. Clients can
reconnect with the Last-Event-ID header to resume the stream after the
last received event, skipping already-processed events.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add updated_at to CreateSessionResponse for consistency with SessionListItem/SessionDetail
- Add format: uuid to session ID fields and parameter across the OpenAPI spec
- Remove SessionEvent discriminated union and SessionEvent* wrapper schemas that conflated
SSE transport-level event names with JSON data payload fields
- Update demo data to use proper UUIDs instead of string IDs
- Add uuid dependency to arc-types crate
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add created_at to AssistantTurn and ToolTurn schemas (matching UserTurn)
- Document SSE event types (assistant_turn, tool_turn, done, error) with
SessionEvent discriminated union schema
- Add title and model to CreateSessionResponse
- Add model and last_message_preview to SessionListItem
- Update demo data with timestamps, model, and preview fields
- Regenerate TypeScript API client
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Mintlify requires all referenced files under docs/, so consolidate to a
single copy and eliminate the symlink and the copy step in the generate
script.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Refactor SessionTurn into discriminated union (UserTurn, AssistantTurn, ToolTurn)
using oneOf + discriminator so invalid states are unrepresentable
- Add format: date-time to all timestamp fields for proper codegen types
- Rename CreateSessionRequest.prompt to .content for consistency with SendMessageRequest
- Change sendSessionMessage from 200 to 202 (async processing via SSE)
- Extract inline response to SendMessageResponse schema
- Add updated_at to SessionListItem for sort-by-activity support
- Extract SessionId parameter to components/parameters (DRY)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add id, is_error, duration_ms to ToolUse; rename args to input
- Add created_at/updated_at timestamps to session schemas; replace time/date display strings
- Add descriptions and examples to all Sessions API fields and endpoints
- Flatten List Sessions response from grouped SessionGroup[] to SessionListItem[]
- Move date grouping (Today/Yesterday/etc.) to React client via groupSessionsByDate()
- Symlink docs/api-reference/arc-api.yaml to canonical openapi/arc-api.yaml
- Update React ToolRow components with duration display and error styling
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add a FeatureFlags config section with a session_sandboxes boolean
(default false) to both the Rust server config and web app config.
Gate the project/branch picker UI behind this flag. Set it to false
in the Docker demo config.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace the server-level `--demo` flag with per-request demo dispatch.
The Rust API builds both a demo and real router; incoming requests with
the `X-Arc-Demo: 1` header hit the demo router (auth disabled, static
data), all others hit the real router with normal auth.
The React web app gets a beaker icon toggle in the top nav bar (next to
the theme toggle) that sets an `arc-demo` cookie. Loaders read the
cookie to decide whether to send the `X-Arc-Demo: 1` header to the API.
The `ARC_DEMO=1` env var still works as a default when no cookie is set.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add three end-to-end tests for the GitHub App Installation Access Token clone
flow (private repo clone, public repo optimization, not-installed error). Also
return an error from is_repo_public on non-404 HTTP failures instead of
attempting to parse the response as JSON.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Share a single reqwest::Client across GitHub API calls in resolve_clone_credentials
to reuse the TLS connection pool
- Accept raw PEM (not just base64-encoded) in build_github_app_credentials, matching
the existing decode_pem_env/decode_pem_value convention
- Move github_app into DaytonaSandbox::new() instead of cloning
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace `gh auth token` with GitHub App IATs scoped to `contents: read` for
Daytona sandbox git cloning. Public repos are auto-detected and cloned without
credentials. Private repos get short-lived, repo-scoped tokens. Clear error
messages for each failure mode (app not installed, suspended, no repo access,
auth failure). Falls back gracefully when no GitHub App is configured.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add both models with aliases (gpt54, gpt54-pro). Neither replaces
gpt-5.2 as the OpenAI default. Update fallback chain test since
GPT-5.4 ($2.50) is now closer to Opus ($15) than GPT-5.2 ($1.75).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Per-node `max_visits` now overrides the graph-level `max_node_visits`
(and the dry-run default of 10) for individual nodes, giving tighter
control over specific loops like fix-and-verify cycles.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Allows literal $ signs in prompts and DOT files by writing $$. Also
refactors VariableExpansionTransform to use expand_vars instead of
string replace, fixing a substring-matching bug with $goal.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add DaytonaNetwork enum (block/allow_all/allow_list) with custom serde
Deserialize for TOML string-or-table syntax
- Wire network config through run_config defaults merging and base_params
- Document network access in sandboxing, environments, and run-configuration
- Fill in execution docs: checkpoints, environments, failures, interviews,
run configuration, observability, retros
- Rename compounding.mdx → retros.mdx, insights.mdx → observability.mdx
- Use DaytonaConfig::default() in tests to reduce boilerplate
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Aligns context key names with the handler name (CommandHandler), which was
previously inconsistent.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add follow_up, steering_before_input, and steering_mid_task parity tests
- Add steering_queue_handle() to Session for test access
- Expand provider_tests! macro to cover all 7 providers (Kimi, Zai, Minimax, Inception)
- Update model names to match current catalog (gpt-5-mini, gemini-3-flash-preview, etc.)
- Fix OpenAI-compatible tool result serialization: emit one ChatMessage per ToolResult
instead of bundling all results into a single message with only the first tool_call_id.
This was causing Kimi to reject requests during parallel tool calls.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Extract shared strip+parse+trailing-check into parser::parse_ast()
so both parser::parse() and parse_command reuse it
- Accept impl Write in parse_command so tests verify actual JSON output
instead of re-parsing independently
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Parses a DOT file and outputs its AST as pretty-printed JSON, useful for
debugging and tooling. Adds Serialize/Deserialize to all AST types.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Accommodates longer model names like gemini-3.1-flash-lite-preview
in both `models list` and `models test` output.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Replace useless format!() with .to_string() in parse_decision
- Derive Default for HookDecision instead of manual impl
- Use contains_key() instead of get().is_none() in semantic parser
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace freeform LLM text generation with schema-constrained
generate_object() for prompt hooks, eliminating the need for
JSON formatting instructions in the system prompt and the
code-fence stripping workaround.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Use thiserror for SlackApiError and ConnectionError (project convention)
- Remove misleading PartialEq/Eq on DispatchAction; use matches!() in tests
- Remove unused _slack_client and _default_channel params from event loop
- Extract check_ok() helper to deduplicate 3 ok-check sites in client.rs
- Make bot_token private on SlackClient; add http() accessor
- Make SlackClient Clone; eliminate duplicate instance in e2e example
- Reuse reqwest::Client in open_socket_url instead of creating a new one
- Filter empty env vars in resolve_credentials; remove redundant is_enabled
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Implements a complete Slack integration for the interviewer system using
Socket Mode (WebSocket-based, no public URL required). Supports all five
question types: YesNo, Confirmation, MultipleChoice, MultiSelect, and
Freeform (via thread replies with @mention).
Modules: config, client, blocks, interaction, socket, dispatch,
connection, threads. 72 unit tests + e2e example.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace flat Vec<Clause> parser with AST-based recursive descent parser
supporting full operator precedence (&& binds tighter than ||, ! is prefix).
New operators: >, <, >=, <= (numeric), contains (substring/array membership),
matches (regex, validated at parse time). Simplify ConditionSyntaxRule to
delegate entirely to parse_condition(). Public API unchanged.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The pattern Some("agent") | Some("agent_loop") | Some("prompt") |
Some("one_shot") was duplicated across preamble.rs (3x) and
validation/rules.rs (1x). Centralizes into a single function in
graph/types.rs. Also fixes stale doc comment on default_registry.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Keep legacy aliases (agent_loop, one_shot) in the handler registry
and validation rules for backwards compatibility. Add codergen_mode
attribute support in the DOT parser, translating legacy values to
the new type names.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Rename handler type strings: codergen → agent_loop, wait.human → human,
script → command, wait.timer → wait
- Rename handler modules/structs to match: AgentHandler, HumanHandler,
CommandHandler, WaitHandler
- Split one_shot into PromptHandler (handler/prompt.rs) with shape=tab mapping
- Remove CodergenMode enum and codergen_mode attribute — one_shot is now its
own handler type, not a mode flag on the agent loop handler
- Update all demo DOT files: codergen_mode="one_shot" → shape=tab
- Update spec, README, validation rules, preamble, and hook tests
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Extract shared prompt/agent hook setup (model resolution, system prompt,
user message, timeout wrapper) into reusable helpers
- Fix blocking I/O: std::process::Command → tokio::process::Command for
host-mode hook execution
- Cache reqwest::Client per TLS mode via OnceLock instead of rebuilding
per HTTP hook call
- Return Cow from resolved_hook_type() to avoid cloning HookType on
every call
- Rename command_executor → executor in HookRunner
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Change HookExecutor trait to take Arc<dyn Sandbox> so agent hooks can
share the sandbox with the ToolRegistry via ToolContext
- Replace hand-rolled 2-tool dispatch with register_core_tools() giving
agent hooks the full tool set (read_file, write_file, shell, grep, glob)
- Add strip_code_fences() to handle LLMs wrapping JSON in markdown
- Set max_tokens(1024) on prompt hooks to avoid exceeding model limits
- Add e2e tests: TOML parsing, prompt proceed/block, agent proceed,
agent with tool use (reads a file via read_file tool)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Prompt hooks make a single-turn LLM call returning {"ok": true/false}.
Agent hooks run a multi-turn LLM tool loop with sandbox access (exec_command, read_file).
Both fail-open on errors/timeouts. Prompt hooks default to 30s timeout,
agent hooks to 60s with max 50 tool rounds. Default model is "haiku".
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Introduce a `tls` field on HTTP hooks with three modes: `verify` (default,
requires https + cert validation), `no_verify` (requires https, skips cert
validation), and `off` (allows http, skips cert validation). This prevents
hooks from accidentally sending credentials over plaintext connections.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
HTTP hooks (type = "http") now actually execute instead of failing with
"no command specified". The executor POSTs the hook context as JSON,
parses HookDecision from the response, and fails open on errors.
Header values support $VAR/${VAR} interpolation gated by an
allowed_env_vars whitelist on the hook definition. Renames
CommandHookExecutor to HookExecutorImpl since it now handles both
command and HTTP hook types.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Replace JSON-based HookEvent::Display with simple match arms
- Use floor_char_boundary for Unicode-safe truncation in effective_name
- Cache compiled regexes in HookRunner instead of recompiling per check
- Simplify run_non_blocking (was run_parallel) to plain sequential loop
- Propagate hook_runner through parallel handler branch services
- Use unique temp file paths for sandbox hook context (avoid collisions)
- Add hook imports to engine.rs, reducing verbose crate:🪝: paths
- Extract duplicate RunFailed hook code into run_failed_hook helper
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Introduce a configurable hook system that triggers user-defined actions
at workflow lifecycle points (RunStart, StageStart, StageComplete,
StageFailed, EdgeSelected, CheckpointSaved, etc). Hooks can block
execution, skip nodes, or override edge routing via JSON decisions.
- New `hook/` module: types, config, executor (command), runner
- Engine instrumented at 8 lifecycle points with HookRunner calls
- TOML config: `[[hooks]]` in server.toml and run config files
- Config cascade: server hooks + run hooks merge, name collisions
resolved by run config winning
- Remove legacy tool_hooks.pre/post from codergen handler (breaking)
- 30 e2e integration tests covering all hook events and behaviors
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Rename the endpoint, schema (RunFiles -> RunCompare), operation ID,
handlers, and frontend route across the full stack.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The question_type field on ApiQuestion was a bare string serialized via
Debug formatting. Define a proper enum in the spec so typify generates a
typed QuestionType, then map from the workflow enum in the handler.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace inline json!() wrappers with a shared ListResponse<T> struct
that serializes directly, avoiding the intermediate serde_json::Value.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Wrap 4 list endpoints that returned bare arrays in the standard
paginated `{ data, meta: { has_more } }` shape so adding real
pagination later is additive rather than a breaking change.
Endpoints: GET /runs/{id}/questions, /runs/{id}/stages,
/runs/{id}/verifications, and /verifications.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Deserialize permissions/output_format as typed enums instead of strings
so invalid values in cli.toml fail at parse time
- Use Option::or/or_else combinators instead of if-is_none pattern
- Load cli.toml only for agent/llm commands, not all CLI invocations
- Standalone arc-agent binary calls apply_cli_defaults for single source
of hardcoded defaults
- Remove redundant #[serde(default)] on Option fields
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Users who always use the same provider/model/permissions no longer need
to pass flags every time. Precedence: CLI flag > cli.toml > hardcoded default.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Define failover_eligible() in terms of retryable() to prevent drift
(only difference: QuotaExceeded is failover-eligible but not retryable)
- Extract spawn_event_forwarder() to eliminate duplicated event-forwarding
spawn blocks and fix missing file-change tracking in failover path
- Replace hardcoded "anthropic" default with self.provider.as_str()
- Refactor create_session into create_session_for(model, provider) to
avoid constructing throwaway AgentApiBackend during failover
- Use &[FallbackTarget] slice instead of cloning Vec on every one_shot
- Remove redundant run_defaults fallback in resolve_fallback_chain
(apply_defaults already merges fallbacks before it's called)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
When an LLM provider returns a transient error (rate limit, server error,
quota exceeded, timeout, network, or stream failure), Arc now automatically
retries on fallback providers using the closest matching model from the
catalog based on capability filters and cost proximity.
Key changes:
- closest_model() and build_fallback_chain() in arc-llm catalog
- failover_eligible() on SdkError to classify transient vs deterministic errors
- fallbacks config field on LlmConfig with task-wins-over-defaults merging
- WorkflowRunEvent::Failover variant for observability
- Failover logic in both one_shot and agent session (run) code paths
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Reuse expand_vars() from run_config to scan for $identifier patterns
in codergen prompt expansion. Unknown variables like $gaol now produce
an ArcError::Validation at runtime instead of silently passing through.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replaces 10 identical copies of the pagination block with a single
shared function. Also eliminates a redundant second collect() by
using truncate() instead.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Apply the same page[limit]/page[offset] pagination pattern from
GET /runs to: workflows, workflow runs, retros, sessions, projects,
branches, saved queries, query history, and stage turns.
Each endpoint now returns { data, meta: { has_more } } instead of
a bare array. Includes OpenAPI spec updates, demo handler changes,
regenerated TS client, updated frontend consumers, and a new
pagination conformance test.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
load_server_config() now accepts an optional explicit path. When
provided, it reads from that path (erroring if missing) instead of
the default ~/.arc/server.toml. The --config flag is wired through
ServeArgs and the hot-reload polling loop.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- AppState.server_config was stored but never read; remove it and revert
create_app_state_with_options back to 5 parameters
- Config polling now compares under a read lock first, only acquiring
the write lock when a change is detected
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Poll ~/.arc/server.toml every 5s and swap run_defaults/git config for new
runs without restarting. CLI overrides (--model, --provider) always win.
On parse error, log a warning and keep the previous config.
Also add a `default` field to the model catalog so default model resolution
uses catalog data instead of hardcoded model names in Rust code.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>