Commit graph

548 commits

Author SHA1 Message Date
Bryan Helmkamp
cf50625824 Switch Daytona git cloning from gh CLI to GitHub App Installation Access Tokens
Replace `gh auth token` with GitHub App IATs scoped to `contents: read` for
Daytona sandbox git cloning. Public repos are auto-detected and cloned without
credentials. Private repos get short-lived, repo-scoped tokens. Clear error
messages for each failure mode (app not installed, suspended, no repo access,
auth failure). Falls back gracefully when no GitHub App is configured.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 22:31:40 -05:00
Bryan Helmkamp
41f72116f2 Update changelog with GPT-5.4, per-node loop limits, and $$ escaping
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 22:20:58 -05:00
Bryan Helmkamp
21f9841623 Fill in tutorials, examples, and reference docs with content and SVGs
- Add SVG workflow diagrams for all tutorials and examples
- Fill in NLSpec Convergence and Semantic Port example content
- Add Solitaire example workflow
- Add error handling sections to tools and subagents docs
- Add context compaction and artifact offloading to context docs
- Add credential redaction note to observability docs
- Add workflow diagram Frame references to tutorials

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 22:14:02 -05:00
Bryan Helmkamp
013101a6f0 Add Clone Substack example workflow adapted from Kilroy
Port Kilroy's substack-spec-v01.dot as a full Arc example with:
- Debate planning (Opus + Gemini Flash independent plans)
- Six-stage verify chain (fmt, build, test, browser, artifacts, fidelity)
- Ensemble review with consensus (two providers)
- Postmortem repair loop with replan/toolchain routing
- Full Kilroy prompts adapted for Arc conventions
- Overview SVG with light/dark mode support

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 22:13:25 -05:00
Bryan Helmkamp
9d1bd1d091 Add GPT-5.4 and GPT-5.4 Pro to model catalog
Add both models with aliases (gpt54, gpt54-pro). Neither replaces
gpt-5.2 as the OpenAI default. Update fallback chain test since
GPT-5.4 ($2.50) is now closer to Opus ($15) than GPT-5.2 ($1.75).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 22:12:21 -05:00
Bryan Helmkamp
1b2a63f6af Implement per-node max_visits override for loop detection
Per-node `max_visits` now overrides the graph-level `max_node_visits`
(and the dry-run default of 10) for individual nodes, giving tighter
control over specific loops like fix-and-verify cycles.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 22:05:11 -05:00
Bryan Helmkamp
69a203c2a8 Add $$ escape mechanism for variable expansion
Allows literal $ signs in prompts and DOT files by writing $$. Also
refactors VariableExpansionTransform to use expand_vars instead of
string replace, fixing a substring-matching bug with $goal.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 21:44:50 -05:00
Bryan Helmkamp
4605f821e8 Add custom TextMate grammar for DOT syntax highlighting in docs
Revert the graphviz language identifier back to dot (Shiki doesn't
bundle either) and register a custom TextMate grammar via Mintlify's
styling.codeblocks.languages.custom config. The grammar (extracted from
apps/arc-web/app/data/dot-grammar.ts) highlights keywords, shape names,
attributes, strings, comments, and operators.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 21:41:54 -05:00
Bryan Helmkamp
6a98dcc246 Use graphviz language identifier for DOT code blocks, add titles to workflow examples
Changes ```dot to ```graphviz across all docs for better syntax highlighting
in Mintlify (Shiki). Full digraph blocks get a title derived from the workflow
name (e.g. title="hello.dot").

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 21:26:45 -05:00
Bryan Helmkamp
598b5b7d19 Update docs: fill in administration, agents, execution, and reference pages; remove placeholder pages
Remove stub pages (bring-your-own-key, costs, monitoring, usage, artifacts,
ai-tools, extensions, hooks, llms) and flesh out remaining pages with real
content covering advanced setup, security, permissions, MCP, observability,
interviews, environments, CLI reference, architecture, and workflow stages.
Add permission-flow diagram and tutorials section.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 20:25:23 -05:00
Bryan Helmkamp
5bdb67c51e Document agent backends (API vs CLI)
Add Backends section to agents page explaining the two execution
strategies, expand backend attribute descriptions in dot-language
and stylesheet references, and add API-only notes to tools and
sub-agents pages.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 20:00:21 -05:00
Bryan Helmkamp
3500d1436b Simplify Skills page, link to Agent Skills spec instead of duplicating
Remove the "Skill file format" and "Creating skills" sections which
duplicated content from the Agent Skills specification. Link to
agentskills.io for the format spec and introduction instead.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 19:19:29 -05:00
Bryan Helmkamp
24f0fa7ca0 Add Daytona network access control and fill in execution docs
- Add DaytonaNetwork enum (block/allow_all/allow_list) with custom serde
  Deserialize for TOML string-or-table syntax
- Wire network config through run_config defaults merging and base_params
- Document network access in sandboxing, environments, and run-configuration
- Fill in execution docs: checkpoints, environments, failures, interviews,
  run configuration, observability, retros
- Rename compounding.mdx → retros.mdx, insights.mdx → observability.mdx
- Use DaytonaConfig::default() in tests to reduce boilerplate

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 18:31:19 -05:00
Bryan Helmkamp
735b4648a0 Rename script.output/script.stderr context keys to command.output/command.stderr
Aligns context key names with the handler name (CommandHandler), which was
previously inconsistent.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 18:09:42 -05:00
Bryan Helmkamp
7505cb4ce1 Fill in Agents, Human-in-the-Loop, Stylesheets, Variables pages
- Agents: agent loop, built-in tools, prompts, sub-agents, skills, hooks
- Human-in-the-Loop: gates, accelerators, freeform input, auto-approve
- Stylesheets: selectors, specificity, properties, cascading, full examples
- Variables: run config vars, $goal expansion, variable merging
- Remove model stylesheets section from core workflows page (now dedicated page)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 17:42:19 -05:00
Bryan Helmkamp
dd797cca1a Fill in core docs: How Arc Works, Models, Workflows, Nodes & Stages, Transitions
- How Arc Works: end-to-end architecture with diagram, execution loop, retries,
  sandboxes, observability, and resume
- Models: multi-provider rationale, full catalog table, stylesheets, overrides
  via CLI and TOML, ensemble SVG
- Workflows: anatomy SVG, key node types, branching/loops SVG, parallel SVG,
  model stylesheets, goal gates
- Nodes & Stages: all node types with attributes, fidelity/thread_id tables,
  join/error/retry policy tables
- Transitions: edge selection priority, condition expression language, agent
  transitions, human gates, weight tiebreaking
- Remove workflows/ingestion page

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 17:29:57 -05:00
Bryan Helmkamp
6970a7c0f6 Draft Quick Start page: install, configure, CLI, API server, and web frontend
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 16:56:14 -05:00
Bryan Helmkamp
d91e0ea6fd Restructure Getting Started docs: add Introduction, rename Overview to Why Arc, add Models placeholder
- Add Introduction page with four card links (Why Arc, Quick Start, Workflows, Agents)
- Rename Overview to Why Arc with updated frontmatter
- Add rendered SVG workflow graph to Why Arc page
- Add Core Concepts > Models placeholder page
- Remove getting-started/workflows page

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 16:53:53 -05:00
Bryan Helmkamp
ba1f37d57d Add parity tests for follow_up and mid-task steering, expand to all providers
- Add follow_up, steering_before_input, and steering_mid_task parity tests
- Add steering_queue_handle() to Session for test access
- Expand provider_tests! macro to cover all 7 providers (Kimi, Zai, Minimax, Inception)
- Update model names to match current catalog (gpt-5-mini, gemini-3-flash-preview, etc.)
- Fix OpenAI-compatible tool result serialization: emit one ChatMessage per ToolResult
  instead of bundling all results into a single message with only the first tool_call_id.
  This was causing Kimi to reject requests during parallel tool calls.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 16:44:30 -05:00
Bryan Helmkamp
1d4af91aab Remove root README (arc-llm content already lives in crates/arc-llm/)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 16:43:52 -05:00
Bryan Helmkamp
9ebc84db67 Draft Getting Started overview page for docs
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 16:43:21 -05:00
Bryan Helmkamp
da5752bfc8 Fix Mintlify dev server: valid navbar URL and ignore AGENTS.md
The navbar href must be a full URL (not a relative path) and AGENTS.md
contains HTML comments that are invalid in MDX.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 16:29:50 -05:00
Bryan Helmkamp
dc348d553e Extract parse_ast() and use Write sink in parse_command
- Extract shared strip+parse+trailing-check into parser::parse_ast()
  so both parser::parse() and parse_command reuse it
- Accept impl Write in parse_command so tests verify actual JSON output
  instead of re-parsing independently

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 16:19:34 -05:00
Bryan Helmkamp
57dd00a1a5 Add arc parse FILE.dot subcommand to print raw AST as JSON
Parses a DOT file and outputs its AST as pretty-printed JSON, useful for
debugging and tooling. Adds Serialize/Deserialize to all AST types.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 16:16:04 -05:00
Bryan Helmkamp
872107e4d6 Fix Docker demo build and verifications API response parsing
- Switch mintlify install from bun to npm with katex version patch
- Add arm64 platform and pull_policy: never to docker-compose.demo.yaml
- Unwrap data envelope in verifications loader response
- Add name-gen script

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 12:03:42 -05:00
Bryan Helmkamp
75a6e2f297 Widen MODEL column from 24 to 30 chars in CLI tables
Accommodates longer model names like gemini-3.1-flash-lite-preview
in both `models list` and `models test` output.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 10:43:29 -05:00
Bryan Helmkamp
f8d96cb7ae Fix server-side error message leaking to browser in production
error.message was displayed to users in all environments, while
stack traces were correctly gated behind import.meta.env.DEV.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 10:25:07 -05:00
Bryan Helmkamp
41ed815a19 Fix apiStages.map crash by using PaginatedRunStageList type
Three routes (run-overview, run-configuration, run-graph) were typing
the stages API response as RunStage[] but the endpoint returns a
paginated wrapper { data, meta }. Destructure .data to get the array.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 10:24:23 -05:00
Bryan Helmkamp
65e3a0bcee Fix pre-existing clippy warnings in arc-workflows
- Replace useless format!() with .to_string() in parse_decision
- Derive Default for HookDecision instead of manual impl
- Use contains_key() instead of get().is_none() in semantic parser

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 10:11:44 -05:00
Bryan Helmkamp
3a7e7deb5a Use structured output for prompt hooks via generate_object()
Replace freeform LLM text generation with schema-constrained
generate_object() for prompt hooks, eliminating the need for
JSON formatting instructions in the system prompt and the
code-fence stripping workaround.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 09:55:16 -05:00
Bryan Helmkamp
e5354c3028 Simplify arc-slack: use thiserror, remove dead params, extract helpers
- Use thiserror for SlackApiError and ConnectionError (project convention)
- Remove misleading PartialEq/Eq on DispatchAction; use matches!() in tests
- Remove unused _slack_client and _default_channel params from event loop
- Extract check_ok() helper to deduplicate 3 ok-check sites in client.rs
- Make bot_token private on SlackClient; add http() accessor
- Make SlackClient Clone; eliminate duplicate instance in e2e example
- Reuse reqwest::Client in open_socket_url instead of creating a new one
- Filter empty env vars in resolve_credentials; remove redundant is_enabled

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 09:24:22 -05:00
Bryan Helmkamp
fb76c0ec3e Add arc-slack crate: Slack Socket Mode integration for interviewer
Implements a complete Slack integration for the interviewer system using
Socket Mode (WebSocket-based, no public URL required). Supports all five
question types: YesNo, Confirmation, MultipleChoice, MultiSelect, and
Freeform (via thread replies with @mention).

Modules: config, client, blocks, interaction, socket, dispatch,
connection, threads. 72 unit tests + e2e example.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 09:16:22 -05:00
Bryan Helmkamp
8a1d1a3ff4 Extend condition expression language with ||, !, numeric comparisons, contains, and matches
Replace flat Vec<Clause> parser with AST-based recursive descent parser
supporting full operator precedence (&& binds tighter than ||, ! is prefix).
New operators: >, <, >=, <= (numeric), contains (substring/array membership),
matches (regex, validated at parse time). Simplify ConditionSyntaxRule to
delegate entirely to parse_condition(). Public API unchanged.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 08:20:07 -05:00
Bryan Helmkamp
2d3ae34059 Extract is_llm_handler_type() to deduplicate 4 match-arm sites
The pattern Some("agent") | Some("agent_loop") | Some("prompt") |
Some("one_shot") was duplicated across preamble.rs (3x) and
validation/rules.rs (1x). Centralizes into a single function in
graph/types.rs. Also fixes stale doc comment on default_registry.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 07:52:13 -05:00
Bryan Helmkamp
059a50e1b3 Rename handler types: agent_loop → agent, one_shot → prompt
Keep legacy aliases (agent_loop, one_shot) in the handler registry
and validation rules for backwards compatibility. Add codergen_mode
attribute support in the DOT parser, translating legacy values to
the new type names.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 07:44:33 -05:00
Bryan Helmkamp
4bb072884c Rename handler types and split one_shot into its own handler
- Rename handler type strings: codergen → agent_loop, wait.human → human,
  script → command, wait.timer → wait
- Rename handler modules/structs to match: AgentHandler, HumanHandler,
  CommandHandler, WaitHandler
- Split one_shot into PromptHandler (handler/prompt.rs) with shape=tab mapping
- Remove CodergenMode enum and codergen_mode attribute — one_shot is now its
  own handler type, not a mode flag on the agent loop handler
- Update all demo DOT files: codergen_mode="one_shot" → shape=tab
- Update spec, README, validation rules, preamble, and hook tests

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 07:22:17 -05:00
Bryan Helmkamp
1b6be5b0c9 Update changelog skill to use narrative H2-per-feature format
Replaces category-header bullet lists with individual H2 sections
per feature, narrative writing with before/after framing, and code
examples. Minor items go in a flat list at the bottom.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 03:14:44 -05:00
Bryan Helmkamp
e7d64136b3 Rewrite changelog entries with improved style and structure
Each major feature gets its own H2 heading with narrative depth
and code examples instead of dense bullet lists under category
headers. Minor improvements and fixes go at the bottom as a flat
list. Style inspired by Qlty, Linear, Vercel, and Resend changelogs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 03:13:42 -05:00
Bryan Helmkamp
45a7762a17 changelog 2026-03-05 03:00:37 -05:00
Bryan Helmkamp
4f4f1f98a3 Simplify hook implementation: deduplicate LLM setup, fix async I/O, cache HTTP clients
- Extract shared prompt/agent hook setup (model resolution, system prompt,
  user message, timeout wrapper) into reusable helpers
- Fix blocking I/O: std::process::Command → tokio::process::Command for
  host-mode hook execution
- Cache reqwest::Client per TLS mode via OnceLock instead of rebuilding
  per HTTP hook call
- Return Cow from resolved_hook_type() to avoid cloning HookType on
  every call
- Rename command_executor → executor in HookRunner

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 02:49:06 -05:00
Bryan Helmkamp
c7e7b30906 Reuse core tool registry in agent hooks and add e2e tests
- Change HookExecutor trait to take Arc<dyn Sandbox> so agent hooks can
  share the sandbox with the ToolRegistry via ToolContext
- Replace hand-rolled 2-tool dispatch with register_core_tools() giving
  agent hooks the full tool set (read_file, write_file, shell, grep, glob)
- Add strip_code_fences() to handle LLMs wrapping JSON in markdown
- Set max_tokens(1024) on prompt hooks to avoid exceeding model limits
- Add e2e tests: TOML parsing, prompt proceed/block, agent proceed,
  agent with tool use (reads a file via read_file tool)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 02:33:10 -05:00
Bryan Helmkamp
fe42189dc9 Add prompt and agent hook types for LLM-based hook evaluation
Prompt hooks make a single-turn LLM call returning {"ok": true/false}.
Agent hooks run a multi-turn LLM tool loop with sandbox access (exec_command, read_file).
Both fail-open on errors/timeouts. Prompt hooks default to 30s timeout,
agent hooks to 60s with max 50 tool rounds. Default model is "haiku".

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 02:19:09 -05:00
Bryan Helmkamp
721224159d Add TLS mode config for HTTP hooks
Introduce a `tls` field on HTTP hooks with three modes: `verify` (default,
requires https + cert validation), `no_verify` (requires https, skips cert
validation), and `off` (allows http, skips cert validation). This prevents
hooks from accidentally sending credentials over plaintext connections.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 02:06:19 -05:00
Bryan Helmkamp
97594b3ab0 Add HTTP hook executor with env var interpolation
HTTP hooks (type = "http") now actually execute instead of failing with
"no command specified". The executor POSTs the hook context as JSON,
parses HookDecision from the response, and fails open on errors.

Header values support $VAR/${VAR} interpolation gated by an
allowed_env_vars whitelist on the hook definition. Renames
CommandHookExecutor to HookExecutorImpl since it now handles both
command and HTTP hook types.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 01:59:38 -05:00
Bryan Helmkamp
fd70d41fca Simplify hook system after code review
- Replace JSON-based HookEvent::Display with simple match arms
- Use floor_char_boundary for Unicode-safe truncation in effective_name
- Cache compiled regexes in HookRunner instead of recompiling per check
- Simplify run_non_blocking (was run_parallel) to plain sequential loop
- Propagate hook_runner through parallel handler branch services
- Use unique temp file paths for sandbox hook context (avoid collisions)
- Add hook imports to engine.rs, reducing verbose crate:🪝: paths
- Extract duplicate RunFailed hook code into run_failed_hook helper

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 01:49:53 -05:00
Bryan Helmkamp
55d44f7488 Add lifecycle hook system replacing legacy tool_hooks
Introduce a configurable hook system that triggers user-defined actions
at workflow lifecycle points (RunStart, StageStart, StageComplete,
StageFailed, EdgeSelected, CheckpointSaved, etc). Hooks can block
execution, skip nodes, or override edge routing via JSON decisions.

- New `hook/` module: types, config, executor (command), runner
- Engine instrumented at 8 lifecycle points with HookRunner calls
- TOML config: `[[hooks]]` in server.toml and run config files
- Config cascade: server hooks + run hooks merge, name collisions
  resolved by run config winning
- Remove legacy tool_hooks.pre/post from codergen handler (breaking)
- 30 e2e integration tests covering all hook events and behaviors

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 01:42:19 -05:00
Bryan Helmkamp
b57e30efad Rename run-files-changed.tsx to run-compare.tsx to match route
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 00:57:52 -05:00
Bryan Helmkamp
26f08a904e Rename GET /runs/{id}/files to GET /runs/{id}/compare
Rename the endpoint, schema (RunFiles -> RunCompare), operation ID,
handlers, and frontend route across the full stack.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 00:53:45 -05:00
Bryan Helmkamp
a40c6bdd29 Add QuestionType enum to OpenAPI spec replacing stringly-typed field
The question_type field on ApiQuestion was a bare string serialized via
Debug formatting. Define a proper enum in the spec so typify generates a
typed QuestionType, then map from the workflow enum in the handler.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 00:49:03 -05:00
Bryan Helmkamp
4f1297572d Extract typed ListResponse struct to eliminate double serialization
Replace inline json!() wrappers with a shared ListResponse<T> struct
that serializes directly, avoiding the intermediate serde_json::Value.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 00:46:11 -05:00