From 21f9841623e1dc2638695d04090e96288d9ff0da Mon Sep 17 00:00:00 2001 From: Bryan Helmkamp Date: Thu, 5 Mar 2026 22:14:02 -0500 Subject: [PATCH] Fill in tutorials, examples, and reference docs with content and SVGs - Add SVG workflow diagrams for all tutorials and examples - Fill in NLSpec Convergence and Semantic Port example content - Add Solitaire example workflow - Add error handling sections to tools and subagents docs - Add context compaction and artifact offloading to context docs - Add credential redaction note to observability docs - Add workflow diagram Frame references to tutorials Co-Authored-By: Claude Opus 4.6 (1M context) --- docs/agents/subagents.mdx | 10 + docs/agents/tools.mdx | 14 + docs/examples/nlspec-convergence.mdx | 173 ++++++++++++- docs/examples/semantic-port.mdx | 269 ++++++++++++++++++- docs/examples/solitaire.mdx | 313 +++++++++++++++++++++++ docs/execution/context.mdx | 45 ++++ docs/execution/observability.mdx | 4 +- docs/images/example-semantic-port.svg | 116 +++++++++ docs/images/example-solitaire.svg | 168 ++++++++++++ docs/images/nlspec-convergence.svg | 148 +++++++++++ docs/images/tutorial-branch-loop.svg | 103 ++++++++ docs/images/tutorial-ensemble.svg | 161 ++++++++++++ docs/images/tutorial-hello.svg | 59 +++++ docs/images/tutorial-multi-model.svg | 100 ++++++++ docs/images/tutorial-parallel-review.svg | 138 ++++++++++ docs/images/tutorial-plan-implement.svg | 102 ++++++++ docs/images/tutorial-subagent.svg | 58 +++++ docs/images/tutorial-tool-use.svg | 58 +++++ docs/tutorials/branch-loop.mdx | 126 +++++++++ docs/tutorials/ensemble.mdx | 125 +++++++++ docs/tutorials/hello-world.mdx | 129 ++++++++++ docs/tutorials/multi-model.mdx | 122 +++++++++ docs/tutorials/parallel-review.mdx | 115 +++++++++ docs/tutorials/plan-implement.mdx | 107 ++++++++ 24 files changed, 2760 insertions(+), 3 deletions(-) create mode 100644 docs/examples/solitaire.mdx create mode 100644 docs/images/example-semantic-port.svg create mode 100644 docs/images/example-solitaire.svg create mode 100644 docs/images/nlspec-convergence.svg create mode 100644 docs/images/tutorial-branch-loop.svg create mode 100644 docs/images/tutorial-ensemble.svg create mode 100644 docs/images/tutorial-hello.svg create mode 100644 docs/images/tutorial-multi-model.svg create mode 100644 docs/images/tutorial-parallel-review.svg create mode 100644 docs/images/tutorial-plan-implement.svg create mode 100644 docs/images/tutorial-subagent.svg create mode 100644 docs/images/tutorial-tool-use.svg diff --git a/docs/agents/subagents.mdx b/docs/agents/subagents.mdx index a62b47575..69551a8e0 100644 --- a/docs/agents/subagents.mdx +++ b/docs/agents/subagents.mdx @@ -98,6 +98,16 @@ Maximum subagent depth (1) reached The depth limit is set via the `max_subagent_depth` field in the session configuration. +## Error handling + +When a sub-agent fails, the error is captured and returned to the parent via the `wait` tool — the parent stage does **not** automatically fail. The parent LLM sees the error message and decides how to respond: retry, try a different approach, or report the failure. + +Specific failure modes: + +- **Sub-agent hits `max_turns`** — The sub-agent stops naturally and returns its last output as a successful result. The parent sees a normal completion with the final assistant message. +- **Sub-agent panics or errors** — The error is captured and returned through `wait` as a failure result. The parent can inspect the error and decide what to do. +- **`spawn_agent` fails** (e.g. depth limit exceeded) — The error is returned immediately as a tool result. The parent can adjust its approach without waiting. + ## Event forwarding Sub-agent events (tool calls, assistant messages, errors) are forwarded to the parent session's event stream as `SubAgentEvent` wrappers. This means the parent's progress log captures the full activity of all children, giving you visibility into what sub-agents are doing. diff --git a/docs/agents/tools.mdx b/docs/agents/tools.mdx index d1a65bbf1..67c4aeda2 100644 --- a/docs/agents/tools.mdx +++ b/docs/agents/tools.mdx @@ -215,6 +215,20 @@ Tool output is truncated before being stored in conversation history to prevent Limits can be overridden per-tool via `SessionConfig.tool_output_limits`. +### Error handling + +When a tool call fails, the error is returned to the agent as a tool result with an error flag — the agent loop continues. Tool errors do **not** fail the stage. The LLM sees the error message and decides how to respond: retry the call, try an alternative approach, or move on. + +Common error cases: + +| Error | Cause | +|---|---| +| Unknown tool | The agent called a tool that doesn't exist | +| Argument validation failure | Arguments don't match the tool's JSON Schema | +| File not found | The target file doesn't exist | +| Command timeout | A shell command exceeded its timeout | +| Read-before-write | The agent tried to write to a file it hasn't read (see [guardrail](#read-before-write-guardrail)) | + ### Timeouts Shell commands have two timeout settings: diff --git a/docs/examples/nlspec-convergence.mdx b/docs/examples/nlspec-convergence.mdx index 36a8e3b98..48d331745 100644 --- a/docs/examples/nlspec-convergence.mdx +++ b/docs/examples/nlspec-convergence.mdx @@ -1,5 +1,176 @@ --- title: "NLSpec Convergence" -description: "Example workflow for NLSpec convergence" +description: "Implement a system from a natural language specification and iterate until conformance tests pass" --- +The NLSpec Convergence pattern gives an agent a detailed specification document, has it build an implementation, and then loops on automated conformance tests until the implementation converges on full compliance. This is the same pattern used by benchmarks like [AttractorBench](https://github.com/strongdm/attractorbench) to measure how well agents follow complex specs. + +## When to use this + +- You have a detailed specification (API contract, RFC, design doc) and want an agent to implement it +- You have automated tests that can verify conformance +- The implementation is too large to get right in one pass and benefits from iterative repair + +## The workflow + + + NLSpec Convergence workflow: Start → Plan → Implement → Quick Tests → Quick passing? → Full Tests → All passing? → Exit, with Fix Failures loop + + +```dot title="n-l-spec-convergence.dot" +digraph NLSpecConvergence { + graph [ + goal="Implement a conformant system from a natural language specification", + model_stylesheet=" + * { llm_model: claude-haiku-4-5; llm_provider: anthropic; } + .impl { llm_model: claude-sonnet-4-5; reasoning_effort: high; } + " + ] + rankdir=LR + + start [shape=Mdiamond, label="Start"] + exit [shape=Msquare, label="Exit"] + + // Phase 1: Read spec and plan + plan [label="Plan", class="impl", prompt="@prompts/plan.md"] + + // Phase 2: Build the initial implementation + subgraph cluster_impl { + label = "Implement & Converge" + node [thread_id="impl", fidelity="full"] + + implement [label="Implement", class="impl", prompt="@prompts/implement.md"] + fix [label="Fix Failures", class="impl", prompt="@prompts/fix.md", max_visits=5] + } + + // Phase 3: Quick conformance loop + test_quick [label="Quick Conformance", shape=parallelogram, script="make conformance-quick 2>&1 || true"] + gate_quick [shape=diamond, label="Quick suite passing?"] + + // Phase 4: Full conformance + test_full [label="Full Conformance", shape=parallelogram, script="make conformance-full 2>&1 || true", goal_gate=true] + gate_full [shape=diamond, label="All tests passing?"] + + // Wiring + start -> plan -> implement -> test_quick -> gate_quick + + gate_quick -> test_full [label="Pass", condition="outcome=success"] + gate_quick -> fix [label="Fix"] + + fix -> test_quick + + test_full -> gate_full + + gate_full -> exit [label="Pass", condition="outcome=success"] + gate_full -> fix [label="Fix"] +} +``` + +```bash +arc run start workflows/nlspec-convergence.dot +``` + +## How it works + +### Spec reading and planning + +The `plan` node reads the specification and produces an implementation plan. Because the spec may be long (thousands of lines), the prompt tells the agent to read it from a file rather than trying to include it inline: + +```markdown + +Read the specification at `docs/spec.md` in full. + +Produce a step-by-step implementation plan in `plan.md` that covers: +1. The key abstractions and data types to define +2. The public API surface (functions, CLI commands, endpoints) +3. The conformance contract (what `make conformance-quick` and `make conformance-full` will test) +4. Implementation order — start with the smallest slice that passes at least one conformance test +``` + +### Implementation with a shared thread + +The `implement` and `fix` nodes share a `thread_id="impl"`, so they accumulate context across loop iterations. When the agent returns to `fix` after a failed conformance run, it sees the full history of what it built and what broke. Combined with `fidelity="full"`, the agent retains the detail it needs to make targeted repairs. + +### The convergence loop + +The core of this pattern is the test-fix loop: + +``` +implement → test_quick → gate_quick → [Fix] → fix → test_quick → gate_quick → [Pass] → test_full → ... +``` + +The **quick conformance** suite runs a subset of tests that execute fast (seconds, not minutes). The agent iterates against this subset first, fixing one failure class at a time. Only after the quick suite passes does it run the **full conformance** suite. + +This two-tier approach mirrors how developers work: run the fast tests while iterating, then run the complete suite before calling it done. + +### Fix node prompt + +The `fix` prompt reads conformance output and targets specific failures: + +```markdown + +The conformance tests found failures. Read the test output from the +previous command node and fix the issues. + +Strategy: +1. Read the failing test names and error messages +2. Identify the root cause — is it a missing feature, wrong format, or integration bug? +3. Fix one failure class at a time (e.g. all JSON schema errors, then all routing errors) +4. After fixing, the workflow will re-run conformance automatically + +Do not rewrite working code. Make targeted fixes to the specific failures. +``` + +### Max visits as a safety valve + +`max_visits=5` on the `fix` node prevents infinite loops. If the agent can't converge in 5 iterations, the workflow moves on with the best result so far. Tune this based on spec complexity: a 30-line spec might need 2 iterations, a 2,000-line spec might need 10. + +### Goal gate on full conformance + +The `test_full` node has `goal_gate=true`. If the full conformance suite never passes, the workflow is marked as failed even if execution reaches the exit node. This makes the workflow's success criteria explicit: partial conformance is not a passing result. + +## Model assignment + +The `model_stylesheet` assigns a cheaper model as the default and routes implementation work to a more capable model: + +```dot +graph [model_stylesheet=" + * { llm_model: claude-haiku-4-5; llm_provider: anthropic; } + .impl { llm_model: claude-sonnet-4-5; reasoning_effort: high; } +"] +``` + +The `.impl` class targets both the `implement` and `fix` nodes (both have `class="impl"`). Planning and implementation get the stronger model with high reasoning effort; any lightweight nodes you add later (summaries, notifications) default to the faster model. + +## Adding a human approval gate + +For high-stakes specs, add a human gate after planning: + +```dot +approve [shape=hexagon, label="Approve Plan"] + +plan -> approve +approve -> implement [label="[A] Approve"] +approve -> plan [label="[R] Revise"] +``` + +The agent writes its plan to `plan.md`, the human reviews it, and either approves (proceeding to implementation) or sends it back for revision. + +## Adapting for your project + +To use this pattern: + +1. **Write your spec** as a Markdown file in the repo (e.g. `docs/spec.md`) +2. **Write conformance tests** that exercise the spec's requirements via a CLI or test runner. Split them into quick (core paths) and full (everything) suites. +3. **Wire the Makefile** so `make conformance-quick` and `make conformance-full` run the suites and output results +4. **Customize the prompts** to reference your spec file, your project's language and conventions, and your conformance contract + +The pattern works for any spec that has automated verification: API contracts with integration tests, protocol implementations with compliance suites, or library specs with unit tests. + +## What you've learned + +- The **convergence loop** (implement, test, fix, repeat) is the core pattern for spec-driven development +- **Two-tier conformance** (quick then full) keeps iteration fast +- **Shared threads** (`thread_id`) give the fix node context from prior iterations +- **`max_visits`** prevents infinite loops when the agent can't converge +- **`goal_gate`** makes conformance a hard requirement for workflow success diff --git a/docs/examples/semantic-port.mdx b/docs/examples/semantic-port.mdx index 492e918fd..4517ee6e5 100644 --- a/docs/examples/semantic-port.mdx +++ b/docs/examples/semantic-port.mdx @@ -1,5 +1,272 @@ --- title: "Semantic Port" -description: "Example workflow for semantic porting" +description: "Continuously port upstream changes from one language to another using a ledger-driven loop" --- +A semantic port workflow tracks commits in an upstream repository, analyzes each one for relevance, and either ports the change to a different codebase or acknowledges it as not applicable — then loops back for the next commit. It's an autonomous maintenance loop, not a one-shot build task. + +This pattern is useful when you maintain a downstream implementation (e.g., a Go SDK) that tracks an upstream reference (e.g., a Python SDK). Instead of manually reviewing every upstream commit, the workflow processes the backlog commit-by-commit, making intelligent port-or-skip decisions. + +## The workflow + + + Semantic Port workflow: Start → Fetch → Analyze → Plan → Implement → Validate → Tests pass? → Finalize → loops back to Fetch, with Skip shortcut from Analyze back to Fetch, Fix loop from gate back to Validate, and Done exit from Fetch + + +```dot title="semantic-port.dot" +digraph SemanticPort { + graph [ + goal="Port semantic changes from upstream Python repository to our Go implementation", + rankdir=LR, + default_max_retry=3, + model_stylesheet=" + * { llm_model: claude-sonnet-4-5; llm_provider: anthropic; } + .hard { llm_model: claude-opus-4-6; llm_provider: anthropic; } + .analyze { llm_model: gemini-3.1-pro-preview; llm_provider: gemini; } + " + ] + + start [shape=Mdiamond, label="Start"] + exit [shape=Msquare, label="Exit"] + + // Phase 1: Find the next unprocessed commit + fetch [ + label="Fetch & Identify", + prompt="Find the next unprocessed upstream commit.\n\n\ + 1. Run `python3 ledger/manage.py earliest` to get the oldest commit with status=new\n\ + 2. If found, write the commit details to .arc/current_commit.md and respond with:\n\ + {\"preferred_next_label\": \"process\"}\n\ + 3. If no new commits exist:\n\ + a. Fetch latest from upstream: cd upstream/ && git fetch && git pull\n\ + b. Find commits newer than the latest in ledger.tsv\n\ + c. Add them with `python3 ledger/manage.py add `\n\ + d. Try `earliest` again\n\ + e. If still none, respond with: {\"preferred_next_label\": \"done\"}\n\n\ + Respond with exactly one of: process or done." + ] + + // Phase 2: Analyze the commit and decide port vs. skip + analyze [ + label="Analyze & Decide", + class="analyze", + prompt="Read .arc/current_commit.md for the commit to process.\n\ + Examine it with `git show ` in the upstream/ directory.\n\n\ + Analyze the semantic changes — what functionality changed, not just syntax.\n\ + Decide if this change is relevant to our Go implementation or if it is\n\ + Python-specific, docs-only, or not applicable.\n\n\ + Write .arc/analysis.md with sections:\n\ + - Commit summary\n\ + - Semantic analysis\n\ + - Decision: PORT or ACKNOWLEDGE (with reasoning)\n\ + - Port plan (if porting): concrete tasks with file:line references\n\n\ + If decision is ACKNOWLEDGE:\n\ + 1. Update ledger: `python3 ledger/manage.py update acknowledged`\n\ + 2. Commit: `git add ledger/ && git commit -m \"semport: acknowledge - \"`\n\ + 3. Respond with: {\"preferred_next_label\": \"skip\"}\n\n\ + If decision is PORT:\n\ + Respond with: {\"preferred_next_label\": \"port\"}" + ] + + // Phase 3: Refine the plan + plan [ + label="Finalize Plan", + prompt="Read .arc/analysis.md. Perform a final editorial pass.\n\ + Write .arc/plan.md ensuring each task has:\n\ + - Concrete file:line references in our Go code\n\ + - Clear acceptance criteria\n\ + - Directly executable instructions\n\n\ + Remove vague language. The plan must be actionable." + ] + + // Phase 4: Implement the port + implement [ + label="Implement Port", + class="hard", + prompt="Follow the plan in .arc/plan.md.\n\ + Port the semantic changes to the Go codebase.\n\ + Focus on semantic equivalence, not literal translation.\n\ + Use Go idioms and respect existing architecture.\n\ + Log all changes to .arc/implementation_log.md." + ] + + // Phase 5: Validate + validate [ + label="Validate", + shape=parallelogram, + script="cd go-sdk && go build ./... && go test ./... -v 2>&1 || true" + ] + + gate [shape=diamond, label="Tests pass?"] + + // Phase 6: Fix failures + fix [ + label="Analyze & Fix", + class="hard", + max_visits=3, + prompt="Tests or build failed. Read the test output from the prior stage.\n\ + Read .arc/plan.md and .arc/implementation_log.md.\n\ + Diagnose the root cause, fix the issue, and log the fix." + ] + + // Phase 7: Update ledger and commit + finalize [ + label="Finalize", + prompt="All tests pass. Finalize this port:\n\ + 1. Update ledger: `python3 ledger/manage.py update implemented`\n\ + 2. Commit all changes:\n\ + `git add -A && git commit -m \"semport: implement - \"`\n\ + 3. Write a brief summary to .arc/implementation_summary.md" + ] + + // Wiring + start -> fetch + + fetch -> analyze [label="Process", condition="preferred_label=process"] + fetch -> exit [label="Done", condition="preferred_label=done"] + + analyze -> plan [label="Port", condition="preferred_label=port"] + analyze -> fetch [label="Skip", condition="preferred_label=skip"] + + plan -> implement -> validate -> gate + + gate -> finalize [label="Pass", condition="outcome=success"] + gate -> fix [label="Fail"] + + fix -> validate + + finalize -> fetch +} +``` + +## Key patterns + +### The commit-processing loop + +The core of this workflow is a loop: `fetch → analyze → (port or skip) → fetch`. Each iteration processes exactly one upstream commit, then loops back for the next. The loop terminates when `fetch` finds no more unprocessed commits and routes to `exit`. + +``` +fetch → analyze → [Skip] → fetch → analyze → [Port] → plan → implement → +validate → gate → [Pass] → finalize → fetch → ... → [Done] → exit +``` + +This is fundamentally different from a build workflow that runs once and exits. The semantic port workflow is designed to process an entire backlog autonomously, handling dozens of commits in a single run. + +### Ledger-driven state + +The workflow tracks disposition in an external ledger file (`ledger.tsv`) with three states: + +| Status | Meaning | +|---|---| +| `new` | Unprocessed — the workflow hasn't looked at this commit yet | +| `acknowledged` | Reviewed and determined to be irrelevant (docs-only, Python-specific, etc.) | +| `implemented` | Semantic changes ported to the Go codebase | + +The ledger is the source of truth for what's been processed. Because it's a plain file committed to Git, it survives across runs — you can stop and resume the workflow and it picks up where it left off. + +### Semantic analysis, not literal translation + +The `analyze` node is the decision point. It examines each upstream commit for *what changed functionally*, not just what code was modified. A commit that refactors Python type hints has no semantic impact on a Go port. A commit that changes retry behavior in the HTTP client does. + +This distinction is critical — routing a different model (Gemini) to the analysis node via the `.analyze` class brings a fresh perspective to the port/skip decision: + +``` +.analyze { llm_model: gemini-3.1-pro-preview; llm_provider: gemini; } +``` + +### The fix loop + +When ported code fails tests, the workflow enters a bounded fix loop: + +```dot +gate -> fix [label="Fail"] +fix -> validate +``` + +The `fix` node has `max_visits=3`, preventing infinite retry cycles. If the fix can't be resolved in 3 attempts, the run terminates rather than looping forever. + +### Skip vs. port branching + +The `analyze` node uses [routing directives](/workflows/transitions#agent-transitions) to choose between two paths: + +```dot +analyze -> plan [label="Port", condition="preferred_label=port"] +analyze -> fetch [label="Skip", condition="preferred_label=skip"] +``` + +When the agent decides a commit is irrelevant, it updates the ledger, commits the acknowledgment, and loops back to `fetch` immediately — no planning or implementation needed. This keeps the workflow efficient: trivial commits (typo fixes, CI config changes, docs updates) are processed in seconds. + +## Multi-model routing + +The stylesheet assigns three tiers of models: + +``` +* { llm_model: claude-sonnet-4-5; } // Default: plan, finalize +.hard { llm_model: claude-opus-4-6; } // Implementation, fixing +.analyze { llm_model: gemini-3.1-pro-preview; } // Analysis: fresh eyes +``` + +- **Sonnet** handles routine tasks: fetching commits, finalizing plans, updating the ledger +- **Opus** handles the hard work: implementing ports and diagnosing test failures +- **Gemini Pro** handles analysis: a different provider brings independent judgment to the port/skip decision, reducing the risk of a single model's blind spots + +## Run configuration + +Pair the workflow with a run config TOML for repeatable execution: + +```toml +version = 1 +goal = "Port semantic changes from upstream openai-agents-python to our Go SDK" +graph = "semport.dot" + +[llm] +model = "claude-sonnet-4-5" +provider = "anthropic" + +[llm.fallbacks] +anthropic = ["openai"] +gemini = ["anthropic"] + +[setup] +commands = [ + "git clone https://github.com/openai/openai-agents-python upstream || (cd upstream && git pull)", + "pip install -r ledger/requirements.txt" +] + +[vars] +upstream_repo = "openai/openai-agents-python" +downstream_lang = "go" +``` + +Launch with: + +```bash +arc run start semport.toml +``` + +## Adapting this pattern + +The semantic port pattern generalizes beyond language porting: + +- **Spec tracking** — monitor an upstream specification (OpenAPI, protobuf) and propagate changes to client libraries +- **Dependency updates** — process a queue of dependency version bumps, testing and committing each one +- **Issue triage** — pull issues from a tracker, classify them, and route to the appropriate workflow +- **Log analysis** — process a backlog of alerts or log entries, investigating each one + +The core structure is always the same: fetch the next item, analyze it, decide on an action, execute, record the disposition, loop. + +## Further reading + + + + Edge conditions, routing directives, and how agents control flow. + + + CSS-like rules for assigning models to workflow nodes. + + + Retry policies, loop detection, and max_visits. + + + TOML configs for repeatable, parameterized runs. + + diff --git a/docs/examples/solitaire.mdx b/docs/examples/solitaire.mdx new file mode 100644 index 000000000..e034fc758 --- /dev/null +++ b/docs/examples/solitaire.mdx @@ -0,0 +1,313 @@ +--- +title: "Build Solitaire" +description: "Build a complete application from a spec using phased implementation with verification gates" +--- + +A build-from-scratch workflow takes a high-level goal, expands it into a spec, then implements the application in phases — each phase followed by an independent verification node and a conditional gate that either advances or retries. A final review with `goal_gate=true` ensures the result meets the spec before the workflow can succeed. + +This pattern is useful when you want an agent to build something non-trivial from zero — a CLI tool, a game, a library — where the implementation naturally decomposes into layers that build on each other. + +## The workflow + + + Build Solitaire workflow: Start → Spec → Setup → OK? → Data → OK? → Logic → OK? → UI → OK? → Integrate → OK? → Review → OK? → Exit, with Retry arcs from each gate back to its phase, and a Fix arc from the review gate back to UI + + +```dot title="build-solitaire.dot" +digraph BuildSolitaire { + graph [ + goal="Build a terminal-based solitaire (Klondike) game in Python", + rankdir=LR, + default_max_retry=3, + retry_target="impl_setup", + fallback_retry_target="impl_game_logic", + model_stylesheet=" + * { llm_model: claude-sonnet-4-5; llm_provider: anthropic; } + .hard { llm_model: claude-opus-4-6; llm_provider: anthropic; } + .verify { llm_model: claude-haiku-4-5; llm_provider: anthropic; } + " + ] + + start [shape=Mdiamond, label="Start"] + exit [shape=Msquare, label="Exit"] + + // Phase 0: Expand the goal into a detailed spec + expand_spec [ + label="Expand Spec", + prompt="Expand the goal into a detailed spec covering:\n\ + - Game rules and data structures (Card, Deck, Pile types)\n\ + - Terminal rendering approach (curses library)\n\ + - Input handling and move validation\n\ + - Win/loss detection\n\ + - UI layout\n\ + - Test strategy\n\n\ + Write the spec to spec.md." + ] + + // Phase 1: Project setup + impl_setup [ + label="Setup Project", + prompt="Read spec.md. Create the Python project structure:\n\ + pyproject.toml, src/ directory, tests/ directory, main.py stub.\n\ + Run: python3 -m py_compile src/*.py" + ] + + verify_setup [label="Verify Setup", class="verify", + prompt="Verify project setup: check pyproject.toml exists,\n\ + source directories exist, and files compile without errors.\n\ + Run: python3 -m py_compile src/*.py" + ] + + check_setup [shape=diamond, label="Setup OK?"] + + // Phase 2: Core data structures + impl_data [ + label="Data Structures", + prompt="Read spec.md. Implement Card, Deck, and Pile types\n\ + with unit tests. Run: python3 -m pytest tests/ -v" + ] + + verify_data [label="Verify Data", class="verify", + prompt="Verify data structures: build, run tests, check that\n\ + Card, Deck, and Pile types are defined and basic operations work.\n\ + Run: python3 -m pytest tests/ -v" + ] + + check_data [shape=diamond, label="Data OK?"] + + // Phase 3: Game logic (hardest phase) + impl_logic [ + label="Game Logic", + class="hard", + max_retries=2, + prompt="Read spec.md and the data structure files.\n\ + Implement Klondike rules: initial deal, move validation,\n\ + auto-complete detection, win condition, undo.\n\ + Write tests for legal/illegal moves, win detection, edge cases.\n\ + Run: python3 -m pytest tests/ -v" + ] + + verify_logic [label="Verify Logic", class="verify", + prompt="Verify game logic: run all tests, check move validation,\n\ + win detection, and undo.\n\ + Run: python3 -m pytest tests/ -v" + ] + + check_logic [shape=diamond, label="Logic OK?"] + + // Phase 4: Terminal UI + impl_ui [ + label="Terminal UI", + class="hard", + max_retries=2, + prompt="Read spec.md and game logic files.\n\ + Implement terminal UI with curses: card rendering (ASCII art),\n\ + board layout, keyboard input, move selection, help text.\n\ + Run: python3 -m pytest tests/ && python3 -m py_compile src/*.py" + ] + + verify_ui [label="Verify UI", class="verify", + prompt="Verify terminal UI: build, run tests, check that\n\ + renderer and input handler exist, game can be instantiated.\n\ + Run: python3 -m pytest tests/" + ] + + check_ui [shape=diamond, label="UI OK?"] + + // Phase 5: Integration + impl_integration [ + label="Integrate", + prompt="Wire up main.py to start the game loop.\n\ + Connect UI input to game logic. Add game over screen,\n\ + help menu, and README with build/run instructions.\n\ + Run: python3 -m pytest tests/" + ] + + verify_integration [label="Verify Integration", class="verify", + prompt="Verify integration: build, run all tests, check README\n\ + exists, verify the game starts without errors.\n\ + Run: python3 -m pytest tests/" + ] + + check_integration [shape=diamond, label="Integration OK?"] + + // Phase 6: Final review (goal gate) + review [ + label="Final Review", + class="hard", + goal_gate=true, + prompt="Read spec.md in full. Review the complete implementation:\n\ + - All Klondike rules correctly implemented\n\ + - Terminal UI works and is intuitive\n\ + - Tests comprehensive and passing\n\ + - README clear and accurate\n\n\ + Run the full test suite. Write a review to review.md.\n\ + Run: python3 -m pytest tests/ -v" + ] + + check_review [shape=diamond, label="Review OK?"] + + // Wiring: linear phases with verify-gate loops + start -> expand_spec -> impl_setup -> verify_setup -> check_setup + + check_setup -> impl_data [condition="outcome=success"] + check_setup -> impl_setup [condition="outcome=fail", label="Retry"] + check_setup -> impl_setup + + impl_data -> verify_data -> check_data + + check_data -> impl_logic [condition="outcome=success"] + check_data -> impl_data [condition="outcome=fail", label="Retry"] + check_data -> impl_data + + impl_logic -> verify_logic -> check_logic + + check_logic -> impl_ui [condition="outcome=success"] + check_logic -> impl_logic [condition="outcome=fail", label="Retry"] + check_logic -> impl_logic + + impl_ui -> verify_ui -> check_ui + + check_ui -> impl_integration [condition="outcome=success"] + check_ui -> impl_ui [condition="outcome=fail", label="Retry"] + check_ui -> impl_ui + + impl_integration -> verify_integration -> check_integration + + check_integration -> review [condition="outcome=success"] + check_integration -> impl_integration [condition="outcome=fail", label="Retry"] + check_integration -> impl_integration + + review -> check_review + + check_review -> exit [condition="outcome=success"] + check_review -> impl_ui [condition="outcome=fail", label="Fix"] + check_review -> impl_ui +} +``` + +## Key patterns + +### Phased implementation with verification gates + +The workflow decomposes the build into six phases, each building on the previous one: + +``` +Spec → Setup → Data Structures → Game Logic → Terminal UI → Integration → Review +``` + +Each phase follows the same three-node pattern: **implement**, **verify**, **gate**. The implement node builds the code, a separate verify node checks it independently, and a conditional gate routes to the next phase on success or back to retry on failure. + +This separation matters. The verify node runs with a different prompt and (via the `.verify` class) a cheaper model. It acts as an independent check — not just "did the implementation node think it succeeded?" but "does an independent evaluation confirm the phase is complete?" + +### Graph-level retry targets + +The graph sets two levels of retry targets: + +```dot +graph [ + retry_target="impl_setup", + fallback_retry_target="impl_game_logic" +] +``` + +If a node fails and has no local retry target, Arc jumps back to `impl_setup` to re-attempt from project setup. If that target itself can't recover, Arc falls back further to `impl_game_logic`. This creates a cascading recovery strategy without cluttering every node with retry configuration. + +### Three-tier model routing + +The stylesheet assigns models by role: + +``` +* { llm_model: claude-sonnet-4-5; } // Default: spec, setup, integration +.hard { llm_model: claude-opus-4-6; } // Hard work: game logic, UI, review +.verify { llm_model: claude-haiku-4-5; } // Verification: fast, cheap checks +``` + +- **Sonnet** handles routine phases: expanding the spec, setting up the project, wiring integration +- **Opus** handles the hard phases: implementing game logic and terminal UI where correctness and complexity demand the strongest model +- **Haiku** handles all verification gates: these are straightforward "run the tests, check the output" tasks that don't need a frontier model + +This keeps costs down. Verification runs after every phase — using Haiku instead of Opus for those checks saves significant tokens across a full run. + +### Goal gate on final review + +The `review` node has `goal_gate=true`: + +```dot +review [label="Final Review", class="hard", goal_gate=true, ...] +``` + +This means the workflow **cannot succeed** unless the review passes. Even if execution reaches the exit node through some edge routing path, Arc checks all goal gates and fails the run if any are unsatisfied. The review is the quality bar — it reads the original spec and verifies the implementation against it. + +If the review fails, execution routes back to `impl_ui` rather than to the beginning. The assumption is that by the time you reach review, the foundation (data structures, game logic) is solid and only the UI or integration needs fixing. + +### Progressive layering + +Each phase reads the spec and builds on the artifacts from prior phases. The data structures phase defines Card, Deck, and Pile types. The game logic phase reads those types and implements rules on top of them. The UI phase reads the game logic and renders it. Each layer has its own tests, so failures are caught at the right level of abstraction. + +This is more reliable than a single "implement everything" node because: +- Earlier phases are validated before later phases begin +- Failures are localized — a broken data structure is caught before game logic tries to use it +- Retries target the right phase, not the entire build + +## Run configuration + +Pair the workflow with a run config for repeatable execution: + +```toml +version = 1 +goal = "Build a terminal-based solitaire (Klondike) game in Python" +graph = "build-solitaire.dot" + +[llm] +model = "claude-sonnet-4-5" +provider = "anthropic" + +[llm.fallbacks] +anthropic = ["openai", "gemini"] + +[setup] +commands = ["python3 -m venv .venv && . .venv/bin/activate && pip install pytest curses"] + +[sandbox] +provider = "daytona" + +[sandbox.daytona.snapshot] +name = "python-dev" +cpu = 4 +memory = 8 +disk = 20 +dockerfile = "FROM python:3.12-slim\nRUN apt-get update && apt-get install -y git libncurses-dev" +``` + +```bash +arc run start build-solitaire.toml +``` + +## Adapting this pattern + +The phased build pattern generalizes to any project that decomposes into layers: + +- **CLI tool** — argument parsing → core logic → output formatting → integration tests +- **REST API** — data models → route handlers → middleware → end-to-end tests +- **Library** — type definitions → core algorithms → public API → documentation +- **Compiler** — lexer → parser → type checker → code generator → test suite + +The structure is always the same: decompose into phases that build on each other, verify each one independently, gate advancement on verification, and enforce overall quality with a goal gate on the final review. + +## Further reading + + + + Retry policies, goal gates, retry targets, and circuit breakers. + + + CSS-like rules for assigning models to workflow nodes. + + + All node types, shapes, and their attributes. + + + TOML configs for repeatable, parameterized runs. + + diff --git a/docs/execution/context.mdx b/docs/execution/context.mdx index 62882adda..0baf760c9 100644 --- a/docs/execution/context.mdx +++ b/docs/execution/context.mdx @@ -19,6 +19,13 @@ Start → Plan → Implement → Test → Exit Context is thread-safe and shared across the entire run. Parallel branches receive an isolated **deep copy** of the context at the point of fan-out, so branches can't interfere with each other. When branches merge, the fan-in handler records the results under `parallel.fan_in.*` keys. +## How agents access context + +Agents do not have a tool to query the context store directly. Instead, context is made available through two mechanisms: + +- **Preamble injection** — When a node starts, Arc assembles a preamble from the current context and prepends it to the node's prompt. The [fidelity setting](#fidelity-controlling-agent-context) controls how detailed this preamble is. +- **Context updates** — Agents can write to the context by including a `context_updates` object in a JSON response. See [Transitions](/workflows/transitions#agent-transitions) for the response format. + ## Keys set by handlers Each handler type writes specific keys into the context after execution: @@ -191,3 +198,41 @@ response.plan → file:///tmp/logs/artifacts/values/response.plan.json The preamble renderer resolves these pointers and displays a reference to the file path. For remote sandboxes (Docker, Daytona), Arc syncs artifact files to the sandbox at `{working_directory}/.arc/artifacts/` so agents can read them. This keeps the context lean — large LLM responses, test output, and file listings don't bloat checkpoint files or overwhelm preamble summaries. + +## Context compaction + +During long-running agent sessions, the conversation history can grow large enough to exceed the LLM's context window. **Context compaction** automatically summarizes older turns and replaces them with a structured summary, keeping the session running without manual intervention. + +Compaction is always enabled and runs with hardcoded defaults — there are no user-facing configuration options. + +### When it triggers + +After every assistant turn, Arc estimates the total token usage of the system prompt and conversation history (using a rough heuristic of 1 token per 4 characters). If the estimate exceeds **80%** of the model's context window, compaction runs. + +### How it works + +1. **Split history** — The conversation is divided into older turns (to be summarized) and the most recent **6 turns** (preserved verbatim). +2. **Render old turns** — Older turns are serialized to a human-readable text format. Tool call arguments and tool results are truncated to 500 characters each. +3. **Summarize via LLM** — Arc makes a non-streaming LLM call (using the same model and provider as the agent session) with a structured summarization prompt. The summary uses these sections: + - **Goal** — what the user asked for + - **Progress** — what was accomplished, with file paths and key decisions + - **Key Decisions** — important choices and their rationale + - **Failed Approaches** — what was tried and didn't work + - **Open Issues** — remaining bugs, edge cases, or TODOs + - **Next Steps** — what should happen next + - **File Operations** — if the agent has tracked file modifications, this section is copied verbatim into the summary +4. **Replace history** — All old turns are removed and replaced with a single `[Context Summary]` system message containing the structured summary. The preserved recent turns remain unchanged. + +### Failure behavior + +Compaction failures are **non-fatal**. If the summarization LLM call fails, Arc emits an `Agent.Error` event and the session continues with the original uncompacted history. The next assistant turn will trigger another compaction attempt. + +### Events + +Compaction emits three events to the [event stream](/execution/observability#event-stream): + +| Event | When | Key fields | +|---|---|---| +| `Agent.ContextWindowWarning` | Token estimate exceeds 80% of context window | `estimated_tokens`, `context_window_size`, `usage_percent` | +| `Agent.CompactionStarted` | Compaction begins | `estimated_tokens`, `context_window_size` | +| `Agent.CompactionCompleted` | Summary generated and history replaced | `original_turn_count`, `preserved_turn_count`, `summary_token_estimate`, `tracked_file_count` | diff --git a/docs/execution/observability.mdx b/docs/execution/observability.mdx index 36cff0cdd..b9172a4ae 100644 --- a/docs/execution/observability.mdx +++ b/docs/execution/observability.mdx @@ -45,7 +45,9 @@ Events fall into several categories: | `Agent.Error` | `stage`, `error` | Agent-level error | | `Agent.LoopDetected` | `stage` | Repeated tool call pattern detected | | `Agent.SteeringInjected` | `stage`, `text` | Human steering message injected | -| `Agent.CompactionStarted` | `stage`, `estimated_tokens` | Context compaction triggered | +| `Agent.ContextWindowWarning` | `stage`, `estimated_tokens`, `context_window_size`, `usage_percent` | Token usage exceeds context window threshold | +| `Agent.CompactionStarted` | `stage`, `estimated_tokens`, `context_window_size` | Context compaction triggered | +| `Agent.CompactionCompleted` | `stage`, `original_turn_count`, `preserved_turn_count`, `summary_token_estimate`, `tracked_file_count` | Context compaction finished | | `Agent.LlmRetry` | `stage`, `provider`, `model`, `attempt`, `delay_secs` | LLM API call retried | | `Agent.SubAgentSpawned` | `stage`, `agent_id`, `task` | Sub-agent launched | | `Agent.SubAgentCompleted` | `stage`, `agent_id`, `success`, `turns_used` | Sub-agent finished | diff --git a/docs/images/example-semantic-port.svg b/docs/images/example-semantic-port.svg new file mode 100644 index 000000000..2dd9009a9 --- /dev/null +++ b/docs/images/example-semantic-port.svg @@ -0,0 +1,116 @@ + + + + + + + + + + + Start + + + + Fetch + + + + Analyze + + + + Plan + + + + Implement + + + + Validate + + + + Tests pass? + + + + Finalize + + + + + + + + Exit + + + + Fix + + + + + + + + + + + Process + + + + + Done + + + + + Port + + + + + Skip + + + + + + + + + + + + + + + + + Pass + + + + + Fail + + + + + + + + + + diff --git a/docs/images/example-solitaire.svg b/docs/images/example-solitaire.svg new file mode 100644 index 000000000..56ea7d46e --- /dev/null +++ b/docs/images/example-solitaire.svg @@ -0,0 +1,168 @@ + + + + + + + + + + + Start + + + + Spec + + + + Setup + + + + OK? + + + + Data + + + + OK? + + + + Logic + + + + OK? + + + + UI + + + + OK? + + + + Integrate + + + + OK? + + + + Review + + + + OK? + + + + + + + + Exit + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + Retry + + + + + + + + + + + + + + + + + + + + + + + Fix + + diff --git a/docs/images/nlspec-convergence.svg b/docs/images/nlspec-convergence.svg new file mode 100644 index 000000000..bb237b699 --- /dev/null +++ b/docs/images/nlspec-convergence.svg @@ -0,0 +1,148 @@ + + + + + + + +NLSpecConvergence + + +start + + + + + +Start + + + +plan + +Plan + + + +start->plan + + + + + +exit + + + + + +Exit + + + +implement + +Implement + + + +plan->implement + + + + + +test_quick + +Quick Tests + + + +implement->test_quick + + + + + +gate_quick + +Quick +passing? + + + +test_quick->gate_quick + + + + + +test_full + +Full Tests + + + +gate_quick->test_full + + +Pass + + + +fix + +Fix Failures + + + +gate_quick->fix + + +Fix + + + +gate_full + +All +passing? + + + +test_full->gate_full + + + + + +gate_full->exit + + +Pass + + + +gate_full->fix + + +Fix + + + +fix->test_quick + + + + + diff --git a/docs/images/tutorial-branch-loop.svg b/docs/images/tutorial-branch-loop.svg new file mode 100644 index 000000000..0baa7070e --- /dev/null +++ b/docs/images/tutorial-branch-loop.svg @@ -0,0 +1,103 @@ + + + + + + + +BranchLoop + + +start + + + + + +Start + + + +plan + + +Plan + + + +start->plan + + + + + +exit + + + + + +Exit + + + +implement + +Implement + + + +plan->implement + + + + + +validate + +Validate + + + +implement->validate + + + + + +gate + +Tests passing? + + + +validate->gate + + + + + +gate->exit + + +Pass + + + +gate->implement + + +Fix + + + diff --git a/docs/images/tutorial-ensemble.svg b/docs/images/tutorial-ensemble.svg new file mode 100644 index 000000000..967579db9 --- /dev/null +++ b/docs/images/tutorial-ensemble.svg @@ -0,0 +1,161 @@ + + + + + + + +Ensemble + + +start + + + + + +Start + + + +fork + + + +Fan Out + + + +start->fork + + + + + +exit + + + + + +Exit + + + +opus + + +Opus +(Anthropic) + + + +fork->opus + + + + + +gemini + + +Gemini +(Google) + + + +fork->gemini + + + + + +codex + + +Codex +(OpenAI) + + + +fork->codex + + + + + +mercury + + +Mercury +(Inception) + + + +fork->mercury + + + + + +merge + + + +Merge + + + +opus->merge + + + + + +gemini->merge + + + + + +codex->merge + + + + + +mercury->merge + + + + + +synth + + +Synthesize + + + +merge->synth + + + + + +synth->exit + + + + + diff --git a/docs/images/tutorial-hello.svg b/docs/images/tutorial-hello.svg new file mode 100644 index 000000000..1677ed8f8 --- /dev/null +++ b/docs/images/tutorial-hello.svg @@ -0,0 +1,59 @@ + + + + + + + +Hello + + +start + + + + + +Start + + + +compose + + +Compose + + + +start->compose + + + + + +exit + + + + + +Exit + + + +compose->exit + + + + + diff --git a/docs/images/tutorial-multi-model.svg b/docs/images/tutorial-multi-model.svg new file mode 100644 index 000000000..271322496 --- /dev/null +++ b/docs/images/tutorial-multi-model.svg @@ -0,0 +1,100 @@ + + + + + + + +MultiModel + + +start + + + + + +Start + + + +spec + + +Write Spec +(Haiku) + + + +start->spec + + + + + +exit + + + + + +Exit + + + +implement + +Implement +(Sonnet) + + + +spec->implement + + + + + +test + +Write Tests +(Sonnet) + + + +implement->test + + + + + +review + + +Code Review +(Sonnet) + + + +test->review + + + + + +review->exit + + + + + diff --git a/docs/images/tutorial-parallel-review.svg b/docs/images/tutorial-parallel-review.svg new file mode 100644 index 000000000..777498537 --- /dev/null +++ b/docs/images/tutorial-parallel-review.svg @@ -0,0 +1,138 @@ + + + + + + + +Parallel + + +start + + + + + +Start + + + +fork + + + +Fan Out + + + +start->fork + + + + + +exit + + + + + +Exit + + + +security + + +Security Audit + + + +fork->security + + + + + +architecture + + +Architecture Review + + + +fork->architecture + + + + + +quality + + +Code Quality + + + +fork->quality + + + + + +merge + + + +Merge + + + +security->merge + + + + + +architecture->merge + + + + + +quality->merge + + + + + +report + + +Final Report + + + +merge->report + + + + + +report->exit + + + + + diff --git a/docs/images/tutorial-plan-implement.svg b/docs/images/tutorial-plan-implement.svg new file mode 100644 index 000000000..f11bd861c --- /dev/null +++ b/docs/images/tutorial-plan-implement.svg @@ -0,0 +1,102 @@ + + + + + + + +PlanImplement + + +start + + + + + +Start + + + +plan + +Plan + + + +start->plan + + + + + +exit + + + + + +Exit + + + +approve + +Approve Plan + + + +plan->approve + + + + + +approve->plan + + +Revise + + + +implement + +Implement + + + +approve->implement + + +Approve + + + +simplify + +Simplify + + + +implement->simplify + + + + + +simplify->exit + + + + + diff --git a/docs/images/tutorial-subagent.svg b/docs/images/tutorial-subagent.svg new file mode 100644 index 000000000..48cbd788d --- /dev/null +++ b/docs/images/tutorial-subagent.svg @@ -0,0 +1,58 @@ + + + + + + + +SubAgent + + +start + + + + + +Start + + + +research + +Research + + + +start->research + + + + + +exit + + + + + +Exit + + + +research->exit + + + + + diff --git a/docs/images/tutorial-tool-use.svg b/docs/images/tutorial-tool-use.svg new file mode 100644 index 000000000..578bdd301 --- /dev/null +++ b/docs/images/tutorial-tool-use.svg @@ -0,0 +1,58 @@ + + + + + + + +ToolUse + + +start + + + + + +Start + + + +explore + +Explore + + + +start->explore + + + + + +exit + + + + + +Exit + + + +explore->exit + + + + + diff --git a/docs/tutorials/branch-loop.mdx b/docs/tutorials/branch-loop.mdx index 789f4dd52..25526bec1 100644 --- a/docs/tutorials/branch-loop.mdx +++ b/docs/tutorials/branch-loop.mdx @@ -2,3 +2,129 @@ title: "Branch & Loop" description: "Conditionals, test validation loops, and max_visits" --- + +This tutorial builds an implement-test-fix loop — the agent writes code, runs tests, and if they fail, fixes the code and tries again. No human intervention needed. + +## The workflow + + + Branch-Loop workflow: Start → Plan → Implement → Validate → Tests passing? → Pass to Exit or Fix back to Implement + + +```dot title="branch-loop.dot" +digraph BranchLoop { + graph [goal="Create a Python script that passes its test suite"] + rankdir=LR + + start [shape=Mdiamond, label="Start"] + exit [shape=Msquare, label="Exit"] + + plan [label="Plan", prompt="Plan a small Python script (fizzbuzz.py) and a test file (test_fizzbuzz.py) using pytest. Describe what you will create.", shape=tab, reasoning_effort="low"] + implement [label="Implement", prompt="Create fizzbuzz.py and test_fizzbuzz.py as planned. Write the files to disk."] + validate [label="Validate", shape=parallelogram, script="python -m pytest test_fizzbuzz.py -v 2>&1 || true"] + gate [shape=diamond, label="Tests passing?"] + + start -> plan -> implement -> validate -> gate + gate -> exit [label="Pass", condition="outcome=success"] + gate -> implement [label="Fix"] +} +``` + +```bash +arc run start demo/05-branch-loop.dot +``` + +## Command nodes + +The `validate` node has `shape=parallelogram`, making it a **command node**. It runs a shell script and captures the output: + +```dot +validate [label="Validate", shape=parallelogram, script="python -m pytest test_fizzbuzz.py -v 2>&1 || true"] +``` + +The `|| true` ensures the command always exits successfully — this way the node itself doesn't fail even when tests fail. The test output is captured as `command.output` in the [run context](/execution/context) for downstream nodes to use. + +## Conditional branching + +The `gate` node has `shape=diamond`, making it a **conditional node**. It evaluates outgoing edge conditions against the current run context: + +```dot +gate [shape=diamond, label="Tests passing?"] + +gate -> exit [label="Pass", condition="outcome=success"] +gate -> implement [label="Fix"] +``` + +- If the previous stage (`validate`) succeeded → take the "Pass" edge to `exit` +- Otherwise → take the unconditional "Fix" edge back to `implement` + +The `condition="outcome=success"` checks the status of the most recently completed stage. An edge without a `condition` acts as the default fallback. + +### Condition expressions + +Conditions support more than just equality checks: + +```dot +// Numeric comparisons +gate -> fast [condition="context.score > 80"] + +// Substring matching +gate -> alert [condition="context.log contains error"] + +// Boolean logic +gate -> deploy [condition="outcome=success && context.tests_passed=true"] + +// Negation +gate -> retry [condition="!outcome=success"] +``` + +See [Transitions](/workflows/transitions) for the full condition grammar. + +## The fix loop + +When tests fail, execution loops back to `implement`. The agent receives context about what happened — the test output and failure details — so it can fix the issues: + +``` +start → plan → implement → validate → gate → [Fix] → implement → validate → gate → [Pass] → exit +``` + +Each time `implement` runs, it sees the previous test output in its preamble, guiding the fix. + +## Preventing infinite loops + +This workflow has no explicit loop limit, but in production you should add one. Use `max_visits` to cap how many times a node can execute: + +```dot +implement [label="Implement", max_visits=3, prompt="..."] +``` + +After 3 visits, Arc terminates the run rather than looping forever. You can also set a graph-level limit: + +```dot +graph [max_node_visits="20"] +``` + +## Node type summary + +This workflow uses four node types: + +| Node | Shape | What it does | +|---|---|---| +| `plan` | `tab` | Single LLM call, no tools | +| `implement` | `box` (default) | Agent with tool access | +| `validate` | `parallelogram` | Runs a shell script | +| `gate` | `diamond` | Routes based on conditions | + +## What you've learned + +- **Command nodes** (`shape=parallelogram`) run shell scripts and capture output +- **Conditional nodes** (`shape=diamond`) route execution based on edge conditions +- **Loops** are just edges that point backward — the agent receives prior context on each visit +- **`max_visits`** prevents infinite loops +- Unconditional edges act as default fallbacks + +## Next + + + Fan out to concurrent branches and merge the results. + diff --git a/docs/tutorials/ensemble.mdx b/docs/tutorials/ensemble.mdx index 500b10214..464336b9e 100644 --- a/docs/tutorials/ensemble.mdx +++ b/docs/tutorials/ensemble.mdx @@ -2,3 +2,128 @@ title: "Ensemble" description: "Multi-provider fan-out, error policies, and result synthesis" --- + +This tutorial combines parallel execution with multi-model routing to get independent opinions from four different LLM providers, then synthesizes the results. This is the ensemble pattern — useful when you want diverse perspectives, consensus-based decisions, or protection against any single model's blind spots. + +## The workflow + + + Ensemble workflow: Start → Fan Out → Opus, Gemini, Codex, Mercury → Merge → Synthesize → Exit + + +```dot title="ensemble.dot" +digraph Ensemble { + graph [ + goal="Get independent opinions from multiple providers, then synthesize", + model_stylesheet=" + #opus { llm_model: claude-opus-4-6; llm_provider: anthropic; } + #gemini { llm_model: gemini-3.1-pro-preview; llm_provider: gemini; } + #codex { llm_model: gpt-5.3-codex; llm_provider: openai; } + #mercury { llm_model: mercury-2; llm_provider: inception; } + #synth { llm_model: claude-opus-4-6; llm_provider: anthropic; reasoning_effort: high; } + " + ] + rankdir=LR + + start [shape=Mdiamond, label="Start"] + exit [shape=Msquare, label="Exit"] + + fork [label="Fan Out", shape=component, join_policy="wait_all", error_policy="continue"] + + opus [label="Opus", prompt="Analyze the goal. Provide your independent assessment, recommendations, and any code or prose needed. Be thorough.", shape=tab] + gemini [label="Gemini", prompt="Analyze the goal. Provide your independent assessment, recommendations, and any code or prose needed. Be thorough.", shape=tab] + codex [label="Codex", prompt="Analyze the goal. Provide your independent assessment, recommendations, and any code or prose needed. Be thorough.", shape=tab] + mercury [label="Mercury", prompt="Analyze the goal. Provide your independent assessment, recommendations, and any code or prose needed. Be thorough.", shape=tab] + + merge [label="Merge", shape=tripleoctagon] + synth [label="Synthesize", prompt="You have received independent analyses from four different models (Opus, Gemini, Codex, Mercury). Compare their perspectives: identify consensus, highlight disagreements, and synthesize the strongest ideas into a single coherent recommendation. Note where models agreed and where they diverged.", shape=tab] + + start -> fork + fork -> opus + fork -> gemini + fork -> codex + fork -> mercury + opus -> merge + gemini -> merge + codex -> merge + mercury -> merge + merge -> synth -> exit +} +``` + +```bash +arc run start demo/11-ensemble.dot +``` + + +This workflow requires API keys for all four providers (`ANTHROPIC_API_KEY`, `GEMINI_API_KEY`, `OPENAI_API_KEY`, `INCEPTION_API_KEY`). If a provider key is missing, that branch will fail — but `error_policy="continue"` ensures the other branches still complete. + + +## How it works + +The workflow has three phases: + +### 1. Fan-out to four providers + +The `fork` node spawns four parallel branches, each assigned to a different provider via the stylesheet: + +``` +#opus { llm_model: claude-opus-4-6; llm_provider: anthropic; } +#gemini { llm_model: gemini-3.1-pro-preview; llm_provider: gemini; } +#codex { llm_model: gpt-5.3-codex; llm_provider: openai; } +#mercury { llm_model: mercury-2; llm_provider: inception; } +``` + +Each branch receives the same prompt but runs on a completely different model. The branches execute concurrently and have no knowledge of each other's responses. + +### 2. Merge results + +The `merge` node collects all four responses. With `error_policy="continue"`, it waits for every branch — even if some fail. A missing API key or provider outage doesn't cancel the entire workflow. + +### 3. Synthesize + +The `synth` node receives all four perspectives in its preamble and produces a unified recommendation. It uses `reasoning_effort: high` because comparing and synthesizing multiple viewpoints is a harder task than generating any single one. + +## Combining patterns + +This workflow combines two patterns from earlier tutorials: + +- **Parallel execution** from [Parallel Review](/tutorials/parallel-review) — fan-out/fan-in with join and error policies +- **Model routing** from [Multi-Model Routing](/tutorials/multi-model) — stylesheet selectors assigning different providers to each node + +The key difference from the parallel review tutorial is that here each branch uses a _different provider_, not just a different prompt. This gives you genuinely independent perspectives — each model has different training data, different reasoning patterns, and different blind spots. + +## When to use ensembles + +The ensemble pattern is most valuable when: + +- **Correctness matters more than speed** — e.g., security audits, architectural decisions, spec reviews +- **You want to detect model-specific blind spots** — if three models agree and one disagrees, the disagreement is worth investigating +- **You need confidence in a judgment call** — consensus across models is stronger than any single model's opinion + +The tradeoff is cost and latency — you're making 4x the LLM calls. Use single-model workflows for routine tasks and ensembles for high-stakes decisions. + +## What you've learned + +- **Ensemble workflows** fan out the same task to multiple providers +- **`error_policy="continue"`** keeps the workflow running even when some branches fail +- **ID selectors** (`#opus`, `#gemini`) assign each branch to a specific model +- A **synthesis node** compares perspectives and produces a unified result +- Combine parallel execution and model routing for diverse, independent analysis + +## Further reading + + + + Available models, providers, and fallback configuration. + + + Full stylesheet syntax and specificity rules. + + + Complete reference for all node types. + + + Edge conditions, routing directives, and tiebreaking. + + diff --git a/docs/tutorials/hello-world.mdx b/docs/tutorials/hello-world.mdx index 86674b857..7153bf4a7 100644 --- a/docs/tutorials/hello-world.mdx +++ b/docs/tutorials/hello-world.mdx @@ -2,3 +2,132 @@ title: "Hello World" description: "Your first workflow: prompt nodes, tool use, and sub-agents" --- + +This tutorial walks through three minimal workflows that introduce the building blocks of Arc: a one-shot prompt, an agent with tool access, and a sub-agent delegation pattern. + +## Prerequisites + +Complete the [Quick Start](/getting-started/quick-start) so you have a working `arc` binary and at least one LLM API key configured. + +## 1. One-shot prompt + +The simplest possible workflow has one node that sends a prompt to an LLM and exits. + + + Hello World workflow: Start → Compose → Exit + + +```dot title="hello.dot" +digraph Hello { + graph [goal="Write a haiku about software workflows"] + rankdir=LR + + start [shape=Mdiamond, label="Start"] + exit [shape=Msquare, label="Exit"] + + compose [label="Compose", prompt="Write a haiku (5-7-5 syllable) about software workflows. Output only the haiku, nothing else.", shape=tab, reasoning_effort="low"] + + start -> compose -> exit +} +``` + +Run it: + +```bash +arc run start demo/01-hello.dot +``` + +### What's happening + +- `shape=tab` makes this a **prompt node** — a single LLM call with no tool access. Good for generation, summarization, and classification. +- `reasoning_effort="low"` tells the model to think less. This is a simple task that doesn't need deep reasoning. +- `graph [goal="..."]` describes the workflow's purpose. Arc uses it in preambles and retrospectives. + +Every workflow needs exactly one `start` node (`shape=Mdiamond`) and one `exit` node (`shape=Msquare`). + +## 2. Agent with tools + +An **agent node** (the default `box` shape) runs an LLM in a loop with access to tools — bash, file reading, file editing, grep, and glob. The agent calls tools autonomously until it decides the task is complete. + + + Tool Use workflow: Start → Explore → Exit + + +```dot title="tool-use.dot" +digraph ToolUse { + graph [goal="Explore the current directory using shell tools"] + rankdir=LR + + start [shape=Mdiamond, label="Start"] + exit [shape=Msquare, label="Exit"] + + explore [label="Explore", prompt="Use bash to list the files in the current directory, then read the first 5 lines of any README or CLAUDE.md file you find. Summarize what this project is about in 2-3 sentences."] + + start -> explore -> exit +} +``` + +```bash +arc run start demo/02-tool-use.dot +``` + +### What's happening + +- No `shape` attribute means the default `box` — an **agent node**. +- The agent has access to [built-in tools](/agents/tools): `shell`, `read_file`, `write_file`, `edit_file`, `grep`, `glob`, `web_search`, and `web_fetch`. +- The agent decides which tools to call and when to stop. Arc handles the tool loop automatically. + +### Prompt vs. agent nodes + +| | Prompt node (`tab`) | Agent node (`box`) | +|---|---|---| +| LLM calls | Single call | Multi-turn loop | +| Tool access | None | Full toolset | +| Use case | Analysis, generation | Tasks requiring file I/O and commands | + +## 3. Sub-agents + +An agent can spawn **sub-agents** to delegate work. Sub-agents run in their own session and return results to the parent. + + + Sub-agent workflow: Start → Research → Exit + + +```dot title="sub-agent.dot" +digraph SubAgent { + graph [goal="Research and summarize using a sub-agent"] + rankdir=LR + + start [shape=Mdiamond, label="Start"] + exit [shape=Msquare, label="Exit"] + + research [label="Research", prompt="You have a sub-agent available via the spawn_agent tool. Spawn a sub-agent to list the files in the current directory and read the first 10 lines of any README or CLAUDE.md. Then, using the sub-agent's findings, write a 2-sentence summary of the project."] + + start -> research -> exit +} +``` + +```bash +arc run start demo/03-subagent.dot +``` + +### What's happening + +- The parent agent uses `spawn_agent` to create a child session, then `wait` to collect the result. +- Sub-agents have their own tool access and conversation history — they don't see the parent's context. +- This pattern is useful for parallelizing research, isolating risky operations, or keeping the parent's context window lean. + +See [Sub-agents](/agents/subagents) for the full tool reference. + +## What you've learned + +- **Prompt nodes** (`shape=tab`) make a single LLM call — no tools +- **Agent nodes** (default `box`) run a multi-turn tool loop +- **Sub-agents** let an agent delegate to independent child sessions +- Every workflow needs a `start` node, an `exit` node, and a `goal` + +## Next + + + Add human gates and revision loops to a multi-step workflow. + diff --git a/docs/tutorials/multi-model.mdx b/docs/tutorials/multi-model.mdx index bd968e9da..dd717b7d9 100644 --- a/docs/tutorials/multi-model.mdx +++ b/docs/tutorials/multi-model.mdx @@ -2,3 +2,125 @@ title: "Multi-Model Routing" description: "Model stylesheets, CSS selectors, and per-node reasoning effort" --- + +This tutorial assigns different models to different tasks in a single workflow — a cheap, fast model for the spec, a capable model for coding, and a different model for review. The routing is controlled by a CSS-like stylesheet. + +## The workflow + + + Multi-Model workflow: Start → Write Spec (Haiku) → Implement (Sonnet) → Write Tests (Sonnet) → Code Review (Sonnet) → Exit + + +```dot title="multi-model.dot" +digraph MultiModel { + graph [ + goal="Build and review a utility function using multiple models", + model_stylesheet=" + * { llm_model: claude-haiku-4-5; llm_provider: anthropic; reasoning_effort: low; } + .coding { llm_model: claude-sonnet-4-5; llm_provider: anthropic; reasoning_effort: high; } + #review { llm_model: claude-sonnet-4-5; llm_provider: anthropic; reasoning_effort: high; } + " + ] + rankdir=LR + + start [shape=Mdiamond, label="Start"] + exit [shape=Msquare, label="Exit"] + + spec [label="Write Spec", prompt="Write a brief spec for a TypeScript string utility module with 3 functions: slugify, truncate, and capitalize. Output the spec only.", shape=tab] + implement [label="Implement", prompt="Implement the TypeScript string utility module from the spec. Write it to string-utils.ts.", class="coding"] + test [label="Write Tests", prompt="Write tests for the string utility module using Bun's test runner. Write to string-utils.test.ts.", class="coding"] + review [label="Code Review", prompt="Review the implementation and tests. Check for edge cases, type safety, and correctness. Provide a brief verdict.", shape=tab] + + start -> spec -> implement -> test -> review -> exit +} +``` + +```bash +arc run start demo/08-multi-model.dot +``` + +## Model stylesheets + +The `model_stylesheet` graph attribute contains CSS-like rules that assign models to nodes: + +``` +* { llm_model: claude-haiku-4-5; llm_provider: anthropic; reasoning_effort: low; } +.coding { llm_model: claude-sonnet-4-5; llm_provider: anthropic; reasoning_effort: high; } +#review { llm_model: claude-sonnet-4-5; llm_provider: anthropic; reasoning_effort: high; } +``` + +### Selectors + +| Selector | Syntax | Matches | Specificity | +|---|---|---|---| +| Universal | `*` | All nodes | 0 | +| Shape | `box`, `tab`, etc. | Nodes with that shape | 1 | +| Class | `.classname` | Nodes with `class="classname"` | 2 | +| ID | `#nodeid` | A specific node by ID | 3 | + +Higher specificity wins. If two rules have the same specificity, the last one in the stylesheet wins. + +### How this workflow routes + +| Node | Matches | Model | Why | +|---|---|---|---| +| `spec` | `*` (universal) | Haiku | Simple generation task — fast and cheap | +| `implement` | `.coding` (class) | Sonnet | Coding requires a capable model | +| `test` | `.coding` (class) | Sonnet | Test writing also needs coding capability | +| `review` | `#review` (ID) | Sonnet | Review needs careful analysis | + +### Assigning classes + +Set the `class` attribute on a node to target it with class selectors: + +```dot +implement [label="Implement", class="coding"] +``` + +Multiple classes are space-separated: `class="coding critical"`. + +## Properties + +Stylesheets support four properties: + +| Property | Description | +|---|---| +| `llm_model` | Model ID or alias (e.g. `claude-sonnet-4-5`, `opus`, `gemini-pro`) | +| `llm_provider` | Provider name (`anthropic`, `openai`, `gemini`, etc.) | +| `reasoning_effort` | `low`, `medium`, or `high` | +| `backend` | `api` (default) or `cli` | + +## Why route models? + +Not every task needs a frontier model: + +- **Spec writing, classification, summarization** — use a fast, cheap model (Haiku, Flash Lite) +- **Code implementation, complex reasoning** — use a capable model (Sonnet, Opus, GPT-5.2) +- **Cross-critique** — use a _different provider_ so the reviewer brings fresh eyes + +Model routing lets you optimize cost and latency without changing the workflow structure. Swap `claude-haiku-4-5` to `gemini-3-flash-preview` in the stylesheet and the workflow behaves the same — just with a different model underneath. + +## Explicit overrides + +A model set directly on a node attribute always beats the stylesheet: + +```dot +implement [label="Implement", class="coding", llm_model="claude-opus-4-6"] +``` + +This node uses Opus regardless of what `.coding` says. + +See [Model Stylesheets](/workflows/stylesheets) for the full reference and [Models](/core-concepts/models) for available model IDs. + +## What you've learned + +- **Model stylesheets** use CSS-like rules to assign models to nodes +- **Selectors** match by universal (`*`), shape, class (`.name`), or ID (`#name`) +- **Specificity** determines which rule wins when multiple match +- Route cheap models to simple tasks and capable models to hard ones + +## Next + + + Fan out to multiple providers and synthesize their independent opinions. + diff --git a/docs/tutorials/parallel-review.mdx b/docs/tutorials/parallel-review.mdx index f564235bb..09ad82286 100644 --- a/docs/tutorials/parallel-review.mdx +++ b/docs/tutorials/parallel-review.mdx @@ -2,3 +2,118 @@ title: "Parallel Review" description: "Fan-out, fan-in, join policies, and merge nodes" --- + +This tutorial runs three code review perspectives in parallel — security, architecture, and quality — then merges the results into a single report. + +## The workflow + + + Parallel Review workflow: Start → Fan Out → Security Audit, Architecture Review, Code Quality → Merge → Final Report → Exit + + +```dot title="parallel.dot" +digraph Parallel { + graph [goal="Perform a multi-perspective code review"] + rankdir=LR + + start [shape=Mdiamond, label="Start"] + exit [shape=Msquare, label="Exit"] + + fork [label="Fork Analysis", shape=component, join_policy="wait_all", error_policy="continue"] + + security [label="Security Audit", prompt="Examine the codebase for security concerns: hardcoded secrets, injection risks, unsafe dependencies. List findings as bullet points.", shape=tab, reasoning_effort="low"] + architecture [label="Architecture Review", prompt="Assess the codebase architecture: separation of concerns, dependency structure, modularity. List findings as bullet points.", shape=tab, reasoning_effort="low"] + quality [label="Code Quality", prompt="Check code quality: naming conventions, dead code, test coverage gaps, error handling. List findings as bullet points.", shape=tab, reasoning_effort="low"] + + merge [label="Merge Findings", shape=tripleoctagon] + report [label="Final Report", prompt="Synthesize the security, architecture, and code quality findings into a prioritized summary report with top 5 action items.", shape=tab] + + start -> fork + fork -> security + fork -> architecture + fork -> quality + security -> merge + architecture -> merge + quality -> merge + merge -> report -> exit +} +``` + +```bash +arc run start demo/06-parallel.dot +``` + +## Fan-out with the fork node + +The `fork` node has `shape=component`, making it a **parallel fan-out node**. Every outgoing edge becomes a concurrent branch: + +```dot +fork [label="Fork Analysis", shape=component, join_policy="wait_all", error_policy="continue"] + +fork -> security +fork -> architecture +fork -> quality +``` + +All three branches start at the same time. Each gets an isolated copy of the run context, so branches can't interfere with each other. + +### Join policies + +The `join_policy` controls when execution can proceed past the merge: + +| Policy | Behavior | +|---|---| +| `wait_all` | Wait for every branch to finish (default) | +| `first_success` | Proceed as soon as one branch succeeds | +| `k_of_n(N)` | Proceed after N branches succeed | +| `quorum(0.5)` | Proceed after a fraction of branches succeed | + +### Error policies + +The `error_policy` controls what happens when a branch fails: + +| Policy | Behavior | +|---|---| +| `continue` | Run all branches even if some fail (default) | +| `fail_fast` | Cancel remaining branches as soon as one fails | +| `ignore` | Treat all branch failures as successes | + +This workflow uses `error_policy="continue"` so that a failure in one review perspective doesn't cancel the others. + +## Fan-in with the merge node + +The `merge` node has `shape=tripleoctagon`, making it a **merge (fan-in) node**. It collects results from all branches into a single context: + +```dot +merge [label="Merge Findings", shape=tripleoctagon] + +security -> merge +architecture -> merge +quality -> merge +``` + +The merged branch results are available to downstream nodes. The `report` node receives all three perspectives in its preamble and synthesizes them. + +## Concurrency control + +By default, Arc runs up to 4 parallel branches simultaneously. Control this with `max_parallel`: + +```dot +fork [shape=component, max_parallel=2] +``` + +This is useful when branches are resource-intensive (e.g., each running a full agent session with tool calls) and you want to limit concurrency. + +## What you've learned + +- **Fan-out nodes** (`shape=component`) spawn concurrent branches +- **Merge nodes** (`shape=tripleoctagon`) collect branch results +- **Join policies** control when execution can proceed past the merge +- **Error policies** control how branch failures are handled +- Each branch gets an isolated copy of the context + +## Next + + + Assign different models to different workflow nodes using stylesheets. + diff --git a/docs/tutorials/plan-implement.mdx b/docs/tutorials/plan-implement.mdx index 50611f3bc..705086422 100644 --- a/docs/tutorials/plan-implement.mdx +++ b/docs/tutorials/plan-implement.mdx @@ -2,3 +2,110 @@ title: "Plan & Implement" description: "Human gates, revision loops, and prompt file references" --- + +This tutorial builds a plan-approve-implement workflow where a human reviews the plan before the agent writes code. If the plan isn't right, the human sends it back for revision — creating a loop. + +## The workflow + + + Plan-Implement workflow: Start → Plan → Approve Plan → Implement → Simplify → Exit, with Revise loop back to Plan + + +```dot title="plan-implement.dot" +digraph PlanImplement { + graph [goal="Plan, approve, implement, and simplify a change"] + rankdir=LR + + start [shape=Mdiamond, label="Start"] + exit [shape=Msquare, label="Exit"] + + plan [label="Plan", prompt="Analyze the goal and codebase. Write a clear, step-by-step implementation plan to a Markdown file called plan.md. Include what files will change and why.", reasoning_effort="high"] + approve [shape=hexagon, label="Approve Plan"] + implement [label="Implement", prompt="Read plan.md and implement every step. Make all the code changes described in the plan."] + simplify [label="Simplify", prompt="@docs-internal/prompts/simplify.md"] + + start -> plan -> approve + approve -> implement [label="[A] Approve"] + approve -> plan [label="[R] Revise"] + implement -> simplify -> exit +} +``` + +```bash +arc run start demo/10-plan-implement.dot +``` + +## Human gates + +The `approve` node has `shape=hexagon`, which makes it a **human gate** — the workflow pauses and waits for a person to choose a path. + +```dot +approve [shape=hexagon, label="Approve Plan"] + +approve -> implement [label="[A] Approve"] +approve -> plan [label="[R] Revise"] +``` + +The outgoing edge labels define the options. In the CLI, you'll see: + +``` +? Approve Plan + [1] A - [A] Approve + [2] R - [R] Revise +Select: +``` + +The `[A]` and `[R]` prefixes are keyboard accelerators — type the letter to select. + +## The revision loop + +If you choose **Revise**, execution goes back to the `plan` node. The agent runs again with context about what happened — it knows its previous plan was rejected and can improve it. This cycle repeats until you approve. + +``` +start → plan → approve → [Revise] → plan → approve → [Approve] → implement → simplify → exit +``` + +Loops are natural in Arc — just point an edge back to an earlier node. For safety, you can set `max_visits` on a node to prevent infinite loops: + +```dot +plan [label="Plan", max_visits=5, ...] +``` + +## Reasoning effort + +The `plan` node sets `reasoning_effort="high"`. This tells the model to think harder — useful for planning tasks that require careful analysis. The default is `high`, but you can set it to `low` or `medium` for simpler tasks to save cost and time. + +## Prompt file references + +The `simplify` node uses `@docs-internal/prompts/simplify.md` instead of an inline prompt string: + +```dot +simplify [label="Simplify", prompt="@docs-internal/prompts/simplify.md"] +``` + +The `@` prefix tells Arc to load the prompt from a Markdown file, resolved relative to the DOT file's location. This keeps DOT files concise and lets you version prompts as standalone files. See [Prompts](/agents/prompts) for details. + +## Context flow between nodes + +Each node receives a **preamble** summarizing what happened in prior stages. When the `implement` node runs, it knows that a plan was written and approved. The preamble includes: + +- The workflow goal +- A summary of completed stages with their outcomes +- Files touched by prior stages +- Run context values + +The agent reads `plan.md` (as instructed by its prompt), but the preamble gives it additional context about the overall workflow state. See [Context](/execution/context) for the full reference. + +## What you've learned + +- **Human gates** (`shape=hexagon`) pause for human input with edge labels as options +- **Revision loops** are just edges that point back to earlier nodes +- **Prompt file references** (`@path/to/file.md`) keep DOT files clean +- **`reasoning_effort`** controls how hard the model thinks +- Nodes receive preambles summarizing prior stages + +## Next + + + Add conditional branching and automated test validation loops. +