Commit graph

78 commits

Author SHA1 Message Date
Bryan Helmkamp
242e82fffd Use Claude Opus 4.6 with 1M context window as default Anthropic model
Switch default model from claude-sonnet-4-5 to claude-opus-4-6 across
run and serve commands. Enable the 1M token context window via the
context-1m-2025-08-07 beta header for opus-4-6 requests.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 14:23:45 -05:00
Bryan Helmkamp
4f31a84ee6 Skip OpenAI function calls with empty names to avoid model-internal items
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 14:07:23 -05:00
Bryan Helmkamp
e52b19d1ce Clarify spec-dod prompts to respond with JSON inline instead of writing files
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 14:05:30 -05:00
Bryan Helmkamp
f4f58078de Update default Gemini model to gemini-3.1-pro-preview and print model name on startup
Updates the default Gemini model from gemini-3-pro-preview to
gemini-3.1-pro-preview across agent, attractor, and ullm CLIs.
Adds "Using model:" output to stderr in both agent and ullm CLIs
so users can see which model is being used.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 14:05:16 -05:00
Bryan Helmkamp
366643d6ef Add attractor-web React frontend for pipeline dashboard
Bun-based React app with pipeline dashboard UI including event log,
graph view, context/checkpoint panels, question panel, and status bar.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 13:47:03 -05:00
Bryan Helmkamp
b4f65e2555 Increase stream_read timeout from 30s to 120s
Prevents premature timeouts on slower streaming responses.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 13:46:32 -05:00
Bryan Helmkamp
e64c0a3328 Switch default model to claude-opus-4-6 and use adaptive thinking
Updates the default Anthropic model from claude-sonnet-4-5 to
claude-opus-4-6 and replaces beta headers with adaptive thinking
provider options for 4.6 models.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 13:46:31 -05:00
Bryan Helmkamp
379e90c11d Add --docker flag to attractor run and serve commands
When --docker is passed, agent tools execute inside a Docker container
via DockerExecutionEnvironment instead of running on the host. The host
working directory is bind-mounted into the container at /workspace.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 13:42:10 -05:00
Bryan Helmkamp
ffa92cd23d Add GET /pipelines endpoint and pipeline list UI
Adds a list endpoint that returns all in-memory pipelines with their
status, and a "Recent Pipelines" section in the start form that polls
every 3s so users can navigate to any pipeline started since server launch.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 13:39:28 -05:00
Bryan Helmkamp
e0f33e072e Fix Dockerfile.agent to use apt ripgrep for multi-arch support
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 13:31:19 -05:00
Bryan Helmkamp
60dad3c1cd Add DockerExecutionEnvironment for sandboxed agent tool execution
Implements ExecutionEnvironment trait backed by Docker containers via
bollard. Host working directory is bind-mounted; all file ops, commands,
grep, and glob execute inside the container via docker exec. Extracts
shared format_lines_numbered() helper from LocalExecutionEnvironment.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 13:29:38 -05:00
Bryan Helmkamp
86935bbb2b Expose MultipleChoice options in API and support selected answers
ApiQuestion now includes options and allow_freeform fields so API clients
can render multiple-choice questions. submit_answer accepts an optional
selected_option_key to produce AnswerValue::Selected instead of always
creating AnswerValue::Text.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 13:28:36 -05:00
Bryan Helmkamp
e4fc345912 Switch agent session from complete() to stream() to avoid request timeouts
Long Opus responses hit the 120s request timeout with complete(). With
stream(), the request timeout only covers initial connection + first
chunk, then the 30s stream_read timeout guards against stalls between
chunks. Also emits AssistantTextDelta events during streaming.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 13:10:54 -05:00
Bryan Helmkamp
dba48f866e Fix CodergenHandler swallowing retryable errors, preventing engine retries
The handler was converting all backend errors into Ok(Outcome::fail(...)),
which the engine treats as a terminal result. Retryable errors (timeouts,
network failures) are now propagated as Err so the engine retry loop kicks in.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 12:43:48 -05:00
Bryan Helmkamp
5573e0e8a2 Fix truncated LLM responses by populating catalog max_output
Requests were always sent with max_tokens: None, which defaulted to
4096 in the Anthropic provider, causing large outputs to be truncated.
Now build_request() looks up the model's max_output from the catalog
and passes it through, giving each model its full output capacity.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 12:30:25 -05:00
Bryan Helmkamp
3a70e6711c serve 2026-02-23 12:20:48 -05:00
Bryan Helmkamp
dddc6df8c8 Extract terminal crate and prettify attractor CLI output
Move ANSI Styles struct from agent/cli.rs into a shared terminal crate
so both binaries can use it. Add green and yellow color codes. Prettify
all attractor CLI output: bold headers, colored diagnostics by severity,
green/red status, yellow warnings, dimmed event details, and styled
interviewer prompts. Move pipeline status output from stdout to stderr.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 12:06:24 -05:00
Bryan Helmkamp
b4397564ec Prettify agent CLI stderr output with ANSI colors
Add a Styles struct that pre-resolves ANSI escape codes based on whether
stderr is a TTY. Tool calls show as "● tool_name(args)" with bold+cyan
names, errors use red "✗", summaries are dimmed, and debug/verbose
middleware output is styled. No ANSI codes emitted when piped.

Also passes tool call arguments through to EventData::ToolCall so the
event handler can format them inline.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 11:45:13 -05:00
Bryan Helmkamp
df8502d83b Merge agent-cli crate into agent as cli module
The agent-cli binary was a thin wrapper over the agent library with
nothing else depending on it. Moving it into the agent crate as a
`pub mod cli` with a `[[bin]]` entry reduces workspace complexity
and follows the attractor crate pattern.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 11:14:25 -05:00
Bryan Helmkamp
eaf934f355 Add tests for tool_approval and agent-cli
- 5 session tests exercising ToolApprovalFn callback (deny, allow,
  arg capture, None passthrough, error event emission)
- 18 unit tests for agent-cli pure functions (tool_category,
  is_auto_approved, default_model, validate_api_key,
  build_tool_approval, build_profile)
- 4 integration tests for the agent binary (usage, help, missing
  API key, invalid permissions)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 11:05:44 -05:00
Bryan Helmkamp
0edc93e1c3 Add CLI tests for validate and dry-run on all test workflows
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:58:54 -05:00
Bryan Helmkamp
f6bf30de3e Disable empty doc-tests and add terse test output alias
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:57:26 -05:00
Bryan Helmkamp
4915574f96 Preserve provider-specific content parts across agent turns
The agent's Turn::Assistant decomposed responses into text, tool_calls,
and reasoning fields, discarding ContentPart::Other items. This lost
OpenAI reasoning items needed for Responses API round-tripping.

Add provider_parts field to Turn::Assistant to carry opaque content
parts through the history, and emit them first in convert_to_messages
so they precede function_call items as the API requires.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:54:35 -05:00
Bryan Helmkamp
b7883d6bef Add agent CLI with tool approval callback
Introduce `ToolApprovalFn` callback in `SessionConfig` to gate tool
execution by permission level. Create `agent-cli` crate as a thin CLI
binary wrapping `Session` with provider/model resolution, permission
model (read-only/read-write/full), interactive approval prompts,
real-time event rendering, debug middleware, and SIGINT handling.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:50:40 -05:00
Bryan Helmkamp
17a42d19f2 Preserve OpenAI reasoning items for Responses API round-trip
The Responses API requires that reasoning items (type: "reasoning",
id: "rs_xxx") are included alongside their associated function_call
items when replaying conversation history. Without them:
"function_call was provided without its required 'reasoning' item"

Store reasoning output items as ContentPart::Other and replay them
as raw input items in translate_input. Handles both non-streaming
(parse_output) and streaming (handle_output_item_done) paths.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:30:39 -05:00
Bryan Helmkamp
6e15d602d0 Fix OpenAI Responses API function call ID prefix mismatch
The Responses API returns two distinct IDs on function_call items:
- `id` (item-level, starts with `fc_`)
- `call_id` (call-level, starts with `call_`)

When sending function calls back as input, the `id` field must start
with `fc_`. Previously we used `call_id` for both fields, causing:
"Invalid 'input[1].id': Expected an ID that begins with 'fc'."

Now we preserve the item-level `id` in provider_metadata and use it
for the `id` field, while `call_id` continues to link tool results.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:27:24 -05:00
Bryan Helmkamp
e992b57ced Add regression tests for Anthropic beta headers and Gemini thought_signature
Covers two recent runtime failures that lacked test coverage:
- Anthropic: assert deprecated beta header values are not sent
- Gemini: test function call parsing/translation with and without thoughtSignature

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:22:14 -05:00
Bryan Helmkamp
94326f5438 docs/specs 2026-02-23 10:17:40 -05:00
Bryan Helmkamp
3b1411a3ac Preserve Gemini thought_signature on function calls
Gemini 3 models include a thoughtSignature field on function call
parts that must be returned in subsequent conversation turns. Add
provider_metadata to ToolCall to carry this through, and extract/emit
it in both the non-streaming and streaming Gemini code paths.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:14:19 -05:00
Bryan Helmkamp
7dddc49121 Remove deprecated anthropic-beta header values
The `extended-thinking-2025-04-14` and `max-tokens-3-5-sonnet-2025-04-14`
beta headers are no longer accepted by the Anthropic API.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:05:43 -05:00
Bryan Helmkamp
d3e396753c Remove stale review docs, add spec DoD pipeline definitions
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:47:50 -05:00
Bryan Helmkamp
4bb2de5a49 Rename unified-llm crate to llm
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:22:14 -05:00
Bryan Helmkamp
023cc54fed Merge unified-llm-cli crate into unified-llm
Move the ullm binary from a separate crate into unified-llm as
src/bin/ullm.rs, consolidating CLI dependencies into the library crate.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:18:41 -05:00
Bryan Helmkamp
30cf3a6787 Rename coding-agent-loop crate to agent
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:17:48 -05:00
Bryan Helmkamp
5932b500e7 Implement GET /pipelines/{id}/graph endpoint for SVG rendering
Stores DOT source in ManagedPipeline and pipes it through `dot -Tsvg`
on request. Returns image/svg+xml on success, 502 if graphviz is
unavailable, 404 if pipeline not found. Resolves spec gap #1.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:13:36 -05:00
Bryan Helmkamp
55b6000374 Fix OpenAI streaming returning no output for reasoning models
The SSE stream terminated prematurely when a network chunk contained
only unhandled event types (e.g. response.in_progress, reasoning
output_item.added). dispatch_sse_messages returned an empty vec, which
the unfold interpreted as stream-end. Now continues reading when
dispatch produces no StreamEvents.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:12:40 -05:00
Bryan Helmkamp
433a7c6247 Implement attractor CLI binary with run and validate subcommands
Add a [[bin]] target to the attractor crate with two subcommands:
- `attractor validate <pipeline.dot>` -- parse and validate only
- `attractor run <pipeline.dot>` -- full pipeline execution with LLM backend

The CLI supports --dry-run, --auto-approve, --resume, --model, --provider,
and two-level verbosity (-v one-line summaries, -vv full event details).

Extracts a shared `default_registry()` function in handler/mod.rs so both
the CLI and server can build a fully-wired HandlerRegistry without
duplicating handler registration boilerplate.

The AgentBackend in cli/backend.rs implements CodergenBackend by creating
a coding-agent-loop Session per node invocation, giving LLM nodes access
to file and shell tools.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:11:42 -05:00
Bryan Helmkamp
54d0c887dc Add end-to-end test for subgraph class derivation, update spec gap analysis
Closes the subgraph class derivation test gap by adding a parse() pipeline
test that verifies DOT subgraph labels produce correct CSS-like classes on
contained nodes. Updates gap analysis to remove resolved items, add spec
contradictions section, and narrow remaining gaps to SVG rendering and
retry predicate customization.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 08:39:42 -05:00
Bryan Helmkamp
a917a8d615 Implement 6 spec gaps: preamble synthesis, thread_id plumbing, cancellation, recording replay, retry presets, inform()
- Gap 1: Real preamble synthesis per fidelity mode (truncate, compact,
  summary:low/medium/high) in new preamble.rs module, replacing placeholder
- Gap 2: Pass thread_id to CodergenBackend.run() so backends can reuse
  LLM sessions across nodes sharing the same thread
- Gap 7: Engine cancellation via AtomicBool token checked between nodes,
  wired to server cancel endpoint, new Cancelled error variant
- Gap 8: RecordingInterviewer serialization (to_json/from_json, file I/O)
  and new ReplayInterviewer for replaying recorded Q&A sessions
- Gap 12: Preset retry policies selectable from DOT via retry_policy attr
  (none, standard, aggressive, linear, patient)
- Gap 15: Engine calls Interviewer.inform() at pipeline start, stage
  start, and stage complete lifecycle points

37 new tests (450 unit + 49 integration, all passing).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 08:23:27 -05:00
Bryan Helmkamp
08d76b93de Add 8 integration tests: HTTP lifecycle, SSE events, sub-pipeline, manager loop, graph merge, real LLM
- Change server registry_factory to accept Arc<dyn Interviewer> so WaitHumanHandler
  shares the same WebInterviewer as the REST API endpoints
- Add full HTTP lifecycle tests: approve-and-complete flow + cancel flow
- Add SSE event stream content parsing test with frame-level verification
- Add sub-pipeline E2E test through the engine with context propagation
- Add manager loop E2E test with SimulatingChildObserver
- Add graph merge E2E test verifying module prefixing and execution ordering
- Add 3 real LLM tests (#[ignore]) using claude-haiku via AnthropicAdapter
- Add dotenvy and http-body-util dev dependencies

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 13:07:26 -04:00
Bryan Helmkamp
a6d5579d70 Merge branch 'worktree-agent-a0f72fec'
# Conflicts:
#	crates/coding-agent-loop/src/provider_profile.rs
#	crates/coding-agent-loop/src/test_support.rs
2026-02-22 12:50:15 -04:00
Bryan Helmkamp
a66e8c4657 Merge branch 'worktree-agent-a224811b'
# Conflicts:
#	crates/coding-agent-loop/src/subagent.rs
2026-02-22 12:49:29 -04:00
Bryan Helmkamp
fc5f9b096b Merge branch 'worktree-agent-a3e10c40' 2026-02-22 12:49:03 -04:00
Bryan Helmkamp
858f5622d5 Merge branch 'worktree-agent-afb2df8a' 2026-02-22 12:49:00 -04:00
Bryan Helmkamp
c9f28631b6 Consolidate test mocks and profiles in coding-agent-loop
Eliminate ~300 lines of duplicated test infrastructure:

- Add MockExecutionEnvironment::linux() constructor, replacing 4
  identical linux_env() helpers across profile test modules
- Extend MockExecutionEnvironment with written_files, captured_timeout,
  and apply_read_offset_limit fields to replace 4 specialized mocks
  (ReadFileEnv, WriteFileEnv, EditFileEnv, ShellCapturingEnv) in tools.rs
- Add MutableMockExecutionEnvironment for apply_patch tests that need
  writes visible to subsequent reads, replacing MockFileEnv in openai.rs
- Merge ParallelTestProfile into TestProfile with configurable
  parallel_tool_calls and context_window fields
- Replace ProviderTestProfile with TestProfile in provider_profile.rs
  tests, updating TestProfile::build_system_prompt to include user
  instructions like real profiles
- Extract shared CapturingLlmProvider into test_support.rs, replacing
  two inline capturing provider structs in session.rs

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 12:41:04 -04:00
Bryan Helmkamp
b2b98be522 Simplify coding-agent-loop: extract required_str helper, simplify spawn, restrict test-only accessors
- Proposal 7.1: Introduce required_str() helper in tools.rs to replace ~15
  repetitions of the args.get("param").and_then(|v| v.as_str()).ok_or_else(...)
  pattern across tools.rs and subagent.rs. Provides consistent error messages.

- Proposal 4.2: Simplify SubAgentManager::spawn success path by using ? on
  process_input() result directly, eliminating the redundant success variable
  (always true on the Ok path) and the if-let-Err-return pattern.

- Proposal 8.2: Mark SubAgent::depth() and SubAgentManager::get() as
  #[cfg(test)] since they are only used in tests. Remove SubAgent::id()
  entirely as it was unused even in tests.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 12:37:08 -04:00
Bryan Helmkamp
e57257bd55 Simplify coding-agent-loop: BaseProfile, &str returns, EnvContext renames, typed errors
- 2.2: Introduce BaseProfile struct with shared id/model/registry fields;
  AnthropicProfile, GeminiProfile, OpenAiProfile now delegate to it
- 5.4: ProviderProfile::id() and ::model() return &str instead of owned String,
  eliminating unnecessary heap allocations on every call
- 6.1: Rename EnvContext fields: date -> current_date, model_name -> model
  for consistency with trait method names
- 8.1: Mark build_env_context_block (no-context variant) as #[cfg(test)]
  since it is only used in one test
- 9.1: Convert SubAgentManager methods (spawn, send_input, wait, close) from
  Result<T, String> to Result<T, AgentError>; tool executors convert at boundary
  via .map_err(|e| e.to_string())

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 12:36:50 -04:00
Bryan Helmkamp
23cb9d8e58 Simplify coding-agent-loop: remove redundant constructors, use let-else, add doc comments
- Remove History::new() (redundant with #[derive(Default)]), replace all callers with History::default()
- Replace match with let-else in extract_signatures_from_assistant for clearer happy path
- Add doc comments to Turn::System and Turn::Steering explaining their LLM role mapping
- Verified: Arc import in openai.rs is used in production code, mutex .expect() messages are consistent

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 12:33:53 -04:00
Bryan Helmkamp
69473dbc20 Simplify session.rs: followup loop, schema check, and cached system prompt
- Replace loop/match/break with clearer let-else pattern for followup
  queue processing in process_input (proposal 4.1)
- Simplify validate_tool_args empty schema check by separating null
  and empty-object checks into distinct if-blocks (proposal 4.3)
- Cache system prompt once per input cycle in run_single_input and pass
  it to build_request and check_context_usage/estimate_token_count,
  avoiding redundant string construction on every tool round (5.2, 5.3)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 12:32:17 -04:00
Bryan Helmkamp
c15b532bcf Implement three missing spec features: SubPipelineHandler, GraphMergeTransform, HTTP Server
Add SubPipelineHandler (handler/sub_pipeline.rs) that inline-executes a parsed
sub-graph within the same engine, reading DOT source from node attributes and
propagating context diffs back to the parent pipeline.

Add GraphMergeTransform (transform.rs) that merges nodes and edges from secondary
graphs into a primary graph with namespace-prefixed IDs to avoid collisions.

Add WebInterviewer (interviewer/web.rs) backed by oneshot channels for async
question/answer flow, and HTTP server (server.rs) with 8 axum endpoints behind
a "server" feature flag for pipeline management and human-in-the-loop via web.

Fix tempfile dev-dependency usage in server production code by using std::env::temp_dir.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 12:07:44 -04:00