Commit graph

24 commits

Author SHA1 Message Date
Bryan Helmkamp
b6d285a79c Remove dead LLM crate code and add --schema to ullm prompt
Remove unused items that have no callers outside their own tests:
middleware wrap_stream_with_middleware/process_stream_event,
common ApiMessage/send_and_read_body, types ProviderEvent variant,
lib CancellationToken re-export, tools execute_all_tools, and
catalog get_latest_model.

Add --schema/-S flag to `ullm prompt` so generate_object() and
stream_object() are exercisable end-to-end. Includes unit test for
invalid JSON rejection and two ignored integration tests.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 18:31:11 -05:00
Bryan Helmkamp
52842e21a7 Add Claude Haiku 4.5 to model catalog
Makes Haiku discoverable via list_models and get_model_info with
aliases "haiku" and "claude-haiku".

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Entire-Checkpoint: ed7df801c2a4
2026-02-25 14:21:58 -05:00
Bryan Helmkamp
778bf45a0b Add cache tokens, reasoning tokens, and file change tracking to pipeline logs
Embed the full Usage struct (with cache_read_tokens, cache_write_tokens,
reasoning_tokens) in AssistantMessage events instead of bare input/output
token fields. Add skip_serializing_if annotations to keep NDJSON clean.
Extend StageUsage with cache/reasoning aggregation. Track files touched
via write_file/edit_file tool call correlation in the backend bridge.
Update format functions, cost accumulator, and TypeScript types.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Entire-Checkpoint: 2c15cd247556
2026-02-25 13:01:45 -05:00
Bryan Helmkamp
b74d837b95 Add ullm models sync command to download OpenRouter model metadata
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Entire-Checkpoint: 6908e4dfbe5e
2026-02-24 14:18:43 -05:00
Bryan Helmkamp
fa828dd8b1 Increase connect timeout to 30s, disable request/stream_read timeouts
LLM API calls should not have HTTP-level request or stream-read timeouts.
These are better controlled at the application level via TimeoutConfig
(total/per_step). The connect timeout is increased from 10s to 30s since
network conditions vary.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-24 09:46:50 -05:00
Bryan Helmkamp
9665fad133 Revert clippy config to defaults, remove all pedantic/nursery/cargo lint suppressions
Removed the workspace-level clippy lint config that enabled all, pedantic, nursery,
and cargo lint groups. Removed all #[allow(clippy::...)] annotations that were only
needed to suppress those extra lints, and fixed the few default clippy warnings that
were uncovered.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-24 09:42:28 -05:00
Bryan Helmkamp
46d4659633 Fix OpenAI error field, add Anthropic provider options pass-through, and improve parity tests
- Use `status: "incomplete"` instead of `is_error` for OpenAI tool results (fixes rejection)
- Add merge_provider_options to forward unknown anthropic provider options to API body
- Derive Clone on Client to enable subagent session factory
- Enable error_recovery scenario for all providers now that OpenAI is fixed
- Improve subagent_spawn test to actually exercise spawn/wait/read workflow
- Adjust multi-turn cache test temperature to 0.5

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 21:21:45 -05:00
Bryan Helmkamp
b4068b6364 Add missing integration tests for multi-turn caching, cross-provider parity, and attractor E2E
Test 1 (llm crate): Multi-turn cache verification runs 6 conversation turns
with a large system prompt (~5460 tokens) and verifies cache_read_tokens on
the final turn. Anthropic threshold 0.5, OpenAI/Gemini 0.0 (automatic
caching not guaranteed).

Test 2 (agent crate): Cross-provider parity matrix with 15 scenarios
(file CRUD, shell, grep/glob, editing, steering, reasoning effort, loop
detection, error recovery, etc.) across Anthropic, OpenAI, and Gemini.
41 total tests. Some scenarios excluded for OpenAI due to gpt-4o-mini
limitations (no reasoning.effort, is_error rejection, weak editing).

Test 3 (attractor crate): E2E pipeline with real LLM using AgentBackend,
AutoApproveInterviewer, and default_registry. Verifies pipeline success,
artifact files, goal gate outcomes, and checkpoint state.

All tests are #[ignore] and require API keys to run.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 18:24:34 -05:00
Bryan Helmkamp
3dfa0f9044 Add shape selectors, LLM stream timeouts, MultiSelect questions, and test coverage
- Stylesheet: add bare-word Shape selector (specificity between Universal and Class)
- LLM: apply per_step timeout to connection and total timeout to stream (Section 4.7)
- Interviewer: add MultiSelect question type alongside MultipleChoice
- Session: move SessionStart/SessionEnd to initialize()/close(), deduplicate close logic
- Engine: return Ok(fail outcome) instead of error when goal gate unsatisfied with no retry_target
- Docker: mark Docker-dependent tests with #[ignore]
- Validation: add extensive unit test coverage for all rule types
- Integration: update tests to match engine/fidelity changes

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 17:43:05 -05:00
Bryan Helmkamp
a3a7f9fb5c Add --no-dotenv flag to ullm binary and use it in tests instead of env_clear()
Prevents tests from loading .env and making real API calls without
manipulating environment state.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 17:24:03 -05:00
Bryan Helmkamp
a3605ec0d4 Speed up slow tests: drop SSE broadcast on completion, switch to rustls-tls, replace hardcoded sleeps with poll loops, and prevent real API calls in ullm tests
- Drop event_tx from ManagedPipeline when pipeline completes/cancels/fails so
  SSE streams end promptly instead of blocking until timeout (3.5s → 0.02s)
- Switch reqwest from native-tls to rustls-tls to avoid 500ms macOS cert store
  load per process (0.67s → 0.005s per OpenAI adapter test)
- Replace hardcoded sleep(500ms)/sleep(200ms)/sleep(100ms) in server and
  integration tests with 10ms poll loops (0.2-0.5s → 0.02-0.03s each)
- Add env_clear() to ullm prompt tests to prevent .env from triggering real
  Anthropic API calls (0.45s → 0.15s)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 17:19:09 -05:00
Bryan Helmkamp
ceaa9bf560 Improve session abort handling, emit tool output deltas, and tighten validation
- Break out of streaming loop on abort and drop the stream before emitting
  SessionEnd to properly cancel the HTTP connection
- Emit ToolCallOutputDelta events for tool call results in both sequential
  and parallel execution paths
- Retry on StageStatus::Fail in addition to Retry in pipeline engine
- Set preferred_label on WaitHumanHandler choice outcomes
- Enforce exactly one terminal node in pipeline validation
- Downgrade unreachable node diagnostic from Error to Warning
- Update context window test to match 1M token limit

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 15:39:04 -05:00
Bryan Helmkamp
6071bb0aa7 Handle signature_delta streaming event for Anthropic thinking blocks
The Anthropic streaming API sends thinking block signatures via a
signature_delta event, not in content_block_start or content_block_stop.
The parser was ignoring this event type, falling back to the empty
placeholder signature from content_block_start, causing "Invalid
signature in thinking block" errors on multi-turn conversations.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 14:39:42 -05:00
Bryan Helmkamp
242e82fffd Use Claude Opus 4.6 with 1M context window as default Anthropic model
Switch default model from claude-sonnet-4-5 to claude-opus-4-6 across
run and serve commands. Enable the 1M token context window via the
context-1m-2025-08-07 beta header for opus-4-6 requests.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 14:23:45 -05:00
Bryan Helmkamp
4f31a84ee6 Skip OpenAI function calls with empty names to avoid model-internal items
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 14:07:23 -05:00
Bryan Helmkamp
f4f58078de Update default Gemini model to gemini-3.1-pro-preview and print model name on startup
Updates the default Gemini model from gemini-3-pro-preview to
gemini-3.1-pro-preview across agent, attractor, and ullm CLIs.
Adds "Using model:" output to stderr in both agent and ullm CLIs
so users can see which model is being used.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 14:05:16 -05:00
Bryan Helmkamp
b4f65e2555 Increase stream_read timeout from 30s to 120s
Prevents premature timeouts on slower streaming responses.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 13:46:32 -05:00
Bryan Helmkamp
5573e0e8a2 Fix truncated LLM responses by populating catalog max_output
Requests were always sent with max_tokens: None, which defaulted to
4096 in the Anthropic provider, causing large outputs to be truncated.
Now build_request() looks up the model's max_output from the catalog
and passes it through, giving each model its full output capacity.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 12:30:25 -05:00
Bryan Helmkamp
f6bf30de3e Disable empty doc-tests and add terse test output alias
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:57:26 -05:00
Bryan Helmkamp
17a42d19f2 Preserve OpenAI reasoning items for Responses API round-trip
The Responses API requires that reasoning items (type: "reasoning",
id: "rs_xxx") are included alongside their associated function_call
items when replaying conversation history. Without them:
"function_call was provided without its required 'reasoning' item"

Store reasoning output items as ContentPart::Other and replay them
as raw input items in translate_input. Handles both non-streaming
(parse_output) and streaming (handle_output_item_done) paths.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:30:39 -05:00
Bryan Helmkamp
6e15d602d0 Fix OpenAI Responses API function call ID prefix mismatch
The Responses API returns two distinct IDs on function_call items:
- `id` (item-level, starts with `fc_`)
- `call_id` (call-level, starts with `call_`)

When sending function calls back as input, the `id` field must start
with `fc_`. Previously we used `call_id` for both fields, causing:
"Invalid 'input[1].id': Expected an ID that begins with 'fc'."

Now we preserve the item-level `id` in provider_metadata and use it
for the `id` field, while `call_id` continues to link tool results.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:27:24 -05:00
Bryan Helmkamp
e992b57ced Add regression tests for Anthropic beta headers and Gemini thought_signature
Covers two recent runtime failures that lacked test coverage:
- Anthropic: assert deprecated beta header values are not sent
- Gemini: test function call parsing/translation with and without thoughtSignature

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:22:14 -05:00
Bryan Helmkamp
3b1411a3ac Preserve Gemini thought_signature on function calls
Gemini 3 models include a thoughtSignature field on function call
parts that must be returned in subsequent conversation turns. Add
provider_metadata to ToolCall to carry this through, and extract/emit
it in both the non-streaming and streaming Gemini code paths.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:14:19 -05:00
Bryan Helmkamp
4bb2de5a49 Rename unified-llm crate to llm
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:22:14 -05:00