From cc71f748e8bc56da3230001fc8083872d8354406 Mon Sep 17 00:00:00 2001 From: Fabro Date: Sat, 23 May 2026 16:42:41 -0400 Subject: [PATCH] =?UTF-8?q?checkpoint=20=E2=9A=92=EF=B8=8F=20Generated=20w?= =?UTF-8?q?ith=20[Fabro](https://fabro.sh)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- run.json | 343 ++- stages/005-implement@1/diff.patch | 2432 +++++++++++++++++ stages/005-implement@1/status.json | 6 + stages/006-simplify_opus@1/prompt.md | 306 +++ stages/006-simplify_opus@1/provider_used.json | 5 + stages/006-simplify_opus@1/response.md | 19 + 6 files changed, 3090 insertions(+), 21 deletions(-) create mode 100644 stages/005-implement@1/diff.patch create mode 100644 stages/005-implement@1/status.json create mode 100644 stages/006-simplify_opus@1/prompt.md create mode 100644 stages/006-simplify_opus@1/provider_used.json create mode 100644 stages/006-simplify_opus@1/response.md diff --git a/run.json b/run.json index a6e078a7b..0f844e0cc 100644 --- a/run.json +++ b/run.json @@ -492,7 +492,7 @@ "kind": "running" }, "status_updated_at": "2026-05-23T19:56:01.493793Z", - "last_event_at": "2026-05-23T20:26:00.219582Z", + "last_event_at": "2026-05-23T20:42:40.964151Z", "pending_control": null, "checkpoints": [ { @@ -740,9 +740,9 @@ } }, { - "seq": 0, + "seq": 736, "checkpoint": { - "timestamp": "2026-05-23T20:26:00.320170Z", + "timestamp": "2026-05-23T20:26:04.061726Z", "current_node": "implement", "completed_nodes": [ "start", @@ -753,31 +753,157 @@ ], "node_retries": {}, "context_values": { - "internal.thread_id": "preflight_lint", - "failure_class": "", - "graph.goal": "# Output Schema Validation Implementation Plan\n\n> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.\n\n**Goal:** Add `output_schema` validation for agent and prompt nodes, with context-preserving repair turns when structured output does not validate.\n\n**Architecture:** Introduce a small structured-output layer in `fabro-workflow` that resolves node-level schema declarations, extracts JSON output, validates it, and produces either routing side effects or a parsed custom output context update. Agent and prompt execution must perform schema repair inside the active LLM conversation instead of using the workflow executor retry path.\n\n**Tech Stack:** Rust, Graphviz workflow attrs, `serde_json`, workspace `jsonschema`, existing `fabro-llm::ResponseFormat`, Fabro agent sessions, `cargo nextest`.\n\n---\n\n## Public Interface\n\nWorkflow authors can opt in on agent and prompt nodes:\n\n```dot\nreview [\n shape=tab,\n output_schema=\"routing\",\n output_retries=2\n]\n\naudit [\n shape=tab,\n output_schema=\"@schemas/audit-result.schema.json\",\n output_retries=2\n]\n```\n\n- `output_schema=\"routing\"` uses Fabro's built-in routing directive schema.\n- `output_schema=\"@path/to/schema.json\"` loads a JSON Schema file through existing workflow file-reference rules.\n- `output_retries` controls corrective turns inside the same node execution. Default: `2`. `0` means validate once and fail without a repair turn.\n- Schema failures are terminal node failures after `output_retries` is exhausted. They are not `retry_requested` outcomes and do not consume `max_retries`.\n- `backend=\"acp\"` with `output_schema` is unsupported in v1 and returns a clear validation error.\n\n## Implementation Tasks\n\n### Task 1: Node Attributes And File Reference Resolution\n\n**Files:**\n- Modify: `lib/crates/fabro-types/src/graph.rs`\n- Modify: `lib/crates/fabro-workflow/src/static_reference.rs`\n- Modify: `lib/crates/fabro-workflow/src/transforms/file_inlining.rs`\n- Test: existing unit tests in those files\n\n- [ ] Add `Node::output_schema(&self) -> Option<&str>` next to other agent/prompt attrs.\n- [ ] Add `Node::output_retries(&self) -> i64` returning `self.int_attr(\"output_retries\").unwrap_or(2).max(0)`.\n- [ ] Teach static reference validation that node attr `output_schema` values starting with `@` are file inline references.\n- [ ] Extend file inlining so `output_schema=\"@schemas/foo.json\"` is replaced with the schema file contents before execution, while `output_schema=\"routing\"` stays unchanged.\n- [ ] Add tests for absent attrs, default retries, zero retries, file inlining, and unresolved schema reference diagnostics.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-types -p fabro-workflow graph:: file_inlining static_reference\n```\n\nExpected: targeted tests pass.\n\n### Task 2: Structured Output Module\n\n**Files:**\n- Create: `lib/crates/fabro-workflow/src/handler/structured_output.rs`\n- Modify: `lib/crates/fabro-workflow/src/handler/mod.rs`\n- Modify: `lib/crates/fabro-workflow/Cargo.toml`\n- Test: unit tests in `structured_output.rs`\n\n- [ ] Add `jsonschema.workspace = true` to `fabro-workflow` dependencies.\n- [ ] Define `OutputSchemaKind` with `Routing` and `JsonSchema { schema: serde_json::Value }`.\n- [ ] Parse `node.output_schema()` into `None`, `Routing`, or custom JSON Schema. Treat literal `routing` as the only built-in keyword.\n- [ ] Add a built-in routing schema requiring an object with at least one recognized field: `preferred_next_label`, `outcome`, `failure_reason`, `suggested_next_ids`, or `context_updates`.\n- [ ] Reuse balanced-object scanning semantics for response text: validate the last JSON object that is relevant to the selected schema.\n- [ ] Return a structured validation result containing the parsed JSON object, concise error messages, and enough information to build a repair prompt.\n- [ ] Add tests for valid routing JSON, missing routing fields, wrong routing field types, valid custom schema, invalid custom schema, invalid JSON, and no JSON object.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow structured_output\n```\n\nExpected: structured-output unit tests pass.\n\n### Task 3: Routing Extraction Compatibility\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/agent.rs`\n- Test: existing agent handler unit tests\n\n- [ ] Keep the loose default unchanged when `output_schema` is absent.\n- [ ] Move current `STATUS_FIELDS`, balanced JSON scanning, and routing-field application behind reusable functions in `structured_output.rs` or call the new module from `agent.rs`.\n- [ ] For `output_schema=\"routing\"`, require schema-valid routing JSON and surface validation failures for repair instead of silently ignoring bad candidates.\n- [ ] Preserve existing routing fallback priority for agent nodes: response text first, then `status.json`, then last file touched.\n- [ ] Keep prompt-node routing behavior response-only unless later tasks explicitly add prompt `status.json` support.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::agent\n```\n\nExpected: existing loose routing tests still pass, plus new strict routing tests pass.\n\n### Task 4: Prompt Node Same-Context Repair\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/llm/api.rs`\n- Modify: `lib/crates/fabro-workflow/src/handler/prompt.rs`\n- Test: prompt/API backend tests in those files\n\n- [ ] In `AgentApiBackend::one_shot`, keep `messages` mutable across attempts.\n- [ ] When a prompt node has a custom JSON Schema, set `response_format=JsonSchema` on the initial and repair LLM requests. For `routing`, use `JsonObject` or no provider-native schema if provider behavior would conflict with Fabro's routing extraction.\n- [ ] After each LLM response, validate according to `output_schema`.\n- [ ] On validation failure with repair attempts remaining, append `Message::assistant(response.text())`, then append a corrective `Message::user(repair_message)`, and call `client.complete` again with the same messages.\n- [ ] On success, return the validated response text and aggregate usage across all attempts.\n- [ ] On exhaustion, return a terminal failed outcome with failure reason `output schema validation failed after N repair attempt(s)`.\n- [ ] Update `PromptHandler` so validated custom output is added to `context_updates[\"output.{node_id}\"]`; routing output still updates outcome routing fields.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::prompt handler::llm::api\n```\n\nExpected: prompt repair keeps previous assistant output in the message list and succeeds after a corrective response.\n\n### Task 5: Agent Node Same-Session Repair\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/llm/api.rs`\n- Modify: `lib/crates/fabro-workflow/src/handler/agent.rs`\n- Test: agent/API backend tests in those files\n\n- [ ] In `AgentApiBackend::run`, validate the final assistant response before releasing, closing, or caching the session.\n- [ ] On validation failure with repair attempts remaining, call `session.process_input(repair_message)` on the same `Session`.\n- [ ] Recompute the final assistant response after each repair turn from `session.history()`.\n- [ ] Aggregate usage across all new assistant turns, including repair turns, without double-counting reused session history.\n- [ ] Do not set provider-native `response_format` for agent sessions in v1, because agent sessions may need normal tool-use messages before final output.\n- [ ] Return terminal failure after exhaustion; do not return a retryable backend error and do not request workflow node retry.\n- [ ] Update `AgentHandler` to apply validated routing/custom output to the final `Outcome`.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::agent handler::llm::api\n```\n\nExpected: agent repair sends a second `process_input` to the same session and final validated output drives outcome/context updates.\n\n### Task 6: ACP Guardrail\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/llm/acp.rs`\n- Test: ACP backend tests in that file\n\n- [ ] At the start of `AgentAcpBackend::run`, reject nodes where `node.output_schema().is_some()`.\n- [ ] Use a clear error message: `output_schema is not supported with backend=\"acp\" in this release`.\n- [ ] Add a test proving the ACP backend does not launch a process when `output_schema` is present.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::llm::acp\n```\n\nExpected: ACP guardrail test passes.\n\n### Task 7: Docs\n\n**Files:**\n- Modify: `docs/public/agents/outputs.mdx`\n- Modify: `docs/public/reference/dot-language.mdx`\n\n- [ ] Document `output_schema=\"routing\"` and `output_schema=\"@schema.json\"` under routing/structured outputs.\n- [ ] Document same-context repair behavior explicitly: Fabro sends validation feedback to the same agent/prompt context before failing.\n- [ ] Document `output_retries`, default `2`, and distinction from `max_retries`.\n- [ ] Document v1 scope: agent/prompt nodes only; ACP unsupported; custom schema output stored at `output.{node_id}`.\n\nRun:\n\n```bash\nrg -n \"output_schema|output_retries|output\\\\.\" docs/public/agents/outputs.mdx docs/public/reference/dot-language.mdx\n```\n\nExpected: docs mention the new attrs and storage behavior.\n\n### Task 8: Full Verification\n\n**Files:**\n- No new files beyond prior tasks\n\n- [ ] Run focused workflow tests:\n\n```bash\ncargo nextest run -p fabro-workflow\n```\n\n- [ ] Run formatting check:\n\n```bash\ncargo +nightly-2026-04-14 fmt --check --all\n```\n\n- [ ] Run clippy:\n\n```bash\ncargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings\n```\n\n- [ ] If snapshots change, inspect before accepting:\n\n```bash\ncargo insta pending-snapshots\n```\n\nOnly run `cargo insta accept` after verifying every pending snapshot is expected.\n\n## Acceptance Criteria\n\n- Existing workflows without `output_schema` behave exactly as before.\n- `output_schema=\"routing\"` prevents malformed/missing routing JSON from silently falling through to normal edge selection.\n- Invalid structured output results in a corrective LLM turn in the same context window.\n- Prompt repair preserves previous assistant output in the message list.\n- Agent repair preserves the same live session and does not re-run the node from scratch.\n- Exhausted output repair attempts produce a clear terminal failure.\n- Custom schema output is available to downstream nodes at `output.{node_id}`.\n- Docs clearly distinguish `output_retries` from `max_retries`.\n\n## Assumptions\n\n- `output_retries=2` is the default.\n- Custom schema validation targets the final JSON object in the response text.\n- `status.json` fallback remains routing-specific.\n- Provider-native response schema is used for prompt nodes only where it is safe.\n- ACP support can be added later after there is a guaranteed context-preserving repair mechanism.\n", - "failure_signature": "", - "internal.retry_count.toolchain": 0, - "internal.retry_count.implement": 0, - "internal.run_id": "01KSB6GTZ00T5V6BNMXN3SPKZF", "internal.retry_count.start": 0, - "graph.rankdir": "LR", - "current_node": "implement", - "thread.start.current_node": "toolchain", - "thread.toolchain.current_node": "preflight_compile", - "internal.node_visit_count": 1, - "command.output": "blob://sha256/12ae32cb1ec02d01eda3581b127c1fee3b0dc53572ed6baf239721a03d82e126", "internal.retry_count.preflight_compile": 0, "thread.preflight_compile.current_node": "preflight_lint", - "internal.retry_count.preflight_lint": 0, + "command.output": "blob://sha256/12ae32cb1ec02d01eda3581b127c1fee3b0dc53572ed6baf239721a03d82e126", + "thread.start.current_node": "toolchain", + "failure_class": "", + "current_node": "implement", + "failure_signature": "", + "graph.model_stylesheet": "\n * { model: claude-opus-4-7; }\n ", + "internal.node_visit_count": 1, + "internal.retry_count.implement": 0, + "internal.retry_count.toolchain": 0, + "last_response": "Implemented output schema validation end-to-end.\n\nSummary:\n- Added `output_schema` / `output_retries` node attrs.\n- Added `output_schema` static/file reference handling and inlining.\n- Added `handler:", "thread.preflight_lint.current_node": "implement", + "internal.thread_id": "preflight_lint", + "response.implement": "Implemented output schema validation end-to-end.\n\nSummary:\n- Added `output_schema` / `output_retries` node attrs.\n- Added `output_schema` static/file reference handling and inlining.\n- Added `handler::structured_output` for:\n - routing/custom schema parsing\n - balanced JSON extraction\n - JSON Schema validation\n - repair prompts\n - routing/custom context application\n- Integrated strict routing/custom output into agent and prompt handlers.\n- Added same-context repair for:\n - prompt API calls via mutable message history\n - agent API sessions via the same live `Session`\n- Added ACP guardrail:\n - `output_schema is not supported with backend=\"acp\" in this release`\n - test verifies no process launch\n- Added custom output storage at `output.{node_id}`.\n- Updated docs for routing/custom schemas, `output_retries`, repair behavior, ACP limitation.\n\nValidation run:\n- `cargo nextest run -p fabro-types -p fabro-workflow graph:: file_inlining static_reference` ✅\n- `cargo nextest run -p fabro-workflow` ✅\n- `cargo +nightly-2026-04-14 fmt --check --all` ✅\n- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` ✅\n- `cargo insta pending-snapshots` unavailable in this environment (`cargo-insta` not installed); checked for `**/*.snap.new` and found none.", + "internal.work_dir": "/home/daytona/workspace/fabro", + "internal.retry_count.preflight_lint": 0, + "graph.rankdir": "LR", + "internal.run_id": "01KSB6GTZ00T5V6BNMXN3SPKZF", + "outcome": "succeeded", + "internal.fidelity": "compact", + "thread.toolchain.current_node": "preflight_compile", + "last_stage": "implement", + "graph.goal": "# Output Schema Validation Implementation Plan\n\n> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.\n\n**Goal:** Add `output_schema` validation for agent and prompt nodes, with context-preserving repair turns when structured output does not validate.\n\n**Architecture:** Introduce a small structured-output layer in `fabro-workflow` that resolves node-level schema declarations, extracts JSON output, validates it, and produces either routing side effects or a parsed custom output context update. Agent and prompt execution must perform schema repair inside the active LLM conversation instead of using the workflow executor retry path.\n\n**Tech Stack:** Rust, Graphviz workflow attrs, `serde_json`, workspace `jsonschema`, existing `fabro-llm::ResponseFormat`, Fabro agent sessions, `cargo nextest`.\n\n---\n\n## Public Interface\n\nWorkflow authors can opt in on agent and prompt nodes:\n\n```dot\nreview [\n shape=tab,\n output_schema=\"routing\",\n output_retries=2\n]\n\naudit [\n shape=tab,\n output_schema=\"@schemas/audit-result.schema.json\",\n output_retries=2\n]\n```\n\n- `output_schema=\"routing\"` uses Fabro's built-in routing directive schema.\n- `output_schema=\"@path/to/schema.json\"` loads a JSON Schema file through existing workflow file-reference rules.\n- `output_retries` controls corrective turns inside the same node execution. Default: `2`. `0` means validate once and fail without a repair turn.\n- Schema failures are terminal node failures after `output_retries` is exhausted. They are not `retry_requested` outcomes and do not consume `max_retries`.\n- `backend=\"acp\"` with `output_schema` is unsupported in v1 and returns a clear validation error.\n\n## Implementation Tasks\n\n### Task 1: Node Attributes And File Reference Resolution\n\n**Files:**\n- Modify: `lib/crates/fabro-types/src/graph.rs`\n- Modify: `lib/crates/fabro-workflow/src/static_reference.rs`\n- Modify: `lib/crates/fabro-workflow/src/transforms/file_inlining.rs`\n- Test: existing unit tests in those files\n\n- [ ] Add `Node::output_schema(&self) -> Option<&str>` next to other agent/prompt attrs.\n- [ ] Add `Node::output_retries(&self) -> i64` returning `self.int_attr(\"output_retries\").unwrap_or(2).max(0)`.\n- [ ] Teach static reference validation that node attr `output_schema` values starting with `@` are file inline references.\n- [ ] Extend file inlining so `output_schema=\"@schemas/foo.json\"` is replaced with the schema file contents before execution, while `output_schema=\"routing\"` stays unchanged.\n- [ ] Add tests for absent attrs, default retries, zero retries, file inlining, and unresolved schema reference diagnostics.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-types -p fabro-workflow graph:: file_inlining static_reference\n```\n\nExpected: targeted tests pass.\n\n### Task 2: Structured Output Module\n\n**Files:**\n- Create: `lib/crates/fabro-workflow/src/handler/structured_output.rs`\n- Modify: `lib/crates/fabro-workflow/src/handler/mod.rs`\n- Modify: `lib/crates/fabro-workflow/Cargo.toml`\n- Test: unit tests in `structured_output.rs`\n\n- [ ] Add `jsonschema.workspace = true` to `fabro-workflow` dependencies.\n- [ ] Define `OutputSchemaKind` with `Routing` and `JsonSchema { schema: serde_json::Value }`.\n- [ ] Parse `node.output_schema()` into `None`, `Routing`, or custom JSON Schema. Treat literal `routing` as the only built-in keyword.\n- [ ] Add a built-in routing schema requiring an object with at least one recognized field: `preferred_next_label`, `outcome`, `failure_reason`, `suggested_next_ids`, or `context_updates`.\n- [ ] Reuse balanced-object scanning semantics for response text: validate the last JSON object that is relevant to the selected schema.\n- [ ] Return a structured validation result containing the parsed JSON object, concise error messages, and enough information to build a repair prompt.\n- [ ] Add tests for valid routing JSON, missing routing fields, wrong routing field types, valid custom schema, invalid custom schema, invalid JSON, and no JSON object.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow structured_output\n```\n\nExpected: structured-output unit tests pass.\n\n### Task 3: Routing Extraction Compatibility\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/agent.rs`\n- Test: existing agent handler unit tests\n\n- [ ] Keep the loose default unchanged when `output_schema` is absent.\n- [ ] Move current `STATUS_FIELDS`, balanced JSON scanning, and routing-field application behind reusable functions in `structured_output.rs` or call the new module from `agent.rs`.\n- [ ] For `output_schema=\"routing\"`, require schema-valid routing JSON and surface validation failures for repair instead of silently ignoring bad candidates.\n- [ ] Preserve existing routing fallback priority for agent nodes: response text first, then `status.json`, then last file touched.\n- [ ] Keep prompt-node routing behavior response-only unless later tasks explicitly add prompt `status.json` support.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::agent\n```\n\nExpected: existing loose routing tests still pass, plus new strict routing tests pass.\n\n### Task 4: Prompt Node Same-Context Repair\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/llm/api.rs`\n- Modify: `lib/crates/fabro-workflow/src/handler/prompt.rs`\n- Test: prompt/API backend tests in those files\n\n- [ ] In `AgentApiBackend::one_shot`, keep `messages` mutable across attempts.\n- [ ] When a prompt node has a custom JSON Schema, set `response_format=JsonSchema` on the initial and repair LLM requests. For `routing`, use `JsonObject` or no provider-native schema if provider behavior would conflict with Fabro's routing extraction.\n- [ ] After each LLM response, validate according to `output_schema`.\n- [ ] On validation failure with repair attempts remaining, append `Message::assistant(response.text())`, then append a corrective `Message::user(repair_message)`, and call `client.complete` again with the same messages.\n- [ ] On success, return the validated response text and aggregate usage across all attempts.\n- [ ] On exhaustion, return a terminal failed outcome with failure reason `output schema validation failed after N repair attempt(s)`.\n- [ ] Update `PromptHandler` so validated custom output is added to `context_updates[\"output.{node_id}\"]`; routing output still updates outcome routing fields.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::prompt handler::llm::api\n```\n\nExpected: prompt repair keeps previous assistant output in the message list and succeeds after a corrective response.\n\n### Task 5: Agent Node Same-Session Repair\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/llm/api.rs`\n- Modify: `lib/crates/fabro-workflow/src/handler/agent.rs`\n- Test: agent/API backend tests in those files\n\n- [ ] In `AgentApiBackend::run`, validate the final assistant response before releasing, closing, or caching the session.\n- [ ] On validation failure with repair attempts remaining, call `session.process_input(repair_message)` on the same `Session`.\n- [ ] Recompute the final assistant response after each repair turn from `session.history()`.\n- [ ] Aggregate usage across all new assistant turns, including repair turns, without double-counting reused session history.\n- [ ] Do not set provider-native `response_format` for agent sessions in v1, because agent sessions may need normal tool-use messages before final output.\n- [ ] Return terminal failure after exhaustion; do not return a retryable backend error and do not request workflow node retry.\n- [ ] Update `AgentHandler` to apply validated routing/custom output to the final `Outcome`.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::agent handler::llm::api\n```\n\nExpected: agent repair sends a second `process_input` to the same session and final validated output drives outcome/context updates.\n\n### Task 6: ACP Guardrail\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/llm/acp.rs`\n- Test: ACP backend tests in that file\n\n- [ ] At the start of `AgentAcpBackend::run`, reject nodes where `node.output_schema().is_some()`.\n- [ ] Use a clear error message: `output_schema is not supported with backend=\"acp\" in this release`.\n- [ ] Add a test proving the ACP backend does not launch a process when `output_schema` is present.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::llm::acp\n```\n\nExpected: ACP guardrail test passes.\n\n### Task 7: Docs\n\n**Files:**\n- Modify: `docs/public/agents/outputs.mdx`\n- Modify: `docs/public/reference/dot-language.mdx`\n\n- [ ] Document `output_schema=\"routing\"` and `output_schema=\"@schema.json\"` under routing/structured outputs.\n- [ ] Document same-context repair behavior explicitly: Fabro sends validation feedback to the same agent/prompt context before failing.\n- [ ] Document `output_retries`, default `2`, and distinction from `max_retries`.\n- [ ] Document v1 scope: agent/prompt nodes only; ACP unsupported; custom schema output stored at `output.{node_id}`.\n\nRun:\n\n```bash\nrg -n \"output_schema|output_retries|output\\\\.\" docs/public/agents/outputs.mdx docs/public/reference/dot-language.mdx\n```\n\nExpected: docs mention the new attrs and storage behavior.\n\n### Task 8: Full Verification\n\n**Files:**\n- No new files beyond prior tasks\n\n- [ ] Run focused workflow tests:\n\n```bash\ncargo nextest run -p fabro-workflow\n```\n\n- [ ] Run formatting check:\n\n```bash\ncargo +nightly-2026-04-14 fmt --check --all\n```\n\n- [ ] Run clippy:\n\n```bash\ncargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings\n```\n\n- [ ] If snapshots change, inspect before accepting:\n\n```bash\ncargo insta pending-snapshots\n```\n\nOnly run `cargo insta accept` after verifying every pending snapshot is expected.\n\n## Acceptance Criteria\n\n- Existing workflows without `output_schema` behave exactly as before.\n- `output_schema=\"routing\"` prevents malformed/missing routing JSON from silently falling through to normal edge selection.\n- Invalid structured output results in a corrective LLM turn in the same context window.\n- Prompt repair preserves previous assistant output in the message list.\n- Agent repair preserves the same live session and does not re-run the node from scratch.\n- Exhausted output repair attempts produce a clear terminal failure.\n- Custom schema output is available to downstream nodes at `output.{node_id}`.\n- Docs clearly distinguish `output_retries` from `max_retries`.\n\n## Assumptions\n\n- `output_retries=2` is the default.\n- Custom schema validation targets the final JSON object in the response text.\n- `status.json` fallback remains routing-specific.\n- Provider-native response schema is used for prompt nodes only where it is safe.\n- ACP support can be added later after there is a guaranteed context-preserving repair mechanism.\n" + }, + "node_outcomes": { + "toolchain": { + "status": "succeeded", + "context_updates": { + "command.output": "blob://sha256/fc14b2ba2d770e5cd3169df7a29525c962adfc4cfa3097b9098c63ebd61a748c" + }, + "notes": "Script completed: command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", + "usage": null + }, + "implement": { + "status": "succeeded", + "context_updates": { + "last_response": "Implemented output schema validation end-to-end.\n\nSummary:\n- Added `output_schema` / `output_retries` node attrs.\n- Added `output_schema` static/file reference handling and inlining.\n- Added `handler:", + "response.implement": "Implemented output schema validation end-to-end.\n\nSummary:\n- Added `output_schema` / `output_retries` node attrs.\n- Added `output_schema` static/file reference handling and inlining.\n- Added `handler::structured_output` for:\n - routing/custom schema parsing\n - balanced JSON extraction\n - JSON Schema validation\n - repair prompts\n - routing/custom context application\n- Integrated strict routing/custom output into agent and prompt handlers.\n- Added same-context repair for:\n - prompt API calls via mutable message history\n - agent API sessions via the same live `Session`\n- Added ACP guardrail:\n - `output_schema is not supported with backend=\"acp\" in this release`\n - test verifies no process launch\n- Added custom output storage at `output.{node_id}`.\n- Updated docs for routing/custom schemas, `output_retries`, repair behavior, ACP limitation.\n\nValidation run:\n- `cargo nextest run -p fabro-types -p fabro-workflow graph:: file_inlining static_reference` ✅\n- `cargo nextest run -p fabro-workflow` ✅\n- `cargo +nightly-2026-04-14 fmt --check --all` ✅\n- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` ✅\n- `cargo insta pending-snapshots` unavailable in this environment (`cargo-insta` not installed); checked for `**/*.snap.new` and found none.", + "last_stage": "implement" + }, + "notes": "Stage completed: implement", + "usage": { + "input": { + "usage": { + "model": { + "provider": "openai", + "model_id": "gpt-5.5" + }, + "tokens": { + "input_tokens": 329828, + "output_tokens": 33087, + "reasoning_tokens": 21712, + "cache_read_tokens": 29338112, + "cache_write_tokens": 0 + } + }, + "facts": { + "algorithm": "openai" + } + }, + "total_usd_micros": 17962166 + }, + "files_touched": [ + "/home/daytona/workspace/fabro/lib/crates/fabro-workflow/src/handler/structured_output.rs" + ] + }, + "start": { + "status": "succeeded", + "usage": null + }, + "preflight_compile": { + "status": "succeeded", + "context_updates": { + "command.output": "blob://sha256/12ae32cb1ec02d01eda3581b127c1fee3b0dc53572ed6baf239721a03d82e126" + }, + "notes": "Script completed: cargo check -q --workspace 2>&1", + "usage": null + }, + "preflight_lint": { + "status": "succeeded", + "context_updates": { + "command.output": "blob://sha256/12ae32cb1ec02d01eda3581b127c1fee3b0dc53572ed6baf239721a03d82e126" + }, + "notes": "Script completed: cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", + "usage": null + } + }, + "next_node_id": "simplify_opus", + "git_commit_sha": "43b8310e6051d435c815d935333e893833fb72a9", + "node_visits": { + "preflight_lint": 1, + "preflight_compile": 1, + "implement": 1, + "start": 1, + "toolchain": 1 + } + }, + "diff": { + "patch": "diff --git a/Cargo.lock b/Cargo.lock\nindex c50b6c665..28b5f81e8 100644\n--- a/Cargo.lock\n+++ b/Cargo.lock\n@@ -2603,6 +2603,7 @@ dependencies = [\n \"git2\",\n \"hex\",\n \"httpmock\",\n+ \"jsonschema\",\n \"md5\",\n \"miette\",\n \"mime_guess\",\ndiff --git a/docs/public/agents/outputs.mdx b/docs/public/agents/outputs.mdx\nindex 1407f4db6..7467f050b 100644\n--- a/docs/public/agents/outputs.mdx\n+++ b/docs/public/agents/outputs.mdx\n@@ -64,6 +64,23 @@ The test coverage is below the threshold.\n \n JSON objects without recognized fields are ignored.\n \n+### Validated routing output\n+\n+Set `output_schema=\"routing\"` on an agent or prompt node to require Fabro's built-in routing directive schema:\n+\n+```dot\n+review [\n+ shape=tab,\n+ prompt=\"Review the implementation and return routing JSON.\",\n+ output_schema=\"routing\",\n+ output_retries=2\n+]\n+```\n+\n+With `output_schema=\"routing\"`, the routing JSON must be an object with at least one recognized routing field (`preferred_next_label`, `outcome`, `failure_reason`, `suggested_next_ids`, or `context_updates`) and those fields must have the expected types. Malformed routing JSON fails validation instead of being silently ignored.\n+\n+Fabro repairs invalid structured output inside the same LLM context before failing the node. For prompt nodes, Fabro appends the invalid assistant response and a corrective user message to the same message list. For agent nodes using the API backend, Fabro sends the corrective message to the same live agent session. `output_retries` controls these repair turns and defaults to `2`; `output_retries=0` validates once and fails without a repair turn. These repair turns are separate from workflow `max_retries` and do not consume node retry attempts.\n+\n ### Fallback: status.json file\n \n If no routing directives are found in the response text, Fabro checks whether the agent wrote a `status.json` file into the sandbox working directory. If the file exists, Fabro extracts routing directives from it using the same logic. This is useful for agents that write structured output to files rather than including JSON in their response text.\n@@ -90,6 +107,35 @@ review -> fix [label=\"Fix\"]\n review -> approve [label=\"Approve\"]\n ```\n \n+## Custom structured outputs\n+\n+Agent and prompt nodes can also validate their final JSON object against a JSON Schema file:\n+\n+```dot\n+audit [\n+ shape=tab,\n+ prompt=\"Audit the change and return JSON that matches the schema.\",\n+ output_schema=\"@schemas/audit-result.schema.json\",\n+ output_retries=2\n+]\n+```\n+\n+`output_schema=\"@path/to/schema.json\"` uses the same workflow file-reference rules as prompt files: the schema is loaded relative to the workflow file and inlined before execution. The final JSON object in the LLM response is validated with `jsonschema`.\n+\n+When custom schema validation succeeds, Fabro stores the parsed JSON value in context at:\n+\n+| Key | Value |\n+|---|---|\n+| `output.{node_id}` | The parsed JSON object that matched the custom schema |\n+\n+For example, node `audit` writes its parsed custom output to `output.audit`. Fabro still stores the raw response text at `response.audit`.\n+\n+If custom schema validation fails, Fabro sends concise validation feedback to the same prompt conversation or agent session and asks for corrected JSON. After `output_retries` repair turns are exhausted, the node fails terminally with `output schema validation failed after N repair attempt(s)`.\n+\n+\n+Structured output validation currently applies to agent and prompt nodes. `backend=\"acp\"` does not support `output_schema` in this release. Custom schemas update `output.{node_id}`; routing schemas update routing fields and `context_updates` instead.\n+\n+\n ## Output logging\n \n Fabro writes several files per stage to `stages/{rank:03}-{node_id}@{visit}/` in metadata snapshots and `fabro dump` output:\ndiff --git a/docs/public/reference/dot-language.mdx b/docs/public/reference/dot-language.mdx\nindex 1c1656057..74a366632 100644\n--- a/docs/public/reference/dot-language.mdx\n+++ b/docs/public/reference/dot-language.mdx\n@@ -206,10 +206,37 @@ Start nodes can also be identified by ID (`start` or `Start`). Exit nodes can be\n | `model` | String | Explicit model ID (overrides stylesheet) |\n | `provider` | String | Explicit provider name (overrides stylesheet). Auto-inferred from the model catalog when omitted. |\n | `project_memory` | Boolean | When `true` (default), prompt nodes discover and include project docs (`AGENTS.md`, `CLAUDE.md`, etc.) as a system prompt. Set to `false` to disable. |\n+| `output_schema` | String | Optional structured output validation. Use `routing` for Fabro's built-in routing directive schema, or `@path/to/schema.json` for a JSON Schema file. Supported on agent and prompt nodes. |\n+| `output_retries` | Integer | Corrective structured-output turns inside the same prompt conversation or agent session. Default `2`; `0` validates once and fails without repair. Separate from `max_retries`. |\n | `backend` | String | Agent execution backend: `api` (default) or `acp`. `api` runs Fabro's tool loop through provider APIs; `acp` runs an Agent Client Protocol stdio agent inside the active sandbox. Prompt nodes are API-only. See [Agents — Backends](/core-concepts/agents#backends). |\n | `acp.command` | String | Shell command for nodes with `backend=\"acp\"`. Mutually exclusive with `acp.config`. The value is always parsed as a command string, not JSON. |\n | `acp.config` | String | JSON stdio ACP config for nodes with `backend=\"acp\"`. Mutually exclusive with `acp.command`. |\n \n+#### Structured output validation\n+\n+`output_schema` opts an agent or prompt node into strict JSON validation:\n+\n+```dot\n+review [\n+ shape=tab,\n+ output_schema=\"routing\",\n+ output_retries=2\n+]\n+\n+audit [\n+ shape=tab,\n+ output_schema=\"@schemas/audit-result.schema.json\",\n+ output_retries=2\n+]\n+```\n+\n+- `output_schema=\"routing\"` requires a JSON object with at least one recognized routing field: `preferred_next_label`, `outcome`, `failure_reason`, `suggested_next_ids`, or `context_updates`.\n+- `output_schema=\"@schemas/audit-result.schema.json\"` loads a JSON Schema file using workflow file-reference rules and validates the final JSON object in the response text.\n+- On validation failure, Fabro sends validation feedback to the same active context before failing: prompt nodes keep the prior assistant response in the message list, and API-backed agent nodes repair in the same live session.\n+- `output_retries` defaults to `2` and controls only these corrective structured-output turns. It is not the same as `max_retries` and does not consume workflow retry attempts.\n+- Custom schema output is stored in context at `output.{node_id}`. Routing schema output updates routing fields and any `context_updates`.\n+- `backend=\"acp\"` with `output_schema` is unsupported in this release.\n+\n ### Command nodes\n \n | Attribute | Type | Description |\n@@ -331,15 +358,16 @@ gate -> v2 [condition=\"context.version matches ^v2\\\\.\"]\n gate -> slow_path\n ```\n \n-## Prompt file references\n+## Prompt and schema file references\n \n-Instead of inlining long prompts, reference an external file:\n+Instead of inlining long prompts or JSON Schemas, reference an external file:\n \n ```dot\n simplify [label=\"Simplify\", prompt=\"@prompts/simplify.md\"]\n+audit [shape=tab, output_schema=\"@schemas/audit-result.schema.json\"]\n ```\n \n-The `@` prefix tells Fabro to load the prompt from a file path relative to the workflow file. Paths support `~` (home directory) and `..` (parent directory):\n+The `@` prefix tells Fabro to load the referenced file relative to the workflow file. Paths support `~` (home directory) and `..` (parent directory):\n \n ```dot\n shared [prompt=\"@~/shared-prompts/review.md\"]\ndiff --git a/lib/crates/fabro-types/src/graph.rs b/lib/crates/fabro-types/src/graph.rs\nindex 27344f8c9..7ef9bad99 100644\n--- a/lib/crates/fabro-types/src/graph.rs\n+++ b/lib/crates/fabro-types/src/graph.rs\n@@ -168,6 +168,16 @@ impl Node {\n self.str_attr(\"prompt\")\n }\n \n+ #[must_use]\n+ pub fn output_schema(&self) -> Option<&str> {\n+ self.str_attr(\"output_schema\")\n+ }\n+\n+ #[must_use]\n+ pub fn output_retries(&self) -> i64 {\n+ self.int_attr(\"output_retries\").unwrap_or(2).max(0)\n+ }\n+\n #[must_use]\n pub fn max_retries(&self) -> Option {\n self.int_attr(\"max_retries\")\n@@ -576,6 +586,8 @@ mod tests {\n assert_eq!(node.shape(), \"box\");\n assert_eq!(node.node_type(), None);\n assert_eq!(node.prompt(), None);\n+ assert_eq!(node.output_schema(), None);\n+ assert_eq!(node.output_retries(), 2);\n assert_eq!(node.max_retries(), None);\n assert!(!node.goal_gate());\n assert_eq!(node.retry_target(), None);\n@@ -602,6 +614,31 @@ mod tests {\n assert!(!node.project_memory());\n }\n \n+ #[test]\n+ fn node_output_retries_defaults_and_clamps_to_zero() {\n+ let mut node = Node::new(\"x\");\n+ assert_eq!(node.output_retries(), 2);\n+\n+ node.attrs\n+ .insert(\"output_retries\".to_string(), AttrValue::Integer(0));\n+ assert_eq!(node.output_retries(), 0);\n+\n+ node.attrs\n+ .insert(\"output_retries\".to_string(), AttrValue::Integer(-3));\n+ assert_eq!(node.output_retries(), 0);\n+ }\n+\n+ #[test]\n+ fn node_output_schema_returns_string_attr() {\n+ let mut node = Node::new(\"x\");\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(\"routing\".to_string()),\n+ );\n+\n+ assert_eq!(node.output_schema(), Some(\"routing\"));\n+ }\n+\n #[test]\n fn node_with_attrs() {\n let mut node = Node::new(\"plan\");\ndiff --git a/lib/crates/fabro-workflow/Cargo.toml b/lib/crates/fabro-workflow/Cargo.toml\nindex d5025f1bd..6021bca4c 100644\n--- a/lib/crates/fabro-workflow/Cargo.toml\n+++ b/lib/crates/fabro-workflow/Cargo.toml\n@@ -46,6 +46,7 @@ fabro-http.workspace = true\n thiserror.workspace = true\n serde.workspace = true\n serde_json.workspace = true\n+jsonschema.workspace = true\n tokio.workspace = true\n bytes.workspace = true\n object_store.workspace = true\ndiff --git a/lib/crates/fabro-workflow/src/error.rs b/lib/crates/fabro-workflow/src/error.rs\nindex 6fcd565e3..4373a16db 100644\n--- a/lib/crates/fabro-workflow/src/error.rs\n+++ b/lib/crates/fabro-workflow/src/error.rs\n@@ -304,6 +304,9 @@ pub enum Error {\n #[error(\"Unsupported operation: {0}\")]\n Unsupported(String),\n \n+ #[error(\"{0}\")]\n+ OutputSchemaValidation(String),\n+\n #[error(\"Pipeline cancelled\")]\n Cancelled,\n }\n@@ -428,8 +431,8 @@ impl Error {\n /// Retryable: Handler (transient handler failures), Engine (could be\n /// transient), Io (network/disk issues are often transient),\n /// Llm (delegates to SdkError). Terminal: Parse, Validation,\n- /// Stylesheet (configuration errors), Checkpoint (storage\n- /// integrity), Cancelled (explicit cancellation).\n+ /// OutputSchemaValidation, Stylesheet (configuration errors), Checkpoint\n+ /// (storage integrity), Cancelled (explicit cancellation).\n #[must_use]\n pub fn is_retryable(&self) -> bool {\n match self {\n@@ -444,6 +447,7 @@ impl Error {\n | Self::Precondition(_)\n | Self::RunNotFound(_)\n | Self::Unsupported(_)\n+ | Self::OutputSchemaValidation(_)\n | Self::Cancelled => false,\n }\n }\n@@ -461,7 +465,8 @@ impl Error {\n | Self::Template { .. }\n | Self::Stylesheet(_)\n | Self::Checkpoint(_)\n- | Self::Unsupported(_) => FailureCategory::Deterministic,\n+ | Self::Unsupported(_)\n+ | Self::OutputSchemaValidation(_) => FailureCategory::Deterministic,\n Self::Precondition(_) | Self::RunNotFound(_) => FailureCategory::Structural,\n Self::Handler { failure_class, .. } | Self::Engine { failure_class, .. } => {\n *failure_class\ndiff --git a/lib/crates/fabro-workflow/src/handler/agent.rs b/lib/crates/fabro-workflow/src/handler/agent.rs\nindex 0788ee1e5..0f2118df8 100644\n--- a/lib/crates/fabro-workflow/src/handler/agent.rs\n+++ b/lib/crates/fabro-workflow/src/handler/agent.rs\n@@ -2,20 +2,21 @@ use std::path::Path;\n use std::sync::Arc;\n \n use async_trait::async_trait;\n-use fabro_agent::Sandbox;\n+use fabro_agent::{Sandbox, shell_quote};\n use fabro_graphviz::graph::{Graph, Node};\n use fabro_types::{RunId, StageModelUsage};\n use tokio_util::sync::CancellationToken;\n \n use super::llm::api::EffectiveRequestControls;\n+use super::structured_output::{\n+ self, OutputSchemaKind, StructuredOutputError, ValidatedStructuredOutput,\n+};\n use super::{EngineServices, Handler, NodeTimeoutPolicy};\n use crate::context::{Context, WorkflowContext, keys};\n use crate::error::Error;\n use crate::event::{Emitter, Event, StageScope};\n use crate::interview_runtime::WorkflowAgentQuestionRuntime;\n-use crate::outcome::{\n- BilledModelUsage, FailureCategory, FailureDetail, Outcome, OutcomeExt, StageOutcome,\n-};\n+use crate::outcome::{BilledModelUsage, Outcome, OutcomeExt};\n \n /// Result from a `CodergenBackend` invocation.\n pub enum CodergenResult {\n@@ -132,110 +133,65 @@ impl AgentHandler {\n }\n }\n \n-/// Status fields that indicate a JSON object contains routing directives.\n-const STATUS_FIELDS: &[&str] = &[\n- \"preferred_next_label\",\n- \"outcome\",\n- \"failure_reason\",\n- \"suggested_next_ids\",\n- \"context_updates\",\n-];\n-\n-/// Find all balanced `{...}` JSON object substrings in the text.\n-fn find_json_objects(text: &str) -> Vec<&str> {\n- let mut results = Vec::new();\n- let bytes = text.as_bytes();\n- let mut i = 0;\n- while i < bytes.len() {\n- if bytes[i] == b'{' {\n- let start = i;\n- let mut depth = 0;\n- let mut in_string = false;\n- let mut escape = false;\n- let mut j = i;\n- while j < bytes.len() {\n- let c = bytes[j];\n- if escape {\n- escape = false;\n- } else if c == b'\\\\' && in_string {\n- escape = true;\n- } else if c == b'\"' {\n- in_string = !in_string;\n- } else if !in_string {\n- if c == b'{' {\n- depth += 1;\n- } else if c == b'}' {\n- depth -= 1;\n- if depth == 0 {\n- results.push(&text[start..=j]);\n- break;\n- }\n- }\n- }\n- j += 1;\n- }\n- }\n- i += 1;\n- }\n- results\n-}\n-\n /// Extract routing directives from LLM response text.\n ///\n /// Searches for the last JSON object in the response that contains at least\n /// one status field (`preferred_next_label`, `outcome`, `suggested_next_ids`,\n /// `context_updates`). Merges extracted fields into the outcome.\n pub(crate) fn extract_status_fields(text: &str, outcome: &mut Outcome) -> bool {\n- let candidates = find_json_objects(text);\n-\n- let parsed = candidates.iter().rev().find_map(|candidate| {\n- let value: serde_json::Value = serde_json::from_str(candidate).ok()?;\n- if let Some(obj) = value.as_object() {\n- if STATUS_FIELDS.iter().any(|f| obj.contains_key(*f)) {\n- return Some(value);\n- }\n- }\n- None\n- });\n-\n- let Some(value) = parsed else { return false };\n- let Some(obj) = value.as_object() else {\n- return false;\n- };\n+ structured_output::extract_status_fields_loose(text, outcome)\n+}\n \n- if let Some(label) = obj.get(\"preferred_next_label\").and_then(|v| v.as_str()) {\n- outcome.preferred_label = Some(label.to_string());\n+pub(crate) async fn validate_agent_output_sources(\n+ schema: &OutputSchemaKind,\n+ response_text: &str,\n+ sandbox: &Arc,\n+ last_file_touched: Option<&str>,\n+) -> Result {\n+ if !matches!(schema, OutputSchemaKind::Routing) {\n+ return structured_output::validate_response_text(schema, response_text);\n }\n \n- if let Some(ids) = obj.get(\"suggested_next_ids\").and_then(|v| v.as_array()) {\n- let string_ids: Vec = ids\n- .iter()\n- .filter_map(|v| v.as_str().map(String::from))\n- .collect();\n- if !string_ids.is_empty() {\n- outcome.suggested_next_ids = string_ids;\n- }\n+ match structured_output::validate_response_text(schema, response_text) {\n+ Ok(validated) => return Ok(validated),\n+ Err(error) if error.allows_routing_fallback() => {}\n+ Err(error) => return Err(error),\n }\n \n- if let Some(status_str) = obj.get(\"outcome\").and_then(|v| v.as_str()) {\n- if let Ok(status) = status_str.parse::() {\n- outcome.status = status;\n- if outcome.status.is_failure() {\n- if let Some(reason) = obj.get(\"failure_reason\").and_then(|v| v.as_str()) {\n- outcome.failure =\n- Some(FailureDetail::new(reason, FailureCategory::Deterministic));\n- }\n+ let mut fallback_error = None;\n+ if let Some(status_json) = read_sandbox_file(sandbox, \"status.json\").await {\n+ match structured_output::validate_response_text(schema, &status_json) {\n+ Ok(validated) => return Ok(validated),\n+ Err(error) if error.allows_routing_fallback() => {\n+ fallback_error = Some(error);\n }\n+ Err(error) => return Err(error),\n }\n }\n \n- if let Some(updates) = obj.get(\"context_updates\").and_then(|v| v.as_object()) {\n- for (key, val) in updates {\n- outcome.context_updates.insert(key.clone(), val.clone());\n+ if let Some(path) = last_file_touched {\n+ if let Some(contents) = read_sandbox_file(sandbox, path).await {\n+ return structured_output::validate_response_text(schema, &contents);\n }\n }\n \n- true\n+ Err(fallback_error.unwrap_or_else(|| {\n+ structured_output::validate_response_text(schema, response_text)\n+ .expect_err(\"response text should have failed routing validation\")\n+ }))\n+}\n+\n+async fn read_sandbox_file(sandbox: &Arc, path: &str) -> Option {\n+ let cmd = format!(\"cat {}\", shell_quote(path));\n+ let result = sandbox\n+ .exec_command(&cmd, 5_000, None, None, None)\n+ .await\n+ .ok()?;\n+ if result.is_success() {\n+ Some(result.stdout)\n+ } else {\n+ None\n+ }\n }\n \n /// Truncate a string to at most `max_chars` characters (char-boundary safe).\n@@ -416,34 +372,46 @@ impl Handler for AgentHandler {\n serde_json::json!(&response_text),\n );\n \n- // 7b. Parse routing directives from response text, falling back to\n- // status.json written by the agent into the sandbox CWD, then to\n- // the last file the agent wrote.\n- let found_in_response = extract_status_fields(&response_text, &mut outcome);\n- if !found_in_response {\n- let mut found_in_status_json = false;\n- if let Ok(result) = services\n- .run\n- .sandbox\n- .exec_command(\"cat status.json\", 5_000, None, None, None)\n- .await\n+ if let Some(schema) = structured_output::parse_node_output_schema(node)? {\n+ match validate_agent_output_sources(\n+ &schema,\n+ &response_text,\n+ &services.run.sandbox,\n+ last_file_touched.as_deref(),\n+ )\n+ .await\n {\n- if result.is_success() {\n- found_in_status_json = extract_status_fields(&result.stdout, &mut outcome);\n+ Ok(validated) => {\n+ structured_output::apply_validated_output(\n+ node,\n+ &schema,\n+ &validated,\n+ &mut outcome,\n+ );\n+ }\n+ Err(_) => {\n+ return Ok(structured_output::exhausted_failure_outcome(\n+ node.output_retries(),\n+ ));\n }\n }\n- if !found_in_status_json {\n- if let Some(ref path) = last_file_touched {\n- let quoted = shlex::try_quote(path).unwrap_or_else(|_| path.into());\n- let cmd = format!(\"cat {quoted}\");\n- if let Ok(result) = services\n- .run\n- .sandbox\n- .exec_command(&cmd, 5_000, None, None, None)\n- .await\n- {\n- if result.is_success() {\n- extract_status_fields(&result.stdout, &mut outcome);\n+ } else {\n+ // 7b. Parse routing directives from response text, falling back to\n+ // status.json written by the agent into the sandbox CWD, then to\n+ // the last file the agent wrote.\n+ let found_in_response = extract_status_fields(&response_text, &mut outcome);\n+ if !found_in_response {\n+ let mut found_in_status_json = false;\n+ if let Some(status_json) =\n+ read_sandbox_file(&services.run.sandbox, \"status.json\").await\n+ {\n+ found_in_status_json = extract_status_fields(&status_json, &mut outcome);\n+ }\n+ if !found_in_status_json {\n+ if let Some(ref path) = last_file_touched {\n+ if let Some(contents) = read_sandbox_file(&services.run.sandbox, path).await\n+ {\n+ extract_status_fields(&contents, &mut outcome);\n }\n }\n }\n@@ -795,6 +763,142 @@ mod tests {\n );\n }\n \n+ #[tokio::test]\n+ async fn codergen_handler_output_schema_routing_uses_status_json_fallback_when_response_has_no_json()\n+ {\n+ let sandbox_dir = TempDir::new().unwrap();\n+ std::fs::write(\n+ sandbox_dir.path().join(\"status.json\"),\n+ r#\"{\"preferred_next_label\": \"review\"}\"#,\n+ )\n+ .unwrap();\n+\n+ let handler = AgentHandler::new(None);\n+ let mut node = Node::new(\"step\");\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(\"routing\".to_string()),\n+ );\n+ let context = test_context();\n+ let graph = Graph::new(\"test\");\n+ let tmp = TempDir::new().unwrap();\n+\n+ let mut services = EngineServices::test_default();\n+ services.run =\n+ services\n+ .run\n+ .with_sandbox(std::sync::Arc::new(fabro_agent::LocalSandbox::new(\n+ sandbox_dir.path().to_path_buf(),\n+ )));\n+\n+ let outcome = handler\n+ .execute(&node, &context, &graph, tmp.path(), &services)\n+ .await\n+ .unwrap();\n+\n+ assert_eq!(outcome.status, crate::outcome::StageOutcome::Succeeded);\n+ assert_eq!(outcome.preferred_label.as_deref(), Some(\"review\"));\n+ }\n+\n+ #[tokio::test]\n+ async fn codergen_handler_output_schema_routing_rejects_malformed_response_before_status_json_fallback()\n+ {\n+ struct BadRoutingBackend;\n+\n+ #[async_trait]\n+ impl CodergenBackend for BadRoutingBackend {\n+ async fn run(&self, _request: CodergenRunRequest<'_>) -> Result {\n+ Ok(CodergenResult::Text {\n+ text: r#\"{\"suggested_next_ids\": [1]}\"#.to_string(),\n+ usage: None,\n+ files_touched: Vec::new(),\n+ last_file_touched: None,\n+ })\n+ }\n+ }\n+\n+ let sandbox_dir = TempDir::new().unwrap();\n+ std::fs::write(\n+ sandbox_dir.path().join(\"status.json\"),\n+ r#\"{\"preferred_next_label\": \"should_not_use\"}\"#,\n+ )\n+ .unwrap();\n+\n+ let handler = AgentHandler::new(Some(Box::new(BadRoutingBackend)));\n+ let mut node = Node::new(\"step\");\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(\"routing\".to_string()),\n+ );\n+ node.attrs\n+ .insert(\"output_retries\".to_string(), AttrValue::Integer(0));\n+ let context = test_context();\n+ let graph = Graph::new(\"test\");\n+ let tmp = TempDir::new().unwrap();\n+\n+ let mut services = EngineServices::test_default();\n+ services.run =\n+ services\n+ .run\n+ .with_sandbox(std::sync::Arc::new(fabro_agent::LocalSandbox::new(\n+ sandbox_dir.path().to_path_buf(),\n+ )));\n+\n+ let outcome = handler\n+ .execute(&node, &context, &graph, tmp.path(), &services)\n+ .await\n+ .unwrap();\n+\n+ assert_eq!(outcome.status, crate::outcome::StageOutcome::Failed {\n+ retry_requested: false,\n+ });\n+ assert_eq!(\n+ outcome.failure_reason(),\n+ Some(\"output schema validation failed after 0 repair attempt(s)\")\n+ );\n+ assert!(outcome.preferred_label.is_none());\n+ }\n+\n+ #[tokio::test]\n+ async fn codergen_handler_custom_output_schema_updates_output_context_key() {\n+ struct CustomOutputBackend;\n+\n+ #[async_trait]\n+ impl CodergenBackend for CustomOutputBackend {\n+ async fn run(&self, _request: CodergenRunRequest<'_>) -> Result {\n+ Ok(CodergenResult::Text {\n+ text: r#\"{\"passed\": true}\"#.to_string(),\n+ usage: None,\n+ files_touched: Vec::new(),\n+ last_file_touched: None,\n+ })\n+ }\n+ }\n+\n+ let handler = AgentHandler::new(Some(Box::new(CustomOutputBackend)));\n+ let mut node = Node::new(\"audit\");\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(\n+ r#\"{\"type\":\"object\",\"required\":[\"passed\"],\"properties\":{\"passed\":{\"type\":\"boolean\"}}}\"#\n+ .to_string(),\n+ ),\n+ );\n+ let context = test_context();\n+ let graph = Graph::new(\"test\");\n+ let tmp = TempDir::new().unwrap();\n+\n+ let outcome = handler\n+ .execute(&node, &context, &graph, tmp.path(), &make_services())\n+ .await\n+ .unwrap();\n+\n+ assert_eq!(\n+ outcome.context_updates.get(\"output.audit\"),\n+ Some(&serde_json::json!({\"passed\": true})),\n+ );\n+ }\n+\n #[tokio::test]\n async fn codergen_handler_projects_provider_used_from_agent_session_events() {\n struct ProviderEventBackend;\ndiff --git a/lib/crates/fabro-workflow/src/handler/llm/acp.rs b/lib/crates/fabro-workflow/src/handler/llm/acp.rs\nindex 827f09e8a..3318c64ef 100644\n--- a/lib/crates/fabro-workflow/src/handler/llm/acp.rs\n+++ b/lib/crates/fabro-workflow/src/handler/llm/acp.rs\n@@ -325,6 +325,11 @@ impl Default for AgentAcpBackend {\n #[async_trait]\n impl CodergenBackend for AgentAcpBackend {\n async fn run(&self, request: CodergenRunRequest<'_>) -> Result {\n+ if request.node.output_schema().is_some() {\n+ return Err(Error::Validation(\n+ \"output_schema is not supported with backend=\\\"acp\\\" in this release\".to_string(),\n+ ));\n+ }\n let stage_scope = StageScope::for_handler(request.context, &request.node.id);\n self.run_turn(\n request.node,\n@@ -476,6 +481,56 @@ mod tests {\n assert_eq!(files_touched, vec![\"hello.txt\"]);\n }\n \n+ #[tokio::test]\n+ async fn acp_backend_rejects_output_schema_without_launching_process() {\n+ let tempdir = tempfile::tempdir().unwrap();\n+ let launched_path = tempdir.path().join(\"launched\");\n+\n+ let mut node = Node::new(\"work\");\n+ node.attrs\n+ .insert(\"backend\".to_string(), AttrValue::String(\"acp\".to_string()));\n+ node.attrs.insert(\n+ \"acp.command\".to_string(),\n+ AttrValue::String(\"sh -c 'touch launched'\".to_string()),\n+ );\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(\"routing\".to_string()),\n+ );\n+\n+ let backend = AgentAcpBackend::new();\n+ let sandbox: Arc = Arc::new(LocalSandbox::new(tempdir.path().to_path_buf()));\n+ let emitter = Arc::new(Emitter::default());\n+ let context = Context::new();\n+ let result = backend\n+ .run(CodergenRunRequest {\n+ node: &node,\n+ prompt: \"write hello\",\n+ context: &context,\n+ thread_id: None,\n+ emitter: &emitter,\n+ sandbox: &sandbox,\n+ tool_hooks: None,\n+ cancel_token: CancellationToken::new(),\n+ agent_tool_runtime: fabro_agent::AgentToolRuntime::default(),\n+ })\n+ .await;\n+\n+ let Err(error) = result else {\n+ panic!(\"expected output_schema guardrail error\");\n+ };\n+ assert!(\n+ error\n+ .to_string()\n+ .contains(\"output_schema is not supported with backend=\\\"acp\\\" in this release\"),\n+ \"unexpected error: {error}\",\n+ );\n+ assert!(\n+ !launched_path.exists(),\n+ \"ACP process should not launch when output_schema is present\",\n+ );\n+ }\n+\n #[tokio::test]\n async fn acp_backend_accepts_steer_and_incorporates_followup_result() {\n let tempdir = tempfile::tempdir().unwrap();\ndiff --git a/lib/crates/fabro-workflow/src/handler/llm/api.rs b/lib/crates/fabro-workflow/src/handler/llm/api.rs\nindex 8da02640d..22d08dc4b 100644\n--- a/lib/crates/fabro-workflow/src/handler/llm/api.rs\n+++ b/lib/crates/fabro-workflow/src/handler/llm/api.rs\n@@ -13,7 +13,8 @@ use fabro_auth::{CredentialSource, EnvCredentialSource};\n use fabro_graphviz::graph::{AttrValue, Node};\n use fabro_llm::client::Client;\n use fabro_llm::types::{\n- Message, ReasoningEffort, Request, Speed, TokenCounts, ToolDefinition as LlmToolDefinition,\n+ Message, ReasoningEffort, Request, Response, Speed, TokenCounts,\n+ ToolDefinition as LlmToolDefinition,\n };\n use fabro_mcp::config::McpServerSettings;\n #[cfg(test)]\n@@ -26,7 +27,11 @@ use tokio::sync::Mutex as TokioMutex;\n use tokio::task::JoinHandle;\n use tokio_util::sync::CancellationToken;\n \n-use super::super::agent::{CodergenBackend, CodergenResult, CodergenRunRequest, OneShotRequest};\n+use super::super::agent::{\n+ CodergenBackend, CodergenResult, CodergenRunRequest, OneShotRequest,\n+ validate_agent_output_sources,\n+};\n+use super::super::structured_output;\n use super::activation_lease::{ActivationLease, ActivationLeaseOptions};\n use super::routing;\n use super::routing::ProviderContext;\n@@ -460,6 +465,32 @@ fn track_file_event(event: &AgentEvent, state: &mut FileTracking) {\n }\n }\n \n+fn file_tracking_snapshot(\n+ file_tracking: &Arc>,\n+) -> (Vec, Option) {\n+ let state = file_tracking.lock().unwrap();\n+ let mut files: Vec = state.touched.iter().cloned().collect();\n+ files.sort();\n+ (files, state.last.clone())\n+}\n+\n+fn last_assistant_response(session: &Session) -> String {\n+ session\n+ .history()\n+ .turns()\n+ .iter()\n+ .rev()\n+ .find_map(|turn| {\n+ if let AgentMessage::Assistant { content, .. } = turn {\n+ if !content.is_empty() {\n+ return Some(content.clone());\n+ }\n+ }\n+ None\n+ })\n+ .unwrap_or_default()\n+}\n+\n /// Spawn a task that subscribes to session events and:\n /// 1. Tracks file changes (write_file/edit_file tool calls) into shared state.\n /// 2. Forwards non-streaming agent events to the pipeline emitter.\n@@ -521,6 +552,13 @@ pub struct AgentApiBackend {\n fabro_run_tools: Option,\n }\n \n+struct OneShotCompletion {\n+ response: Response,\n+ actual_model: String,\n+ actual_provider: String,\n+ actual_speed: Option,\n+}\n+\n impl AgentApiBackend {\n #[must_use]\n pub fn new(\n@@ -812,71 +850,18 @@ impl AgentApiBackend {\n }\n }\n }\n-}\n-\n-#[async_trait]\n-impl CodergenBackend for AgentApiBackend {\n- async fn shutdown(&self, emitter: &Arc) {\n- self.shutdown_cached_sessions(emitter);\n- }\n-\n- fn effective_request_controls(&self, node: &Node) -> Result {\n- self.resolve_effective_request_controls(node)\n- }\n-\n- async fn one_shot(&self, request: OneShotRequest<'_>) -> Result {\n- let node = request.node;\n- let prompt = request.prompt;\n- let system_prompt = request.system_prompt;\n- let emitter = request.emitter;\n- let stage_scope = request.stage_scope;\n-\n- let client = Client::from_source(self.source.as_ref(), Arc::clone(&self.catalog))\n- .await\n- .map_err(|e| Error::handler_with_source(\"Failed to create LLM client\", e))?;\n-\n- let model = node.model().unwrap_or(&self.model);\n- let provider = self.resolve_provider_context(model, node.provider())?;\n- let provider_id = provider.provider_id.to_string();\n- let controls = self.resolve_effective_request_controls(node)?;\n-\n- let max_tokens = node\n- .max_tokens()\n- .or_else(|| self.catalog.get(model).and_then(|m| m.limits.max_output));\n-\n- let mut messages = Vec::new();\n- if let Some(sys) = system_prompt {\n- messages.push(Message::system(sys));\n- }\n- messages.push(Message::user(prompt));\n-\n- let request = Request {\n- model: model.to_string(),\n- messages,\n- provider: Some(provider_id),\n- reasoning_effort: controls.reasoning_effort,\n- speed: controls.speed,\n- tools: None,\n- tool_choice: None,\n- response_format: None,\n- temperature: None,\n- top_p: None,\n- max_tokens,\n- stop_sequences: None,\n- metadata: None,\n- provider_options: None,\n- };\n-\n- // Build per-request fallback chain: if the node overrides the provider,\n- // no failover is available; otherwise use the backend's.\n- let fallback_chain: &[FallbackTarget] = if node.provider().is_some() {\n- &[]\n- } else {\n- &self.fallback_chain\n- };\n-\n- let result = client.complete(&request).await;\n \n+ async fn complete_one_shot_request(\n+ &self,\n+ client: &Client,\n+ node: &Node,\n+ emitter: &Arc,\n+ stage_scope: &StageScope,\n+ request: &Request,\n+ controls: EffectiveRequestControls,\n+ fallback_chain: &[FallbackTarget],\n+ ) -> Result {\n+ let result = client.complete(request).await;\n let default_provider = self.provider_id.to_string();\n \n let (response, actual_model, actual_provider, actual_speed) = match result {\n@@ -916,7 +901,7 @@ impl CodergenBackend for AgentApiBackend {\n let max_tokens = node.max_tokens().or_else(|| {\n self.catalog\n .get(&target.model)\n- .and_then(|m| m.limits.max_output)\n+ .and_then(|model| model.limits.max_output)\n });\n \n let fallback_request = Request {\n@@ -930,12 +915,12 @@ impl CodergenBackend for AgentApiBackend {\n \n match client.complete(&fallback_request).await {\n Ok(resp) => {\n- found = Some((\n- resp,\n- target.model.clone(),\n- target.provider.clone(),\n- controls.speed,\n- ));\n+ found = Some(OneShotCompletion {\n+ response: resp,\n+ actual_model: target.model.clone(),\n+ actual_provider: target.provider.clone(),\n+ actual_speed: controls.speed,\n+ });\n break;\n }\n Err(err) if err.failover_eligible() => {\n@@ -946,30 +931,144 @@ impl CodergenBackend for AgentApiBackend {\n }\n \n match found {\n- Some(triple) => triple,\n+ Some(completion) => return Ok(completion),\n None => return Err(Error::Llm(last_err)),\n }\n }\n Err(sdk_err) => return Err(Error::Llm(sdk_err)),\n };\n \n- let stage_usage = billed_model_usage_from_llm(\n- self.catalog.as_ref(),\n- &ModelRef {\n- provider: ProviderId::from(actual_provider),\n- model_id: actual_model,\n- speed: actual_speed,\n- },\n- &response.usage,\n- )?;\n-\n- Ok(CodergenResult::Text {\n- text: response.text(),\n- usage: Some(stage_usage),\n- files_touched: Vec::new(),\n- last_file_touched: None,\n+ Ok(OneShotCompletion {\n+ response,\n+ actual_model,\n+ actual_provider,\n+ actual_speed,\n })\n }\n+}\n+\n+#[async_trait]\n+impl CodergenBackend for AgentApiBackend {\n+ async fn shutdown(&self, emitter: &Arc) {\n+ self.shutdown_cached_sessions(emitter);\n+ }\n+\n+ fn effective_request_controls(&self, node: &Node) -> Result {\n+ self.resolve_effective_request_controls(node)\n+ }\n+\n+ async fn one_shot(&self, request: OneShotRequest<'_>) -> Result {\n+ let node = request.node;\n+ let prompt = request.prompt;\n+ let system_prompt = request.system_prompt;\n+ let emitter = request.emitter;\n+ let stage_scope = request.stage_scope;\n+\n+ let client = Client::from_source(self.source.as_ref(), Arc::clone(&self.catalog))\n+ .await\n+ .map_err(|e| Error::handler_with_source(\"Failed to create LLM client\", e))?;\n+\n+ let model = node.model().unwrap_or(&self.model);\n+ let provider = self.resolve_provider_context(model, node.provider())?;\n+ let provider_id = provider.provider_id.to_string();\n+ let controls = self.resolve_effective_request_controls(node)?;\n+\n+ let max_tokens = node\n+ .max_tokens()\n+ .or_else(|| self.catalog.get(model).and_then(|m| m.limits.max_output));\n+\n+ let mut messages = Vec::new();\n+ if let Some(sys) = system_prompt {\n+ messages.push(Message::system(sys));\n+ }\n+ messages.push(Message::user(prompt));\n+\n+ // Build per-request fallback chain: if the node overrides the provider,\n+ // no failover is available; otherwise use the backend's.\n+ let fallback_chain: &[FallbackTarget] = if node.provider().is_some() {\n+ &[]\n+ } else {\n+ &self.fallback_chain\n+ };\n+\n+ let output_schema = structured_output::parse_node_output_schema(node)?;\n+ let response_format = output_schema\n+ .as_ref()\n+ .map(structured_output::prompt_response_format);\n+ let mut repair_attempts = 0_i64;\n+ let mut total_usage = TokenCounts::default();\n+\n+ loop {\n+ let request = Request {\n+ model: model.to_string(),\n+ messages: messages.clone(),\n+ provider: Some(provider_id.clone()),\n+ reasoning_effort: controls.reasoning_effort,\n+ speed: controls.speed,\n+ tools: None,\n+ tool_choice: None,\n+ response_format: response_format.clone(),\n+ temperature: None,\n+ top_p: None,\n+ max_tokens,\n+ stop_sequences: None,\n+ metadata: None,\n+ provider_options: None,\n+ };\n+\n+ let completion = self\n+ .complete_one_shot_request(\n+ &client,\n+ node,\n+ emitter,\n+ stage_scope,\n+ &request,\n+ controls,\n+ fallback_chain,\n+ )\n+ .await?;\n+ total_usage += completion.response.usage.clone();\n+ let response_text = completion.response.text();\n+\n+ let validation_error = if let Some(schema) = &output_schema {\n+ match structured_output::validate_response_text(schema, &response_text) {\n+ Ok(_) => None,\n+ Err(error) => Some((schema, error)),\n+ }\n+ } else {\n+ None\n+ };\n+\n+ if let Some((schema, error)) = validation_error {\n+ if repair_attempts >= node.output_retries() {\n+ return Err(Error::OutputSchemaValidation(\n+ structured_output::exhausted_failure_reason(node.output_retries()),\n+ ));\n+ }\n+ messages.push(Message::assistant(response_text));\n+ messages.push(Message::user(error.repair_message(schema)));\n+ repair_attempts += 1;\n+ continue;\n+ }\n+\n+ let stage_usage = billed_model_usage_from_llm(\n+ self.catalog.as_ref(),\n+ &ModelRef {\n+ provider: ProviderId::from(completion.actual_provider),\n+ model_id: completion.actual_model,\n+ speed: completion.actual_speed,\n+ },\n+ &total_usage,\n+ )?;\n+\n+ return Ok(CodergenResult::Text {\n+ text: response_text,\n+ usage: Some(stage_usage),\n+ files_touched: Vec::new(),\n+ last_file_touched: None,\n+ });\n+ }\n+ }\n \n async fn run(&self, request: CodergenRunRequest<'_>) -> Result {\n let node = request.node;\n@@ -981,6 +1080,7 @@ impl CodergenBackend for AgentApiBackend {\n let tool_hooks = request.tool_hooks;\n let cancel_token = request.cancel_token;\n let agent_tool_runtime = request.agent_tool_runtime;\n+ let output_schema = structured_output::parse_node_output_schema(node)?;\n \n let fidelity = context.fidelity();\n let reuse_key = if fidelity == Fidelity::Full {\n@@ -1258,8 +1358,59 @@ impl CodergenBackend for AgentApiBackend {\n return Err(err);\n }\n \n+ let mut response = last_assistant_response(&session);\n+ if let Some(schema) = &output_schema {\n+ let mut repair_attempts = 0_i64;\n+ loop {\n+ let (_, last_file_touched) = file_tracking_snapshot(&file_tracking);\n+ match validate_agent_output_sources(\n+ schema,\n+ &response,\n+ sandbox,\n+ last_file_touched.as_deref(),\n+ )\n+ .await\n+ {\n+ Ok(_) => break,\n+ Err(error) => {\n+ if repair_attempts >= node.output_retries() {\n+ bridge.abort();\n+ discard_session(&mut session, &mut lease, emitter);\n+ return Err(Error::OutputSchemaValidation(\n+ structured_output::exhausted_failure_reason(node.output_retries()),\n+ ));\n+ }\n+ let repair_message = error.repair_message(schema);\n+ match session.process_input(&repair_message).await {\n+ Ok(()) => {\n+ repair_attempts += 1;\n+ response = last_assistant_response(&session);\n+ }\n+ Err(err) => match classify_agent_error(err, false) {\n+ AgentApiErrorDisposition::Cancelled => {\n+ bridge.abort();\n+ discard_session(&mut session, &mut lease, emitter);\n+ return Err(Error::Cancelled);\n+ }\n+ AgentApiErrorDisposition::Terminal(err) => {\n+ bridge.abort();\n+ discard_session(&mut session, &mut lease, emitter);\n+ return Err(err);\n+ }\n+ AgentApiErrorDisposition::FailoverEligible(sdk_err) => {\n+ bridge.abort();\n+ discard_session(&mut session, &mut lease, emitter);\n+ return Err(Error::Llm(sdk_err));\n+ }\n+ },\n+ }\n+ }\n+ }\n+ }\n+ }\n+\n // Aggregate token usage only from new turns (prevents double-counting on\n- // reuse).\n+ // reuse), including any output-schema repair turns.\n let mut total_usage = TokenCounts::default();\n for turn in &session.history().turns()[turns_before..] {\n if let AgentMessage::Assistant { usage, .. } = turn {\n@@ -1278,29 +1429,8 @@ impl CodergenBackend for AgentApiBackend {\n &total_usage,\n )?;\n \n- // Extract last assistant response from the session history.\n- let response = session\n- .history()\n- .turns()\n- .iter()\n- .rev()\n- .find_map(|turn| {\n- if let AgentMessage::Assistant { content, .. } = turn {\n- if !content.is_empty() {\n- return Some(content.clone());\n- }\n- }\n- None\n- })\n- .unwrap_or_default();\n-\n // Collect files_touched from the shared tracking state.\n- let (files_touched, last_file_touched) = {\n- let s = file_tracking.lock().unwrap();\n- let mut v: Vec = s.touched.iter().cloned().collect();\n- v.sort();\n- (v, s.last.clone())\n- };\n+ let (files_touched, last_file_touched) = file_tracking_snapshot(&file_tracking);\n \n if let Some(lease) = lease.take() {\n lease.release();\n@@ -1376,10 +1506,13 @@ mod tests {\n };\n use fabro_vault::{SecretType, Vault};\n use futures::stream;\n+ use httpmock::Method::POST;\n+ use httpmock::MockServer;\n use tokio::sync::RwLock as AsyncRwLock;\n use tokio_util::sync::CancellationToken;\n \n use super::*;\n+ use crate::context::Context;\n use crate::services::FabroRunToolServices;\n \n struct ShutdownTestProfile {\n@@ -1447,6 +1580,109 @@ mod tests {\n }\n }\n \n+ fn mock_llm_catalog(server: &MockServer) -> Arc {\n+ let settings: LlmCatalogSettings = toml::from_str(&format!(\n+ r#\"\n+[providers.mock]\n+adapter = \"openai_compatible\"\n+agent_profile = \"openai\"\n+base_url = \"{}\"\n+\n+[providers.mock.auth]\n+credentials = [\"env:MOCK_API_KEY\"]\n+\n+[models.mock-model]\n+provider = \"mock\"\n+display_name = \"Mock Model\"\n+family = \"mock\"\n+default = true\n+\n+[models.mock-model.limits]\n+context_window = 8192\n+max_output = 1024\n+\n+[models.mock-model.features]\n+tools = true\n+vision = false\n+reasoning = false\n+\"#,\n+ server.base_url()\n+ ))\n+ .unwrap();\n+ Arc::new(Catalog::from_builtin_with_overrides(&settings).unwrap())\n+ }\n+\n+ fn mock_api_backend(server: &MockServer) -> AgentApiBackend {\n+ let source = EnvCredentialSource::with_env_lookup(Arc::new(|name| {\n+ if name == \"MOCK_API_KEY\" {\n+ Some(\"sk-test\".to_string())\n+ } else {\n+ None\n+ }\n+ }));\n+ AgentApiBackend::new_with_catalog(\n+ \"mock-model\".to_string(),\n+ ProviderId::from(\"mock\"),\n+ Vec::new(),\n+ Arc::new(source),\n+ SteeringHub::for_tests(),\n+ mock_llm_catalog(server),\n+ )\n+ }\n+\n+ fn chat_completion_response(\n+ text: &str,\n+ input_tokens: i64,\n+ output_tokens: i64,\n+ ) -> serde_json::Value {\n+ serde_json::json!({\n+ \"id\": uuid::Uuid::new_v4().to_string(),\n+ \"model\": \"mock-model\",\n+ \"choices\": [{\n+ \"message\": {\n+ \"content\": text\n+ },\n+ \"finish_reason\": \"stop\"\n+ }],\n+ \"usage\": {\n+ \"prompt_tokens\": input_tokens,\n+ \"completion_tokens\": output_tokens,\n+ \"total_tokens\": input_tokens + output_tokens\n+ }\n+ })\n+ }\n+\n+ fn chat_completion_stream(text: &str, input_tokens: i64, output_tokens: i64) -> String {\n+ let text_chunk = serde_json::json!({\n+ \"id\": uuid::Uuid::new_v4().to_string(),\n+ \"model\": \"mock-model\",\n+ \"choices\": [{\n+ \"delta\": {\n+ \"content\": text\n+ },\n+ \"finish_reason\": null\n+ }]\n+ });\n+ let usage_chunk = serde_json::json!({\n+ \"id\": uuid::Uuid::new_v4().to_string(),\n+ \"model\": \"mock-model\",\n+ \"choices\": [],\n+ \"usage\": {\n+ \"prompt_tokens\": input_tokens,\n+ \"completion_tokens\": output_tokens,\n+ \"total_tokens\": input_tokens + output_tokens\n+ }\n+ });\n+ format!(\"data: {text_chunk}\\n\\ndata: {usage_chunk}\\n\\ndata: [DONE]\\n\\n\")\n+ }\n+\n+ fn custom_output_schema_attr() -> AttrValue {\n+ AttrValue::String(\n+ r#\"{\"type\":\"object\",\"required\":[\"passed\"],\"properties\":{\"passed\":{\"type\":\"boolean\"}}}\"#\n+ .to_string(),\n+ )\n+ }\n+\n #[test]\n fn agent_backend_stores_config() {\n let backend = AgentApiBackend::new_from_env(\n@@ -2212,6 +2448,127 @@ reasoning = false\n assert_eq!(client.provider_names(), vec![\"anthropic\"]);\n }\n \n+ #[tokio::test]\n+ async fn one_shot_repairs_custom_output_schema_with_previous_assistant_message() {\n+ let server = MockServer::start();\n+ let first = server.mock(|when, then| {\n+ when.method(POST)\n+ .path(\"/chat/completions\")\n+ .body_includes(r#\"\"type\":\"json_schema\"\"#)\n+ .body_excludes(r#\"\"role\":\"assistant\"\"#);\n+ then.status(200)\n+ .header(\"content-type\", \"application/json\")\n+ .json_body(chat_completion_response(\"not json\", 10, 1));\n+ });\n+ let repair = server.mock(|when, then| {\n+ when.method(POST)\n+ .path(\"/chat/completions\")\n+ .body_includes(r#\"\"type\":\"json_schema\"\"#)\n+ .body_includes(r#\"\"role\":\"assistant\"\"#)\n+ .body_includes(\"not json\")\n+ .body_includes(\"output_schema\");\n+ then.status(200)\n+ .header(\"content-type\", \"application/json\")\n+ .json_body(chat_completion_response(r#\"{\"passed\":true}\"#, 11, 2));\n+ });\n+ let backend = mock_api_backend(&server);\n+ let mut node = Node::new(\"audit\");\n+ node.attrs\n+ .insert(\"output_schema\".to_string(), custom_output_schema_attr());\n+ node.attrs\n+ .insert(\"output_retries\".to_string(), AttrValue::Integer(1));\n+ let context = Context::new();\n+ let stage_scope = StageScope::for_handler(&context, &node.id);\n+ let emitter = Arc::new(Emitter::new(fabro_types::RunId::new()));\n+ let workspace = tempfile::tempdir().unwrap();\n+ let sandbox: Arc =\n+ Arc::new(LocalSandbox::new(workspace.path().to_path_buf()));\n+\n+ let result = backend\n+ .one_shot(OneShotRequest {\n+ node: &node,\n+ prompt: \"Audit the result\",\n+ system_prompt: None,\n+ emitter: &emitter,\n+ stage_scope: &stage_scope,\n+ sandbox: &sandbox,\n+ cancel_token: CancellationToken::new(),\n+ })\n+ .await\n+ .unwrap();\n+\n+ first.assert_calls(1);\n+ repair.assert_calls(1);\n+ let CodergenResult::Text { text, usage, .. } = result else {\n+ panic!(\"one_shot should return text\");\n+ };\n+ assert_eq!(text, r#\"{\"passed\":true}\"#);\n+ let usage = usage.expect(\"usage should be aggregated\");\n+ assert_eq!(usage.tokens().input_tokens, 21);\n+ assert_eq!(usage.tokens().output_tokens, 3);\n+ }\n+\n+ #[tokio::test]\n+ async fn agent_run_repairs_custom_output_schema_in_same_session() {\n+ let server = MockServer::start();\n+ let first = server.mock(|when, then| {\n+ when.method(POST)\n+ .path(\"/chat/completions\")\n+ .body_includes(r#\"\"stream\":true\"#)\n+ .body_excludes(r#\"\"role\":\"assistant\"\"#);\n+ then.status(200)\n+ .header(\"content-type\", \"text/event-stream\")\n+ .body(chat_completion_stream(\"not json\", 20, 3));\n+ });\n+ let repair = server.mock(|when, then| {\n+ when.method(POST)\n+ .path(\"/chat/completions\")\n+ .body_includes(r#\"\"stream\":true\"#)\n+ .body_includes(r#\"\"role\":\"assistant\"\"#)\n+ .body_includes(\"not json\")\n+ .body_includes(\"output_schema\");\n+ then.status(200)\n+ .header(\"content-type\", \"text/event-stream\")\n+ .body(chat_completion_stream(r#\"{\"passed\":true}\"#, 21, 4));\n+ });\n+ let backend = mock_api_backend(&server);\n+ let mut node = Node::new(\"audit\");\n+ node.attrs\n+ .insert(\"output_schema\".to_string(), custom_output_schema_attr());\n+ node.attrs\n+ .insert(\"output_retries\".to_string(), AttrValue::Integer(1));\n+ let context = Context::new();\n+ let emitter = Arc::new(Emitter::new(fabro_types::RunId::new()));\n+ let workspace = tempfile::tempdir().unwrap();\n+ let sandbox: Arc =\n+ Arc::new(LocalSandbox::new(workspace.path().to_path_buf()));\n+\n+ let result = backend\n+ .run(CodergenRunRequest {\n+ node: &node,\n+ prompt: \"Audit the result\",\n+ context: &context,\n+ thread_id: None,\n+ emitter: &emitter,\n+ sandbox: &sandbox,\n+ tool_hooks: None,\n+ cancel_token: CancellationToken::new(),\n+ agent_tool_runtime: fabro_agent::AgentToolRuntime::default(),\n+ })\n+ .await\n+ .unwrap();\n+\n+ first.assert_calls(1);\n+ repair.assert_calls(1);\n+ let CodergenResult::Text { text, usage, .. } = result else {\n+ panic!(\"run should return text\");\n+ };\n+ assert_eq!(text, r#\"{\"passed\":true}\"#);\n+ let usage = usage.expect(\"usage should be aggregated\");\n+ assert_eq!(usage.tokens().input_tokens, 41);\n+ assert_eq!(usage.tokens().output_tokens, 7);\n+ }\n+\n #[tokio::test]\n async fn api_backend_shutdown_closes_cached_sessions_once() {\n let backend = AgentApiBackend::new_from_env(\ndiff --git a/lib/crates/fabro-workflow/src/handler/mod.rs b/lib/crates/fabro-workflow/src/handler/mod.rs\nindex 675f24b39..bc94af3b8 100644\n--- a/lib/crates/fabro-workflow/src/handler/mod.rs\n+++ b/lib/crates/fabro-workflow/src/handler/mod.rs\n@@ -9,6 +9,7 @@ pub mod manager_loop;\n pub mod parallel;\n pub mod prompt;\n pub mod start;\n+pub mod structured_output;\n pub mod wait;\n \n use std::any::Any;\ndiff --git a/lib/crates/fabro-workflow/src/handler/prompt.rs b/lib/crates/fabro-workflow/src/handler/prompt.rs\nindex 6dbc0576d..27b9fa524 100644\n--- a/lib/crates/fabro-workflow/src/handler/prompt.rs\n+++ b/lib/crates/fabro-workflow/src/handler/prompt.rs\n@@ -10,7 +10,7 @@ use super::agent::{\n truncate,\n };\n use super::llm::routing;\n-use super::{EngineServices, Handler};\n+use super::{EngineServices, Handler, structured_output};\n use crate::context::{Context, WorkflowContext, keys};\n use crate::error::Error;\n use crate::event::{Emitter, Event};\n@@ -191,7 +191,25 @@ impl Handler for PromptHandler {\n serde_json::json!(&response_text),\n );\n \n- extract_status_fields(&response_text, &mut outcome);\n+ if let Some(schema) = structured_output::parse_node_output_schema(node)? {\n+ match structured_output::validate_response_text(&schema, &response_text) {\n+ Ok(validated) => {\n+ structured_output::apply_validated_output(\n+ node,\n+ &schema,\n+ &validated,\n+ &mut outcome,\n+ );\n+ }\n+ Err(_) => {\n+ return Ok(structured_output::exhausted_failure_outcome(\n+ node.output_retries(),\n+ ));\n+ }\n+ }\n+ } else {\n+ extract_status_fields(&response_text, &mut outcome);\n+ }\n outcome.usage = stage_usage;\n outcome.files_touched = backend_files_touched;\n \n@@ -214,6 +232,7 @@ mod tests {\n use super::*;\n use crate::event::Emitter;\n use crate::handler::agent::CodergenRunRequest;\n+ use crate::outcome::OutcomeExt;\n \n fn make_services() -> EngineServices {\n EngineServices::test_default()\n@@ -368,6 +387,102 @@ mod tests {\n );\n }\n \n+ #[tokio::test]\n+ async fn prompt_handler_custom_output_schema_updates_output_context_key() {\n+ struct CustomOutputBackend;\n+\n+ #[async_trait]\n+ impl CodergenBackend for CustomOutputBackend {\n+ async fn run(&self, _request: CodergenRunRequest<'_>) -> Result {\n+ panic!(\"run() should not be called for prompt handler\");\n+ }\n+\n+ async fn one_shot(\n+ &self,\n+ _request: OneShotRequest<'_>,\n+ ) -> Result {\n+ Ok(CodergenResult::Text {\n+ text: r#\"{\"passed\": true}\"#.to_string(),\n+ usage: None,\n+ files_touched: Vec::new(),\n+ last_file_touched: None,\n+ })\n+ }\n+ }\n+\n+ let handler = PromptHandler::new(Some(Box::new(CustomOutputBackend)));\n+ let mut node = Node::new(\"audit\");\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(\n+ r#\"{\"type\":\"object\",\"required\":[\"passed\"],\"properties\":{\"passed\":{\"type\":\"boolean\"}}}\"#\n+ .to_string(),\n+ ),\n+ );\n+ let context = Context::new();\n+ let graph = Graph::new(\"test\");\n+ let tmp = TempDir::new().unwrap();\n+\n+ let outcome = handler\n+ .execute(&node, &context, &graph, tmp.path(), &make_services())\n+ .await\n+ .unwrap();\n+\n+ assert_eq!(\n+ outcome.context_updates.get(\"output.audit\"),\n+ Some(&serde_json::json!({\"passed\": true})),\n+ );\n+ }\n+\n+ #[tokio::test]\n+ async fn prompt_handler_routing_output_schema_requires_valid_routing_json() {\n+ struct BadRoutingBackend;\n+\n+ #[async_trait]\n+ impl CodergenBackend for BadRoutingBackend {\n+ async fn run(&self, _request: CodergenRunRequest<'_>) -> Result {\n+ panic!(\"run() should not be called for prompt handler\");\n+ }\n+\n+ async fn one_shot(\n+ &self,\n+ _request: OneShotRequest<'_>,\n+ ) -> Result {\n+ Ok(CodergenResult::Text {\n+ text: r#\"{\"outcome\": 123}\"#.to_string(),\n+ usage: None,\n+ files_touched: Vec::new(),\n+ last_file_touched: None,\n+ })\n+ }\n+ }\n+\n+ let handler = PromptHandler::new(Some(Box::new(BadRoutingBackend)));\n+ let mut node = Node::new(\"route\");\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(\"routing\".to_string()),\n+ );\n+ node.attrs\n+ .insert(\"output_retries\".to_string(), AttrValue::Integer(0));\n+ let context = Context::new();\n+ let graph = Graph::new(\"test\");\n+ let tmp = TempDir::new().unwrap();\n+\n+ let outcome = handler\n+ .execute(&node, &context, &graph, tmp.path(), &make_services())\n+ .await\n+ .unwrap();\n+\n+ assert_eq!(outcome.status, crate::outcome::StageOutcome::Failed {\n+ retry_requested: false,\n+ });\n+ assert_eq!(\n+ outcome.failure_reason(),\n+ Some(\"output schema validation failed after 0 repair attempt(s)\")\n+ );\n+ }\n+\n #[tokio::test]\n async fn prompt_handler_projects_provider_used_from_prompt_events() {\n struct ProviderOneShotBackend;\ndiff --git a/lib/crates/fabro-workflow/src/handler/structured_output.rs b/lib/crates/fabro-workflow/src/handler/structured_output.rs\nnew file mode 100644\nindex 000000000..d98a42afb\n--- /dev/null\n+++ b/lib/crates/fabro-workflow/src/handler/structured_output.rs\n@@ -0,0 +1,608 @@\n+use fabro_graphviz::graph::Node;\n+use fabro_llm::types::{ResponseFormat, ResponseFormatType};\n+use serde_json::Value;\n+\n+use crate::error::Error;\n+use crate::outcome::{FailureCategory, FailureDetail, Outcome, StageOutcome};\n+\n+pub(crate) const ROUTING_STATUS_FIELDS: &[&str] = &[\n+ \"preferred_next_label\",\n+ \"outcome\",\n+ \"failure_reason\",\n+ \"suggested_next_ids\",\n+ \"context_updates\",\n+];\n+\n+#[derive(Debug, Clone, PartialEq)]\n+pub(crate) enum OutputSchemaKind {\n+ Routing,\n+ JsonSchema { schema: Value },\n+}\n+\n+#[derive(Debug, Clone, Copy, PartialEq, Eq)]\n+pub(crate) enum StructuredOutputErrorKind {\n+ NoJsonObject,\n+ NoRelevantJsonObject,\n+ InvalidJson,\n+ SchemaValidation,\n+}\n+\n+#[derive(Debug, Clone, PartialEq, Eq)]\n+pub(crate) struct StructuredOutputError {\n+ kind: StructuredOutputErrorKind,\n+ messages: Vec,\n+}\n+\n+impl StructuredOutputError {\n+ fn new(kind: StructuredOutputErrorKind, message: impl Into) -> Self {\n+ Self {\n+ kind,\n+ messages: vec![message.into()],\n+ }\n+ }\n+\n+ fn validation(messages: Vec) -> Self {\n+ Self {\n+ kind: StructuredOutputErrorKind::SchemaValidation,\n+ messages,\n+ }\n+ }\n+\n+ #[cfg(test)]\n+ #[must_use]\n+ pub(crate) fn kind(&self) -> StructuredOutputErrorKind {\n+ self.kind\n+ }\n+\n+ #[cfg(test)]\n+ #[must_use]\n+ pub(crate) fn messages(&self) -> &[String] {\n+ &self.messages\n+ }\n+\n+ #[must_use]\n+ pub(crate) fn allows_routing_fallback(&self) -> bool {\n+ matches!(\n+ self.kind,\n+ StructuredOutputErrorKind::NoJsonObject\n+ | StructuredOutputErrorKind::NoRelevantJsonObject\n+ )\n+ }\n+\n+ #[must_use]\n+ pub(crate) fn repair_message(&self, schema: &OutputSchemaKind) -> String {\n+ let expectation = match schema {\n+ OutputSchemaKind::Routing => format!(\n+ \"Return a single JSON object with at least one routing field: {}.\",\n+ ROUTING_STATUS_FIELDS.join(\", \")\n+ ),\n+ OutputSchemaKind::JsonSchema { .. } => {\n+ \"Return a single JSON object that satisfies the configured JSON Schema.\".to_string()\n+ }\n+ };\n+ let errors = self\n+ .messages\n+ .iter()\n+ .map(|message| format!(\"- {message}\"))\n+ .collect::>()\n+ .join(\"\\n\");\n+ format!(\n+ \"Your previous response did not satisfy the node's output_schema.\\n\\n\\\n+ Validation errors:\\n{errors}\\n\\n\\\n+ {expectation}\\n\\\n+ Do not include Markdown fences or explanatory prose; reply only with the corrected JSON object.\"\n+ )\n+ }\n+}\n+\n+#[derive(Debug, Clone, PartialEq)]\n+pub(crate) struct ValidatedStructuredOutput {\n+ pub(crate) value: Value,\n+}\n+\n+#[must_use]\n+pub(crate) fn output_key(node_id: &str) -> String {\n+ format!(\"output.{node_id}\")\n+}\n+\n+#[must_use]\n+pub(crate) fn exhausted_failure_reason(repair_attempts: i64) -> String {\n+ format!(\"output schema validation failed after {repair_attempts} repair attempt(s)\")\n+}\n+\n+#[must_use]\n+pub(crate) fn exhausted_failure_outcome(repair_attempts: i64) -> Outcome {\n+ Outcome {\n+ status: StageOutcome::Failed {\n+ retry_requested: false,\n+ },\n+ failure: Some(FailureDetail::new(\n+ exhausted_failure_reason(repair_attempts),\n+ FailureCategory::Deterministic,\n+ )),\n+ ..Outcome::default()\n+ }\n+}\n+\n+pub(crate) fn parse_node_output_schema(node: &Node) -> Result, Error> {\n+ let Some(raw) = node.output_schema() else {\n+ return Ok(None);\n+ };\n+ let value = raw.trim();\n+ if value.is_empty() {\n+ return Err(Error::Validation(format!(\n+ \"Invalid output_schema for node \\\"{}\\\": value must not be empty\",\n+ node.id\n+ )));\n+ }\n+ if value == \"routing\" {\n+ return Ok(Some(OutputSchemaKind::Routing));\n+ }\n+ if value.starts_with('@') {\n+ return Err(Error::Validation(format!(\n+ \"Invalid output_schema for node \\\"{}\\\": unresolved file reference {value}\",\n+ node.id\n+ )));\n+ }\n+\n+ let schema = serde_json::from_str::(value).map_err(|err| {\n+ Error::Validation(format!(\n+ \"Invalid output_schema for node \\\"{}\\\": expected \\\"routing\\\" or a JSON Schema object: {err}\",\n+ node.id\n+ ))\n+ })?;\n+ compile_schema(&schema).map_err(|err| {\n+ Error::Validation(format!(\n+ \"Invalid output_schema for node \\\"{}\\\": {err}\",\n+ node.id\n+ ))\n+ })?;\n+ Ok(Some(OutputSchemaKind::JsonSchema { schema }))\n+}\n+\n+#[must_use]\n+pub(crate) fn prompt_response_format(schema: &OutputSchemaKind) -> ResponseFormat {\n+ match schema {\n+ OutputSchemaKind::Routing => ResponseFormat {\n+ kind: ResponseFormatType::JsonObject,\n+ json_schema: None,\n+ strict: false,\n+ },\n+ OutputSchemaKind::JsonSchema { schema } => ResponseFormat {\n+ kind: ResponseFormatType::JsonSchema,\n+ json_schema: Some(schema.clone()),\n+ strict: true,\n+ },\n+ }\n+}\n+\n+pub(crate) fn validate_response_text(\n+ schema: &OutputSchemaKind,\n+ text: &str,\n+) -> Result {\n+ match schema {\n+ OutputSchemaKind::Routing => validate_routing_response_text(text),\n+ OutputSchemaKind::JsonSchema { schema } => validate_custom_response_text(schema, text),\n+ }\n+}\n+\n+pub(crate) fn apply_validated_output(\n+ node: &Node,\n+ schema: &OutputSchemaKind,\n+ validated: &ValidatedStructuredOutput,\n+ outcome: &mut Outcome,\n+) {\n+ match schema {\n+ OutputSchemaKind::Routing => apply_routing_fields(&validated.value, outcome),\n+ OutputSchemaKind::JsonSchema { .. } => {\n+ outcome\n+ .context_updates\n+ .insert(output_key(&node.id), validated.value.clone());\n+ }\n+ }\n+}\n+\n+/// Find all balanced `{...}` JSON object substrings in the text.\n+pub(crate) fn find_json_objects(text: &str) -> Vec<&str> {\n+ let mut results = Vec::new();\n+ let bytes = text.as_bytes();\n+ let mut i = 0;\n+ while i < bytes.len() {\n+ if bytes[i] == b'{' {\n+ let start = i;\n+ let mut depth = 0;\n+ let mut in_string = false;\n+ let mut escape = false;\n+ let mut j = i;\n+ while j < bytes.len() {\n+ let c = bytes[j];\n+ if escape {\n+ escape = false;\n+ } else if c == b'\\\\' && in_string {\n+ escape = true;\n+ } else if c == b'\"' {\n+ in_string = !in_string;\n+ } else if !in_string {\n+ if c == b'{' {\n+ depth += 1;\n+ } else if c == b'}' {\n+ depth -= 1;\n+ if depth == 0 {\n+ results.push(&text[start..=j]);\n+ break;\n+ }\n+ }\n+ }\n+ j += 1;\n+ }\n+ }\n+ i += 1;\n+ }\n+ results\n+}\n+\n+pub(crate) fn extract_status_fields_loose(text: &str, outcome: &mut Outcome) -> bool {\n+ let candidates = find_json_objects(text);\n+\n+ let parsed = candidates.iter().rev().find_map(|candidate| {\n+ let value: Value = serde_json::from_str(candidate).ok()?;\n+ if value.as_object().is_some_and(contains_routing_field) {\n+ Some(value)\n+ } else {\n+ None\n+ }\n+ });\n+\n+ let Some(value) = parsed else { return false };\n+ apply_routing_fields(&value, outcome);\n+ true\n+}\n+\n+fn validate_routing_response_text(\n+ text: &str,\n+) -> Result {\n+ let candidates = find_json_objects(text);\n+ if candidates.is_empty() {\n+ return Err(StructuredOutputError::new(\n+ StructuredOutputErrorKind::NoJsonObject,\n+ \"no JSON object found in response\",\n+ ));\n+ }\n+\n+ for candidate in candidates.iter().rev() {\n+ let parsed = match serde_json::from_str::(candidate) {\n+ Ok(value) => value,\n+ Err(err) if raw_mentions_routing_field(candidate) => {\n+ return Err(StructuredOutputError::new(\n+ StructuredOutputErrorKind::InvalidJson,\n+ format!(\"invalid routing JSON object: {err}\"),\n+ ));\n+ }\n+ Err(_) => continue,\n+ };\n+ let Some(obj) = parsed.as_object() else {\n+ continue;\n+ };\n+ if !contains_routing_field(obj) {\n+ continue;\n+ }\n+ validate_value_against_schema(&routing_schema(), &parsed)?;\n+ return Ok(ValidatedStructuredOutput { value: parsed });\n+ }\n+\n+ Err(StructuredOutputError::new(\n+ StructuredOutputErrorKind::NoRelevantJsonObject,\n+ format!(\n+ \"no JSON object contained any recognized routing field ({})\",\n+ ROUTING_STATUS_FIELDS.join(\", \")\n+ ),\n+ ))\n+}\n+\n+fn validate_custom_response_text(\n+ schema: &Value,\n+ text: &str,\n+) -> Result {\n+ let candidates = find_json_objects(text);\n+ let Some(candidate) = candidates.last() else {\n+ return Err(StructuredOutputError::new(\n+ StructuredOutputErrorKind::NoJsonObject,\n+ \"no JSON object found in response\",\n+ ));\n+ };\n+ let parsed = serde_json::from_str::(candidate).map_err(|err| {\n+ StructuredOutputError::new(\n+ StructuredOutputErrorKind::InvalidJson,\n+ format!(\"invalid JSON object: {err}\"),\n+ )\n+ })?;\n+ validate_value_against_schema(schema, &parsed)?;\n+ Ok(ValidatedStructuredOutput { value: parsed })\n+}\n+\n+fn validate_value_against_schema(\n+ schema: &Value,\n+ value: &Value,\n+) -> Result<(), StructuredOutputError> {\n+ let validator = compile_schema(schema).map_err(|err| {\n+ StructuredOutputError::new(\n+ StructuredOutputErrorKind::SchemaValidation,\n+ format!(\"invalid JSON Schema: {err}\"),\n+ )\n+ })?;\n+ let errors = validator\n+ .iter_errors(value)\n+ .map(|error| error.to_string())\n+ .take(5)\n+ .collect::>();\n+ if errors.is_empty() {\n+ Ok(())\n+ } else {\n+ Err(StructuredOutputError::validation(errors))\n+ }\n+}\n+\n+fn compile_schema(\n+ schema: &Value,\n+) -> Result> {\n+ jsonschema::validator_for(schema)\n+}\n+\n+fn contains_routing_field(obj: &serde_json::Map) -> bool {\n+ ROUTING_STATUS_FIELDS\n+ .iter()\n+ .any(|field| obj.contains_key(*field))\n+}\n+\n+fn raw_mentions_routing_field(candidate: &str) -> bool {\n+ ROUTING_STATUS_FIELDS\n+ .iter()\n+ .any(|field| candidate.contains(&format!(\"\\\"{field}\\\"\")))\n+}\n+\n+fn routing_schema() -> Value {\n+ serde_json::json!({\n+ \"type\": \"object\",\n+ \"additionalProperties\": true,\n+ \"properties\": {\n+ \"preferred_next_label\": { \"type\": \"string\" },\n+ \"outcome\": {\n+ \"type\": \"string\",\n+ \"enum\": [\"succeeded\", \"partially_succeeded\", \"failed\", \"skipped\"]\n+ },\n+ \"failure_reason\": { \"type\": \"string\" },\n+ \"suggested_next_ids\": {\n+ \"type\": \"array\",\n+ \"items\": { \"type\": \"string\" }\n+ },\n+ \"context_updates\": { \"type\": \"object\" }\n+ },\n+ \"anyOf\": ROUTING_STATUS_FIELDS\n+ .iter()\n+ .map(|field| serde_json::json!({ \"required\": [field] }))\n+ .collect::>()\n+ })\n+}\n+\n+fn apply_routing_fields(value: &Value, outcome: &mut Outcome) {\n+ let Some(obj) = value.as_object() else {\n+ return;\n+ };\n+\n+ if let Some(label) = obj.get(\"preferred_next_label\").and_then(Value::as_str) {\n+ outcome.preferred_label = Some(label.to_string());\n+ }\n+\n+ if let Some(ids) = obj.get(\"suggested_next_ids\").and_then(Value::as_array) {\n+ let string_ids: Vec = ids\n+ .iter()\n+ .filter_map(|value| value.as_str().map(String::from))\n+ .collect();\n+ if !string_ids.is_empty() {\n+ outcome.suggested_next_ids = string_ids;\n+ }\n+ }\n+\n+ if let Some(status_str) = obj.get(\"outcome\").and_then(Value::as_str) {\n+ if let Ok(status) = status_str.parse::() {\n+ outcome.status = status;\n+ if outcome.status.is_failure() {\n+ if let Some(reason) = obj.get(\"failure_reason\").and_then(Value::as_str) {\n+ outcome.failure =\n+ Some(FailureDetail::new(reason, FailureCategory::Deterministic));\n+ }\n+ }\n+ }\n+ }\n+\n+ if let Some(updates) = obj.get(\"context_updates\").and_then(Value::as_object) {\n+ for (key, value) in updates {\n+ outcome.context_updates.insert(key.clone(), value.clone());\n+ }\n+ }\n+}\n+\n+#[cfg(test)]\n+mod tests {\n+ use fabro_graphviz::graph::{AttrValue, Node};\n+\n+ use super::*;\n+\n+ fn routing() -> OutputSchemaKind {\n+ OutputSchemaKind::Routing\n+ }\n+\n+ fn schema(value: Value) -> OutputSchemaKind {\n+ OutputSchemaKind::JsonSchema { schema: value }\n+ }\n+\n+ #[test]\n+ fn validates_routing_json_and_applies_fields() {\n+ let validated = validate_response_text(\n+ &routing(),\n+ r#\"done {\"outcome\":\"failed\",\"failure_reason\":\"tests failed\",\"preferred_next_label\":\"fix\",\"suggested_next_ids\":[\"a\"],\"context_updates\":{\"verified\":true}}\"#,\n+ )\n+ .unwrap();\n+ let mut outcome = Outcome::success();\n+\n+ apply_routing_fields(&validated.value, &mut outcome);\n+\n+ assert_eq!(outcome.status, StageOutcome::Failed {\n+ retry_requested: false,\n+ });\n+ assert_eq!(\n+ outcome.failure.as_ref().map(|f| f.message.as_str()),\n+ Some(\"tests failed\")\n+ );\n+ assert_eq!(outcome.preferred_label.as_deref(), Some(\"fix\"));\n+ assert_eq!(outcome.suggested_next_ids, vec![\"a\".to_string()]);\n+ assert_eq!(\n+ outcome.context_updates.get(\"verified\"),\n+ Some(&serde_json::json!(true)),\n+ );\n+ }\n+\n+ #[test]\n+ fn routing_json_missing_routing_fields_is_invalid() {\n+ let error = validate_response_text(&routing(), r#\"{\"summary\":\"ok\"}\"#).unwrap_err();\n+\n+ assert_eq!(\n+ error.kind(),\n+ StructuredOutputErrorKind::NoRelevantJsonObject\n+ );\n+ assert!(error.messages()[0].contains(\"recognized routing field\"));\n+ }\n+\n+ #[test]\n+ fn routing_json_with_wrong_field_type_is_invalid() {\n+ let error =\n+ validate_response_text(&routing(), r#\"{\"suggested_next_ids\":[1]}\"#).unwrap_err();\n+\n+ assert_eq!(error.kind(), StructuredOutputErrorKind::SchemaValidation);\n+ assert!(\n+ error\n+ .messages()\n+ .iter()\n+ .any(|message| message.contains(\"string\")),\n+ \"unexpected messages: {:?}\",\n+ error.messages(),\n+ );\n+ }\n+\n+ #[test]\n+ fn validates_custom_schema_against_last_json_object() {\n+ let schema = schema(serde_json::json!({\n+ \"type\": \"object\",\n+ \"required\": [\"passed\"],\n+ \"properties\": {\n+ \"passed\": { \"type\": \"boolean\" }\n+ }\n+ }));\n+\n+ let validated =\n+ validate_response_text(&schema, r#\"ignore {\"other\":1} final {\"passed\":true}\"#).unwrap();\n+\n+ assert_eq!(validated.value, serde_json::json!({\"passed\": true}));\n+ }\n+\n+ #[test]\n+ fn custom_schema_validation_errors_are_reported() {\n+ let schema = schema(serde_json::json!({\n+ \"type\": \"object\",\n+ \"required\": [\"passed\"],\n+ \"properties\": {\n+ \"passed\": { \"type\": \"boolean\" }\n+ }\n+ }));\n+\n+ let error = validate_response_text(&schema, r#\"{\"passed\":\"yes\"}\"#).unwrap_err();\n+\n+ assert_eq!(error.kind(), StructuredOutputErrorKind::SchemaValidation);\n+ assert!(\n+ error\n+ .messages()\n+ .iter()\n+ .any(|message| message.contains(\"boolean\")),\n+ \"unexpected messages: {:?}\",\n+ error.messages(),\n+ );\n+ }\n+\n+ #[test]\n+ fn invalid_custom_schema_is_rejected_when_parsing_node_attr() {\n+ let mut node = Node::new(\"audit\");\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(r#\"{\"type\": 5}\"#.to_string()),\n+ );\n+\n+ let error = parse_node_output_schema(&node).unwrap_err();\n+\n+ assert!(\n+ error.to_string().contains(\"Invalid output_schema\"),\n+ \"unexpected error: {error}\",\n+ );\n+ }\n+\n+ #[test]\n+ fn invalid_json_candidate_is_reported_for_custom_schema() {\n+ let schema = schema(serde_json::json!({\"type\": \"object\"}));\n+\n+ let error = validate_response_text(&schema, r\"{not json}\").unwrap_err();\n+\n+ assert_eq!(error.kind(), StructuredOutputErrorKind::InvalidJson);\n+ assert!(error.messages()[0].contains(\"invalid JSON object\"));\n+ }\n+\n+ #[test]\n+ fn no_json_object_is_reported() {\n+ let error = validate_response_text(&routing(), \"plain text only\").unwrap_err();\n+\n+ assert_eq!(error.kind(), StructuredOutputErrorKind::NoJsonObject);\n+ assert!(error.messages()[0].contains(\"no JSON object\"));\n+ }\n+\n+ #[test]\n+ fn parse_node_output_schema_accepts_builtin_routing_keyword() {\n+ let mut node = Node::new(\"route\");\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(\"routing\".to_string()),\n+ );\n+\n+ let parsed = parse_node_output_schema(&node).unwrap();\n+\n+ assert_eq!(parsed, Some(OutputSchemaKind::Routing));\n+ }\n+\n+ #[test]\n+ fn prompt_response_format_uses_json_schema_for_custom_schema() {\n+ let schema = schema(serde_json::json!({\"type\": \"object\"}));\n+\n+ let format = prompt_response_format(&schema);\n+\n+ assert_eq!(format.kind, ResponseFormatType::JsonSchema);\n+ assert_eq!(\n+ format.json_schema,\n+ Some(serde_json::json!({\"type\": \"object\"}))\n+ );\n+ assert!(format.strict);\n+ }\n+\n+ #[test]\n+ fn apply_validated_custom_output_updates_output_context_key() {\n+ let node = Node::new(\"audit\");\n+ let schema = schema(serde_json::json!({\"type\": \"object\"}));\n+ let validated = ValidatedStructuredOutput {\n+ value: serde_json::json!({\"passed\": true}),\n+ };\n+ let mut outcome = Outcome::success();\n+\n+ apply_validated_output(&node, &schema, &validated, &mut outcome);\n+\n+ assert_eq!(\n+ outcome.context_updates.get(\"output.audit\"),\n+ Some(&serde_json::json!({\"passed\": true})),\n+ );\n+ }\n+}\ndiff --git a/lib/crates/fabro-workflow/src/static_reference.rs b/lib/crates/fabro-workflow/src/static_reference.rs\nindex 276cf37c1..af83539eb 100644\n--- a/lib/crates/fabro-workflow/src/static_reference.rs\n+++ b/lib/crates/fabro-workflow/src/static_reference.rs\n@@ -87,9 +87,57 @@ pub fn reference_kind_for_attribute(\n \"goal\" if matches!(scope, AttributeScope::Graph) && value.starts_with('@') => {\n Some(ReferenceKind::GraphGoalFile)\n }\n- \"prompt\" if matches!(scope, AttributeScope::Node) && value.starts_with('@') => {\n+ \"prompt\" | \"output_schema\"\n+ if matches!(scope, AttributeScope::Node) && value.starts_with('@') =>\n+ {\n Some(ReferenceKind::FileInline)\n }\n _ => None,\n }\n }\n+\n+#[cfg(test)]\n+mod tests {\n+ use super::*;\n+\n+ #[test]\n+ fn output_schema_at_value_is_file_inline_reference() {\n+ assert_eq!(\n+ reference_kind_for_attribute(\n+ AttributeScope::Node,\n+ \"output_schema\",\n+ \"@schemas/result.schema.json\",\n+ ),\n+ Some(ReferenceKind::FileInline),\n+ );\n+ }\n+\n+ #[test]\n+ fn output_schema_builtin_keyword_is_not_file_inline_reference() {\n+ assert_eq!(\n+ reference_kind_for_attribute(AttributeScope::Node, \"output_schema\", \"routing\"),\n+ None,\n+ );\n+ }\n+\n+ #[test]\n+ fn output_schema_reference_rejects_template_syntax() {\n+ let error = reference_kind_for_attribute(\n+ AttributeScope::Node,\n+ \"output_schema\",\n+ \"@schemas/{{ inputs.schema }}.json\",\n+ )\n+ .expect(\"output_schema @ references should be static references\")\n+ .validate(\"@schemas/{{ inputs.schema }}.json\")\n+ .unwrap_err();\n+\n+ assert_eq!(error.kind(), ReferenceKind::FileInline);\n+ assert_eq!(error.value(), \"@schemas/{{ inputs.schema }}.json\");\n+ assert!(\n+ error\n+ .to_string()\n+ .contains(\"templates are not supported in file inline references\"),\n+ \"unexpected error: {error}\",\n+ );\n+ }\n+}\ndiff --git a/lib/crates/fabro-workflow/src/transforms/file_inlining.rs b/lib/crates/fabro-workflow/src/transforms/file_inlining.rs\nindex 1e71699b1..942ba4af6 100644\n--- a/lib/crates/fabro-workflow/src/transforms/file_inlining.rs\n+++ b/lib/crates/fabro-workflow/src/transforms/file_inlining.rs\n@@ -192,33 +192,46 @@ impl FileInliningTransform {\n .with_inputs(self.inputs.clone());\n \n for (node_id, node) in &mut graph.nodes {\n- let Some(AttrValue::String(prompt)) = node.attrs.get(\"prompt\") else {\n- continue;\n- };\n- let target = TemplateRenderTarget::node_attr(\n- self.source_name.clone(),\n- node_id.clone(),\n- \"prompt\",\n- )\n- .with_source_origin(self.source_text.as_deref(), prompt)\n- .with_template_store(template_render_store(\n- &self.current_dir,\n- Arc::clone(&self.resolver),\n- self.source_name.as_deref(),\n- prompt,\n- )?);\n- let rendered = render_template_for_target(\n- prompt,\n- &ctx,\n- self.render_mode,\n- &target,\n- &mut diagnostics,\n- )?;\n- let value = self\n- .render_resolved_file_ref(&rendered, &ctx, target, &mut diagnostics)?\n- .unwrap_or(rendered);\n- node.attrs\n- .insert(\"prompt\".to_string(), AttrValue::String(value));\n+ for attr_name in [\"prompt\", \"output_schema\"] {\n+ let Some(AttrValue::String(attr_value)) = node.attrs.get(attr_name) else {\n+ continue;\n+ };\n+ let target = TemplateRenderTarget::node_attr(\n+ self.source_name.clone(),\n+ node_id.clone(),\n+ attr_name,\n+ )\n+ .with_source_origin(self.source_text.as_deref(), attr_value)\n+ .with_template_store(template_render_store(\n+ &self.current_dir,\n+ Arc::clone(&self.resolver),\n+ self.source_name.as_deref(),\n+ attr_value,\n+ )?);\n+ let rendered = render_template_for_target(\n+ attr_value,\n+ &ctx,\n+ self.render_mode,\n+ &target,\n+ &mut diagnostics,\n+ )?;\n+ let value = match self.render_resolved_file_ref(\n+ &rendered,\n+ &ctx,\n+ target,\n+ &mut diagnostics,\n+ )? {\n+ Some(value) => value,\n+ None if attr_name == \"output_schema\" && rendered.starts_with('@') => {\n+ return Err(Error::Validation(format!(\n+ \"node '{node_id}' output_schema has unresolved file reference: {rendered}\"\n+ )));\n+ }\n+ None => rendered,\n+ };\n+ node.attrs\n+ .insert(attr_name.to_string(), AttrValue::String(value));\n+ }\n }\n \n Ok((graph, diagnostics))\n@@ -465,6 +478,90 @@ mod tests {\n );\n }\n \n+ #[test]\n+ fn file_inlining_transform_inlines_output_schema_reference() {\n+ let dir = tempfile::tempdir().unwrap();\n+ std::fs::create_dir_all(dir.path().join(\"schemas\")).unwrap();\n+ std::fs::write(\n+ dir.path().join(\"schemas/audit-result.schema.json\"),\n+ r#\"{\"type\":\"object\",\"required\":[\"passed\"]}\"#,\n+ )\n+ .unwrap();\n+\n+ let mut graph = Graph::new(\"test\");\n+ let mut node = Node::new(\"audit\");\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(\"@schemas/audit-result.schema.json\".to_string()),\n+ );\n+ graph.nodes.insert(\"audit\".to_string(), node);\n+\n+ let transform = FileInliningTransform::new(\n+ dir.path().to_path_buf(),\n+ Arc::new(FilesystemFileResolver::new(None)),\n+ );\n+ let graph = transform.apply(graph).unwrap();\n+\n+ assert_eq!(\n+ graph.nodes[\"audit\"]\n+ .attrs\n+ .get(\"output_schema\")\n+ .and_then(AttrValue::as_str),\n+ Some(r#\"{\"type\":\"object\",\"required\":[\"passed\"]}\"#)\n+ );\n+ }\n+\n+ #[test]\n+ fn file_inlining_transform_leaves_routing_output_schema_keyword_unchanged() {\n+ let dir = tempfile::tempdir().unwrap();\n+ let mut graph = Graph::new(\"test\");\n+ let mut node = Node::new(\"route\");\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(\"routing\".to_string()),\n+ );\n+ graph.nodes.insert(\"route\".to_string(), node);\n+\n+ let transform = FileInliningTransform::new(\n+ dir.path().to_path_buf(),\n+ Arc::new(FilesystemFileResolver::new(None)),\n+ );\n+ let graph = transform.apply(graph).unwrap();\n+\n+ assert_eq!(\n+ graph.nodes[\"route\"]\n+ .attrs\n+ .get(\"output_schema\")\n+ .and_then(AttrValue::as_str),\n+ Some(\"routing\")\n+ );\n+ }\n+\n+ #[test]\n+ fn file_inlining_transform_reports_unresolved_output_schema_reference() {\n+ let dir = tempfile::tempdir().unwrap();\n+ let mut graph = Graph::new(\"test\");\n+ let mut node = Node::new(\"audit\");\n+ node.attrs.insert(\n+ \"output_schema\".to_string(),\n+ AttrValue::String(\"@schemas/missing.schema.json\".to_string()),\n+ );\n+ graph.nodes.insert(\"audit\".to_string(), node);\n+\n+ let transform = FileInliningTransform::new(\n+ dir.path().to_path_buf(),\n+ Arc::new(FilesystemFileResolver::new(None)),\n+ );\n+ let error = transform.apply(graph).unwrap_err();\n+\n+ assert!(\n+ error.to_string().contains(\n+ \"node 'audit' output_schema has unresolved file reference: @schemas/missing.schema.json\"\n+ ),\n+ \"unexpected error: {error}\",\n+ );\n+ }\n+\n #[test]\n fn file_inlining_transform_resolves_minijinja_includes_for_prompts_and_goal() {\n let dir = tempfile::tempdir().unwrap();\n", + "summary": { + "files_changed": 14, + "additions": 1766, + "deletions": 263 + } + } + }, + { + "seq": 0, + "checkpoint": { + "timestamp": "2026-05-23T20:42:41.068135Z", + "current_node": "simplify_opus", + "completed_nodes": [ + "start", + "toolchain", + "preflight_compile", + "preflight_lint", + "implement", + "simplify_opus" + ], + "node_retries": {}, + "context_values": { + "internal.retry_count.implement": 0, + "failure_class": "", + "internal.run_id": "01KSB6GTZ00T5V6BNMXN3SPKZF", + "thread.toolchain.current_node": "preflight_compile", + "current_node": "simplify_opus", + "thread.preflight_compile.current_node": "preflight_lint", + "thread.preflight_lint.current_node": "implement", + "internal.retry_count.preflight_lint": 0, "response.implement": "Implemented output schema validation end-to-end.\n\nSummary:\n- Added `output_schema` / `output_retries` node attrs.\n- Added `output_schema` static/file reference handling and inlining.\n- Added `handler::structured_output` for:\n - routing/custom schema parsing\n - balanced JSON extraction\n - JSON Schema validation\n - repair prompts\n - routing/custom context application\n- Integrated strict routing/custom output into agent and prompt handlers.\n- Added same-context repair for:\n - prompt API calls via mutable message history\n - agent API sessions via the same live `Session`\n- Added ACP guardrail:\n - `output_schema is not supported with backend=\"acp\" in this release`\n - test verifies no process launch\n- Added custom output storage at `output.{node_id}`.\n- Updated docs for routing/custom schemas, `output_retries`, repair behavior, ACP limitation.\n\nValidation run:\n- `cargo nextest run -p fabro-types -p fabro-workflow graph:: file_inlining static_reference` ✅\n- `cargo nextest run -p fabro-workflow` ✅\n- `cargo +nightly-2026-04-14 fmt --check --all` ✅\n- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` ✅\n- `cargo insta pending-snapshots` unavailable in this environment (`cargo-insta` not installed); checked for `**/*.snap.new` and found none.", "graph.model_stylesheet": "\n * { model: claude-opus-4-7; }\n ", "internal.work_dir": "/home/daytona/workspace/fabro", "internal.fidelity": "compact", - "last_response": "Implemented output schema validation end-to-end.\n\nSummary:\n- Added `output_schema` / `output_retries` node attrs.\n- Added `output_schema` static/file reference handling and inlining.\n- Added `handler:", + "last_response": "## Summary\n\nThree review agents flagged ~25 findings; I applied the high-impact, low-risk fixes (about 12 lines net reduction across the touched files, plus correctness wins on schema compilation).\n\n*", + "last_stage": "simplify_opus", + "internal.thread_id": "implement", + "failure_signature": "", + "graph.goal": "# Output Schema Validation Implementation Plan\n\n> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.\n\n**Goal:** Add `output_schema` validation for agent and prompt nodes, with context-preserving repair turns when structured output does not validate.\n\n**Architecture:** Introduce a small structured-output layer in `fabro-workflow` that resolves node-level schema declarations, extracts JSON output, validates it, and produces either routing side effects or a parsed custom output context update. Agent and prompt execution must perform schema repair inside the active LLM conversation instead of using the workflow executor retry path.\n\n**Tech Stack:** Rust, Graphviz workflow attrs, `serde_json`, workspace `jsonschema`, existing `fabro-llm::ResponseFormat`, Fabro agent sessions, `cargo nextest`.\n\n---\n\n## Public Interface\n\nWorkflow authors can opt in on agent and prompt nodes:\n\n```dot\nreview [\n shape=tab,\n output_schema=\"routing\",\n output_retries=2\n]\n\naudit [\n shape=tab,\n output_schema=\"@schemas/audit-result.schema.json\",\n output_retries=2\n]\n```\n\n- `output_schema=\"routing\"` uses Fabro's built-in routing directive schema.\n- `output_schema=\"@path/to/schema.json\"` loads a JSON Schema file through existing workflow file-reference rules.\n- `output_retries` controls corrective turns inside the same node execution. Default: `2`. `0` means validate once and fail without a repair turn.\n- Schema failures are terminal node failures after `output_retries` is exhausted. They are not `retry_requested` outcomes and do not consume `max_retries`.\n- `backend=\"acp\"` with `output_schema` is unsupported in v1 and returns a clear validation error.\n\n## Implementation Tasks\n\n### Task 1: Node Attributes And File Reference Resolution\n\n**Files:**\n- Modify: `lib/crates/fabro-types/src/graph.rs`\n- Modify: `lib/crates/fabro-workflow/src/static_reference.rs`\n- Modify: `lib/crates/fabro-workflow/src/transforms/file_inlining.rs`\n- Test: existing unit tests in those files\n\n- [ ] Add `Node::output_schema(&self) -> Option<&str>` next to other agent/prompt attrs.\n- [ ] Add `Node::output_retries(&self) -> i64` returning `self.int_attr(\"output_retries\").unwrap_or(2).max(0)`.\n- [ ] Teach static reference validation that node attr `output_schema` values starting with `@` are file inline references.\n- [ ] Extend file inlining so `output_schema=\"@schemas/foo.json\"` is replaced with the schema file contents before execution, while `output_schema=\"routing\"` stays unchanged.\n- [ ] Add tests for absent attrs, default retries, zero retries, file inlining, and unresolved schema reference diagnostics.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-types -p fabro-workflow graph:: file_inlining static_reference\n```\n\nExpected: targeted tests pass.\n\n### Task 2: Structured Output Module\n\n**Files:**\n- Create: `lib/crates/fabro-workflow/src/handler/structured_output.rs`\n- Modify: `lib/crates/fabro-workflow/src/handler/mod.rs`\n- Modify: `lib/crates/fabro-workflow/Cargo.toml`\n- Test: unit tests in `structured_output.rs`\n\n- [ ] Add `jsonschema.workspace = true` to `fabro-workflow` dependencies.\n- [ ] Define `OutputSchemaKind` with `Routing` and `JsonSchema { schema: serde_json::Value }`.\n- [ ] Parse `node.output_schema()` into `None`, `Routing`, or custom JSON Schema. Treat literal `routing` as the only built-in keyword.\n- [ ] Add a built-in routing schema requiring an object with at least one recognized field: `preferred_next_label`, `outcome`, `failure_reason`, `suggested_next_ids`, or `context_updates`.\n- [ ] Reuse balanced-object scanning semantics for response text: validate the last JSON object that is relevant to the selected schema.\n- [ ] Return a structured validation result containing the parsed JSON object, concise error messages, and enough information to build a repair prompt.\n- [ ] Add tests for valid routing JSON, missing routing fields, wrong routing field types, valid custom schema, invalid custom schema, invalid JSON, and no JSON object.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow structured_output\n```\n\nExpected: structured-output unit tests pass.\n\n### Task 3: Routing Extraction Compatibility\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/agent.rs`\n- Test: existing agent handler unit tests\n\n- [ ] Keep the loose default unchanged when `output_schema` is absent.\n- [ ] Move current `STATUS_FIELDS`, balanced JSON scanning, and routing-field application behind reusable functions in `structured_output.rs` or call the new module from `agent.rs`.\n- [ ] For `output_schema=\"routing\"`, require schema-valid routing JSON and surface validation failures for repair instead of silently ignoring bad candidates.\n- [ ] Preserve existing routing fallback priority for agent nodes: response text first, then `status.json`, then last file touched.\n- [ ] Keep prompt-node routing behavior response-only unless later tasks explicitly add prompt `status.json` support.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::agent\n```\n\nExpected: existing loose routing tests still pass, plus new strict routing tests pass.\n\n### Task 4: Prompt Node Same-Context Repair\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/llm/api.rs`\n- Modify: `lib/crates/fabro-workflow/src/handler/prompt.rs`\n- Test: prompt/API backend tests in those files\n\n- [ ] In `AgentApiBackend::one_shot`, keep `messages` mutable across attempts.\n- [ ] When a prompt node has a custom JSON Schema, set `response_format=JsonSchema` on the initial and repair LLM requests. For `routing`, use `JsonObject` or no provider-native schema if provider behavior would conflict with Fabro's routing extraction.\n- [ ] After each LLM response, validate according to `output_schema`.\n- [ ] On validation failure with repair attempts remaining, append `Message::assistant(response.text())`, then append a corrective `Message::user(repair_message)`, and call `client.complete` again with the same messages.\n- [ ] On success, return the validated response text and aggregate usage across all attempts.\n- [ ] On exhaustion, return a terminal failed outcome with failure reason `output schema validation failed after N repair attempt(s)`.\n- [ ] Update `PromptHandler` so validated custom output is added to `context_updates[\"output.{node_id}\"]`; routing output still updates outcome routing fields.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::prompt handler::llm::api\n```\n\nExpected: prompt repair keeps previous assistant output in the message list and succeeds after a corrective response.\n\n### Task 5: Agent Node Same-Session Repair\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/llm/api.rs`\n- Modify: `lib/crates/fabro-workflow/src/handler/agent.rs`\n- Test: agent/API backend tests in those files\n\n- [ ] In `AgentApiBackend::run`, validate the final assistant response before releasing, closing, or caching the session.\n- [ ] On validation failure with repair attempts remaining, call `session.process_input(repair_message)` on the same `Session`.\n- [ ] Recompute the final assistant response after each repair turn from `session.history()`.\n- [ ] Aggregate usage across all new assistant turns, including repair turns, without double-counting reused session history.\n- [ ] Do not set provider-native `response_format` for agent sessions in v1, because agent sessions may need normal tool-use messages before final output.\n- [ ] Return terminal failure after exhaustion; do not return a retryable backend error and do not request workflow node retry.\n- [ ] Update `AgentHandler` to apply validated routing/custom output to the final `Outcome`.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::agent handler::llm::api\n```\n\nExpected: agent repair sends a second `process_input` to the same session and final validated output drives outcome/context updates.\n\n### Task 6: ACP Guardrail\n\n**Files:**\n- Modify: `lib/crates/fabro-workflow/src/handler/llm/acp.rs`\n- Test: ACP backend tests in that file\n\n- [ ] At the start of `AgentAcpBackend::run`, reject nodes where `node.output_schema().is_some()`.\n- [ ] Use a clear error message: `output_schema is not supported with backend=\"acp\" in this release`.\n- [ ] Add a test proving the ACP backend does not launch a process when `output_schema` is present.\n\nRun:\n\n```bash\ncargo nextest run -p fabro-workflow handler::llm::acp\n```\n\nExpected: ACP guardrail test passes.\n\n### Task 7: Docs\n\n**Files:**\n- Modify: `docs/public/agents/outputs.mdx`\n- Modify: `docs/public/reference/dot-language.mdx`\n\n- [ ] Document `output_schema=\"routing\"` and `output_schema=\"@schema.json\"` under routing/structured outputs.\n- [ ] Document same-context repair behavior explicitly: Fabro sends validation feedback to the same agent/prompt context before failing.\n- [ ] Document `output_retries`, default `2`, and distinction from `max_retries`.\n- [ ] Document v1 scope: agent/prompt nodes only; ACP unsupported; custom schema output stored at `output.{node_id}`.\n\nRun:\n\n```bash\nrg -n \"output_schema|output_retries|output\\\\.\" docs/public/agents/outputs.mdx docs/public/reference/dot-language.mdx\n```\n\nExpected: docs mention the new attrs and storage behavior.\n\n### Task 8: Full Verification\n\n**Files:**\n- No new files beyond prior tasks\n\n- [ ] Run focused workflow tests:\n\n```bash\ncargo nextest run -p fabro-workflow\n```\n\n- [ ] Run formatting check:\n\n```bash\ncargo +nightly-2026-04-14 fmt --check --all\n```\n\n- [ ] Run clippy:\n\n```bash\ncargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings\n```\n\n- [ ] If snapshots change, inspect before accepting:\n\n```bash\ncargo insta pending-snapshots\n```\n\nOnly run `cargo insta accept` after verifying every pending snapshot is expected.\n\n## Acceptance Criteria\n\n- Existing workflows without `output_schema` behave exactly as before.\n- `output_schema=\"routing\"` prevents malformed/missing routing JSON from silently falling through to normal edge selection.\n- Invalid structured output results in a corrective LLM turn in the same context window.\n- Prompt repair preserves previous assistant output in the message list.\n- Agent repair preserves the same live session and does not re-run the node from scratch.\n- Exhausted output repair attempts produce a clear terminal failure.\n- Custom schema output is available to downstream nodes at `output.{node_id}`.\n- Docs clearly distinguish `output_retries` from `max_retries`.\n\n## Assumptions\n\n- `output_retries=2` is the default.\n- Custom schema validation targets the final JSON object in the response text.\n- `status.json` fallback remains routing-specific.\n- Provider-native response schema is used for prompt nodes only where it is safe.\n- ACP support can be added later after there is a guaranteed context-preserving repair mechanism.\n", + "internal.retry_count.toolchain": 0, + "thread.implement.current_node": "simplify_opus", + "internal.retry_count.start": 0, + "graph.rankdir": "LR", + "thread.start.current_node": "toolchain", + "internal.node_visit_count": 1, + "command.output": "blob://sha256/12ae32cb1ec02d01eda3581b127c1fee3b0dc53572ed6baf239721a03d82e126", + "internal.retry_count.preflight_compile": 0, + "internal.retry_count.simplify_opus": 0, "outcome": "succeeded", - "last_stage": "implement" + "response.simplify_opus": "## Summary\n\nThree review agents flagged ~25 findings; I applied the high-impact, low-risk fixes (about 12 lines net reduction across the touched files, plus correctness wins on schema compilation).\n\n**Fixed:**\n1. **Schema compiled twice per node** (efficiency #1, reuse #3, quality #1) — `OutputSchemaKind::JsonSchema` now stores an `Arc` compiled once in `parse_node_output_schema`; the routing validator is a `LazyLock`, so repair iterations no longer recompile. Also removed the redundant `compile_schema` helper.\n2. **`read_sandbox_file` reimplemented `Sandbox::read_file_text`** (reuse #2) — replaced the `cat ` exec_command call with the existing trait method; dropped the unused `shell_quote` import and the magic 5s timeout.\n3. **Redundant re-validation** in `validate_agent_output_sources` (efficiency #6) — bound the first error and reused it instead of calling `validate_response_text` again just to `.expect_err(...)`.\n4. **`OneShotCompletion::actual_speed` was always `controls.speed`** (quality #2, reuse #4) — collapsed the three `actual_*` fields into a single `model: ModelRef`, removing the manual reconstruction one line later.\n5. **`extract_status_fields` wrapper deletion** (reuse #1, quality #5) — renamed `extract_status_fields_loose` → `extract_status_fields` in `structured_output.rs`; agent.rs re-exports it so existing call sites (including `prompt.rs`) work unchanged.\n6. **`find_json_objects` visibility** (quality #6) — tightened from `pub(crate)` to `fn`.\n7. **`last_touched_file` cheap accessor** (efficiency #5) — added a tiny helper that grabs only `.last`; the repair loop no longer clones+sorts the full `touched` set each iteration.\n8. **`ROUTING_KEYWORD` constant** (quality #8) — extracted the literal `\"routing\"` to a const used by `parse_node_output_schema`.\n\n**Skipped:**\n- Folding `Error::OutputSchemaValidation` into `Error::Validation` — `Validation` adds a `\"Validation error: \"` prefix that would change user-visible failure messages and downstream event content. Not worth the drift for one extra enum variant.\n- Other suggestions (request structs for multi-param functions, deduping the two repair loops across `one_shot`/`run`, moving `validate_agent_output_sources` into `structured_output.rs`, collapsing `StructuredOutputErrorKind`) — judged as larger refactors with no clear correctness payoff for this iteration.\n\n**Verification:** `cargo nextest run -p fabro-types -p fabro-workflow` → 1461 tests pass; `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` clean; `cargo +nightly-2026-04-14 fmt --check --all` clean." }, "node_outcomes": { "start": { @@ -840,10 +966,48 @@ }, "notes": "Script completed: cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", "usage": null + }, + "simplify_opus": { + "status": "succeeded", + "context_updates": { + "last_stage": "simplify_opus", + "last_response": "## Summary\n\nThree review agents flagged ~25 findings; I applied the high-impact, low-risk fixes (about 12 lines net reduction across the touched files, plus correctness wins on schema compilation).\n\n*", + "response.simplify_opus": "## Summary\n\nThree review agents flagged ~25 findings; I applied the high-impact, low-risk fixes (about 12 lines net reduction across the touched files, plus correctness wins on schema compilation).\n\n**Fixed:**\n1. **Schema compiled twice per node** (efficiency #1, reuse #3, quality #1) — `OutputSchemaKind::JsonSchema` now stores an `Arc` compiled once in `parse_node_output_schema`; the routing validator is a `LazyLock`, so repair iterations no longer recompile. Also removed the redundant `compile_schema` helper.\n2. **`read_sandbox_file` reimplemented `Sandbox::read_file_text`** (reuse #2) — replaced the `cat ` exec_command call with the existing trait method; dropped the unused `shell_quote` import and the magic 5s timeout.\n3. **Redundant re-validation** in `validate_agent_output_sources` (efficiency #6) — bound the first error and reused it instead of calling `validate_response_text` again just to `.expect_err(...)`.\n4. **`OneShotCompletion::actual_speed` was always `controls.speed`** (quality #2, reuse #4) — collapsed the three `actual_*` fields into a single `model: ModelRef`, removing the manual reconstruction one line later.\n5. **`extract_status_fields` wrapper deletion** (reuse #1, quality #5) — renamed `extract_status_fields_loose` → `extract_status_fields` in `structured_output.rs`; agent.rs re-exports it so existing call sites (including `prompt.rs`) work unchanged.\n6. **`find_json_objects` visibility** (quality #6) — tightened from `pub(crate)` to `fn`.\n7. **`last_touched_file` cheap accessor** (efficiency #5) — added a tiny helper that grabs only `.last`; the repair loop no longer clones+sorts the full `touched` set each iteration.\n8. **`ROUTING_KEYWORD` constant** (quality #8) — extracted the literal `\"routing\"` to a const used by `parse_node_output_schema`.\n\n**Skipped:**\n- Folding `Error::OutputSchemaValidation` into `Error::Validation` — `Validation` adds a `\"Validation error: \"` prefix that would change user-visible failure messages and downstream event content. Not worth the drift for one extra enum variant.\n- Other suggestions (request structs for multi-param functions, deduping the two repair loops across `one_shot`/`run`, moving `validate_agent_output_sources` into `structured_output.rs`, collapsing `StructuredOutputErrorKind`) — judged as larger refactors with no clear correctness payoff for this iteration.\n\n**Verification:** `cargo nextest run -p fabro-types -p fabro-workflow` → 1461 tests pass; `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` clean; `cargo +nightly-2026-04-14 fmt --check --all` clean." + }, + "notes": "Stage completed: simplify_opus", + "usage": { + "input": { + "usage": { + "model": { + "provider": "anthropic", + "model_id": "claude-opus-4-7" + }, + "tokens": { + "input_tokens": 141487, + "output_tokens": 35894, + "reasoning_tokens": 0, + "cache_read_tokens": 8125892, + "cache_write_tokens": 684904 + } + }, + "facts": { + "algorithm": "anthropic", + "cache_write_5m_tokens": 684904, + "cache_write_1h_tokens": 0 + } + }, + "total_usd_micros": 9948381 + }, + "files_touched": [ + "/home/daytona/workspace/fabro/lib/crates/fabro-workflow/src/handler/agent.rs", + "/home/daytona/workspace/fabro/lib/crates/fabro-workflow/src/handler/llm/api.rs", + "/home/daytona/workspace/fabro/lib/crates/fabro-workflow/src/handler/structured_output.rs" + ] } }, - "next_node_id": "simplify_opus", + "next_node_id": "simplify_gpt", "node_visits": { + "simplify_opus": 1, "implement": 1, "preflight_compile": 1, "start": 1, @@ -1022,7 +1186,12 @@ "first_event_seq": 51, "prompt": null, "response": null, - "completion": null, + "completion": { + "outcome": "succeeded", + "notes": "Stage completed: implement", + "failure_reason": null, + "timestamp": "2026-05-23T20:26:00.319401Z" + }, "provider_used": { "mode": "agent", "provider": "openai", @@ -1036,6 +1205,12 @@ "output": null, "started_at": "2026-05-23T20:01:00.772435Z", "handler": "agent", + "timing": { + "wall_time_ms": 1499532, + "inference_time_ms": 0, + "tool_time_ms": 0, + "active_time_ms": 0 + }, "usage": { "input_tokens": 329828, "output_tokens": 33087, @@ -1101,6 +1276,132 @@ } } ], + "state": "succeeded" + }, + "simplify_opus@1": { + "first_event_seq": 739, + "prompt": null, + "response": null, + "completion": null, + "provider_used": { + "mode": "agent", + "provider": "anthropic", + "model": "claude-opus-4-7" + }, + "diff": null, + "script_invocation": null, + "script_timing": null, + "parallel_results": null, + "output": null, + "started_at": "2026-05-23T20:26:04.064143Z", + "handler": "agent", + "usage": { + "input_tokens": 141487, + "output_tokens": 35894, + "total_tokens": 8988177, + "reasoning_tokens": 0, + "cache_read_tokens": 8125892, + "cache_write_tokens": 684904, + "total_usd_micros": 9948381 + }, + "model": { + "provider": "anthropic", + "model_id": "claude-opus-4-7" + }, + "todos": { + "kind": "anthropic_tasks", + "list_id": "anthropic_tasks:b0219b28-093f-4ecf-b204-fd0ef5b1a91e", + "items": [ + { + "id": "1", + "status": "completed", + "order": 0, + "subject": "Cache compiled jsonschema::Validator in OutputSchemaKind", + "description": "Today parse_node_output_schema compiles the schema only to discard it, and validate_value_against_schema recompiles it every call (including on each repair turn). Store an Arc in OutputSchemaKind::JsonSchema so compilation happens once per node.", + "active_form": "Caching compiled jsonschema::Validator in OutputSchemaKind" + }, + { + "id": "2", + "status": "completed", + "order": 1, + "subject": "Use Sandbox::read_file_text instead of exec_command(\"cat ...\")", + "description": "Replace read_sandbox_file's exec_command(\"cat …\") implementation in agent.rs with the existing Sandbox::read_file_text trait method. Removes shell quoting concerns, a 5s magic timeout, and an unused fabro_agent::shell_quote import.", + "active_form": "Using Sandbox::read_file_text for status.json / last-file routing fallback" + }, + { + "id": "3", + "status": "completed", + "order": 2, + "subject": "Reuse first error in validate_agent_output_sources fallback", + "description": "Today the fallback chain discards the initial routing error, then re-runs validate_response_text on the response text purely to reconstruct it via .expect_err. Bind the first error and reuse it.", + "active_form": "Reusing first validation error instead of re-running validation" + }, + { + "id": "4", + "status": "completed", + "order": 3, + "subject": "Drop OneShotCompletion::actual_speed and inline through ModelRef", + "description": "actual_speed always equals controls.speed in every branch; the struct's three actual_* fields are reassembled into a ModelRef one line later. Replace the three fields with a single ModelRef and read speed from controls at the call site.", + "active_form": "Simplifying OneShotCompletion to hold ModelRef directly" + }, + { + "id": "5", + "status": "completed", + "order": 4, + "subject": "Make structured_output::find_json_objects private", + "description": "find_json_objects is only used inside structured_output.rs (no external pub(crate) call sites).", + "active_form": "Tightening structured_output::find_json_objects visibility" + }, + { + "id": "6", + "status": "completed", + "order": 5, + "subject": "Replace agent.rs::extract_status_fields wrapper with direct call", + "description": "extract_status_fields is now a one-line forwarder to structured_output::extract_status_fields_loose. Delete the wrapper and have callers (agent.rs, prompt.rs) use the canonical name.", + "active_form": "Deleting redundant extract_status_fields wrapper" + }, + { + "id": "8", + "status": "completed", + "order": 7, + "subject": "Cheap last_touched accessor for repair loop", + "description": "file_tracking_snapshot clones the full touched HashSet and sorts a Vec just to read .last. In the repair loop only .last is needed — add a `last_touched_file` helper that returns Option without copying touched.", + "active_form": "Adding cheap last_touched_file accessor for repair loop" + } + ] + }, + "subagents": [ + { + "agent_id": "2bd89a0f", + "depth": 1, + "task": "You are reviewing a code change in the Fabro repo (Rust). The change adds an `output_schema` validation feature to workflow agent/prompt nodes with same-context repair turns.\n\n**Files changed:**\n- NEW: `lib/crates/fabro-workflow/src/handler/structured_output.rs` (608 lines, validation/parsing/routing application)\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/agent.rs` (refactored to use structured_output, added `validate_agent_output_sources`)\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/llm/api.rs` (added repair loops in `one_shot` and `run`, helpers `file_tracking_snapshot` and `last_assistant_response`, extracted `complete_one_shot_request` and `OneShotCompletion` struct)\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/prompt.rs` (validation/repair on prompt handler)\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/llm/acp.rs` (guardrail)\n- MODIFIED: `lib/crates/fabro-types/src/graph.rs` (added `output_schema()` and `output_retries()`)\n- MODIFIED: `lib/crates/fabro-workflow/src/transforms/file_inlining.rs` (loops over prompt + output_schema)\n- MODIFIED: `lib/crates/fabro-workflow/src/static_reference.rs` (recognizes `output_schema` `@` refs)\n- MODIFIED: `lib/crates/fabro-workflow/src/error.rs` (added `OutputSchemaValidation` variant)\n\nStart by running `git diff HEAD~1 -- ` to inspect each file. The previous commit is `190065dc1`, the new commit is `43b8310e6`.\n\n**Your task — CODE REUSE REVIEW:**\n\nFor each change:\n\n1. Search for existing utilities and helpers that could replace newly written code. Use grep/glob to find similar patterns elsewhere — common locations: `fabro-util`, `fabro-llm`, `fabro-graphviz`, `fabro-agent`, sibling files in `handler/`. In particular check:\n - Is there an existing JSON object extractor / balanced-brace scanner anywhere? `find_json_objects` is hand-rolled in structured_output.rs.\n - Is there already a routing-field application helper? The new `apply_routing_fields` may duplicate logic elsewhere.\n - The existing `extract_status_fields` in `agent.rs` was moved/wrapped — make sure the wrapper is actually necessary, or whether call sites can use the new function directly.\n - Is `shell_quote` from `fabro-agent` (used in `validate_agent_output_sources`) the same thing as the now-deleted `shlex::try_quote` call? Check.\n - Schema compilation: `jsonschema::validator_for(schema)` is called both inside `parse_node_output_schema` (just to validate) and inside `validate_value_against_schema` (to actually use). Are we compiling the schema twice per node run?\n - `read_sandbox_file` helper in agent.rs — is there an existing sandbox helper for reading a file?\n\n2. Flag any new function that duplicates existing functionality. Suggest the existing function to use instead.\n\n3. Flag any inline logic that could use an existing utility.\n\nNote: This is a greenfield app, so focus on maximizing simplicity. Be aggressive in suggesting consolidation. Only report findings; do not edit files.\n\nReport findings as a concise numbered list with file:line references and brief recommendations.\n", + "status": { + "kind": "completed", + "success": true, + "turns_used": 115 + } + }, + { + "agent_id": "64539ee5", + "depth": 1, + "task": "You are reviewing a code change in the Fabro repo (Rust). The change adds an `output_schema` validation feature to workflow agent/prompt nodes with same-context repair turns.\n\n**Files changed:**\n- NEW: `lib/crates/fabro-workflow/src/handler/structured_output.rs` (608 lines)\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/agent.rs` (added `validate_agent_output_sources`, wraps `extract_status_fields_loose`)\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/llm/api.rs` (added repair loops in `one_shot` and `run`, helpers `file_tracking_snapshot` and `last_assistant_response`, extracted `complete_one_shot_request` and `OneShotCompletion` struct)\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/prompt.rs` (validation/repair on prompt handler)\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/llm/acp.rs` (guardrail)\n- MODIFIED: `lib/crates/fabro-types/src/graph.rs` (added `output_schema()` and `output_retries()`)\n- MODIFIED: `lib/crates/fabro-workflow/src/transforms/file_inlining.rs`\n- MODIFIED: `lib/crates/fabro-workflow/src/static_reference.rs`\n- MODIFIED: `lib/crates/fabro-workflow/src/error.rs`\n\nStart by running `git diff HEAD~1 -- ` to inspect each file. Previous commit is `190065dc1`, new commit is `43b8310e6`. Also read the actual files to see the full context.\n\n**Your task — CODE QUALITY REVIEW:**\n\nLook for hacky patterns:\n\n1. **Redundant state**: state that duplicates existing state, cached values that could be derived. E.g.:\n - `StructuredOutputErrorKind` + `StructuredOutputError` — is the kind needed at all? Could the error variants encode the data they need directly? Look at how `allows_routing_fallback()` and `repair_message()` are consumed.\n - `OneShotCompletion` struct has 4 fields including `actual_*` — was the old tuple really worth replacing with a named struct?\n - The `ROUTING_STATUS_FIELDS` constant + the schema's `anyOf` rebuild — could these be derived from a single source?\n\n2. **Parameter sprawl**: adding new parameters to a function instead of generalizing or restructuring.\n - `validate_agent_output_sources` takes 4 params; would a small request struct be cleaner?\n - `complete_one_shot_request` takes 7 params — consider whether this is reasonable.\n\n3. **Copy-paste with slight variation**: \n - The repair loop logic appears twice (`one_shot` and `run`) — could it be unified?\n - The two error-disposition match arms inside the `run` repair loop are nearly identical to other dispositions in the same function.\n - Tests `mock_llm_catalog`, `mock_api_backend`, `chat_completion_response`, `chat_completion_stream` — are there existing helpers? Are the new ones duplicating sibling test fixtures?\n\n4. **Leaky abstractions**: \n - `extract_status_fields_loose` is pub(crate) — but `extract_status_fields` is also pub(crate) and just wraps it. Why both?\n - `find_json_objects` is pub(crate) — is it actually used outside `structured_output.rs`?\n - `validate_agent_output_sources` lives in `agent.rs` but most of its logic is in `structured_output.rs`. Should it move?\n\n5. **Stringly-typed code**: \n - The literal `\"routing\"` keyword appears in multiple places (`structured_output.rs`, `static_reference.rs` test, possibly elsewhere) — should be a constant.\n - Attr names like `\"output_schema\"`, `\"output_retries\"` appear repeatedly — already a pattern, but check if there's a convention to follow.\n\n6. **Other concerns**:\n - Is `repair_message` a method on the error, or should it be a method on `OutputSchemaKind`? It primarily uses the schema; the error's contribution is just the message strings.\n - The `validation_error` Option-with-tuple pattern in `one_shot` repair loop — does that read well?\n - `allows_routing_fallback()` — is the kind/messages split clean here, or could fallback be a separate sentinel?\n\nNote: This is a greenfield app, so be aggressive in suggesting cleanups. Only report findings; do not edit files.\n\nReport findings as a concise numbered list with file:line references and brief recommendations.\n", + "status": { + "kind": "completed", + "success": true, + "turns_used": 41 + } + }, + { + "agent_id": "5b296f13", + "depth": 1, + "task": "You are reviewing a code change in the Fabro repo (Rust). The change adds an `output_schema` validation feature to workflow agent/prompt nodes with same-context repair turns.\n\n**Files changed:**\n- NEW: `lib/crates/fabro-workflow/src/handler/structured_output.rs` (608 lines, validation/parsing/routing application)\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/agent.rs` (added `validate_agent_output_sources`)\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/llm/api.rs` (added repair loops in `one_shot` and `run`)\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/prompt.rs`\n- MODIFIED: `lib/crates/fabro-workflow/src/handler/llm/acp.rs`\n\nStart by running `git diff HEAD~1 -- ` to inspect each file. Previous commit is `190065dc1`, new commit is `43b8310e6`.\n\n**Your task — EFFICIENCY REVIEW:**\n\nLook for:\n\n1. **Unnecessary work / repeated computation**:\n - `parse_node_output_schema` calls `compile_schema` to *validate* the schema, but never caches it. Then `validate_value_against_schema` calls `compile_schema` *again* every time a response is validated, including on each repair attempt. Should we compile once and reuse?\n - Inside the repair loop, `parse_node_output_schema` is called once but `validate_response_text` does fresh JSON object extraction + schema compilation each iteration.\n - `find_json_objects` builds a full Vec for every validation, even when only the last object is needed (custom schema path) — could short-circuit.\n\n2. **Missed concurrency**: \n - Is the file_tracking_snapshot reasonably called?\n\n3. **Hot-path bloat**:\n - The JSON object extraction in `find_json_objects` runs over every response. Acceptable for response size; check it's not double-run.\n - `last_assistant_response` iterates the session history every time it's called. In the repair loop it's called twice per iteration (initial, then after each repair turn).\n\n4. **Memory**:\n - `messages.clone()` in `one_shot` repair loop — clones the entire message history each iteration. Acceptable if Vec is small but check.\n - `validator.iter_errors(value)` collects errors via `.take(5)` — bounded, OK.\n\n5. **TOCTOU / unnecessary existence checks**:\n - `read_sandbox_file` checks `result.is_success()` after running `cat`. Acceptable.\n\n6. **Overly broad operations**:\n - `find_json_objects` returns `Vec<&str>` of all matches when sometimes only the last is needed.\n\nNote: Only report findings; do not edit files. Report findings as a concise numbered list with file:line references and brief recommendations.\n", + "status": { + "kind": "completed", + "success": true, + "turns_used": 41 + } + } + ], "state": "running" }, "start@1": { diff --git a/stages/005-implement@1/diff.patch b/stages/005-implement@1/diff.patch new file mode 100644 index 000000000..b3adc7385 --- /dev/null +++ b/stages/005-implement@1/diff.patch @@ -0,0 +1,2432 @@ +diff --git a/Cargo.lock b/Cargo.lock +index c50b6c665..28b5f81e8 100644 +--- a/Cargo.lock ++++ b/Cargo.lock +@@ -2603,6 +2603,7 @@ dependencies = [ + "git2", + "hex", + "httpmock", ++ "jsonschema", + "md5", + "miette", + "mime_guess", +diff --git a/docs/public/agents/outputs.mdx b/docs/public/agents/outputs.mdx +index 1407f4db6..7467f050b 100644 +--- a/docs/public/agents/outputs.mdx ++++ b/docs/public/agents/outputs.mdx +@@ -64,6 +64,23 @@ The test coverage is below the threshold. + + JSON objects without recognized fields are ignored. + ++### Validated routing output ++ ++Set `output_schema="routing"` on an agent or prompt node to require Fabro's built-in routing directive schema: ++ ++```dot ++review [ ++ shape=tab, ++ prompt="Review the implementation and return routing JSON.", ++ output_schema="routing", ++ output_retries=2 ++] ++``` ++ ++With `output_schema="routing"`, the routing JSON must be an object with at least one recognized routing field (`preferred_next_label`, `outcome`, `failure_reason`, `suggested_next_ids`, or `context_updates`) and those fields must have the expected types. Malformed routing JSON fails validation instead of being silently ignored. ++ ++Fabro repairs invalid structured output inside the same LLM context before failing the node. For prompt nodes, Fabro appends the invalid assistant response and a corrective user message to the same message list. For agent nodes using the API backend, Fabro sends the corrective message to the same live agent session. `output_retries` controls these repair turns and defaults to `2`; `output_retries=0` validates once and fails without a repair turn. These repair turns are separate from workflow `max_retries` and do not consume node retry attempts. ++ + ### Fallback: status.json file + + If no routing directives are found in the response text, Fabro checks whether the agent wrote a `status.json` file into the sandbox working directory. If the file exists, Fabro extracts routing directives from it using the same logic. This is useful for agents that write structured output to files rather than including JSON in their response text. +@@ -90,6 +107,35 @@ review -> fix [label="Fix"] + review -> approve [label="Approve"] + ``` + ++## Custom structured outputs ++ ++Agent and prompt nodes can also validate their final JSON object against a JSON Schema file: ++ ++```dot ++audit [ ++ shape=tab, ++ prompt="Audit the change and return JSON that matches the schema.", ++ output_schema="@schemas/audit-result.schema.json", ++ output_retries=2 ++] ++``` ++ ++`output_schema="@path/to/schema.json"` uses the same workflow file-reference rules as prompt files: the schema is loaded relative to the workflow file and inlined before execution. The final JSON object in the LLM response is validated with `jsonschema`. ++ ++When custom schema validation succeeds, Fabro stores the parsed JSON value in context at: ++ ++| Key | Value | ++|---|---| ++| `output.{node_id}` | The parsed JSON object that matched the custom schema | ++ ++For example, node `audit` writes its parsed custom output to `output.audit`. Fabro still stores the raw response text at `response.audit`. ++ ++If custom schema validation fails, Fabro sends concise validation feedback to the same prompt conversation or agent session and asks for corrected JSON. After `output_retries` repair turns are exhausted, the node fails terminally with `output schema validation failed after N repair attempt(s)`. ++ ++ ++Structured output validation currently applies to agent and prompt nodes. `backend="acp"` does not support `output_schema` in this release. Custom schemas update `output.{node_id}`; routing schemas update routing fields and `context_updates` instead. ++ ++ + ## Output logging + + Fabro writes several files per stage to `stages/{rank:03}-{node_id}@{visit}/` in metadata snapshots and `fabro dump` output: +diff --git a/docs/public/reference/dot-language.mdx b/docs/public/reference/dot-language.mdx +index 1c1656057..74a366632 100644 +--- a/docs/public/reference/dot-language.mdx ++++ b/docs/public/reference/dot-language.mdx +@@ -206,10 +206,37 @@ Start nodes can also be identified by ID (`start` or `Start`). Exit nodes can be + | `model` | String | Explicit model ID (overrides stylesheet) | + | `provider` | String | Explicit provider name (overrides stylesheet). Auto-inferred from the model catalog when omitted. | + | `project_memory` | Boolean | When `true` (default), prompt nodes discover and include project docs (`AGENTS.md`, `CLAUDE.md`, etc.) as a system prompt. Set to `false` to disable. | ++| `output_schema` | String | Optional structured output validation. Use `routing` for Fabro's built-in routing directive schema, or `@path/to/schema.json` for a JSON Schema file. Supported on agent and prompt nodes. | ++| `output_retries` | Integer | Corrective structured-output turns inside the same prompt conversation or agent session. Default `2`; `0` validates once and fails without repair. Separate from `max_retries`. | + | `backend` | String | Agent execution backend: `api` (default) or `acp`. `api` runs Fabro's tool loop through provider APIs; `acp` runs an Agent Client Protocol stdio agent inside the active sandbox. Prompt nodes are API-only. See [Agents — Backends](/core-concepts/agents#backends). | + | `acp.command` | String | Shell command for nodes with `backend="acp"`. Mutually exclusive with `acp.config`. The value is always parsed as a command string, not JSON. | + | `acp.config` | String | JSON stdio ACP config for nodes with `backend="acp"`. Mutually exclusive with `acp.command`. | + ++#### Structured output validation ++ ++`output_schema` opts an agent or prompt node into strict JSON validation: ++ ++```dot ++review [ ++ shape=tab, ++ output_schema="routing", ++ output_retries=2 ++] ++ ++audit [ ++ shape=tab, ++ output_schema="@schemas/audit-result.schema.json", ++ output_retries=2 ++] ++``` ++ ++- `output_schema="routing"` requires a JSON object with at least one recognized routing field: `preferred_next_label`, `outcome`, `failure_reason`, `suggested_next_ids`, or `context_updates`. ++- `output_schema="@schemas/audit-result.schema.json"` loads a JSON Schema file using workflow file-reference rules and validates the final JSON object in the response text. ++- On validation failure, Fabro sends validation feedback to the same active context before failing: prompt nodes keep the prior assistant response in the message list, and API-backed agent nodes repair in the same live session. ++- `output_retries` defaults to `2` and controls only these corrective structured-output turns. It is not the same as `max_retries` and does not consume workflow retry attempts. ++- Custom schema output is stored in context at `output.{node_id}`. Routing schema output updates routing fields and any `context_updates`. ++- `backend="acp"` with `output_schema` is unsupported in this release. ++ + ### Command nodes + + | Attribute | Type | Description | +@@ -331,15 +358,16 @@ gate -> v2 [condition="context.version matches ^v2\\."] + gate -> slow_path + ``` + +-## Prompt file references ++## Prompt and schema file references + +-Instead of inlining long prompts, reference an external file: ++Instead of inlining long prompts or JSON Schemas, reference an external file: + + ```dot + simplify [label="Simplify", prompt="@prompts/simplify.md"] ++audit [shape=tab, output_schema="@schemas/audit-result.schema.json"] + ``` + +-The `@` prefix tells Fabro to load the prompt from a file path relative to the workflow file. Paths support `~` (home directory) and `..` (parent directory): ++The `@` prefix tells Fabro to load the referenced file relative to the workflow file. Paths support `~` (home directory) and `..` (parent directory): + + ```dot + shared [prompt="@~/shared-prompts/review.md"] +diff --git a/lib/crates/fabro-types/src/graph.rs b/lib/crates/fabro-types/src/graph.rs +index 27344f8c9..7ef9bad99 100644 +--- a/lib/crates/fabro-types/src/graph.rs ++++ b/lib/crates/fabro-types/src/graph.rs +@@ -168,6 +168,16 @@ impl Node { + self.str_attr("prompt") + } + ++ #[must_use] ++ pub fn output_schema(&self) -> Option<&str> { ++ self.str_attr("output_schema") ++ } ++ ++ #[must_use] ++ pub fn output_retries(&self) -> i64 { ++ self.int_attr("output_retries").unwrap_or(2).max(0) ++ } ++ + #[must_use] + pub fn max_retries(&self) -> Option { + self.int_attr("max_retries") +@@ -576,6 +586,8 @@ mod tests { + assert_eq!(node.shape(), "box"); + assert_eq!(node.node_type(), None); + assert_eq!(node.prompt(), None); ++ assert_eq!(node.output_schema(), None); ++ assert_eq!(node.output_retries(), 2); + assert_eq!(node.max_retries(), None); + assert!(!node.goal_gate()); + assert_eq!(node.retry_target(), None); +@@ -602,6 +614,31 @@ mod tests { + assert!(!node.project_memory()); + } + ++ #[test] ++ fn node_output_retries_defaults_and_clamps_to_zero() { ++ let mut node = Node::new("x"); ++ assert_eq!(node.output_retries(), 2); ++ ++ node.attrs ++ .insert("output_retries".to_string(), AttrValue::Integer(0)); ++ assert_eq!(node.output_retries(), 0); ++ ++ node.attrs ++ .insert("output_retries".to_string(), AttrValue::Integer(-3)); ++ assert_eq!(node.output_retries(), 0); ++ } ++ ++ #[test] ++ fn node_output_schema_returns_string_attr() { ++ let mut node = Node::new("x"); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String("routing".to_string()), ++ ); ++ ++ assert_eq!(node.output_schema(), Some("routing")); ++ } ++ + #[test] + fn node_with_attrs() { + let mut node = Node::new("plan"); +diff --git a/lib/crates/fabro-workflow/Cargo.toml b/lib/crates/fabro-workflow/Cargo.toml +index d5025f1bd..6021bca4c 100644 +--- a/lib/crates/fabro-workflow/Cargo.toml ++++ b/lib/crates/fabro-workflow/Cargo.toml +@@ -46,6 +46,7 @@ fabro-http.workspace = true + thiserror.workspace = true + serde.workspace = true + serde_json.workspace = true ++jsonschema.workspace = true + tokio.workspace = true + bytes.workspace = true + object_store.workspace = true +diff --git a/lib/crates/fabro-workflow/src/error.rs b/lib/crates/fabro-workflow/src/error.rs +index 6fcd565e3..4373a16db 100644 +--- a/lib/crates/fabro-workflow/src/error.rs ++++ b/lib/crates/fabro-workflow/src/error.rs +@@ -304,6 +304,9 @@ pub enum Error { + #[error("Unsupported operation: {0}")] + Unsupported(String), + ++ #[error("{0}")] ++ OutputSchemaValidation(String), ++ + #[error("Pipeline cancelled")] + Cancelled, + } +@@ -428,8 +431,8 @@ impl Error { + /// Retryable: Handler (transient handler failures), Engine (could be + /// transient), Io (network/disk issues are often transient), + /// Llm (delegates to SdkError). Terminal: Parse, Validation, +- /// Stylesheet (configuration errors), Checkpoint (storage +- /// integrity), Cancelled (explicit cancellation). ++ /// OutputSchemaValidation, Stylesheet (configuration errors), Checkpoint ++ /// (storage integrity), Cancelled (explicit cancellation). + #[must_use] + pub fn is_retryable(&self) -> bool { + match self { +@@ -444,6 +447,7 @@ impl Error { + | Self::Precondition(_) + | Self::RunNotFound(_) + | Self::Unsupported(_) ++ | Self::OutputSchemaValidation(_) + | Self::Cancelled => false, + } + } +@@ -461,7 +465,8 @@ impl Error { + | Self::Template { .. } + | Self::Stylesheet(_) + | Self::Checkpoint(_) +- | Self::Unsupported(_) => FailureCategory::Deterministic, ++ | Self::Unsupported(_) ++ | Self::OutputSchemaValidation(_) => FailureCategory::Deterministic, + Self::Precondition(_) | Self::RunNotFound(_) => FailureCategory::Structural, + Self::Handler { failure_class, .. } | Self::Engine { failure_class, .. } => { + *failure_class +diff --git a/lib/crates/fabro-workflow/src/handler/agent.rs b/lib/crates/fabro-workflow/src/handler/agent.rs +index 0788ee1e5..0f2118df8 100644 +--- a/lib/crates/fabro-workflow/src/handler/agent.rs ++++ b/lib/crates/fabro-workflow/src/handler/agent.rs +@@ -2,20 +2,21 @@ use std::path::Path; + use std::sync::Arc; + + use async_trait::async_trait; +-use fabro_agent::Sandbox; ++use fabro_agent::{Sandbox, shell_quote}; + use fabro_graphviz::graph::{Graph, Node}; + use fabro_types::{RunId, StageModelUsage}; + use tokio_util::sync::CancellationToken; + + use super::llm::api::EffectiveRequestControls; ++use super::structured_output::{ ++ self, OutputSchemaKind, StructuredOutputError, ValidatedStructuredOutput, ++}; + use super::{EngineServices, Handler, NodeTimeoutPolicy}; + use crate::context::{Context, WorkflowContext, keys}; + use crate::error::Error; + use crate::event::{Emitter, Event, StageScope}; + use crate::interview_runtime::WorkflowAgentQuestionRuntime; +-use crate::outcome::{ +- BilledModelUsage, FailureCategory, FailureDetail, Outcome, OutcomeExt, StageOutcome, +-}; ++use crate::outcome::{BilledModelUsage, Outcome, OutcomeExt}; + + /// Result from a `CodergenBackend` invocation. + pub enum CodergenResult { +@@ -132,110 +133,65 @@ impl AgentHandler { + } + } + +-/// Status fields that indicate a JSON object contains routing directives. +-const STATUS_FIELDS: &[&str] = &[ +- "preferred_next_label", +- "outcome", +- "failure_reason", +- "suggested_next_ids", +- "context_updates", +-]; +- +-/// Find all balanced `{...}` JSON object substrings in the text. +-fn find_json_objects(text: &str) -> Vec<&str> { +- let mut results = Vec::new(); +- let bytes = text.as_bytes(); +- let mut i = 0; +- while i < bytes.len() { +- if bytes[i] == b'{' { +- let start = i; +- let mut depth = 0; +- let mut in_string = false; +- let mut escape = false; +- let mut j = i; +- while j < bytes.len() { +- let c = bytes[j]; +- if escape { +- escape = false; +- } else if c == b'\\' && in_string { +- escape = true; +- } else if c == b'"' { +- in_string = !in_string; +- } else if !in_string { +- if c == b'{' { +- depth += 1; +- } else if c == b'}' { +- depth -= 1; +- if depth == 0 { +- results.push(&text[start..=j]); +- break; +- } +- } +- } +- j += 1; +- } +- } +- i += 1; +- } +- results +-} +- + /// Extract routing directives from LLM response text. + /// + /// Searches for the last JSON object in the response that contains at least + /// one status field (`preferred_next_label`, `outcome`, `suggested_next_ids`, + /// `context_updates`). Merges extracted fields into the outcome. + pub(crate) fn extract_status_fields(text: &str, outcome: &mut Outcome) -> bool { +- let candidates = find_json_objects(text); +- +- let parsed = candidates.iter().rev().find_map(|candidate| { +- let value: serde_json::Value = serde_json::from_str(candidate).ok()?; +- if let Some(obj) = value.as_object() { +- if STATUS_FIELDS.iter().any(|f| obj.contains_key(*f)) { +- return Some(value); +- } +- } +- None +- }); +- +- let Some(value) = parsed else { return false }; +- let Some(obj) = value.as_object() else { +- return false; +- }; ++ structured_output::extract_status_fields_loose(text, outcome) ++} + +- if let Some(label) = obj.get("preferred_next_label").and_then(|v| v.as_str()) { +- outcome.preferred_label = Some(label.to_string()); ++pub(crate) async fn validate_agent_output_sources( ++ schema: &OutputSchemaKind, ++ response_text: &str, ++ sandbox: &Arc, ++ last_file_touched: Option<&str>, ++) -> Result { ++ if !matches!(schema, OutputSchemaKind::Routing) { ++ return structured_output::validate_response_text(schema, response_text); + } + +- if let Some(ids) = obj.get("suggested_next_ids").and_then(|v| v.as_array()) { +- let string_ids: Vec = ids +- .iter() +- .filter_map(|v| v.as_str().map(String::from)) +- .collect(); +- if !string_ids.is_empty() { +- outcome.suggested_next_ids = string_ids; +- } ++ match structured_output::validate_response_text(schema, response_text) { ++ Ok(validated) => return Ok(validated), ++ Err(error) if error.allows_routing_fallback() => {} ++ Err(error) => return Err(error), + } + +- if let Some(status_str) = obj.get("outcome").and_then(|v| v.as_str()) { +- if let Ok(status) = status_str.parse::() { +- outcome.status = status; +- if outcome.status.is_failure() { +- if let Some(reason) = obj.get("failure_reason").and_then(|v| v.as_str()) { +- outcome.failure = +- Some(FailureDetail::new(reason, FailureCategory::Deterministic)); +- } ++ let mut fallback_error = None; ++ if let Some(status_json) = read_sandbox_file(sandbox, "status.json").await { ++ match structured_output::validate_response_text(schema, &status_json) { ++ Ok(validated) => return Ok(validated), ++ Err(error) if error.allows_routing_fallback() => { ++ fallback_error = Some(error); + } ++ Err(error) => return Err(error), + } + } + +- if let Some(updates) = obj.get("context_updates").and_then(|v| v.as_object()) { +- for (key, val) in updates { +- outcome.context_updates.insert(key.clone(), val.clone()); ++ if let Some(path) = last_file_touched { ++ if let Some(contents) = read_sandbox_file(sandbox, path).await { ++ return structured_output::validate_response_text(schema, &contents); + } + } + +- true ++ Err(fallback_error.unwrap_or_else(|| { ++ structured_output::validate_response_text(schema, response_text) ++ .expect_err("response text should have failed routing validation") ++ })) ++} ++ ++async fn read_sandbox_file(sandbox: &Arc, path: &str) -> Option { ++ let cmd = format!("cat {}", shell_quote(path)); ++ let result = sandbox ++ .exec_command(&cmd, 5_000, None, None, None) ++ .await ++ .ok()?; ++ if result.is_success() { ++ Some(result.stdout) ++ } else { ++ None ++ } + } + + /// Truncate a string to at most `max_chars` characters (char-boundary safe). +@@ -416,34 +372,46 @@ impl Handler for AgentHandler { + serde_json::json!(&response_text), + ); + +- // 7b. Parse routing directives from response text, falling back to +- // status.json written by the agent into the sandbox CWD, then to +- // the last file the agent wrote. +- let found_in_response = extract_status_fields(&response_text, &mut outcome); +- if !found_in_response { +- let mut found_in_status_json = false; +- if let Ok(result) = services +- .run +- .sandbox +- .exec_command("cat status.json", 5_000, None, None, None) +- .await ++ if let Some(schema) = structured_output::parse_node_output_schema(node)? { ++ match validate_agent_output_sources( ++ &schema, ++ &response_text, ++ &services.run.sandbox, ++ last_file_touched.as_deref(), ++ ) ++ .await + { +- if result.is_success() { +- found_in_status_json = extract_status_fields(&result.stdout, &mut outcome); ++ Ok(validated) => { ++ structured_output::apply_validated_output( ++ node, ++ &schema, ++ &validated, ++ &mut outcome, ++ ); ++ } ++ Err(_) => { ++ return Ok(structured_output::exhausted_failure_outcome( ++ node.output_retries(), ++ )); + } + } +- if !found_in_status_json { +- if let Some(ref path) = last_file_touched { +- let quoted = shlex::try_quote(path).unwrap_or_else(|_| path.into()); +- let cmd = format!("cat {quoted}"); +- if let Ok(result) = services +- .run +- .sandbox +- .exec_command(&cmd, 5_000, None, None, None) +- .await +- { +- if result.is_success() { +- extract_status_fields(&result.stdout, &mut outcome); ++ } else { ++ // 7b. Parse routing directives from response text, falling back to ++ // status.json written by the agent into the sandbox CWD, then to ++ // the last file the agent wrote. ++ let found_in_response = extract_status_fields(&response_text, &mut outcome); ++ if !found_in_response { ++ let mut found_in_status_json = false; ++ if let Some(status_json) = ++ read_sandbox_file(&services.run.sandbox, "status.json").await ++ { ++ found_in_status_json = extract_status_fields(&status_json, &mut outcome); ++ } ++ if !found_in_status_json { ++ if let Some(ref path) = last_file_touched { ++ if let Some(contents) = read_sandbox_file(&services.run.sandbox, path).await ++ { ++ extract_status_fields(&contents, &mut outcome); + } + } + } +@@ -795,6 +763,142 @@ mod tests { + ); + } + ++ #[tokio::test] ++ async fn codergen_handler_output_schema_routing_uses_status_json_fallback_when_response_has_no_json() ++ { ++ let sandbox_dir = TempDir::new().unwrap(); ++ std::fs::write( ++ sandbox_dir.path().join("status.json"), ++ r#"{"preferred_next_label": "review"}"#, ++ ) ++ .unwrap(); ++ ++ let handler = AgentHandler::new(None); ++ let mut node = Node::new("step"); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String("routing".to_string()), ++ ); ++ let context = test_context(); ++ let graph = Graph::new("test"); ++ let tmp = TempDir::new().unwrap(); ++ ++ let mut services = EngineServices::test_default(); ++ services.run = ++ services ++ .run ++ .with_sandbox(std::sync::Arc::new(fabro_agent::LocalSandbox::new( ++ sandbox_dir.path().to_path_buf(), ++ ))); ++ ++ let outcome = handler ++ .execute(&node, &context, &graph, tmp.path(), &services) ++ .await ++ .unwrap(); ++ ++ assert_eq!(outcome.status, crate::outcome::StageOutcome::Succeeded); ++ assert_eq!(outcome.preferred_label.as_deref(), Some("review")); ++ } ++ ++ #[tokio::test] ++ async fn codergen_handler_output_schema_routing_rejects_malformed_response_before_status_json_fallback() ++ { ++ struct BadRoutingBackend; ++ ++ #[async_trait] ++ impl CodergenBackend for BadRoutingBackend { ++ async fn run(&self, _request: CodergenRunRequest<'_>) -> Result { ++ Ok(CodergenResult::Text { ++ text: r#"{"suggested_next_ids": [1]}"#.to_string(), ++ usage: None, ++ files_touched: Vec::new(), ++ last_file_touched: None, ++ }) ++ } ++ } ++ ++ let sandbox_dir = TempDir::new().unwrap(); ++ std::fs::write( ++ sandbox_dir.path().join("status.json"), ++ r#"{"preferred_next_label": "should_not_use"}"#, ++ ) ++ .unwrap(); ++ ++ let handler = AgentHandler::new(Some(Box::new(BadRoutingBackend))); ++ let mut node = Node::new("step"); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String("routing".to_string()), ++ ); ++ node.attrs ++ .insert("output_retries".to_string(), AttrValue::Integer(0)); ++ let context = test_context(); ++ let graph = Graph::new("test"); ++ let tmp = TempDir::new().unwrap(); ++ ++ let mut services = EngineServices::test_default(); ++ services.run = ++ services ++ .run ++ .with_sandbox(std::sync::Arc::new(fabro_agent::LocalSandbox::new( ++ sandbox_dir.path().to_path_buf(), ++ ))); ++ ++ let outcome = handler ++ .execute(&node, &context, &graph, tmp.path(), &services) ++ .await ++ .unwrap(); ++ ++ assert_eq!(outcome.status, crate::outcome::StageOutcome::Failed { ++ retry_requested: false, ++ }); ++ assert_eq!( ++ outcome.failure_reason(), ++ Some("output schema validation failed after 0 repair attempt(s)") ++ ); ++ assert!(outcome.preferred_label.is_none()); ++ } ++ ++ #[tokio::test] ++ async fn codergen_handler_custom_output_schema_updates_output_context_key() { ++ struct CustomOutputBackend; ++ ++ #[async_trait] ++ impl CodergenBackend for CustomOutputBackend { ++ async fn run(&self, _request: CodergenRunRequest<'_>) -> Result { ++ Ok(CodergenResult::Text { ++ text: r#"{"passed": true}"#.to_string(), ++ usage: None, ++ files_touched: Vec::new(), ++ last_file_touched: None, ++ }) ++ } ++ } ++ ++ let handler = AgentHandler::new(Some(Box::new(CustomOutputBackend))); ++ let mut node = Node::new("audit"); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String( ++ r#"{"type":"object","required":["passed"],"properties":{"passed":{"type":"boolean"}}}"# ++ .to_string(), ++ ), ++ ); ++ let context = test_context(); ++ let graph = Graph::new("test"); ++ let tmp = TempDir::new().unwrap(); ++ ++ let outcome = handler ++ .execute(&node, &context, &graph, tmp.path(), &make_services()) ++ .await ++ .unwrap(); ++ ++ assert_eq!( ++ outcome.context_updates.get("output.audit"), ++ Some(&serde_json::json!({"passed": true})), ++ ); ++ } ++ + #[tokio::test] + async fn codergen_handler_projects_provider_used_from_agent_session_events() { + struct ProviderEventBackend; +diff --git a/lib/crates/fabro-workflow/src/handler/llm/acp.rs b/lib/crates/fabro-workflow/src/handler/llm/acp.rs +index 827f09e8a..3318c64ef 100644 +--- a/lib/crates/fabro-workflow/src/handler/llm/acp.rs ++++ b/lib/crates/fabro-workflow/src/handler/llm/acp.rs +@@ -325,6 +325,11 @@ impl Default for AgentAcpBackend { + #[async_trait] + impl CodergenBackend for AgentAcpBackend { + async fn run(&self, request: CodergenRunRequest<'_>) -> Result { ++ if request.node.output_schema().is_some() { ++ return Err(Error::Validation( ++ "output_schema is not supported with backend=\"acp\" in this release".to_string(), ++ )); ++ } + let stage_scope = StageScope::for_handler(request.context, &request.node.id); + self.run_turn( + request.node, +@@ -476,6 +481,56 @@ mod tests { + assert_eq!(files_touched, vec!["hello.txt"]); + } + ++ #[tokio::test] ++ async fn acp_backend_rejects_output_schema_without_launching_process() { ++ let tempdir = tempfile::tempdir().unwrap(); ++ let launched_path = tempdir.path().join("launched"); ++ ++ let mut node = Node::new("work"); ++ node.attrs ++ .insert("backend".to_string(), AttrValue::String("acp".to_string())); ++ node.attrs.insert( ++ "acp.command".to_string(), ++ AttrValue::String("sh -c 'touch launched'".to_string()), ++ ); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String("routing".to_string()), ++ ); ++ ++ let backend = AgentAcpBackend::new(); ++ let sandbox: Arc = Arc::new(LocalSandbox::new(tempdir.path().to_path_buf())); ++ let emitter = Arc::new(Emitter::default()); ++ let context = Context::new(); ++ let result = backend ++ .run(CodergenRunRequest { ++ node: &node, ++ prompt: "write hello", ++ context: &context, ++ thread_id: None, ++ emitter: &emitter, ++ sandbox: &sandbox, ++ tool_hooks: None, ++ cancel_token: CancellationToken::new(), ++ agent_tool_runtime: fabro_agent::AgentToolRuntime::default(), ++ }) ++ .await; ++ ++ let Err(error) = result else { ++ panic!("expected output_schema guardrail error"); ++ }; ++ assert!( ++ error ++ .to_string() ++ .contains("output_schema is not supported with backend=\"acp\" in this release"), ++ "unexpected error: {error}", ++ ); ++ assert!( ++ !launched_path.exists(), ++ "ACP process should not launch when output_schema is present", ++ ); ++ } ++ + #[tokio::test] + async fn acp_backend_accepts_steer_and_incorporates_followup_result() { + let tempdir = tempfile::tempdir().unwrap(); +diff --git a/lib/crates/fabro-workflow/src/handler/llm/api.rs b/lib/crates/fabro-workflow/src/handler/llm/api.rs +index 8da02640d..22d08dc4b 100644 +--- a/lib/crates/fabro-workflow/src/handler/llm/api.rs ++++ b/lib/crates/fabro-workflow/src/handler/llm/api.rs +@@ -13,7 +13,8 @@ use fabro_auth::{CredentialSource, EnvCredentialSource}; + use fabro_graphviz::graph::{AttrValue, Node}; + use fabro_llm::client::Client; + use fabro_llm::types::{ +- Message, ReasoningEffort, Request, Speed, TokenCounts, ToolDefinition as LlmToolDefinition, ++ Message, ReasoningEffort, Request, Response, Speed, TokenCounts, ++ ToolDefinition as LlmToolDefinition, + }; + use fabro_mcp::config::McpServerSettings; + #[cfg(test)] +@@ -26,7 +27,11 @@ use tokio::sync::Mutex as TokioMutex; + use tokio::task::JoinHandle; + use tokio_util::sync::CancellationToken; + +-use super::super::agent::{CodergenBackend, CodergenResult, CodergenRunRequest, OneShotRequest}; ++use super::super::agent::{ ++ CodergenBackend, CodergenResult, CodergenRunRequest, OneShotRequest, ++ validate_agent_output_sources, ++}; ++use super::super::structured_output; + use super::activation_lease::{ActivationLease, ActivationLeaseOptions}; + use super::routing; + use super::routing::ProviderContext; +@@ -460,6 +465,32 @@ fn track_file_event(event: &AgentEvent, state: &mut FileTracking) { + } + } + ++fn file_tracking_snapshot( ++ file_tracking: &Arc>, ++) -> (Vec, Option) { ++ let state = file_tracking.lock().unwrap(); ++ let mut files: Vec = state.touched.iter().cloned().collect(); ++ files.sort(); ++ (files, state.last.clone()) ++} ++ ++fn last_assistant_response(session: &Session) -> String { ++ session ++ .history() ++ .turns() ++ .iter() ++ .rev() ++ .find_map(|turn| { ++ if let AgentMessage::Assistant { content, .. } = turn { ++ if !content.is_empty() { ++ return Some(content.clone()); ++ } ++ } ++ None ++ }) ++ .unwrap_or_default() ++} ++ + /// Spawn a task that subscribes to session events and: + /// 1. Tracks file changes (write_file/edit_file tool calls) into shared state. + /// 2. Forwards non-streaming agent events to the pipeline emitter. +@@ -521,6 +552,13 @@ pub struct AgentApiBackend { + fabro_run_tools: Option, + } + ++struct OneShotCompletion { ++ response: Response, ++ actual_model: String, ++ actual_provider: String, ++ actual_speed: Option, ++} ++ + impl AgentApiBackend { + #[must_use] + pub fn new( +@@ -812,71 +850,18 @@ impl AgentApiBackend { + } + } + } +-} +- +-#[async_trait] +-impl CodergenBackend for AgentApiBackend { +- async fn shutdown(&self, emitter: &Arc) { +- self.shutdown_cached_sessions(emitter); +- } +- +- fn effective_request_controls(&self, node: &Node) -> Result { +- self.resolve_effective_request_controls(node) +- } +- +- async fn one_shot(&self, request: OneShotRequest<'_>) -> Result { +- let node = request.node; +- let prompt = request.prompt; +- let system_prompt = request.system_prompt; +- let emitter = request.emitter; +- let stage_scope = request.stage_scope; +- +- let client = Client::from_source(self.source.as_ref(), Arc::clone(&self.catalog)) +- .await +- .map_err(|e| Error::handler_with_source("Failed to create LLM client", e))?; +- +- let model = node.model().unwrap_or(&self.model); +- let provider = self.resolve_provider_context(model, node.provider())?; +- let provider_id = provider.provider_id.to_string(); +- let controls = self.resolve_effective_request_controls(node)?; +- +- let max_tokens = node +- .max_tokens() +- .or_else(|| self.catalog.get(model).and_then(|m| m.limits.max_output)); +- +- let mut messages = Vec::new(); +- if let Some(sys) = system_prompt { +- messages.push(Message::system(sys)); +- } +- messages.push(Message::user(prompt)); +- +- let request = Request { +- model: model.to_string(), +- messages, +- provider: Some(provider_id), +- reasoning_effort: controls.reasoning_effort, +- speed: controls.speed, +- tools: None, +- tool_choice: None, +- response_format: None, +- temperature: None, +- top_p: None, +- max_tokens, +- stop_sequences: None, +- metadata: None, +- provider_options: None, +- }; +- +- // Build per-request fallback chain: if the node overrides the provider, +- // no failover is available; otherwise use the backend's. +- let fallback_chain: &[FallbackTarget] = if node.provider().is_some() { +- &[] +- } else { +- &self.fallback_chain +- }; +- +- let result = client.complete(&request).await; + ++ async fn complete_one_shot_request( ++ &self, ++ client: &Client, ++ node: &Node, ++ emitter: &Arc, ++ stage_scope: &StageScope, ++ request: &Request, ++ controls: EffectiveRequestControls, ++ fallback_chain: &[FallbackTarget], ++ ) -> Result { ++ let result = client.complete(request).await; + let default_provider = self.provider_id.to_string(); + + let (response, actual_model, actual_provider, actual_speed) = match result { +@@ -916,7 +901,7 @@ impl CodergenBackend for AgentApiBackend { + let max_tokens = node.max_tokens().or_else(|| { + self.catalog + .get(&target.model) +- .and_then(|m| m.limits.max_output) ++ .and_then(|model| model.limits.max_output) + }); + + let fallback_request = Request { +@@ -930,12 +915,12 @@ impl CodergenBackend for AgentApiBackend { + + match client.complete(&fallback_request).await { + Ok(resp) => { +- found = Some(( +- resp, +- target.model.clone(), +- target.provider.clone(), +- controls.speed, +- )); ++ found = Some(OneShotCompletion { ++ response: resp, ++ actual_model: target.model.clone(), ++ actual_provider: target.provider.clone(), ++ actual_speed: controls.speed, ++ }); + break; + } + Err(err) if err.failover_eligible() => { +@@ -946,30 +931,144 @@ impl CodergenBackend for AgentApiBackend { + } + + match found { +- Some(triple) => triple, ++ Some(completion) => return Ok(completion), + None => return Err(Error::Llm(last_err)), + } + } + Err(sdk_err) => return Err(Error::Llm(sdk_err)), + }; + +- let stage_usage = billed_model_usage_from_llm( +- self.catalog.as_ref(), +- &ModelRef { +- provider: ProviderId::from(actual_provider), +- model_id: actual_model, +- speed: actual_speed, +- }, +- &response.usage, +- )?; +- +- Ok(CodergenResult::Text { +- text: response.text(), +- usage: Some(stage_usage), +- files_touched: Vec::new(), +- last_file_touched: None, ++ Ok(OneShotCompletion { ++ response, ++ actual_model, ++ actual_provider, ++ actual_speed, + }) + } ++} ++ ++#[async_trait] ++impl CodergenBackend for AgentApiBackend { ++ async fn shutdown(&self, emitter: &Arc) { ++ self.shutdown_cached_sessions(emitter); ++ } ++ ++ fn effective_request_controls(&self, node: &Node) -> Result { ++ self.resolve_effective_request_controls(node) ++ } ++ ++ async fn one_shot(&self, request: OneShotRequest<'_>) -> Result { ++ let node = request.node; ++ let prompt = request.prompt; ++ let system_prompt = request.system_prompt; ++ let emitter = request.emitter; ++ let stage_scope = request.stage_scope; ++ ++ let client = Client::from_source(self.source.as_ref(), Arc::clone(&self.catalog)) ++ .await ++ .map_err(|e| Error::handler_with_source("Failed to create LLM client", e))?; ++ ++ let model = node.model().unwrap_or(&self.model); ++ let provider = self.resolve_provider_context(model, node.provider())?; ++ let provider_id = provider.provider_id.to_string(); ++ let controls = self.resolve_effective_request_controls(node)?; ++ ++ let max_tokens = node ++ .max_tokens() ++ .or_else(|| self.catalog.get(model).and_then(|m| m.limits.max_output)); ++ ++ let mut messages = Vec::new(); ++ if let Some(sys) = system_prompt { ++ messages.push(Message::system(sys)); ++ } ++ messages.push(Message::user(prompt)); ++ ++ // Build per-request fallback chain: if the node overrides the provider, ++ // no failover is available; otherwise use the backend's. ++ let fallback_chain: &[FallbackTarget] = if node.provider().is_some() { ++ &[] ++ } else { ++ &self.fallback_chain ++ }; ++ ++ let output_schema = structured_output::parse_node_output_schema(node)?; ++ let response_format = output_schema ++ .as_ref() ++ .map(structured_output::prompt_response_format); ++ let mut repair_attempts = 0_i64; ++ let mut total_usage = TokenCounts::default(); ++ ++ loop { ++ let request = Request { ++ model: model.to_string(), ++ messages: messages.clone(), ++ provider: Some(provider_id.clone()), ++ reasoning_effort: controls.reasoning_effort, ++ speed: controls.speed, ++ tools: None, ++ tool_choice: None, ++ response_format: response_format.clone(), ++ temperature: None, ++ top_p: None, ++ max_tokens, ++ stop_sequences: None, ++ metadata: None, ++ provider_options: None, ++ }; ++ ++ let completion = self ++ .complete_one_shot_request( ++ &client, ++ node, ++ emitter, ++ stage_scope, ++ &request, ++ controls, ++ fallback_chain, ++ ) ++ .await?; ++ total_usage += completion.response.usage.clone(); ++ let response_text = completion.response.text(); ++ ++ let validation_error = if let Some(schema) = &output_schema { ++ match structured_output::validate_response_text(schema, &response_text) { ++ Ok(_) => None, ++ Err(error) => Some((schema, error)), ++ } ++ } else { ++ None ++ }; ++ ++ if let Some((schema, error)) = validation_error { ++ if repair_attempts >= node.output_retries() { ++ return Err(Error::OutputSchemaValidation( ++ structured_output::exhausted_failure_reason(node.output_retries()), ++ )); ++ } ++ messages.push(Message::assistant(response_text)); ++ messages.push(Message::user(error.repair_message(schema))); ++ repair_attempts += 1; ++ continue; ++ } ++ ++ let stage_usage = billed_model_usage_from_llm( ++ self.catalog.as_ref(), ++ &ModelRef { ++ provider: ProviderId::from(completion.actual_provider), ++ model_id: completion.actual_model, ++ speed: completion.actual_speed, ++ }, ++ &total_usage, ++ )?; ++ ++ return Ok(CodergenResult::Text { ++ text: response_text, ++ usage: Some(stage_usage), ++ files_touched: Vec::new(), ++ last_file_touched: None, ++ }); ++ } ++ } + + async fn run(&self, request: CodergenRunRequest<'_>) -> Result { + let node = request.node; +@@ -981,6 +1080,7 @@ impl CodergenBackend for AgentApiBackend { + let tool_hooks = request.tool_hooks; + let cancel_token = request.cancel_token; + let agent_tool_runtime = request.agent_tool_runtime; ++ let output_schema = structured_output::parse_node_output_schema(node)?; + + let fidelity = context.fidelity(); + let reuse_key = if fidelity == Fidelity::Full { +@@ -1258,8 +1358,59 @@ impl CodergenBackend for AgentApiBackend { + return Err(err); + } + ++ let mut response = last_assistant_response(&session); ++ if let Some(schema) = &output_schema { ++ let mut repair_attempts = 0_i64; ++ loop { ++ let (_, last_file_touched) = file_tracking_snapshot(&file_tracking); ++ match validate_agent_output_sources( ++ schema, ++ &response, ++ sandbox, ++ last_file_touched.as_deref(), ++ ) ++ .await ++ { ++ Ok(_) => break, ++ Err(error) => { ++ if repair_attempts >= node.output_retries() { ++ bridge.abort(); ++ discard_session(&mut session, &mut lease, emitter); ++ return Err(Error::OutputSchemaValidation( ++ structured_output::exhausted_failure_reason(node.output_retries()), ++ )); ++ } ++ let repair_message = error.repair_message(schema); ++ match session.process_input(&repair_message).await { ++ Ok(()) => { ++ repair_attempts += 1; ++ response = last_assistant_response(&session); ++ } ++ Err(err) => match classify_agent_error(err, false) { ++ AgentApiErrorDisposition::Cancelled => { ++ bridge.abort(); ++ discard_session(&mut session, &mut lease, emitter); ++ return Err(Error::Cancelled); ++ } ++ AgentApiErrorDisposition::Terminal(err) => { ++ bridge.abort(); ++ discard_session(&mut session, &mut lease, emitter); ++ return Err(err); ++ } ++ AgentApiErrorDisposition::FailoverEligible(sdk_err) => { ++ bridge.abort(); ++ discard_session(&mut session, &mut lease, emitter); ++ return Err(Error::Llm(sdk_err)); ++ } ++ }, ++ } ++ } ++ } ++ } ++ } ++ + // Aggregate token usage only from new turns (prevents double-counting on +- // reuse). ++ // reuse), including any output-schema repair turns. + let mut total_usage = TokenCounts::default(); + for turn in &session.history().turns()[turns_before..] { + if let AgentMessage::Assistant { usage, .. } = turn { +@@ -1278,29 +1429,8 @@ impl CodergenBackend for AgentApiBackend { + &total_usage, + )?; + +- // Extract last assistant response from the session history. +- let response = session +- .history() +- .turns() +- .iter() +- .rev() +- .find_map(|turn| { +- if let AgentMessage::Assistant { content, .. } = turn { +- if !content.is_empty() { +- return Some(content.clone()); +- } +- } +- None +- }) +- .unwrap_or_default(); +- + // Collect files_touched from the shared tracking state. +- let (files_touched, last_file_touched) = { +- let s = file_tracking.lock().unwrap(); +- let mut v: Vec = s.touched.iter().cloned().collect(); +- v.sort(); +- (v, s.last.clone()) +- }; ++ let (files_touched, last_file_touched) = file_tracking_snapshot(&file_tracking); + + if let Some(lease) = lease.take() { + lease.release(); +@@ -1376,10 +1506,13 @@ mod tests { + }; + use fabro_vault::{SecretType, Vault}; + use futures::stream; ++ use httpmock::Method::POST; ++ use httpmock::MockServer; + use tokio::sync::RwLock as AsyncRwLock; + use tokio_util::sync::CancellationToken; + + use super::*; ++ use crate::context::Context; + use crate::services::FabroRunToolServices; + + struct ShutdownTestProfile { +@@ -1447,6 +1580,109 @@ mod tests { + } + } + ++ fn mock_llm_catalog(server: &MockServer) -> Arc { ++ let settings: LlmCatalogSettings = toml::from_str(&format!( ++ r#" ++[providers.mock] ++adapter = "openai_compatible" ++agent_profile = "openai" ++base_url = "{}" ++ ++[providers.mock.auth] ++credentials = ["env:MOCK_API_KEY"] ++ ++[models.mock-model] ++provider = "mock" ++display_name = "Mock Model" ++family = "mock" ++default = true ++ ++[models.mock-model.limits] ++context_window = 8192 ++max_output = 1024 ++ ++[models.mock-model.features] ++tools = true ++vision = false ++reasoning = false ++"#, ++ server.base_url() ++ )) ++ .unwrap(); ++ Arc::new(Catalog::from_builtin_with_overrides(&settings).unwrap()) ++ } ++ ++ fn mock_api_backend(server: &MockServer) -> AgentApiBackend { ++ let source = EnvCredentialSource::with_env_lookup(Arc::new(|name| { ++ if name == "MOCK_API_KEY" { ++ Some("sk-test".to_string()) ++ } else { ++ None ++ } ++ })); ++ AgentApiBackend::new_with_catalog( ++ "mock-model".to_string(), ++ ProviderId::from("mock"), ++ Vec::new(), ++ Arc::new(source), ++ SteeringHub::for_tests(), ++ mock_llm_catalog(server), ++ ) ++ } ++ ++ fn chat_completion_response( ++ text: &str, ++ input_tokens: i64, ++ output_tokens: i64, ++ ) -> serde_json::Value { ++ serde_json::json!({ ++ "id": uuid::Uuid::new_v4().to_string(), ++ "model": "mock-model", ++ "choices": [{ ++ "message": { ++ "content": text ++ }, ++ "finish_reason": "stop" ++ }], ++ "usage": { ++ "prompt_tokens": input_tokens, ++ "completion_tokens": output_tokens, ++ "total_tokens": input_tokens + output_tokens ++ } ++ }) ++ } ++ ++ fn chat_completion_stream(text: &str, input_tokens: i64, output_tokens: i64) -> String { ++ let text_chunk = serde_json::json!({ ++ "id": uuid::Uuid::new_v4().to_string(), ++ "model": "mock-model", ++ "choices": [{ ++ "delta": { ++ "content": text ++ }, ++ "finish_reason": null ++ }] ++ }); ++ let usage_chunk = serde_json::json!({ ++ "id": uuid::Uuid::new_v4().to_string(), ++ "model": "mock-model", ++ "choices": [], ++ "usage": { ++ "prompt_tokens": input_tokens, ++ "completion_tokens": output_tokens, ++ "total_tokens": input_tokens + output_tokens ++ } ++ }); ++ format!("data: {text_chunk}\n\ndata: {usage_chunk}\n\ndata: [DONE]\n\n") ++ } ++ ++ fn custom_output_schema_attr() -> AttrValue { ++ AttrValue::String( ++ r#"{"type":"object","required":["passed"],"properties":{"passed":{"type":"boolean"}}}"# ++ .to_string(), ++ ) ++ } ++ + #[test] + fn agent_backend_stores_config() { + let backend = AgentApiBackend::new_from_env( +@@ -2212,6 +2448,127 @@ reasoning = false + assert_eq!(client.provider_names(), vec!["anthropic"]); + } + ++ #[tokio::test] ++ async fn one_shot_repairs_custom_output_schema_with_previous_assistant_message() { ++ let server = MockServer::start(); ++ let first = server.mock(|when, then| { ++ when.method(POST) ++ .path("/chat/completions") ++ .body_includes(r#""type":"json_schema""#) ++ .body_excludes(r#""role":"assistant""#); ++ then.status(200) ++ .header("content-type", "application/json") ++ .json_body(chat_completion_response("not json", 10, 1)); ++ }); ++ let repair = server.mock(|when, then| { ++ when.method(POST) ++ .path("/chat/completions") ++ .body_includes(r#""type":"json_schema""#) ++ .body_includes(r#""role":"assistant""#) ++ .body_includes("not json") ++ .body_includes("output_schema"); ++ then.status(200) ++ .header("content-type", "application/json") ++ .json_body(chat_completion_response(r#"{"passed":true}"#, 11, 2)); ++ }); ++ let backend = mock_api_backend(&server); ++ let mut node = Node::new("audit"); ++ node.attrs ++ .insert("output_schema".to_string(), custom_output_schema_attr()); ++ node.attrs ++ .insert("output_retries".to_string(), AttrValue::Integer(1)); ++ let context = Context::new(); ++ let stage_scope = StageScope::for_handler(&context, &node.id); ++ let emitter = Arc::new(Emitter::new(fabro_types::RunId::new())); ++ let workspace = tempfile::tempdir().unwrap(); ++ let sandbox: Arc = ++ Arc::new(LocalSandbox::new(workspace.path().to_path_buf())); ++ ++ let result = backend ++ .one_shot(OneShotRequest { ++ node: &node, ++ prompt: "Audit the result", ++ system_prompt: None, ++ emitter: &emitter, ++ stage_scope: &stage_scope, ++ sandbox: &sandbox, ++ cancel_token: CancellationToken::new(), ++ }) ++ .await ++ .unwrap(); ++ ++ first.assert_calls(1); ++ repair.assert_calls(1); ++ let CodergenResult::Text { text, usage, .. } = result else { ++ panic!("one_shot should return text"); ++ }; ++ assert_eq!(text, r#"{"passed":true}"#); ++ let usage = usage.expect("usage should be aggregated"); ++ assert_eq!(usage.tokens().input_tokens, 21); ++ assert_eq!(usage.tokens().output_tokens, 3); ++ } ++ ++ #[tokio::test] ++ async fn agent_run_repairs_custom_output_schema_in_same_session() { ++ let server = MockServer::start(); ++ let first = server.mock(|when, then| { ++ when.method(POST) ++ .path("/chat/completions") ++ .body_includes(r#""stream":true"#) ++ .body_excludes(r#""role":"assistant""#); ++ then.status(200) ++ .header("content-type", "text/event-stream") ++ .body(chat_completion_stream("not json", 20, 3)); ++ }); ++ let repair = server.mock(|when, then| { ++ when.method(POST) ++ .path("/chat/completions") ++ .body_includes(r#""stream":true"#) ++ .body_includes(r#""role":"assistant""#) ++ .body_includes("not json") ++ .body_includes("output_schema"); ++ then.status(200) ++ .header("content-type", "text/event-stream") ++ .body(chat_completion_stream(r#"{"passed":true}"#, 21, 4)); ++ }); ++ let backend = mock_api_backend(&server); ++ let mut node = Node::new("audit"); ++ node.attrs ++ .insert("output_schema".to_string(), custom_output_schema_attr()); ++ node.attrs ++ .insert("output_retries".to_string(), AttrValue::Integer(1)); ++ let context = Context::new(); ++ let emitter = Arc::new(Emitter::new(fabro_types::RunId::new())); ++ let workspace = tempfile::tempdir().unwrap(); ++ let sandbox: Arc = ++ Arc::new(LocalSandbox::new(workspace.path().to_path_buf())); ++ ++ let result = backend ++ .run(CodergenRunRequest { ++ node: &node, ++ prompt: "Audit the result", ++ context: &context, ++ thread_id: None, ++ emitter: &emitter, ++ sandbox: &sandbox, ++ tool_hooks: None, ++ cancel_token: CancellationToken::new(), ++ agent_tool_runtime: fabro_agent::AgentToolRuntime::default(), ++ }) ++ .await ++ .unwrap(); ++ ++ first.assert_calls(1); ++ repair.assert_calls(1); ++ let CodergenResult::Text { text, usage, .. } = result else { ++ panic!("run should return text"); ++ }; ++ assert_eq!(text, r#"{"passed":true}"#); ++ let usage = usage.expect("usage should be aggregated"); ++ assert_eq!(usage.tokens().input_tokens, 41); ++ assert_eq!(usage.tokens().output_tokens, 7); ++ } ++ + #[tokio::test] + async fn api_backend_shutdown_closes_cached_sessions_once() { + let backend = AgentApiBackend::new_from_env( +diff --git a/lib/crates/fabro-workflow/src/handler/mod.rs b/lib/crates/fabro-workflow/src/handler/mod.rs +index 675f24b39..bc94af3b8 100644 +--- a/lib/crates/fabro-workflow/src/handler/mod.rs ++++ b/lib/crates/fabro-workflow/src/handler/mod.rs +@@ -9,6 +9,7 @@ pub mod manager_loop; + pub mod parallel; + pub mod prompt; + pub mod start; ++pub mod structured_output; + pub mod wait; + + use std::any::Any; +diff --git a/lib/crates/fabro-workflow/src/handler/prompt.rs b/lib/crates/fabro-workflow/src/handler/prompt.rs +index 6dbc0576d..27b9fa524 100644 +--- a/lib/crates/fabro-workflow/src/handler/prompt.rs ++++ b/lib/crates/fabro-workflow/src/handler/prompt.rs +@@ -10,7 +10,7 @@ use super::agent::{ + truncate, + }; + use super::llm::routing; +-use super::{EngineServices, Handler}; ++use super::{EngineServices, Handler, structured_output}; + use crate::context::{Context, WorkflowContext, keys}; + use crate::error::Error; + use crate::event::{Emitter, Event}; +@@ -191,7 +191,25 @@ impl Handler for PromptHandler { + serde_json::json!(&response_text), + ); + +- extract_status_fields(&response_text, &mut outcome); ++ if let Some(schema) = structured_output::parse_node_output_schema(node)? { ++ match structured_output::validate_response_text(&schema, &response_text) { ++ Ok(validated) => { ++ structured_output::apply_validated_output( ++ node, ++ &schema, ++ &validated, ++ &mut outcome, ++ ); ++ } ++ Err(_) => { ++ return Ok(structured_output::exhausted_failure_outcome( ++ node.output_retries(), ++ )); ++ } ++ } ++ } else { ++ extract_status_fields(&response_text, &mut outcome); ++ } + outcome.usage = stage_usage; + outcome.files_touched = backend_files_touched; + +@@ -214,6 +232,7 @@ mod tests { + use super::*; + use crate::event::Emitter; + use crate::handler::agent::CodergenRunRequest; ++ use crate::outcome::OutcomeExt; + + fn make_services() -> EngineServices { + EngineServices::test_default() +@@ -368,6 +387,102 @@ mod tests { + ); + } + ++ #[tokio::test] ++ async fn prompt_handler_custom_output_schema_updates_output_context_key() { ++ struct CustomOutputBackend; ++ ++ #[async_trait] ++ impl CodergenBackend for CustomOutputBackend { ++ async fn run(&self, _request: CodergenRunRequest<'_>) -> Result { ++ panic!("run() should not be called for prompt handler"); ++ } ++ ++ async fn one_shot( ++ &self, ++ _request: OneShotRequest<'_>, ++ ) -> Result { ++ Ok(CodergenResult::Text { ++ text: r#"{"passed": true}"#.to_string(), ++ usage: None, ++ files_touched: Vec::new(), ++ last_file_touched: None, ++ }) ++ } ++ } ++ ++ let handler = PromptHandler::new(Some(Box::new(CustomOutputBackend))); ++ let mut node = Node::new("audit"); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String( ++ r#"{"type":"object","required":["passed"],"properties":{"passed":{"type":"boolean"}}}"# ++ .to_string(), ++ ), ++ ); ++ let context = Context::new(); ++ let graph = Graph::new("test"); ++ let tmp = TempDir::new().unwrap(); ++ ++ let outcome = handler ++ .execute(&node, &context, &graph, tmp.path(), &make_services()) ++ .await ++ .unwrap(); ++ ++ assert_eq!( ++ outcome.context_updates.get("output.audit"), ++ Some(&serde_json::json!({"passed": true})), ++ ); ++ } ++ ++ #[tokio::test] ++ async fn prompt_handler_routing_output_schema_requires_valid_routing_json() { ++ struct BadRoutingBackend; ++ ++ #[async_trait] ++ impl CodergenBackend for BadRoutingBackend { ++ async fn run(&self, _request: CodergenRunRequest<'_>) -> Result { ++ panic!("run() should not be called for prompt handler"); ++ } ++ ++ async fn one_shot( ++ &self, ++ _request: OneShotRequest<'_>, ++ ) -> Result { ++ Ok(CodergenResult::Text { ++ text: r#"{"outcome": 123}"#.to_string(), ++ usage: None, ++ files_touched: Vec::new(), ++ last_file_touched: None, ++ }) ++ } ++ } ++ ++ let handler = PromptHandler::new(Some(Box::new(BadRoutingBackend))); ++ let mut node = Node::new("route"); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String("routing".to_string()), ++ ); ++ node.attrs ++ .insert("output_retries".to_string(), AttrValue::Integer(0)); ++ let context = Context::new(); ++ let graph = Graph::new("test"); ++ let tmp = TempDir::new().unwrap(); ++ ++ let outcome = handler ++ .execute(&node, &context, &graph, tmp.path(), &make_services()) ++ .await ++ .unwrap(); ++ ++ assert_eq!(outcome.status, crate::outcome::StageOutcome::Failed { ++ retry_requested: false, ++ }); ++ assert_eq!( ++ outcome.failure_reason(), ++ Some("output schema validation failed after 0 repair attempt(s)") ++ ); ++ } ++ + #[tokio::test] + async fn prompt_handler_projects_provider_used_from_prompt_events() { + struct ProviderOneShotBackend; +diff --git a/lib/crates/fabro-workflow/src/handler/structured_output.rs b/lib/crates/fabro-workflow/src/handler/structured_output.rs +new file mode 100644 +index 000000000..d98a42afb +--- /dev/null ++++ b/lib/crates/fabro-workflow/src/handler/structured_output.rs +@@ -0,0 +1,608 @@ ++use fabro_graphviz::graph::Node; ++use fabro_llm::types::{ResponseFormat, ResponseFormatType}; ++use serde_json::Value; ++ ++use crate::error::Error; ++use crate::outcome::{FailureCategory, FailureDetail, Outcome, StageOutcome}; ++ ++pub(crate) const ROUTING_STATUS_FIELDS: &[&str] = &[ ++ "preferred_next_label", ++ "outcome", ++ "failure_reason", ++ "suggested_next_ids", ++ "context_updates", ++]; ++ ++#[derive(Debug, Clone, PartialEq)] ++pub(crate) enum OutputSchemaKind { ++ Routing, ++ JsonSchema { schema: Value }, ++} ++ ++#[derive(Debug, Clone, Copy, PartialEq, Eq)] ++pub(crate) enum StructuredOutputErrorKind { ++ NoJsonObject, ++ NoRelevantJsonObject, ++ InvalidJson, ++ SchemaValidation, ++} ++ ++#[derive(Debug, Clone, PartialEq, Eq)] ++pub(crate) struct StructuredOutputError { ++ kind: StructuredOutputErrorKind, ++ messages: Vec, ++} ++ ++impl StructuredOutputError { ++ fn new(kind: StructuredOutputErrorKind, message: impl Into) -> Self { ++ Self { ++ kind, ++ messages: vec![message.into()], ++ } ++ } ++ ++ fn validation(messages: Vec) -> Self { ++ Self { ++ kind: StructuredOutputErrorKind::SchemaValidation, ++ messages, ++ } ++ } ++ ++ #[cfg(test)] ++ #[must_use] ++ pub(crate) fn kind(&self) -> StructuredOutputErrorKind { ++ self.kind ++ } ++ ++ #[cfg(test)] ++ #[must_use] ++ pub(crate) fn messages(&self) -> &[String] { ++ &self.messages ++ } ++ ++ #[must_use] ++ pub(crate) fn allows_routing_fallback(&self) -> bool { ++ matches!( ++ self.kind, ++ StructuredOutputErrorKind::NoJsonObject ++ | StructuredOutputErrorKind::NoRelevantJsonObject ++ ) ++ } ++ ++ #[must_use] ++ pub(crate) fn repair_message(&self, schema: &OutputSchemaKind) -> String { ++ let expectation = match schema { ++ OutputSchemaKind::Routing => format!( ++ "Return a single JSON object with at least one routing field: {}.", ++ ROUTING_STATUS_FIELDS.join(", ") ++ ), ++ OutputSchemaKind::JsonSchema { .. } => { ++ "Return a single JSON object that satisfies the configured JSON Schema.".to_string() ++ } ++ }; ++ let errors = self ++ .messages ++ .iter() ++ .map(|message| format!("- {message}")) ++ .collect::>() ++ .join("\n"); ++ format!( ++ "Your previous response did not satisfy the node's output_schema.\n\n\ ++ Validation errors:\n{errors}\n\n\ ++ {expectation}\n\ ++ Do not include Markdown fences or explanatory prose; reply only with the corrected JSON object." ++ ) ++ } ++} ++ ++#[derive(Debug, Clone, PartialEq)] ++pub(crate) struct ValidatedStructuredOutput { ++ pub(crate) value: Value, ++} ++ ++#[must_use] ++pub(crate) fn output_key(node_id: &str) -> String { ++ format!("output.{node_id}") ++} ++ ++#[must_use] ++pub(crate) fn exhausted_failure_reason(repair_attempts: i64) -> String { ++ format!("output schema validation failed after {repair_attempts} repair attempt(s)") ++} ++ ++#[must_use] ++pub(crate) fn exhausted_failure_outcome(repair_attempts: i64) -> Outcome { ++ Outcome { ++ status: StageOutcome::Failed { ++ retry_requested: false, ++ }, ++ failure: Some(FailureDetail::new( ++ exhausted_failure_reason(repair_attempts), ++ FailureCategory::Deterministic, ++ )), ++ ..Outcome::default() ++ } ++} ++ ++pub(crate) fn parse_node_output_schema(node: &Node) -> Result, Error> { ++ let Some(raw) = node.output_schema() else { ++ return Ok(None); ++ }; ++ let value = raw.trim(); ++ if value.is_empty() { ++ return Err(Error::Validation(format!( ++ "Invalid output_schema for node \"{}\": value must not be empty", ++ node.id ++ ))); ++ } ++ if value == "routing" { ++ return Ok(Some(OutputSchemaKind::Routing)); ++ } ++ if value.starts_with('@') { ++ return Err(Error::Validation(format!( ++ "Invalid output_schema for node \"{}\": unresolved file reference {value}", ++ node.id ++ ))); ++ } ++ ++ let schema = serde_json::from_str::(value).map_err(|err| { ++ Error::Validation(format!( ++ "Invalid output_schema for node \"{}\": expected \"routing\" or a JSON Schema object: {err}", ++ node.id ++ )) ++ })?; ++ compile_schema(&schema).map_err(|err| { ++ Error::Validation(format!( ++ "Invalid output_schema for node \"{}\": {err}", ++ node.id ++ )) ++ })?; ++ Ok(Some(OutputSchemaKind::JsonSchema { schema })) ++} ++ ++#[must_use] ++pub(crate) fn prompt_response_format(schema: &OutputSchemaKind) -> ResponseFormat { ++ match schema { ++ OutputSchemaKind::Routing => ResponseFormat { ++ kind: ResponseFormatType::JsonObject, ++ json_schema: None, ++ strict: false, ++ }, ++ OutputSchemaKind::JsonSchema { schema } => ResponseFormat { ++ kind: ResponseFormatType::JsonSchema, ++ json_schema: Some(schema.clone()), ++ strict: true, ++ }, ++ } ++} ++ ++pub(crate) fn validate_response_text( ++ schema: &OutputSchemaKind, ++ text: &str, ++) -> Result { ++ match schema { ++ OutputSchemaKind::Routing => validate_routing_response_text(text), ++ OutputSchemaKind::JsonSchema { schema } => validate_custom_response_text(schema, text), ++ } ++} ++ ++pub(crate) fn apply_validated_output( ++ node: &Node, ++ schema: &OutputSchemaKind, ++ validated: &ValidatedStructuredOutput, ++ outcome: &mut Outcome, ++) { ++ match schema { ++ OutputSchemaKind::Routing => apply_routing_fields(&validated.value, outcome), ++ OutputSchemaKind::JsonSchema { .. } => { ++ outcome ++ .context_updates ++ .insert(output_key(&node.id), validated.value.clone()); ++ } ++ } ++} ++ ++/// Find all balanced `{...}` JSON object substrings in the text. ++pub(crate) fn find_json_objects(text: &str) -> Vec<&str> { ++ let mut results = Vec::new(); ++ let bytes = text.as_bytes(); ++ let mut i = 0; ++ while i < bytes.len() { ++ if bytes[i] == b'{' { ++ let start = i; ++ let mut depth = 0; ++ let mut in_string = false; ++ let mut escape = false; ++ let mut j = i; ++ while j < bytes.len() { ++ let c = bytes[j]; ++ if escape { ++ escape = false; ++ } else if c == b'\\' && in_string { ++ escape = true; ++ } else if c == b'"' { ++ in_string = !in_string; ++ } else if !in_string { ++ if c == b'{' { ++ depth += 1; ++ } else if c == b'}' { ++ depth -= 1; ++ if depth == 0 { ++ results.push(&text[start..=j]); ++ break; ++ } ++ } ++ } ++ j += 1; ++ } ++ } ++ i += 1; ++ } ++ results ++} ++ ++pub(crate) fn extract_status_fields_loose(text: &str, outcome: &mut Outcome) -> bool { ++ let candidates = find_json_objects(text); ++ ++ let parsed = candidates.iter().rev().find_map(|candidate| { ++ let value: Value = serde_json::from_str(candidate).ok()?; ++ if value.as_object().is_some_and(contains_routing_field) { ++ Some(value) ++ } else { ++ None ++ } ++ }); ++ ++ let Some(value) = parsed else { return false }; ++ apply_routing_fields(&value, outcome); ++ true ++} ++ ++fn validate_routing_response_text( ++ text: &str, ++) -> Result { ++ let candidates = find_json_objects(text); ++ if candidates.is_empty() { ++ return Err(StructuredOutputError::new( ++ StructuredOutputErrorKind::NoJsonObject, ++ "no JSON object found in response", ++ )); ++ } ++ ++ for candidate in candidates.iter().rev() { ++ let parsed = match serde_json::from_str::(candidate) { ++ Ok(value) => value, ++ Err(err) if raw_mentions_routing_field(candidate) => { ++ return Err(StructuredOutputError::new( ++ StructuredOutputErrorKind::InvalidJson, ++ format!("invalid routing JSON object: {err}"), ++ )); ++ } ++ Err(_) => continue, ++ }; ++ let Some(obj) = parsed.as_object() else { ++ continue; ++ }; ++ if !contains_routing_field(obj) { ++ continue; ++ } ++ validate_value_against_schema(&routing_schema(), &parsed)?; ++ return Ok(ValidatedStructuredOutput { value: parsed }); ++ } ++ ++ Err(StructuredOutputError::new( ++ StructuredOutputErrorKind::NoRelevantJsonObject, ++ format!( ++ "no JSON object contained any recognized routing field ({})", ++ ROUTING_STATUS_FIELDS.join(", ") ++ ), ++ )) ++} ++ ++fn validate_custom_response_text( ++ schema: &Value, ++ text: &str, ++) -> Result { ++ let candidates = find_json_objects(text); ++ let Some(candidate) = candidates.last() else { ++ return Err(StructuredOutputError::new( ++ StructuredOutputErrorKind::NoJsonObject, ++ "no JSON object found in response", ++ )); ++ }; ++ let parsed = serde_json::from_str::(candidate).map_err(|err| { ++ StructuredOutputError::new( ++ StructuredOutputErrorKind::InvalidJson, ++ format!("invalid JSON object: {err}"), ++ ) ++ })?; ++ validate_value_against_schema(schema, &parsed)?; ++ Ok(ValidatedStructuredOutput { value: parsed }) ++} ++ ++fn validate_value_against_schema( ++ schema: &Value, ++ value: &Value, ++) -> Result<(), StructuredOutputError> { ++ let validator = compile_schema(schema).map_err(|err| { ++ StructuredOutputError::new( ++ StructuredOutputErrorKind::SchemaValidation, ++ format!("invalid JSON Schema: {err}"), ++ ) ++ })?; ++ let errors = validator ++ .iter_errors(value) ++ .map(|error| error.to_string()) ++ .take(5) ++ .collect::>(); ++ if errors.is_empty() { ++ Ok(()) ++ } else { ++ Err(StructuredOutputError::validation(errors)) ++ } ++} ++ ++fn compile_schema( ++ schema: &Value, ++) -> Result> { ++ jsonschema::validator_for(schema) ++} ++ ++fn contains_routing_field(obj: &serde_json::Map) -> bool { ++ ROUTING_STATUS_FIELDS ++ .iter() ++ .any(|field| obj.contains_key(*field)) ++} ++ ++fn raw_mentions_routing_field(candidate: &str) -> bool { ++ ROUTING_STATUS_FIELDS ++ .iter() ++ .any(|field| candidate.contains(&format!("\"{field}\""))) ++} ++ ++fn routing_schema() -> Value { ++ serde_json::json!({ ++ "type": "object", ++ "additionalProperties": true, ++ "properties": { ++ "preferred_next_label": { "type": "string" }, ++ "outcome": { ++ "type": "string", ++ "enum": ["succeeded", "partially_succeeded", "failed", "skipped"] ++ }, ++ "failure_reason": { "type": "string" }, ++ "suggested_next_ids": { ++ "type": "array", ++ "items": { "type": "string" } ++ }, ++ "context_updates": { "type": "object" } ++ }, ++ "anyOf": ROUTING_STATUS_FIELDS ++ .iter() ++ .map(|field| serde_json::json!({ "required": [field] })) ++ .collect::>() ++ }) ++} ++ ++fn apply_routing_fields(value: &Value, outcome: &mut Outcome) { ++ let Some(obj) = value.as_object() else { ++ return; ++ }; ++ ++ if let Some(label) = obj.get("preferred_next_label").and_then(Value::as_str) { ++ outcome.preferred_label = Some(label.to_string()); ++ } ++ ++ if let Some(ids) = obj.get("suggested_next_ids").and_then(Value::as_array) { ++ let string_ids: Vec = ids ++ .iter() ++ .filter_map(|value| value.as_str().map(String::from)) ++ .collect(); ++ if !string_ids.is_empty() { ++ outcome.suggested_next_ids = string_ids; ++ } ++ } ++ ++ if let Some(status_str) = obj.get("outcome").and_then(Value::as_str) { ++ if let Ok(status) = status_str.parse::() { ++ outcome.status = status; ++ if outcome.status.is_failure() { ++ if let Some(reason) = obj.get("failure_reason").and_then(Value::as_str) { ++ outcome.failure = ++ Some(FailureDetail::new(reason, FailureCategory::Deterministic)); ++ } ++ } ++ } ++ } ++ ++ if let Some(updates) = obj.get("context_updates").and_then(Value::as_object) { ++ for (key, value) in updates { ++ outcome.context_updates.insert(key.clone(), value.clone()); ++ } ++ } ++} ++ ++#[cfg(test)] ++mod tests { ++ use fabro_graphviz::graph::{AttrValue, Node}; ++ ++ use super::*; ++ ++ fn routing() -> OutputSchemaKind { ++ OutputSchemaKind::Routing ++ } ++ ++ fn schema(value: Value) -> OutputSchemaKind { ++ OutputSchemaKind::JsonSchema { schema: value } ++ } ++ ++ #[test] ++ fn validates_routing_json_and_applies_fields() { ++ let validated = validate_response_text( ++ &routing(), ++ r#"done {"outcome":"failed","failure_reason":"tests failed","preferred_next_label":"fix","suggested_next_ids":["a"],"context_updates":{"verified":true}}"#, ++ ) ++ .unwrap(); ++ let mut outcome = Outcome::success(); ++ ++ apply_routing_fields(&validated.value, &mut outcome); ++ ++ assert_eq!(outcome.status, StageOutcome::Failed { ++ retry_requested: false, ++ }); ++ assert_eq!( ++ outcome.failure.as_ref().map(|f| f.message.as_str()), ++ Some("tests failed") ++ ); ++ assert_eq!(outcome.preferred_label.as_deref(), Some("fix")); ++ assert_eq!(outcome.suggested_next_ids, vec!["a".to_string()]); ++ assert_eq!( ++ outcome.context_updates.get("verified"), ++ Some(&serde_json::json!(true)), ++ ); ++ } ++ ++ #[test] ++ fn routing_json_missing_routing_fields_is_invalid() { ++ let error = validate_response_text(&routing(), r#"{"summary":"ok"}"#).unwrap_err(); ++ ++ assert_eq!( ++ error.kind(), ++ StructuredOutputErrorKind::NoRelevantJsonObject ++ ); ++ assert!(error.messages()[0].contains("recognized routing field")); ++ } ++ ++ #[test] ++ fn routing_json_with_wrong_field_type_is_invalid() { ++ let error = ++ validate_response_text(&routing(), r#"{"suggested_next_ids":[1]}"#).unwrap_err(); ++ ++ assert_eq!(error.kind(), StructuredOutputErrorKind::SchemaValidation); ++ assert!( ++ error ++ .messages() ++ .iter() ++ .any(|message| message.contains("string")), ++ "unexpected messages: {:?}", ++ error.messages(), ++ ); ++ } ++ ++ #[test] ++ fn validates_custom_schema_against_last_json_object() { ++ let schema = schema(serde_json::json!({ ++ "type": "object", ++ "required": ["passed"], ++ "properties": { ++ "passed": { "type": "boolean" } ++ } ++ })); ++ ++ let validated = ++ validate_response_text(&schema, r#"ignore {"other":1} final {"passed":true}"#).unwrap(); ++ ++ assert_eq!(validated.value, serde_json::json!({"passed": true})); ++ } ++ ++ #[test] ++ fn custom_schema_validation_errors_are_reported() { ++ let schema = schema(serde_json::json!({ ++ "type": "object", ++ "required": ["passed"], ++ "properties": { ++ "passed": { "type": "boolean" } ++ } ++ })); ++ ++ let error = validate_response_text(&schema, r#"{"passed":"yes"}"#).unwrap_err(); ++ ++ assert_eq!(error.kind(), StructuredOutputErrorKind::SchemaValidation); ++ assert!( ++ error ++ .messages() ++ .iter() ++ .any(|message| message.contains("boolean")), ++ "unexpected messages: {:?}", ++ error.messages(), ++ ); ++ } ++ ++ #[test] ++ fn invalid_custom_schema_is_rejected_when_parsing_node_attr() { ++ let mut node = Node::new("audit"); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String(r#"{"type": 5}"#.to_string()), ++ ); ++ ++ let error = parse_node_output_schema(&node).unwrap_err(); ++ ++ assert!( ++ error.to_string().contains("Invalid output_schema"), ++ "unexpected error: {error}", ++ ); ++ } ++ ++ #[test] ++ fn invalid_json_candidate_is_reported_for_custom_schema() { ++ let schema = schema(serde_json::json!({"type": "object"})); ++ ++ let error = validate_response_text(&schema, r"{not json}").unwrap_err(); ++ ++ assert_eq!(error.kind(), StructuredOutputErrorKind::InvalidJson); ++ assert!(error.messages()[0].contains("invalid JSON object")); ++ } ++ ++ #[test] ++ fn no_json_object_is_reported() { ++ let error = validate_response_text(&routing(), "plain text only").unwrap_err(); ++ ++ assert_eq!(error.kind(), StructuredOutputErrorKind::NoJsonObject); ++ assert!(error.messages()[0].contains("no JSON object")); ++ } ++ ++ #[test] ++ fn parse_node_output_schema_accepts_builtin_routing_keyword() { ++ let mut node = Node::new("route"); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String("routing".to_string()), ++ ); ++ ++ let parsed = parse_node_output_schema(&node).unwrap(); ++ ++ assert_eq!(parsed, Some(OutputSchemaKind::Routing)); ++ } ++ ++ #[test] ++ fn prompt_response_format_uses_json_schema_for_custom_schema() { ++ let schema = schema(serde_json::json!({"type": "object"})); ++ ++ let format = prompt_response_format(&schema); ++ ++ assert_eq!(format.kind, ResponseFormatType::JsonSchema); ++ assert_eq!( ++ format.json_schema, ++ Some(serde_json::json!({"type": "object"})) ++ ); ++ assert!(format.strict); ++ } ++ ++ #[test] ++ fn apply_validated_custom_output_updates_output_context_key() { ++ let node = Node::new("audit"); ++ let schema = schema(serde_json::json!({"type": "object"})); ++ let validated = ValidatedStructuredOutput { ++ value: serde_json::json!({"passed": true}), ++ }; ++ let mut outcome = Outcome::success(); ++ ++ apply_validated_output(&node, &schema, &validated, &mut outcome); ++ ++ assert_eq!( ++ outcome.context_updates.get("output.audit"), ++ Some(&serde_json::json!({"passed": true})), ++ ); ++ } ++} +diff --git a/lib/crates/fabro-workflow/src/static_reference.rs b/lib/crates/fabro-workflow/src/static_reference.rs +index 276cf37c1..af83539eb 100644 +--- a/lib/crates/fabro-workflow/src/static_reference.rs ++++ b/lib/crates/fabro-workflow/src/static_reference.rs +@@ -87,9 +87,57 @@ pub fn reference_kind_for_attribute( + "goal" if matches!(scope, AttributeScope::Graph) && value.starts_with('@') => { + Some(ReferenceKind::GraphGoalFile) + } +- "prompt" if matches!(scope, AttributeScope::Node) && value.starts_with('@') => { ++ "prompt" | "output_schema" ++ if matches!(scope, AttributeScope::Node) && value.starts_with('@') => ++ { + Some(ReferenceKind::FileInline) + } + _ => None, + } + } ++ ++#[cfg(test)] ++mod tests { ++ use super::*; ++ ++ #[test] ++ fn output_schema_at_value_is_file_inline_reference() { ++ assert_eq!( ++ reference_kind_for_attribute( ++ AttributeScope::Node, ++ "output_schema", ++ "@schemas/result.schema.json", ++ ), ++ Some(ReferenceKind::FileInline), ++ ); ++ } ++ ++ #[test] ++ fn output_schema_builtin_keyword_is_not_file_inline_reference() { ++ assert_eq!( ++ reference_kind_for_attribute(AttributeScope::Node, "output_schema", "routing"), ++ None, ++ ); ++ } ++ ++ #[test] ++ fn output_schema_reference_rejects_template_syntax() { ++ let error = reference_kind_for_attribute( ++ AttributeScope::Node, ++ "output_schema", ++ "@schemas/{{ inputs.schema }}.json", ++ ) ++ .expect("output_schema @ references should be static references") ++ .validate("@schemas/{{ inputs.schema }}.json") ++ .unwrap_err(); ++ ++ assert_eq!(error.kind(), ReferenceKind::FileInline); ++ assert_eq!(error.value(), "@schemas/{{ inputs.schema }}.json"); ++ assert!( ++ error ++ .to_string() ++ .contains("templates are not supported in file inline references"), ++ "unexpected error: {error}", ++ ); ++ } ++} +diff --git a/lib/crates/fabro-workflow/src/transforms/file_inlining.rs b/lib/crates/fabro-workflow/src/transforms/file_inlining.rs +index 1e71699b1..942ba4af6 100644 +--- a/lib/crates/fabro-workflow/src/transforms/file_inlining.rs ++++ b/lib/crates/fabro-workflow/src/transforms/file_inlining.rs +@@ -192,33 +192,46 @@ impl FileInliningTransform { + .with_inputs(self.inputs.clone()); + + for (node_id, node) in &mut graph.nodes { +- let Some(AttrValue::String(prompt)) = node.attrs.get("prompt") else { +- continue; +- }; +- let target = TemplateRenderTarget::node_attr( +- self.source_name.clone(), +- node_id.clone(), +- "prompt", +- ) +- .with_source_origin(self.source_text.as_deref(), prompt) +- .with_template_store(template_render_store( +- &self.current_dir, +- Arc::clone(&self.resolver), +- self.source_name.as_deref(), +- prompt, +- )?); +- let rendered = render_template_for_target( +- prompt, +- &ctx, +- self.render_mode, +- &target, +- &mut diagnostics, +- )?; +- let value = self +- .render_resolved_file_ref(&rendered, &ctx, target, &mut diagnostics)? +- .unwrap_or(rendered); +- node.attrs +- .insert("prompt".to_string(), AttrValue::String(value)); ++ for attr_name in ["prompt", "output_schema"] { ++ let Some(AttrValue::String(attr_value)) = node.attrs.get(attr_name) else { ++ continue; ++ }; ++ let target = TemplateRenderTarget::node_attr( ++ self.source_name.clone(), ++ node_id.clone(), ++ attr_name, ++ ) ++ .with_source_origin(self.source_text.as_deref(), attr_value) ++ .with_template_store(template_render_store( ++ &self.current_dir, ++ Arc::clone(&self.resolver), ++ self.source_name.as_deref(), ++ attr_value, ++ )?); ++ let rendered = render_template_for_target( ++ attr_value, ++ &ctx, ++ self.render_mode, ++ &target, ++ &mut diagnostics, ++ )?; ++ let value = match self.render_resolved_file_ref( ++ &rendered, ++ &ctx, ++ target, ++ &mut diagnostics, ++ )? { ++ Some(value) => value, ++ None if attr_name == "output_schema" && rendered.starts_with('@') => { ++ return Err(Error::Validation(format!( ++ "node '{node_id}' output_schema has unresolved file reference: {rendered}" ++ ))); ++ } ++ None => rendered, ++ }; ++ node.attrs ++ .insert(attr_name.to_string(), AttrValue::String(value)); ++ } + } + + Ok((graph, diagnostics)) +@@ -465,6 +478,90 @@ mod tests { + ); + } + ++ #[test] ++ fn file_inlining_transform_inlines_output_schema_reference() { ++ let dir = tempfile::tempdir().unwrap(); ++ std::fs::create_dir_all(dir.path().join("schemas")).unwrap(); ++ std::fs::write( ++ dir.path().join("schemas/audit-result.schema.json"), ++ r#"{"type":"object","required":["passed"]}"#, ++ ) ++ .unwrap(); ++ ++ let mut graph = Graph::new("test"); ++ let mut node = Node::new("audit"); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String("@schemas/audit-result.schema.json".to_string()), ++ ); ++ graph.nodes.insert("audit".to_string(), node); ++ ++ let transform = FileInliningTransform::new( ++ dir.path().to_path_buf(), ++ Arc::new(FilesystemFileResolver::new(None)), ++ ); ++ let graph = transform.apply(graph).unwrap(); ++ ++ assert_eq!( ++ graph.nodes["audit"] ++ .attrs ++ .get("output_schema") ++ .and_then(AttrValue::as_str), ++ Some(r#"{"type":"object","required":["passed"]}"#) ++ ); ++ } ++ ++ #[test] ++ fn file_inlining_transform_leaves_routing_output_schema_keyword_unchanged() { ++ let dir = tempfile::tempdir().unwrap(); ++ let mut graph = Graph::new("test"); ++ let mut node = Node::new("route"); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String("routing".to_string()), ++ ); ++ graph.nodes.insert("route".to_string(), node); ++ ++ let transform = FileInliningTransform::new( ++ dir.path().to_path_buf(), ++ Arc::new(FilesystemFileResolver::new(None)), ++ ); ++ let graph = transform.apply(graph).unwrap(); ++ ++ assert_eq!( ++ graph.nodes["route"] ++ .attrs ++ .get("output_schema") ++ .and_then(AttrValue::as_str), ++ Some("routing") ++ ); ++ } ++ ++ #[test] ++ fn file_inlining_transform_reports_unresolved_output_schema_reference() { ++ let dir = tempfile::tempdir().unwrap(); ++ let mut graph = Graph::new("test"); ++ let mut node = Node::new("audit"); ++ node.attrs.insert( ++ "output_schema".to_string(), ++ AttrValue::String("@schemas/missing.schema.json".to_string()), ++ ); ++ graph.nodes.insert("audit".to_string(), node); ++ ++ let transform = FileInliningTransform::new( ++ dir.path().to_path_buf(), ++ Arc::new(FilesystemFileResolver::new(None)), ++ ); ++ let error = transform.apply(graph).unwrap_err(); ++ ++ assert!( ++ error.to_string().contains( ++ "node 'audit' output_schema has unresolved file reference: @schemas/missing.schema.json" ++ ), ++ "unexpected error: {error}", ++ ); ++ } ++ + #[test] + fn file_inlining_transform_resolves_minijinja_includes_for_prompts_and_goal() { + let dir = tempfile::tempdir().unwrap(); diff --git a/stages/005-implement@1/status.json b/stages/005-implement@1/status.json new file mode 100644 index 000000000..d12d9ea32 --- /dev/null +++ b/stages/005-implement@1/status.json @@ -0,0 +1,6 @@ +{ + "outcome": "succeeded", + "notes": "Stage completed: implement", + "failure_reason": null, + "timestamp": "2026-05-23T20:26:00.319401Z" +} \ No newline at end of file diff --git a/stages/006-simplify_opus@1/prompt.md b/stages/006-simplify_opus@1/prompt.md new file mode 100644 index 000000000..584ed6a26 --- /dev/null +++ b/stages/006-simplify_opus@1/prompt.md @@ -0,0 +1,306 @@ +Goal: # Output Schema Validation Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Add `output_schema` validation for agent and prompt nodes, with context-preserving repair turns when structured output does not validate. + +**Architecture:** Introduce a small structured-output layer in `fabro-workflow` that resolves node-level schema declarations, extracts JSON output, validates it, and produces either routing side effects or a parsed custom output context update. Agent and prompt execution must perform schema repair inside the active LLM conversation instead of using the workflow executor retry path. + +**Tech Stack:** Rust, Graphviz workflow attrs, `serde_json`, workspace `jsonschema`, existing `fabro-llm::ResponseFormat`, Fabro agent sessions, `cargo nextest`. + +--- + +## Public Interface + +Workflow authors can opt in on agent and prompt nodes: + +```dot +review [ + shape=tab, + output_schema="routing", + output_retries=2 +] + +audit [ + shape=tab, + output_schema="@schemas/audit-result.schema.json", + output_retries=2 +] +``` + +- `output_schema="routing"` uses Fabro's built-in routing directive schema. +- `output_schema="@path/to/schema.json"` loads a JSON Schema file through existing workflow file-reference rules. +- `output_retries` controls corrective turns inside the same node execution. Default: `2`. `0` means validate once and fail without a repair turn. +- Schema failures are terminal node failures after `output_retries` is exhausted. They are not `retry_requested` outcomes and do not consume `max_retries`. +- `backend="acp"` with `output_schema` is unsupported in v1 and returns a clear validation error. + +## Implementation Tasks + +### Task 1: Node Attributes And File Reference Resolution + +**Files:** +- Modify: `lib/crates/fabro-types/src/graph.rs` +- Modify: `lib/crates/fabro-workflow/src/static_reference.rs` +- Modify: `lib/crates/fabro-workflow/src/transforms/file_inlining.rs` +- Test: existing unit tests in those files + +- [ ] Add `Node::output_schema(&self) -> Option<&str>` next to other agent/prompt attrs. +- [ ] Add `Node::output_retries(&self) -> i64` returning `self.int_attr("output_retries").unwrap_or(2).max(0)`. +- [ ] Teach static reference validation that node attr `output_schema` values starting with `@` are file inline references. +- [ ] Extend file inlining so `output_schema="@schemas/foo.json"` is replaced with the schema file contents before execution, while `output_schema="routing"` stays unchanged. +- [ ] Add tests for absent attrs, default retries, zero retries, file inlining, and unresolved schema reference diagnostics. + +Run: + +```bash +cargo nextest run -p fabro-types -p fabro-workflow graph:: file_inlining static_reference +``` + +Expected: targeted tests pass. + +### Task 2: Structured Output Module + +**Files:** +- Create: `lib/crates/fabro-workflow/src/handler/structured_output.rs` +- Modify: `lib/crates/fabro-workflow/src/handler/mod.rs` +- Modify: `lib/crates/fabro-workflow/Cargo.toml` +- Test: unit tests in `structured_output.rs` + +- [ ] Add `jsonschema.workspace = true` to `fabro-workflow` dependencies. +- [ ] Define `OutputSchemaKind` with `Routing` and `JsonSchema { schema: serde_json::Value }`. +- [ ] Parse `node.output_schema()` into `None`, `Routing`, or custom JSON Schema. Treat literal `routing` as the only built-in keyword. +- [ ] Add a built-in routing schema requiring an object with at least one recognized field: `preferred_next_label`, `outcome`, `failure_reason`, `suggested_next_ids`, or `context_updates`. +- [ ] Reuse balanced-object scanning semantics for response text: validate the last JSON object that is relevant to the selected schema. +- [ ] Return a structured validation result containing the parsed JSON object, concise error messages, and enough information to build a repair prompt. +- [ ] Add tests for valid routing JSON, missing routing fields, wrong routing field types, valid custom schema, invalid custom schema, invalid JSON, and no JSON object. + +Run: + +```bash +cargo nextest run -p fabro-workflow structured_output +``` + +Expected: structured-output unit tests pass. + +### Task 3: Routing Extraction Compatibility + +**Files:** +- Modify: `lib/crates/fabro-workflow/src/handler/agent.rs` +- Test: existing agent handler unit tests + +- [ ] Keep the loose default unchanged when `output_schema` is absent. +- [ ] Move current `STATUS_FIELDS`, balanced JSON scanning, and routing-field application behind reusable functions in `structured_output.rs` or call the new module from `agent.rs`. +- [ ] For `output_schema="routing"`, require schema-valid routing JSON and surface validation failures for repair instead of silently ignoring bad candidates. +- [ ] Preserve existing routing fallback priority for agent nodes: response text first, then `status.json`, then last file touched. +- [ ] Keep prompt-node routing behavior response-only unless later tasks explicitly add prompt `status.json` support. + +Run: + +```bash +cargo nextest run -p fabro-workflow handler::agent +``` + +Expected: existing loose routing tests still pass, plus new strict routing tests pass. + +### Task 4: Prompt Node Same-Context Repair + +**Files:** +- Modify: `lib/crates/fabro-workflow/src/handler/llm/api.rs` +- Modify: `lib/crates/fabro-workflow/src/handler/prompt.rs` +- Test: prompt/API backend tests in those files + +- [ ] In `AgentApiBackend::one_shot`, keep `messages` mutable across attempts. +- [ ] When a prompt node has a custom JSON Schema, set `response_format=JsonSchema` on the initial and repair LLM requests. For `routing`, use `JsonObject` or no provider-native schema if provider behavior would conflict with Fabro's routing extraction. +- [ ] After each LLM response, validate according to `output_schema`. +- [ ] On validation failure with repair attempts remaining, append `Message::assistant(response.text())`, then append a corrective `Message::user(repair_message)`, and call `client.complete` again with the same messages. +- [ ] On success, return the validated response text and aggregate usage across all attempts. +- [ ] On exhaustion, return a terminal failed outcome with failure reason `output schema validation failed after N repair attempt(s)`. +- [ ] Update `PromptHandler` so validated custom output is added to `context_updates["output.{node_id}"]`; routing output still updates outcome routing fields. + +Run: + +```bash +cargo nextest run -p fabro-workflow handler::prompt handler::llm::api +``` + +Expected: prompt repair keeps previous assistant output in the message list and succeeds after a corrective response. + +### Task 5: Agent Node Same-Session Repair + +**Files:** +- Modify: `lib/crates/fabro-workflow/src/handler/llm/api.rs` +- Modify: `lib/crates/fabro-workflow/src/handler/agent.rs` +- Test: agent/API backend tests in those files + +- [ ] In `AgentApiBackend::run`, validate the final assistant response before releasing, closing, or caching the session. +- [ ] On validation failure with repair attempts remaining, call `session.process_input(repair_message)` on the same `Session`. +- [ ] Recompute the final assistant response after each repair turn from `session.history()`. +- [ ] Aggregate usage across all new assistant turns, including repair turns, without double-counting reused session history. +- [ ] Do not set provider-native `response_format` for agent sessions in v1, because agent sessions may need normal tool-use messages before final output. +- [ ] Return terminal failure after exhaustion; do not return a retryable backend error and do not request workflow node retry. +- [ ] Update `AgentHandler` to apply validated routing/custom output to the final `Outcome`. + +Run: + +```bash +cargo nextest run -p fabro-workflow handler::agent handler::llm::api +``` + +Expected: agent repair sends a second `process_input` to the same session and final validated output drives outcome/context updates. + +### Task 6: ACP Guardrail + +**Files:** +- Modify: `lib/crates/fabro-workflow/src/handler/llm/acp.rs` +- Test: ACP backend tests in that file + +- [ ] At the start of `AgentAcpBackend::run`, reject nodes where `node.output_schema().is_some()`. +- [ ] Use a clear error message: `output_schema is not supported with backend="acp" in this release`. +- [ ] Add a test proving the ACP backend does not launch a process when `output_schema` is present. + +Run: + +```bash +cargo nextest run -p fabro-workflow handler::llm::acp +``` + +Expected: ACP guardrail test passes. + +### Task 7: Docs + +**Files:** +- Modify: `docs/public/agents/outputs.mdx` +- Modify: `docs/public/reference/dot-language.mdx` + +- [ ] Document `output_schema="routing"` and `output_schema="@schema.json"` under routing/structured outputs. +- [ ] Document same-context repair behavior explicitly: Fabro sends validation feedback to the same agent/prompt context before failing. +- [ ] Document `output_retries`, default `2`, and distinction from `max_retries`. +- [ ] Document v1 scope: agent/prompt nodes only; ACP unsupported; custom schema output stored at `output.{node_id}`. + +Run: + +```bash +rg -n "output_schema|output_retries|output\\." docs/public/agents/outputs.mdx docs/public/reference/dot-language.mdx +``` + +Expected: docs mention the new attrs and storage behavior. + +### Task 8: Full Verification + +**Files:** +- No new files beyond prior tasks + +- [ ] Run focused workflow tests: + +```bash +cargo nextest run -p fabro-workflow +``` + +- [ ] Run formatting check: + +```bash +cargo +nightly-2026-04-14 fmt --check --all +``` + +- [ ] Run clippy: + +```bash +cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings +``` + +- [ ] If snapshots change, inspect before accepting: + +```bash +cargo insta pending-snapshots +``` + +Only run `cargo insta accept` after verifying every pending snapshot is expected. + +## Acceptance Criteria + +- Existing workflows without `output_schema` behave exactly as before. +- `output_schema="routing"` prevents malformed/missing routing JSON from silently falling through to normal edge selection. +- Invalid structured output results in a corrective LLM turn in the same context window. +- Prompt repair preserves previous assistant output in the message list. +- Agent repair preserves the same live session and does not re-run the node from scratch. +- Exhausted output repair attempts produce a clear terminal failure. +- Custom schema output is available to downstream nodes at `output.{node_id}`. +- Docs clearly distinguish `output_retries` from `max_retries`. + +## Assumptions + +- `output_retries=2` is the default. +- Custom schema validation targets the final JSON object in the response text. +- `status.json` fallback remains routing-specific. +- Provider-native response schema is used for prompt nodes only where it is safe. +- ACP support can be added later after there is a guaranteed context-preserving repair mechanism. + + +## Completed stages +- **toolchain**: succeeded + - Script: `command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1` + - Output: + ``` + cargo 1.95.0 (f2d3ce0bd 2026-03-21) + ``` +- **preflight_compile**: succeeded + - Script: `cargo check -q --workspace 2>&1` + - Output: (empty) +- **preflight_lint**: succeeded + - Script: `cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1` + - Output: (empty) +- **implement**: succeeded + - Model: gpt-5.5, 329.8k tokens in / 54.8k out + - Files: /home/daytona/workspace/fabro/lib/crates/fabro-workflow/src/handler/structured_output.rs + + +# Simplify: Code Review and Cleanup + +Review changes vs. origin for reuse, quality, and efficiency. Fix any issues found. + +## Phase 1: Identify Changes + +Run git diff (or git diff HEAD if there are staged changes) to see what changed. If there are no git changes, review the most recently modified files that the user mentioned or that you edited earlier in this conversation. + +## Phase 2: Launch Three Review Agents in Parallel + +Use the Agent tool to launch all three agents concurrently in a single message. Pass each agent the full diff so it has the complete context. + +### Agent 1: Code Reuse Review + +For each change: + +1. Search for existing utilities and helpers that could replace newly written code. Use Grep to find similar patterns elsewhere in the codebase — common locations are utility directories, shared modules, and files adjacent to the changed ones. +2. Flag any new function that duplicates existing functionality. Suggest the existing function to use instead. +3. Flag any inline logic that could use an existing utility — hand-rolled string manipulation, manual path handling, custom environment checks, ad-hoc type guards, and similar patterns are common candidates. + +Note: This is a greenfield app, so focus on maximizing simplicity and don't worry about changing things to achieve it. + +### Agent 2: Code Quality Review + +Review the same changes for hacky patterns: + +1. Redundant state: state that duplicates existing state, cached values that could be derived, observers/effects that could be direct calls +2. Parameter sprawl: adding new parameters to a function instead of generalizing or restructuring existing ones +3. Copy-paste with slight variation: near-duplicate code blocks that should be unified with a shared abstraction +4. Leaky abstractions: exposing internal details that should be encapsulated, or breaking existing abstraction boundaries +5. Stringly-typed code: using raw strings where constants, enums (string unions), or branded types already exist in the codebase + +Note: This is a greenfield app, so be aggressive in optimizing quality. + +### Agent 3: Efficiency Review + +Review the same changes for efficiency: + +1. Unnecessary work: redundant computations, repeated file reads, duplicate network/API calls, N+1 patterns +2. Missed concurrency: independent operations run sequentially when they could run in parallel +3. Hot-path bloat: new blocking work added to startup or per-request/per-render hot paths +4. Unnecessary existence checks: pre-checking file/resource existence before operating (TOCTOU anti-pattern) — operate directly and handle the error +5. Memory: unbounded data structures, missing cleanup, event listener leaks +6. Overly broad operations: reading entire files when only a portion is needed, loading all items when filtering for one + +## Phase 3: Fix Issues + +Wait for all three agents to complete. Aggregate their findings and fix each issue directly. If a finding is a false positive or not worth addressing, note it and move on — do not argue with the finding, just skip it. + +When done, briefly summarize what was fixed (or confirm the code was already clean). \ No newline at end of file diff --git a/stages/006-simplify_opus@1/provider_used.json b/stages/006-simplify_opus@1/provider_used.json new file mode 100644 index 000000000..672cc50e5 --- /dev/null +++ b/stages/006-simplify_opus@1/provider_used.json @@ -0,0 +1,5 @@ +{ + "mode": "agent", + "provider": "anthropic", + "model": "claude-opus-4-7" +} \ No newline at end of file diff --git a/stages/006-simplify_opus@1/response.md b/stages/006-simplify_opus@1/response.md new file mode 100644 index 000000000..b075cdb89 --- /dev/null +++ b/stages/006-simplify_opus@1/response.md @@ -0,0 +1,19 @@ +## Summary + +Three review agents flagged ~25 findings; I applied the high-impact, low-risk fixes (about 12 lines net reduction across the touched files, plus correctness wins on schema compilation). + +**Fixed:** +1. **Schema compiled twice per node** (efficiency #1, reuse #3, quality #1) — `OutputSchemaKind::JsonSchema` now stores an `Arc` compiled once in `parse_node_output_schema`; the routing validator is a `LazyLock`, so repair iterations no longer recompile. Also removed the redundant `compile_schema` helper. +2. **`read_sandbox_file` reimplemented `Sandbox::read_file_text`** (reuse #2) — replaced the `cat ` exec_command call with the existing trait method; dropped the unused `shell_quote` import and the magic 5s timeout. +3. **Redundant re-validation** in `validate_agent_output_sources` (efficiency #6) — bound the first error and reused it instead of calling `validate_response_text` again just to `.expect_err(...)`. +4. **`OneShotCompletion::actual_speed` was always `controls.speed`** (quality #2, reuse #4) — collapsed the three `actual_*` fields into a single `model: ModelRef`, removing the manual reconstruction one line later. +5. **`extract_status_fields` wrapper deletion** (reuse #1, quality #5) — renamed `extract_status_fields_loose` → `extract_status_fields` in `structured_output.rs`; agent.rs re-exports it so existing call sites (including `prompt.rs`) work unchanged. +6. **`find_json_objects` visibility** (quality #6) — tightened from `pub(crate)` to `fn`. +7. **`last_touched_file` cheap accessor** (efficiency #5) — added a tiny helper that grabs only `.last`; the repair loop no longer clones+sorts the full `touched` set each iteration. +8. **`ROUTING_KEYWORD` constant** (quality #8) — extracted the literal `"routing"` to a const used by `parse_node_output_schema`. + +**Skipped:** +- Folding `Error::OutputSchemaValidation` into `Error::Validation` — `Validation` adds a `"Validation error: "` prefix that would change user-visible failure messages and downstream event content. Not worth the drift for one extra enum variant. +- Other suggestions (request structs for multi-param functions, deduping the two repair loops across `one_shot`/`run`, moving `validate_agent_output_sources` into `structured_output.rs`, collapsing `StructuredOutputErrorKind`) — judged as larger refactors with no clear correctness payoff for this iteration. + +**Verification:** `cargo nextest run -p fabro-types -p fabro-workflow` → 1461 tests pass; `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` clean; `cargo +nightly-2026-04-14 fmt --check --all` clean. \ No newline at end of file